Forecast Value Added in Workforce Management

From WFM Labs

Forecast value added (FVA) is the change in a forecast performance measure that is attributable to one step of a forecasting process, measured against the step before it or against a simple benchmark. A step adds value when the forecast after it is more accurate than the forecast before it, and destroys value when it is less accurate. The idea comes from demand planning, where it was developed and popularized to test whether the layers of review, override and consensus that organizations add to a statistical forecast earn the effort they cost.[1] In workforce management the same question applies to every layer a volume and handle-time forecast passes through before staff are scheduled against it, and an agentic planning function needs the answer before it decides which layers to automate. This page sets out the method, the choices that matter for staffing forecasts, the evidence on human adjustment, a worked example on a synthetic estate, and the use of FVA as the condition that releases automation.

The stairstep report

FVA is reported as a stairstep: the process is written down as an ordered list of layers, each layer's forecast is scored against the same actuals over the same period, and each layer's value is the difference between its error and the error of the layer before it. With WAPE (weighted absolute percentage error) as the measure, FVA for a layer is the earlier layer's WAPE minus the layer's own, in percentage points, so a positive figure means the layer helped.

A typical short-term workforce forecast has five layers:

  1. The benchmark — a naive forecast (the same day last week or last year) or, in a running operation, the forecast of record before today's changes.
  2. The statistical update — the model's revision on new actuals, including any level or trend correction learned from recent error.
  3. Overlays for known change — explicit adjustments with a stated source, such as a client migration's contact ratio or a confirmed calendar event.
  4. Signals — adjustments taken from intelligence: emails, meeting notes and stand-up remarks that announce a change.
  5. Human edits — the forecaster's or owner's final override at sign-off.

The order is the order in which the process applies the layers, not an order of importance. A layer is only measurable if the process keeps it: a forecast published as one number, with the contributions of the update, the overlays, the signals and the edit merged into it, cannot later be decomposed. The design rule that follows is that the published forecast stores each layer's contribution beside the final figure, so that every layer can be scored when actuals arrive, weeks after the decision.

Choosing the benchmark and the measure

The benchmark. Demand-planning practice compares every step with a naive forecast (see Naive and Seasonal Naive Forecasting), which tests whether the forecasting process as a whole beats doing nothing clever.[1] Ratio measures such as MASE and Theil's U express the same comparison against a naive forecast as a single number. A workforce function that already runs a forecast also needs a second comparison, against its own forecast of record, because the operational question is whether today's change should be made at all. The two benchmarks answer different questions and both are worth reporting.

The measure. Mean absolute percentage error is unstable for the low-volume intervals and small queues that workforce forecasts contain, because percentage errors are undefined or extreme where actuals are near zero;[2] WAPE, which divides total absolute error by total actual volume rather than dividing interval by interval, is far less exposed to that defect. For staffing, the more decision-relevant quantity is workload, volume multiplied by handle time, since a volume forecast that is right while its handle time is wrong still staffs the floor wrongly. Bias, the signed error, is reported beside WAPE: two layers can have the same WAPE and opposite bias, and a persistently negative bias understaffs systematically. The measures and their trade-offs are covered on MAPE WAPE and Forecast Bias and Forecast Accuracy Metrics.

The horizon. A layer is scored at the horizon at which its forecast was used. A daily reforecast that sets tomorrow's intraday plan is scored on tomorrow; one that feeds a four-week schedule is scored over those four weeks. Scoring a layer at a horizon it was not used for flatters or penalizes it for decisions it did not inform; the bookkeeping that keeps each forecast version with its decision date is described on Forecast Vintages.

Walk-forward and forward checks

Two scoring regimes are needed, and they answer different questions.

Walk-forward backtest. To ask whether a layer would have added value in the past, the process is replayed from a series of forecast origins, each using only the data available at that origin, and scored on what followed. This is rolling-origin evaluation, the standard out-of-sample test for forecasting methods (see Model Evaluation and Validation); it avoids the in-sample optimism of fitting and scoring on the same period.[3][4] A backtest can score layers that a machine applies, but not human edits that were never made in the past.

Forward check. To ask whether the forecast actually published added value, the stored layers of a published version are scored once actuals arrive. The forward check is the only regime that can score human edits and the signals accepted at review, and it is the one an owner should be shown, because it reports on decisions the owner actually took.

Is the difference real? Layer-to-layer differences in WAPE are often a few hundredths of a point, and FVA results over short periods may be due simply to randomness.[1] A difference should be tested before it is acted on, for example with a paired test on the daily difference in the two layers' losses, aggregated across queues because their errors are correlated, of which the Diebold–Mariano test is the standard form.[5] A finding that a layer adds nothing measurable is itself a result, and frequently the most useful one.

What the evidence says about judgmental adjustment

Field studies of company forecasts, discussed more broadly on Judgmental Forecasting, find that people adjust statistical forecasts often, that many adjustments are small, and that small adjustments add little value or reduce accuracy; larger adjustments and downward adjustments were more often beneficial, while upward adjustments, frequently optimistic, were more often harmful.[6] The recommendations that follow from that work, that each adjustment carry a recorded reason and that adjustments be measured against the forecast they replaced, are the same two requirements an FVA report imposes.[7] The evidence is from product demand forecasting; whether workforce forecasters behave the same way is plausible but has been studied far less, and FVA is the instrument by which a function can find out for itself.

Worked example: a synthetic estate

The following figures come from a synthetic estate built to test a daily forecasting cycle: 60 queues in four product lines, voice, chat and email channels, two years of daily history with recurring storms, outages and absences, and planted effects: a client migration in three waves, a level shift on one queue, a hurricane, a chat outage, an absence that was never logged, three scheduled changes announced through intake, and a rumor that is false. They are illustrative of the method, not of any operation. The same estate runs end to end, from intake to this learning step, on The Agentic WFM Cycle: a Worked Demo.

Backtest. 44 weekly origins, a 28-day horizon, workload WAPE on every queue, channel and day:

Layer Workload WAPE FVA against the record FVA against the previous layer Bias
Forecast of record 11.01% — — −1.89%
+ learned volume level, applied everywhere 11.04% −0.03 pts −0.03 pts −2.36%
+ learned handle time, applied everywhere 11.02% −0.01 pts +0.02 pts −2.37%
+ migration contact ratio 11.00% +0.01 pts +0.02 pts −2.34%
Alternative: learning only where earned, plus migration ratio 10.97% +0.04 pts +0.03 pts against the row above −1.89%

Forward check. The forecast published on the first day of the test, scored against the nine weeks of actuals that followed:

Layer of the published forecast Volume WAPE Bias
Forecast of record 10.75% −2.36%
+ learned level 10.75% −2.35%
+ migration overlay 10.69% −2.13%
+ applied signals 10.50% −2.13%
+ human edits 10.50% −2.13%

No significance test was run on any of these figures.

Signals not applied. The forward check also scores what was kept off the forecast. On the days and queues it would have touched, the false rumor, graded as unattributed and therefore not applied, would have raised WAPE from 9.4 to 18.3 percent. A signal implied by a stand-up remark, graded as unconfirmed and held back pending the owner's confirmation, would have lowered WAPE from 11.1 to 9.95 percent: the gate cost 1.1 points on those days. A gate that keeps bad signals off also keeps some good ones off, and the forward check is how that cost is seen.

Reading the example. Three conclusions follow. First, the forecast of record was hard to beat: on an estate where most queues had nothing real to learn, no learned layer moved accuracy by more than a few hundredths of a point in either direction. Second, the only layer that moved the published forecast by more than a tenth of a point was the applied signals (10.69 to 10.50 percent); the migration overlay's 0.06 points is the same order as the differences treated as noise above. Third, the human edit layer added nothing because no edits were made; a scripted reviewer approved every queue, which is the case FVA is designed to expose, not to praise.

Earned automation

FVA becomes an operating rule when it is used to decide, queue by queue, which layers run and which proposals may pass without review.

Earned learning. A queue uses its learned corrections only after they have beaten the forecast of record in earlier walk-forward windows on that queue, or after a confirmed shift in its level; elsewhere the record stands. In the example, the walk-forward part of this rule scored 10.97 percent against 11.00 percent for the same corrections applied everywhere; with no significance test, a difference of three hundredths should not be read as a gain. The rule's value is that it left the record in place on 55 of the 60 queues.

Earned auto-approval. A queue's proposal may be approved without individual review when the proposal moves the forecast by less than 1 percent and has beaten the record, or come within half a point of it, in at least 80 percent of its last twelve windows, never losing one by more than 2 points. In the example, 54 of 60 queues met the rule. Most of those queues qualified by staying within that tolerance rather than by beating the record, so auto-approval releases review of small changes whose cost has been measured and bounded, and keeps every large change in front of a person.

Both rules apply the principle set out on Human Gates and Number Grades, that a gate is released per class of action on published evidence rather than at a maturity level, with FVA as the evidence.

Common failure modes

  • Scoring only the total. An estate-level gain can hide losses on many small queues offset by a large gain on one; FVA is reported by queue or segment as well as in total.
  • Scoring at the wrong horizon. See above; the horizon is the one at which the forecast was used.
  • Not storing the layers. Once the contributions are merged into one published number, the forward check is impossible.
  • Scoring people only when they act. Human edits are scored on the queues and days they touched, and the absence of edits is reported as such, not as a neutral layer.
  • Declaring wins inside noise. A layer kept on the strength of a difference no larger than the period-to-period variation will be removed by the next period's numbers.
  • Ignoring cost. A layer that adds a tenth of a point but takes a forecaster's morning may not be worth keeping; FVA measures value, not net value.

What would change this

The worked example is synthetic and its estate is deliberately quiet; an operation with frequent real shifts in demand would expect learned corrections to earn their place on more queues, and the earned-by-queue rule is designed to find that out rather than assume it. The thresholds in the auto-approval rule (1 percent, half a point, 80 percent of twelve windows, 2 points) are design choices that a function should set from its own forward checks. Evidence that workforce forecasters' adjustments behave differently from those in the demand-planning studies would change the expectation, though not the method.

How this connects

FVA is step 13 of the Agentic WFM Protocol, the learning step that closes the loop on what was published. It scores the reforecast proposed by The Short-Term Forecasting Loop with an Agent Team, the events of The Intelligence Feed and the edits made at the planner gate of Human Gates and Number Grades. The bookkeeping it depends on is that of Forecast Vintages, and its accuracy measures are those of Forecast Accuracy Metrics and MAPE WAPE and Forecast Bias. Systematic bias found by the report is treated on Forecast Bias Detection and Correction.

Maturity Model Position

Measuring forecast accuracy at all is Level 2 practice on the WFM Labs Maturity Model™. Keeping versions and scoring each layer of the process against the one before it is Level 3. Using FVA, scored walk-forward and forward, to decide which layers run and which proposals pass without review, queue by queue, is Level 4 practice.

See Also

References

  1. ↑ 1.0 1.1 1.2 Gilliland, M. (2015). Forecast Value Added Analysis: Step by Step. SAS Institute white paper. sas.com.
  2. ↑ Hyndman, R. J., & Koehler, A. B. (2006). "Another look at measures of forecast accuracy". International Journal of Forecasting 22(4), 679–688. doi:10.1016/j.ijforecast.2006.03.001.
  3. ↑ Tashman, L. J. (2000). "Out-of-sample tests of forecasting accuracy: an analysis and review". International Journal of Forecasting 16(4), 437–450. doi:10.1016/S0169-2070(00)00065-0.
  4. ↑ Hyndman, R. J., & Athanasopoulos, G. (2021). Forecasting: Principles and Practice (3rd ed.), section 5.10, "Time series cross-validation". OTexts. otexts.com/fpp3/tscv.html.
  5. ↑ Diebold, F. X., & Mariano, R. S. (1995). "Comparing predictive accuracy". Journal of Business & Economic Statistics 13(3), 253–263. doi:10.1080/07350015.1995.10524599.
  6. ↑ Fildes, R., Goodwin, P., Lawrence, M., & Nikolopoulos, K. (2009). "Effective forecasting and judgmental adjustments: an empirical evaluation and strategies for improvement in supply-chain planning". International Journal of Forecasting 25(1), 3–23. doi:10.1016/j.ijforecast.2008.11.010.
  7. ↑ Fildes, R., & Goodwin, P. (2007). "Against your better judgment? How organizations can improve their use of management judgment in forecasting". Interfaces 37(6), 570–576. doi:10.1287/inte.1070.0309.