Forecast Vintages

From WFM Labs
A forecast table with intervals, actuals and forecasts, and one empty gold column: the date each forecast was made.

Forecast vintages are the successive versions of a forecast for the same future interval, each stamped with the date on which it was produced. A forecast table that stores one forecast per interval — overwriting the previous version each time the forecast is refreshed — cannot say how good the forecast was at any horizon, cannot separate forecast error by the layer that produced it, and cannot resolve the variance decomposition by horizon. The missing column is small and its absence is total: without it the forecast can be reported but not audited.

The distinction the page turns on is between two kinds of evaluation. A backtest re-runs a model over history from many origins and establishes what the model can do; it needs no archive. A vintage record establishes decision-time performance — how good the number was that the hiring, scheduling or intraday decision was actually made against — and that cannot be reconstructed after the fact. Most forecast evaluation in service operations is of the first kind, and it silently answers a different question from the one the operation is asking.

This page defines the vintage, lists what depends on it, gives the storage grain and retention that hold it, and places the borrowing in its statistical lineage. The horizon-resolution step of the variance decomposition is at Doubly Stochastic Arrivals and Demand Variance Decomposition; the evaluation of probabilistic forecasts at Probabilistic Forecasting; horizon-specific accuracy measurement at Forecast Accuracy Metrics; the governance of forecast revisions at Reforecast and Rolling Forecast Methodology.

What a vintage is

A forecast for the 09:00 interval next Tuesday is made many times: in the annual plan, in the monthly capacity cycle, in the weekly schedule build, on the day before, and intraday. Each is a different number, produced by a different method on different information, and each is the forecast that some decision was made against — hiring against the monthly one, shifts against the weekly one, breaks against the intraday one. A vintage is one of those versions, identified by the interval it forecasts and the date on which it was made. The as-of date is the vintage's key.

Most forecast tables do not hold vintages. They hold the current forecast for each future interval, and when the forecast is refreshed the previous value is overwritten. Once the interval has passed, the table holds the actual and the last forecast made before it — typically the intraday or day-before value — and nothing else. Every earlier version, including the ones the expensive decisions were made against, is gone.

What depends on it

  • Error by horizon. Whether the forecast was good enough for hiring (a months-out horizon), for scheduling (weeks), or for intraday (hours) are three different questions with three different answers, and a table with one forecast per interval can answer only the last.
  • Error by layer. A monthly capacity forecast, a weekly schedule forecast and an intraday reforecast are produced by different people and methods. Attributing error to the layer that produced it — the discipline the Capacity Planning Cycle describes at Level 4 — requires each layer's forecast to survive.
  • The variance decomposition by horizon. The three-way split described at Doubly Stochastic Arrivals and Demand Variance Decomposition runs on a single locked-horizon forecast and is usually available without an archive; what vintages add is its final step, mapping variance to horizon — the resolution curve that shows at which lead time forecast investment stops buying anything. Without them the split can be run at a single horizon and the curve cannot be drawn.
  • Calibration drift. A probabilistic forecast is evaluated by whether its stated bands contained the actual at the stated rate — coverage — over many intervals; Probabilistic Forecasting gives the scoring. Coverage on held-out data tests the model's bands; coverage of the bands actually issued at each horizon tests the forecast the operation planned against, and only the second needs vintages.
  • Model cards. Level 4: Planning in Distributions describes backtesting a model on holdout periods, which needs no archive. A model card built only from holdout backtests documents what the model can do, not what the published forecast did; the calibration history that distinguishes them is a vintage record.

The consequence is stronger than "the forecast cannot be improved." An operation without vintages can know its forecast's quality at one horizon but not how that quality decays with lead time — and it is the decay curve, not the single-horizon figure, that says whether investment belongs in the forecast or in supply elasticity. That decision is made on a single-horizon figure when the decay curve is what bears on it.[1]

The schema

The fix is a table keyed on three things rather than one.

Forecast table with vintages
Key Meaning
Target interval The interval being forecast
As-of The date and time the forecast was produced — equivalently, the horizon
Layer or method Which planning layer or model produced this version

Each row holds the forecast value — and, for probabilistic forecasts, its bands — and the refresh appends rather than replaces. The actual is joined at the interval key when it arrives. Not every refresh needs keeping: the practical set is one snapshot per planning lead time — the horizons Doubly Stochastic Arrivals and Demand Variance Decomposition uses are thirteen weeks, four weeks, one week and one day — which is a handful of rows per interval rather than one per refresh. Retention follows the forecast-version row of the table at WFM Data Governance and Quality: current plus ninety days hot, two years warm for accuracy trending. Long-horizon vintages are therefore evaluated from the warm tier: a thirteen-week vintage outlives the hot window before its actual arrives.

Where a platform cannot store vintages natively, a nightly extract of the current forecast table, stamped and appended to an archive, is sufficient.[1] Reforecast and Rolling Forecast Methodology gives the governance form of this record — revision trigger, approver, forecast lock; this page gives the storage grain the record has to be kept at to remain evaluable.

The lineage

The word is borrowed from macroeconomic forecasting, where two archives are kept. Real-time data sets hold every vintage of every data series as it stood on each date, because the data a forecaster had at the time is itself revised later and evaluating a forecast against final data rather than the data available when it was made is a recognized error;[2] survey archives such as the Survey of Professional Forecasters, maintained by the same institution, hold every vintage of the forecasts.[3] A service operation's forecast archive is the second object at a smaller scale, and this page's use of vintage for successive forecasts of one interval is a borrowing of the term rather than its original sense. The first archive's lesson applies too: where actuals are corrected after the fact, the forecast should be judged against the actual as first recorded.

Failure modes

The following recur in planning environments.[1]

  • Overwriting on refresh. The design default of most planning tables, and the whole problem.
  • Keeping only the last forecast before the interval. The data trap Doubly Stochastic Arrivals and Demand Variance Decomposition names as forecast vintage drift: accuracy is reported at the horizon where it matters least.
  • Storing vintages without the layer. Versions survive but cannot be attributed, so error by layer is still impossible.
  • Evaluating against revised actuals. Where actuals are themselves corrected later — reclassified contacts, late-arriving data — the forecast is judged against a number that did not exist when it was made; this is the error the real-time data literature was built to prevent.[2]
  • Deferring the column until the new platform. A stamped nightly extract costs nothing and starts the history now; a platform migration that arrives in two years arrives with two years of history if the extract exists and none if it does not.

Maturity Model Position

Forecast accuracy is measured from Level 2 on the WFM Labs Maturity Model™, but only at the last horizon; keeping vintages is what Level 3 adds — Reforecast and Rolling Forecast Methodology places version tracking there — and it is what every Level 4 practice reads from: the horizon-resolution curve, error by layer, calibration drift and the calibration history in a model card. It is the cheapest column in a planning environment and the one most do not have.

See Also

References

  1. 1.0 1.1 1.2 Practitioner observation from capacity-planning environments in service operations; a consistent pattern rather than a measured result.
  2. 2.0 2.1 Croushore, D., & Stark, T. (2001). A real-time data set for macroeconomists. Journal of Econometrics, 105(1), 111–130.
  3. Federal Reserve Bank of Philadelphia. Survey of Professional Forecasters. Begun in 1968 by the American Statistical Association and the National Bureau of Economic Research; conducted by the Philadelphia Fed since 1990. https://www.philadelphiafed.org/surveys-and-data/real-time-data-research/survey-of-professional-forecasters