The Agentic WFM Cycle: a Worked Demo

From WFM Labs

The Agentic WFM Cycle: a Worked Demo is a public, runnable demonstration of the Agentic WFM Protocol on synthetic data. Twelve pieces of material arrive as the files people send, pass through intake and triage, feed a daily forecasting cycle that a person signs, regenerate the views leaders read, and are scored against what happened nine weeks later. Nothing in it is mocked: every number on its pages is produced by running code on a synthetic estate of 60 queues with two years of history and effects planted in the data. Two steps that a person performs in practice, triage and the forecast owner's signature, are scripted so that the demo runs unattended, and are labeled as such. The demo is at agentic-demo.wfmlabs.com; its animated map, which plays the day and follows single pieces of information end to end, is at agentic-demo.wfmlabs.com/flow.html.

What the demo shows

The protocol runs a planning cycle as a chain of fixed steps, each with a defined input and output, with agents proposing and people signing at the points where a plan would change; a fixed chain of this kind is the form agent builders recommend wherever the path through a task is known in advance.[1] A description of such a chain is easy to agree with and hard to picture. The demo exists to make it concrete: it shows what each step consumes and produces, where a person decides, what a leader receives each morning, and how the cycle is held to account for what it published.

The estate is a synthetic contact operation: 60 queues across four product lines and three channels (voice, chat and email), two years of daily history, recurring storms, outages and absences, and planted effects including a client migration in three waves, a level shift on one queue, an unlogged absence, three scheduled changes announced through intake, and a rumor that is false. Because the planted effects are known, the demo can check whether each step recovered them, which is the protocol's standard of evidence rather than whether a step ran.

The day, step by step

# Step What happened in this run (Day 1, as of 4 October) The method
1 Material arrives 12 items: five emails (one with a wave-plan spreadsheet), a client-review deck, Word notes, a PDF ops note, a stand-up transcript and three chat exports Work Intake for Planning and Analytics Teams
2 Intake Each file read (tables to CSV), scrubbed against a list of real names, and written as an ask with a suggested route; nothing routed automatically Work Intake for Planning and Analytics Teams
3 Triage (scripted) 8 asks to the forecast, 1 urgent request for a view, 3 closed because they change no decision Work Intake for Planning and Analytics Teams
4 Bridge and signals The 8 forecast asks become the cycle's intake; 8 signals graded by source: 2 applied, 1 corroborating, 1 rumor held, 1 waiting on its owner, 2 explaining past days, 1 already in the plan Source grades on intake signals
5 The daily cycle, review and publish (scripted signature) 93,091 queue-channel-days validated against the contracts; yesterday's service level 80.6 percent against an 82 percent target; 4 queues moved by 1 percent or more; the daily brief is written with the proposal (11 decisions, 3 due today); 16 queues to look at and 44 approvable on their earned record; the signed forecast published as 23,114 queue-channel-days, each layer kept for later scoring The Short-Term Forecasting Loop with an Agent Team, Human Gates and Number Grades
6 Views regenerate A Staffing Outlook rebuilt from the signed record, the weekly and monthly plans (47 hiring asks) and an export for the workforce management platform The Shape File Bridge, Capacity Planning Cycle
7 Handoffs Every ask records what it became: the signal's grade and fate, or the view that answered it Work Intake for Planning and Analytics Teams
8 Nine weeks later The forecast published on Day 1 scored 10.50 percent volume WAPE against actuals; the forecast of record alone would have scored 10.75 percent Forecast Value Added in Workforce Management
9 Guard Every text file and path in the demo scanned against a list of real names before it is published: none found Agentic WFM Protocol

Three threads, followed end to end

A client onboarding. A client-review deck from an account lead announces a new division of about 1,200 users onboarding from 9 November, with volume expected up about 15 percent on three queues. The signal is graded as coming from a named owner, which on its own would wait for confirmation. A chat message from a different role, in a different item, says the same thing; the second, independent source makes it applicable, and the cycle adds a dated overlay of +15 percent on the three queues from 9 November. Those queues are the day's largest movers on the review screen (about +9 percent over the 26-week horizon), the brief lists the change under risks ahead, and the learning step later scores the applied-signals layer of the published forecast, which carries this change and a one-day 40 percent holiday reduction on eight queues together and moved error from 10.69 to 10.50 percent.

A rumor. An email forwarded with no author claims that a client will move 3,000 users to self-service on 12 October, cutting one queue's voice volume by 15 percent. With no accountable source it is graded unattributed and held as a rumor; the brief asks the forecast owner to find someone who will stand behind it or drop it. Nobody does. Once its date passes the brief closes it, and the learning step shows what applying it would have done: on the 55 queue-days it named, error would have risen from 9.4 to 18.3 percent. The source grade is what kept that hit off the forecast.

The same rule has a cost. A platform change raised at a stand-up, graded as unconfirmed and held for its owner, would have lowered error on the 282 queue-days it named from 11.1 to 9.95 percent had it been applied. A gate that keeps rumors off also delays true changes until someone stands behind them, and the learning step reports both. No significance test was run on these differences.

An executive's question. An executive asks for one page by Monday showing staffing against requirement for every product through the wave plan, then weekly. Triage marks it urgent. It is not a forecast signal, so it waits for the signed record; once that is published, a Staffing Outlook is regenerated from it with no hands, turning the monthly lock's hiring asks into hiring waves. The answer: the tightest week is the week of 8 February, right after the final migration wave, short by about 58 FTE (5.3 percent) even with the hiring asks in place. The ask is handed off with the view and an update note written for the slide notes; next week the view rebuilds from the next record.

The daily brief

Each run also writes a one-page daily brief: five headline lines, every decision the cycle needs from a person with an owner, a due date and its evidence, the risks ahead, what the cycle did on its own, and three trend charts. It is built only from the run's outputs and decides nothing itself.

The two briefs in the demo show the loop over time. On Day 1 the brief carries 11 decisions, three due that day, and reports the tightest week ahead as 80 FTE short at P50 on the supply plan alone; the Staffing Outlook's figure of about 58 FTE is smaller because it counts the hiring asks that would land by then. On Day 63 the synthetic world has acted on none of the hiring asks, so three of them are overdue and lead the list, while the accuracy line reports that the forecast published on Day 1 scored 10.50 percent against 10.75 percent for the record alone. Two signals that reached their start dates unconfirmed have left the decision list and are scored by the learning step instead.

What is scripted and what is not

  • Scripted: triage applies a fixed table, the forecast owner's signature approves every queue, and Day 63 reuses Day 1's intake because the demo sends no new mail. In practice a person triages and signs, and both steps exist precisely because a person must produce something, a route or a decision per queue.[2]
  • Not scripted: extraction from the files (by fixed rules here; a deployment would use a language model to fill the same schema, and every step after extraction is unchanged), the signal grades, the reforecast, the review packet and its refusal rules, the regenerated views, the brief, and the learning step. Their numbers are what the code produced.
  • Not shown: requirement in intervals, schedules, real-time actions and placement. The demo covers the forecasting half of the protocol.

What would change this

The estate is synthetic and deliberately quiet, so the small size of the measured gains, a quarter of a point between the record and the published forecast, says as much about the estate as about the method; an operation with frequent real changes would expect the source-graded signals to matter more. The scripted steps understate the cost of the human gates: in practice the review takes a forecaster's time, and forecast value added is the measure that decides whether it earns it.[3]

How this connects

The demo is the worked example for the Agentic WFM Protocol and runs its forecasting steps at a daily cadence, as The Short-Term Forecasting Loop with an Agent Team describes. Its intake follows Work Intake for Planning and Analytics Teams; its signal grades, review and guarded publish follow Human Gates and Number Grades; its views are regenerated from the published record as on The Shape File Bridge; its monthly lock and hiring asks belong to the Capacity Planning Cycle; and its learning step is the measurement set out on Forecast Value Added in Workforce Management, whose worked example is this demo's synthetic estate.

Maturity Model Position

The demo runs at the practices the WFM Labs Maturity Model™ places at Level 3 (versioned forecasts, graded numbers, signed versions, views regenerated from files) and Level 4 (planted-truth testing, scoring each layer of what was published, earning autonomy on that evidence).

See Also

References

  1. ↑ Anthropic (2024). "Building effective agents". Anthropic Engineering. anthropic.com/engineering/building-effective-agents.
  2. ↑ Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). "A model for types and levels of human interaction with automation". IEEE Transactions on Systems, Man, and Cybernetics — Part A 30(3), 286–297. doi:10.1109/3468.844354.
  3. ↑ Gilliland, M. (2015). Forecast Value Added Analysis: Step by Step. SAS Institute white paper. sas.com.