Agentic WFM Protocol
The Agentic WFM Protocol is a design for running a workforce management function's planning cycle with software agents that propose and people who sign. It sets out thirteen steps from intake to learning, gives each a role, a mode of computation, a grade on what it produces and, where a plan would change, a gate at which a named person signs. It extends the AI Agent Teams for Workforce Management series from a team pattern to an operating protocol, and adds three things the team pattern leaves open: the contracts that let a tool built in one place run on data that may not leave another, a first worked implementation in the daily forecast cycle, and a measurement of value by layer that decides where automation has been earned rather than assumed. This page is the hub of the series: it lists the steps, states the principles, and points to the page on which each step is documented.
The thirteen steps
The protocol is a chain, not a single agent. Each step has a fixed input and output, so that a step can be tested, replaced or handed to a person without rebuilding the others; this is the "workflow" rather than "autonomous agent" end of the design space that agent builders recommend wherever the path through a task is known in advance.[1] The mode column matters as much as the order: arithmetic is done by deterministic code, ranges by seeded simulation, and only reading and writing by a language model.
| # | Step | Input → output | Mode | Documented on |
|---|---|---|---|---|
| 1 | Trigger and intake | A question, an event or a cadence tick → an intake entry with an owner and the decision it changes | Language model classifies; a person routes | Work Intake for Planning and Analytics Teams, The Question Register and Knowledge Base |
| 2 | Data preparation and contract validation | Raw extracts in any shape, one mapping per source → canonical files and a gap list | Language model drafts mappings; validation is deterministic | The Shape File Bridge, WFM Data Governance and Quality |
| 3 | Baseline forecast | History, calendar and events → volume and handle time by interval and channel | Statistical, deterministic given a seed | The Short-Term Forecasting Loop with an Agent Team, Three-Step Forecast Build, The Intelligence Feed |
| 4 | Assumption register | Gaps and planner inputs → one row per assumption with value, grade, owner and sensitivity | Language model drafts; a person owns | The Assumption Register |
| 5 | Requirement | Forecast and service targets → required productive hours and FTE | Deterministic queueing model per channel | Erlang-A, Erlang Sensitivity and the Staffing Cliff |
| 6 | Supply and mobilization | Roster, hiring waves, ramp and attrition → productive hours available by week | Deterministic | Capacity Planning Cycle, Recruiting Pipeline and Capacity Planning |
| 7 | Scenarios and simulation | Central case and the register's ranges → a P10 to P90 band and a sensitivity ranking | Probabilistic, seeded and reproducible | Level 4: Planning in Distributions, The Assumption Register |
| 8 | Outlook and report | Model outputs → an answer-first report, a platform-loadable file, and "what would change the answer" | Language model writes; numbers are injected from files, never retyped | The Question Register and Knowledge Base, Long-Term Planning Agents and the Plan of Record |
| 9 | Critic and ground truth | The report, and a synthetic world with planted effects → pass or fail per check and a catch rate | Deterministic checks and a separate evaluator | The Agent Overseer, Hypothesis Testing with Agent Teams |
| 10 | Human sign-off | A proposal → a signed version | A person | Human Gates and Number Grades |
| 11 | Publish | A signed version → the plan of record and the platform of record, read-only first | Deterministic | Forecast Lock Process, Long-Term Planning Agents and the Plan of Record |
| 12 | Monitor actuals | Actuals against the plan → variance and drift flags | Deterministic, with statistical process control | Forecast Accuracy Metrics, Forecast Bias Detection and Correction |
| 13 | Learn | Actuals → updated parameters and a score for every layer of the published plan | Statistical update | Forecast Value Added in Workforce Management, Forecast Vintages |
Two rules bind the chain together. A drift flag at step 12 never rewrites a plan; it raises an entry at step 1, so that every change to a published number passes through the same gate as the number itself. And a step whose output is a number attaches a grade to it, under the scale set out on Human Gates and Number Grades; a computed figure inherits the weakest grade among its inputs.
Ten operating principles
- Agents propose, people sign. The sign-off is built into the file format, so an unsigned plan cannot load (Human Gates and Number Grades).
- Every number carries a grade. The scale is the one on Human Gates and Number Grades.
- Deterministic arithmetic, probabilistic range, language for language. One implementation of each calculation, used by every tool and every step.
- Questions pull analysis; schedules do not push it. An intake entry that names no decision it changes waits; it is not worked.
- Validate the contract before modeling. Definitional qualifiers, such as whether a handle time is worked or elapsed, are data fields rather than footnotes.
- A supply break is not a demand change. Shrinkage is applied once, on the supply side.
- Test against planted truth. The question is whether a step recovered an effect deliberately placed in synthetic data, not whether it ran.
- When a check fails, suspect the instrument first. A surprising result is more often a defect in the measurement than a fact about the operation.
- Persist and learn before simulating. State that is not written back is not learned.
- Release autonomy per class of action on a published catch rate, read before write. The evidence releases a gate; a maturity label does not.
The first and last principles restate, for a planning function, the human-factors finding that the level of automation should be set by the consequences of error and the measured reliability of the automation.[2] The fifth principle is the data-validation step that machine-learning operations practice requires before an automated pipeline retrains.[3]
Two zones: method crosses, data stays
Methods, contracts and single-file tools are built and tested on synthetic data with planted effects, then carried to the operating zone, where they run on real data on the function's own machines; nothing crosses back except files that follow an agreed contract and carry no identifying names. The contracts, the mapping files that adapt a platform's export to them, and the one-file tools are described on The Shape File Bridge. The protocol adds one device that page leaves implicit:
- A name guard. Every file produced for the build zone is scanned against a machine-local list of real names before it is committed or published, and the published surfaces are scanned again on a schedule. A missing list fails the check rather than passing it.
The daily forecast cycle: the first worked implementation
The protocol's first implementation runs the chain on the short-term forecast. Intake, preparation, reforecast, review, publication, monitoring and learning run daily; requirement, supply and range are refreshed weekly and monthly; a planted-truth critic (step 9) is not yet part of it. The daily order is the one a forecaster's morning already follows, and it implements the loop described on The Short-Term Forecasting Loop with an Agent Team: prior-day performance; logged events; signals extracted from email, notes and transcripts and graded by their source; a reforecast that proposes changes to the forecast of record; an owner's review with a signature; publication; and a learning step that scores what was published once actuals arrive. A weekly step refreshes a 26-week outlook and states what changed and why, and a monthly step locks the next month and turns the staffing runway, after attrition, into hiring asks dated by when hiring must start.
Built and tested on a synthetic estate of 60 queues with two years of history and planted effects, the cycle's most useful result was negative. On a quiet estate, the forecast of record was hard to beat: corrections learned from recent error and applied everywhere chased noise and did not improve accuracy, while the same corrections applied only where they had already beaten the record in earlier windows did no harm and occasionally helped. The largest gains, still small, came from explicit changes with a source: in the forward check, a migration overlay and applied intake signals moved volume error from 10.75 to 10.50 percent, while a rumor kept off the forecast would have worsened its queues by almost nine points. The measurement behind that result, and the rule it produced, that automation is earned queue by queue on its own record, are the subject of Forecast Value Added in Workforce Management.
See it run
A worked run of the protocol on synthetic data is described on The Agentic WFM Cycle: a Worked Demo: twelve items of material through intake and triage, a signed daily forecast, views regenerated from the record, a daily brief of decisions with owners and due dates, and a learning step that scores the published forecast nine weeks later. The run itself can be explored at agentic-demo.wfmlabs.com, including an animated map that follows single pieces of information from the moment they are mentioned to the staffing decision they inform.
Maturity path
| Stage | What runs | Evidence that unlocks the next stage |
|---|---|---|
| 1. Assist | An assistant helps one analyst; calculations move from spreadsheets into parameterized, graded, versioned notebooks | Analyst hours moved off spreadsheet assembly; a definitions register with identifiers |
| 2. Chain | The fixed chain runs from intake to report on each book; the human gates are exercised; the planted-truth tests run on every change | Around thirty consecutive cycles on cadence; tests green; grades moving from [A] toward [C] and [M] (as on Human Gates and Number Grades); a first catch rate published |
| 3. Supervised autonomy | Scheduled runs with learning persisted; read connections to the data warehouse and the platform of record; drafts published for a person's signature | Forecast intervals that cover actuals at their stated rate; a catch rate per class of action above threshold; a complete audit log |
| 4. Gated write-back | Low-risk classes of action written back to the platform of record, one class at a time | The per-class catch rate sustained; rollback tested; an independent validator's sign-off |
The stages are ordered by autonomy; which work is handed over first is a separate question, answered on A Roadmap Pattern for Agent Teams in WFM and Agent Team Readiness for a Planning Process. The stages are consistent with the AI Risk Management Framework's functions, under which risks are mapped and measured before they are managed,[4] and the independent validation required at stage 4 follows model-risk supervisory practice.[5]
What would change this
The thirteen steps are a design validated against prototypes and synthetic estates, not yet a record of many functions in production. Evidence that a step's separate existence adds nothing, for example a critic that never catches what the human gate would, would argue for merging it. The daily cycle's negative result is specific to a quiet estate; an operation whose demand shifts frequently could find learned corrections earning their place on most queues, and the earned-by-queue rule is designed to discover that rather than assume either answer.
How this connects
The roles that run the steps are those of The Agent Team Model, over the ledgers of Living Ledgers; the planner gate and the grade scale are on Human Gates and Number Grades; the daily forecast cycle above is an implementation of the daily clock described on The Short-Term Forecasting Loop with an Agent Team. The decision to hand a process to agents at all is a separate gate, The Agentic Handover Gate, and the broader path is mapped on The Agentic Journey Map and Agentic AI Workforce Planning.
Maturity Model Position
Running the chain's steps as separate, tested programs with graded numbers and signed versions is Level 3 practice on the WFM Labs Maturity Model™. Planting effects in synthetic data to measure a catch rate and scoring every layer of a published plan are Level 4 practice; releasing autonomy per class of action follows the condition set out on Human Gates and Number Grades.
See Also
- AI Agent Teams for Workforce Management — the team pattern this protocol extends
- The Agentic WFM Cycle: a Worked Demo — the forecasting steps of the protocol, run end to end on synthetic data
- Forecast Value Added in Workforce Management — the measurement that decides where automation is earned
- Human Gates and Number Grades — the grades and the gates
- The Shape File Bridge — files that cross the data boundary
- The Short-Term Forecasting Loop with an Agent Team — the daily clock
- Work Intake for Planning and Analytics Teams — step 1
- The Agentic Journey Map — the wider path
References
- ↑ Anthropic (2024). "Building effective agents". Anthropic Engineering. anthropic.com/engineering/building-effective-agents.
- ↑ Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). "A model for types and levels of human interaction with automation". IEEE Transactions on Systems, Man, and Cybernetics — Part A 30(3), 286–297. doi:10.1109/3468.844354.
- ↑ Google Cloud (2024). "MLOps: Continuous delivery and automation pipelines in machine learning". Cloud Architecture Center. docs.cloud.google.com.
- ↑ National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. doi:10.6028/NIST.AI.100-1.
- ↑ Board of Governors of the Federal Reserve System (2011). SR 11-7: Supervisory Guidance on Model Risk Management. federalreserve.gov.
