The Agent Team Ladder: Alpha to Production
Part of the Planning Week chain · previous: AI Agent Program for a Resource Optimization Center · next: Anatomy of a ROC Standard
The agent team ladder is the sequence of four rungs — alpha in shadow, beta under a gate, pilot on the whole clock, production released by class — that an agent team climbs on one book of business before any class of its actions runs without a person signing. The page's claim is that a rung is defined by the evidence that exits it, not by the time spent on it, and that the clone pattern — what a team carries to the next book and what it never carries — turns one pilot into a consolidation instrument. The page produces the agent ladder card, template AL (AL-001 onward), and the clone checklist, and is worked in a session with Wiki:Packs/Agent Program Readiness (CP-OPS-005).
A card is opened per team per book, not per function: the three rosters of an agent program run different clocks, earn different evidence, and will normally sit on different rungs at once. Rungs are not phases of a plan. A later card may open before an earlier card's exit evidence is complete, and when it does the acceptance is written down with its price rather than the opening deferred.
The ladder is the program's route through The Agentic Handover Gate, not a second gate: the gate's five tests decide whether a process may be run by agents, and the ladder says in what order a team earns the evidence those tests read. The run-time gates of Human Gates and Number Grades are a third object and stay on at every rung, the last included.
Why rungs, and why evidence
Staged release is the ordinary discipline for a system whose reliability is not yet measured: expose it to a small, reversible population, measure, and widen only on the measurement.[1] Readiness scales in engineering make the same move, defining each level by the environment in which a component has been demonstrated rather than by the effort spent.[2] The ladder applies both, with the adjustment human-factors research requires: the person at the gate stays practiced at the work for as long as the gate stands.[3] Evidence rather than time defines a rung for a second reason: headline agent accuracy is inflated by retrying and by benchmark overfitting, and is rarely cost-controlled or repeated enough to carry an error bar, so a published score is a poor guide to deployable reliability.[4] Thirty consecutive daily runs is a count of measurements, not a waiting period. The counts on this page are design choices validated against one prototype; a function that measures a different count is entitled to use it.
The four rungs
| Rung | What the team does | Entry condition | Exit evidence (recorded on the card) | Who signs the exit |
|---|---|---|---|---|
| Alpha (shadow) | Runs its clock beside the planner; writes every ledger; publishes nothing; the planner's output stays the version of record | The check on Agent Team Readiness for a Planning Process passed for this process on this book; every ledger column cites a definition; an owner for every agent per Agent Identity and Custody | Twenty consecutive runs with no unexplained block; the team's output reconciled against the planner's each time; every carried assumption labeled; zero ungraded cells at load | The planner who owns the book, with the program owner |
| Beta (gated) | The gate is on: the team's proposal is the version a person confirms or amends, and the version of record is the team's; the overseer runs error injection and sampled deep verification | Alpha exited; the overseer named; the injection rate and sample size published before the first injected error | A first catch rate per The Agent Overseer, with its injection rate and sample size; thirty consecutive runs with the gate exercised on each; the grade profile measured against the planner-hour log | The person at the gate and the overseer; the program owner records it |
| Pilot (the whole clock) | The clock runs end to end on one real book as the version of record, including the long-cycle steps the earlier rungs only sampled; the clone onto a second book is run | Beta exited, or opened with a recorded acceptance; the clone checklist run on the second book | A record traceable to a real catch — a stale forecast version, an unmatched event, a contract-rule breach; the long-cycle step run once under its own gate; the second book's alpha exited on its own definitions with no carried benchmark | The seat that owns the clock, with the overseer |
| Production (released, by class) | For one named class of action the format no longer requires a signature, and the overseer's sampling replaces the approver for that class only | Pilot exited; the class named and its catch rate published per The Agent Overseer, on the condition stated once on Human Gates and Number Grades; the format change versioned; a reversal condition and a re-verification date logged | Continued catch rate per period; the reversal never silently skipped; the re-verification date met | The seat that holds the method, on the overseer's evidence |
One departure is deliberate, and is the ladder's own entry condition, not a second version of the release condition. Human Gates and Number Grades releases a class of action when a catch rate has been published for it. A function whose output is a plan others commit money against may want that catch rate held a second period before the signature goes; where it does, the stricter condition goes on the card, named as the function's.
Three properties matter more than any cell. The exit evidence is the same kind of thing at every rung — a count of runs, a catch rate with its denominator, a record that caught something — so a rung cannot be exited on a demonstration. The person at the gate stays practiced: amending assumptions at beta, signing the plan at pilot, signing everything outside the released class at production. And production is per class of action, never per team, so one team may hold a released class and a gated class at once. Durations are absent by design: a function entering alpha with clean definitions can exit in a month, and one entering with two definitions of handle time will spend that month on the register, which is the correct use of it.
Reconciling the rungs to the handover gate
The Agentic Handover Gate gives five cumulative tests. The ladder adds none; it distributes the evidence for them across the rungs.
| Handover-gate test | Where the ladder produces its evidence |
|---|---|
| 1 Documented | Entry to alpha: the readiness check's second question reads the L2 table, whose tools column is the permission map |
| 2 Instrumented | Alpha: every objective the process serves has a graded column in the ledgers; a blank with an owner is the honest state of an objective with no instrument |
| 3 Measured here | Beta: the catch rate is measured on this book's work with this book's definitions; anything carried from another channel, platform, cohort or book is [A] on the card |
| 4 Overseen first | Beta, then pilot: error injection and sampled verification run through beta; each team publishes a catch rate for its own clock before its pilot card closes |
| 5 Reversible and logged | Production: a versioned format change with a reversal condition and a re-verification date; a reversal at any rung returns the previous rung's format and is recorded on the card |
The gate admits an agent to act without a signature, so a team in shadow or under a gate has not yet been handed anything: the five tests accumulate as the team climbs and are closed at production, not at alpha. A card may therefore open on a process whose acceptance-gate response has not yet returned — the package the gate's first test asks for is what production waits on, not what alpha waits on.
The five outcomes of the placement check, which Placement Engine Architecture defines and The Agentic Handover Gate applies, apply at every rung exit too; the ladder borrows them rather than defining a set of its own, and an exit that cannot be tested is recorded with the definitional or data gap that stopped it.
The clone pattern
A clone is the same team started on a second book — where a program discovers whether it built a team or a demonstration. What a clone must not carry is the carried-assumption rule, made mechanical.
| A clone carries | A clone never carries |
|---|---|
| The scaffolding: every agent's identity and rules, the coordinator's clock order, the evaluator's rules, the ledger schemas | The previous book's definitions: a clone's columns cite the second book's own register entries |
| The standards: the grade rules, the signature-block format, the intake door's fields and routes, the answer-card format | The previous book's benchmarks: a shrinkage rate, a contacts-per-transaction ratio or a handle-time baseline carried from the first book is [A] on the second |
| The clocks and gates: daily, weekly, monthly and the intake door; every gate on | The previous book's events: the event ledger starts empty; a first-book go-live window explains nothing on the second |
| The adapters: mapping-file format and rejection log, the second book's own mappings filled | The previous book's ladder position: a clone enters at alpha, the readiness check re-run first |
Three consequences follow. The readiness check is per book: the four questions are asked again with the second book's ledgers open. The carried-benchmark test is the clone's first real test, which is why the example's second book is one whose own migration window opens during the pilot: a clone that inherited the first book's post-go-live handle time as its baseline would report the second book's go-live as normal. And a book that cannot be cloned onto is a book whose definitions are not in the register: the checklist stops at its first line, which is the program's most useful output, because a program of agent teams guides consolidation not by choosing which books to standardize but by reporting which ones cannot yet be run on the standard. The mechanics rest on a working prototype in which the team is specified in files rather than in a platform, so cloning copies the specification and not the data.
Worked example
The series' example function does not start its ladder at the planning week. It arrives with three cards already open, and the week's job is to say what the next rung is. The dates below are the example's and are illustrative.
The planning-loop team, card AL-001, is at pilot. Alpha ran in February 2026 and was exited before any card existed: the readiness check passed on the short-term voice reforecast with one narrow fail on definitions (a second contacts-per-transaction record, CPT-02, retired the following week), and question 4 failed outright, remediated by opening the question register with Q-001 on Mon 23 Feb. Beta opened Mon 2 Mar 2026, the day phase 1 of the migration went live; the gate was exercised on every run, and by Fri 3 Apr the loop had run 25 gated business days with a first catch rate published — 9 of 10 known errors detected, from 6 injected and 4 found by deep verification of a 5 percent sample [C from the overseer's log]. Beta's exit asks for thirty runs; the thirtieth fell Fri 10 Apr 2026. Pilot opened before it, on Fri 27 Mar 2026, when the monthly clock signed the phase 2 plan of record with one move declined and priced. The acceptance was recorded with its price: for two weeks a plan of record was signed on a loop whose catch rate was one period old.
The scheduling and real-time teams, cards AL-002 and AL-003, are at beta. The checker first ran Thu 26 Mar 2026 under a publication gate and caught a schedule built on a superseded forecast version; the issuer wrote its first incident Wed 8 Apr 2026 under an action gate, carrying the reoptimization proposal to the gate rather than executing it. Both are earning the catch rates their pilot exits will need.
No card is at production, and the week does not put one there. At D-13 the room takes the next rung per card — the planning loop's pilot exit set on the clone, the scheduling team's pilot entry on a catch rate for the checker, the real-time team's on an outcome record at the action gate — and the second book. That book is chosen because its own migration window opens Mon 14 Sep 2026, which is the carried-benchmark test; card AL-004 opens its alpha inside the front-load closing Tue 30 Jun 2026, the readiness check re-run on it first. One production candidate is named and left undated: the reforecast proposal on unchanged assumptions, whose release waits on its own published catch rate.
The artifact this page produces
Agent ladder card (AL), one row per team per book per rung. One filled example row:
| ID | Book | Team | Rung | Entered | Exit evidence required | Exit evidence observed | Catch rate (proportion, injection rate, n) | Days-to-detection | Clone source |
|---|---|---|---|---|---|---|---|---|---|
| AL-001 | the migrating corporate client book | planning loop | pilot | Fri 27 Mar 2026, with the acceptance recorded that beta's thirtieth run would fall Fri 10 Apr | the monthly clock run under its own gate with a decline priced; a record traceable to a real catch; the clone run on the second book | the phase 2 plan of record signed Fri 27 Mar with one move declined and priced; beta's thirty gated runs complete Fri 10 Apr; clone open | 9 of 10 (6 injected, 4 by deep verification of a 5 percent sample) [C from the overseer's log] | open | none; source for AL-004 |
Figures measured on this book before the card was opened keep the grade they were published at. Human Gates and Number Grades puts a figure at [A] when it is carried across a channel, platform or cohort change; the ladder adds one condition of its own, named here as the ladder's: a figure carried from another book, or from a prototype, is [A] on the card it arrives on until this book measures it.
Produced in a working session with Wiki:Packs/Agent Program Readiness (CP-OPS-005), with the clone checklist; the filled set is part of blueprint v0.1.
What would change this
The claim is that evidence, not elapsed time, should define a rung; that a card belongs to a team on a book rather than to a function; and that a clone must carry scaffolding and never definitions. A function that ran a clone with the first book's definitions and benchmarks carried across and held forecast accuracy and plan quality through the second book's own go-live would weaken the clone rule. A function that let its second and third rosters wait for the first roster's exit evidence, and reached the same catch rates no later than one that opened them early with a recorded acceptance, would show the acceptance device to be unnecessary bookkeeping. Evidence that people at a gate amended fewer assumptions over time without the catch rate falling would argue for relaxing the gate's "must produce" rule on unchanged assumptions. None of these observations exists at the time of writing.
How this connects
- Previous in the chain: AI Agent Program for a Resource Optimization Center — the charter the cards are opened under
- Next in the chain: Anatomy of a ROC Standard — the Standard whose §3 sections are the permission maps
- Defers to: The Agentic Handover Gate (the five tests, reconciled above, never restated) · Human Gates and Number Grades (the release condition, stated there once) · Agent Team Readiness for a Planning Process (the check, here per book) · The Agent Overseer (catch rate) · The Short-Term Forecasting Loop with an Agent Team, Scheduling Agents and Real-Time Agents (the three clocks the cards sit on) · Living Ledgers (what a clone's ledgers are)
- Reads: Process Standardization Lifecycle (highly specified) · Migrating a Book of Business (the carried-assumption rule the clone test mirrors) · Process Shells for a Workforce Standard (the shells marking processes that cannot yet enter alpha)
Maturity Model Position
On the WFM Labs Maturity Model™ alpha and beta are Level 3 work made visible: definitions cited, a clock kept, a register open, a person signing. The pilot rung, with a team's whole clock running as the version of record and a second book cloned onto the standard, is the Level 4 form of a planning function's own work. Production per class of action is released on a published catch rate rather than at a level, the condition stated once on Human Gates and Number Grades; the wiki's level pages agree on it since 17 September 2026, and this page places the rung on evidence and stops there by design. Four scales on this wiki use the word "level"; Planning Week for a Workforce Function states which is which. This page uses the maturity Levels 1 to 5, and its rungs are not levels on any of the four scales.
See Also
- Planning Week for a Workforce Function — Day 4 morning
- AI Agent Program for a Resource Optimization Center — the charter the ladder serves
- The Agentic Handover Gate — the five tests
- Human Gates and Number Grades — the gates that stay on, and the release condition
- Agent Team Readiness for a Planning Process — re-run per book
- The Agent Overseer — the catch-rate program
- A Roadmap Pattern for Agent Teams in WFM — the order, and the published record this page's rungs are read from
References
- ↑ Humble, J., & Farley, D. (2010). Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation. Addison-Wesley. ISBN 978-0-321-60191-9.
- ↑ Mankins, J. C. (1995). "Technology Readiness Levels: A White Paper". NASA Office of Space Access and Technology, Advanced Concepts Office.
- ↑ Bainbridge, L. (1983). "Ironies of Automation". Automatica 19(6), 775–779. doi:10.1016/0005-1098(83)90046-8.
- ↑ Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). "AI Agents That Matter". arXiv:2407.01502.
