Hypothesis Testing with Agent Teams

From WFM Labs

Hypothesis testing with agent teams is the discipline by which a planning agent team turns a question about the operation into a set of competing claims, states for each what evidence supports and contradicts it, names the test that would settle it and the data the test needs, runs the test when the data arrives, and tags the surviving driver as structural or transitional because that tag decides steady-state staffing. The page's claim is that the method is not new, it is strong inference applied to planning data, and that what an agent team adds is the discipline of writing the table before looking for the answer and the rule about which agent may say what. This page gives the table, the tag, two traps as worked examples, and the rung rule.

The table

Every register row on The Question Register and Knowledge Base carries a hypothesis table. Its columns are fixed.

Column Holds Rule
Claim One candidate explanation, stated so that it could be false A claim that no observation could contradict is not admitted
Evidence for Graded findings that support it, each cited to a ledger version Numbers carry [M], [C], [E] or [A]
Evidence against Graded findings that contradict it An empty column is a finding: nobody has looked
Settling test The observation or analysis that would decide between this claim and its rivals Written before the test is run
Data needed The ledger, version and cut the test requires; whether it exists If it does not exist, the row says who obtains it and by when
Grade Established · Inferred · Asserted · Open, per Data Synthesis Before Decision Set by a person; the causal analyst proposes
Tag Structural or transitional, once the claim survives Required on every driver that enters a forecast or plan

The method is the one the physical sciences call strong inference: devise alternative hypotheses, devise the experiment that excludes one or more of them, run it, and repeat with what remains.[1] Its planning form has one addition: the settling test is written into the row before anyone looks at the data, so that the test cannot be chosen to flatter the favored claim.

The tag: structural or transitional

Every driver that survives its test is tagged. Structural means the effect persists at steady state and the plan must carry it: a handle-time level shift that does not decay, a mix change that a product change made permanent. Transitional means the effect will pass and the plan must not carry it beyond its window: a learning curve after a go-live, a two-day training pull, a backlog being worked down. The tag is what the capacity planner reads when it sizes the next phase, and it is the single most consequential field in the register, because a transitional driver carried as structural overstaffs every month until someone notices, and a structural driver carried as transitional understaffs until service breaks.

The tag is a claim like any other and carries a grade. A level shift observed for three weeks is structural at Inferred; it becomes Established when the settling test, a cohort-level curve that is flat, has run over a period long enough for a curve to have shown. The row says how long that is.

The rung rule

Pearl's ladder of causation has three rungs: association (seeing), intervention (doing), counterfactual (imagining), and each answers a question the rung below cannot.[2] The Ladder of Causation in WFM applies it to planning data. This series applies it as a permission.

Rung Question Who may write it What it looks like in a ledger
1, association What co-occurs with what? Post-analyst, scout, monitor, detector "Handle time rose from 2 March; the rise is out of control on the chart; the go-live's window covers it"
2, intervention What would change if X were changed? Causal analyst, after a causal diagram and a test "Holding cohort and channel fixed, the go-live moved partner-cohort handle time by 40 seconds (Inferred)"
3, counterfactual Had X not happened, what would Y have been? Causal analyst, only where the diagram identifies it "Had the training pull not occurred, 8 April staffing would have met schedule (Inferred; the counterfactual is the schedule itself)"

The evaluator blocks any run in which a rung-1 agent has written a rung-2 sentence. The causal analyst draws a diagram per question, names the confounders explicitly, and follows the standard method for isolating them: decompose first, test what can be tested, label the residual, state what would change the answer. Causal Diagrams (DAGs) in WFM gives the diagram rules; this page assumes them.

Two traps, worked

Both examples use the series' corporate client book: three channels, an in-house and a partner cohort, a phased migration live from Monday 2 March 2026.

The composition trap

Phase 2 of the migration on Monday 30 March raises the partner cohort's share of voice handled from 30 percent to 55 percent. The blended voice handle time reads 442 seconds in the week of 23 March, the last week before phase 2, and 441 in the last week of April [M]: flat across the phase change. The in-house cohort's curve had already pulled the blend from the 452 seconds measured on 3 March and signed into the plan on 27 March down to 442 by late March; that is the movement the blend then stopped showing. A leader concludes handle time has stabilized. The register row's first claim is "handle time has stabilized"; the settling test, written before looking, is the cohort split. The split shows the in-house cohort fell from 430 to 405 seconds [M], a learning curve, while the partner cohort held at 470 [M], a level shift with no curve. The blend was flat because the in-house improvement of 25 seconds was offset, almost exactly, by the share moving toward the slower cohort: 0.7 × 430 + 0.3 × 470 = 442; 0.45 × 405 + 0.55 × 470 = 441 [C]. Decomposing the change into a within-cohort component and a mix component, the method demographers formalized for exactly this problem, gives the answer the blend hid: within-cohort, handle time fell; mix, it rose.[3] Where more than two drivers interact, the Shapley decomposition gives an order-independent attribution.[4]

Tags: the in-house curve is transitional; the partner level shift is structural at Inferred, becoming Established if the partner cohort shows no curve by the phase 3 review. The plan for phase 3 carries 470 seconds for the partner cohort and a decaying curve for in-house, not 441 for everyone. Simpson's Paradox in Contact Center Metrics and Mix Effects in Blended Quality Targets describe the same arithmetic in quality and service measures; the register row is where a planning function catches it in handle time.

The two-definitions trap

In the fourth week after phase 2, a register row opens on "contacts per transaction has risen 35 percent since the migration." The causal analyst's first act is not a diagram but a definitions check, because the claim compares two numbers. The demand ledger's definitions block shows two records in circulation: definition CPT-01 counts handled contacts over booked transactions and reads 0.42 [M]; definition CPT-02, from the legacy platform's report, counted handled contacts excluding follow-ups on an open case over all transactions including cancellations, and read 0.31 [M] on the same days. The rise is the difference between two definitions, not a change in customer behavior, and the row closes at Established with the finding "no change in contacts per transaction on one definition; the comparison mixed CPT-01 and CPT-02." The definitions ledger gains a pinned note; the legacy report is retired. Without a definitions ledger, this row would have produced a forecast adjustment and a plan built on it. Deming's insistence that a measurement has no meaning without its operational definition is the whole of the trap and the whole of the cure.[5]

What would change this

The table's columns and the rung rule are design choices resting on the causal-inference and measurement literatures; the two traps are generic reconstructions, not measured cases. A function that found its causal analyst's Inferred tags were overturned on later evidence more often than its planners' untagged judgments would have grounds to reconsider the machinery, though not the table. The structural/transitional dichotomy is coarse by design; a function with enough history could grade the decay rate instead.

How this connects

The table lives inside the register row on The Question Register and Knowledge Base and its findings enter the assumption registers of Living Ledgers. Its rung rule is a permission the roster on The Agent Team Model enforces and the evaluator checks; its tags are what Long-Term Planning Agents and the Plan of Record reads at the requirement stage. The causal method is the wiki's: The Ladder of Causation in WFM, Causal Diagrams (DAGs) in WFM, Correlation and Causation in WFM, and the decomposition traps are on Simpson's Paradox in Contact Center Metrics and Mix Effects in Blended Quality Targets.

Maturity Model Position

A planning function at Level 2 on the WFM Labs Maturity Model™ explains misses in prose from memory. Writing the table by hand, with grades and a settling test, is a Level 3 discipline that needs no agent. An agent team that runs the settling test the day the data arrives, and a plan that reads the structural tag, is Level 4, where every driver in the distribution has a grade and a decay assumption.

See Also

References

  1. Platt, J. R. (1964). "Strong Inference". Science 146(3642), 347–353. doi:10.1126/science.146.3642.347.
  2. Pearl, J., & Mackenzie, D. (2018). The Book of Why: The New Science of Cause and Effect. Basic Books. ISBN 978-0-465-09760-9.
  3. Kitagawa, E. M. (1955). "Components of a Difference Between Two Rates". Journal of the American Statistical Association 50(272), 1168–1194. doi:10.1080/01621459.1955.10501299.
  4. Shorrocks, A. F. (2013). "Decomposition procedures for distributional analysis: a unified framework based on the Shapley value". Journal of Economic Inequality 11(1), 99–126. doi:10.1007/s10888-011-9214-z.
  5. Deming, W. E. (1986). Out of the Crisis. MIT Center for Advanced Engineering Study. ISBN 0-911379-01-0.