Planning for Agentic AI Is Not a Carve-Out

From WFM Labs
The old carve-out treated deflected volume as gone. Under agentic automation the boundary moves, and four effects (rebound, escalation, a harder residual, new capability) return to the plan.

Planning for agentic AI is not a carve-out is a planning proposition about how automated work should enter a staffing plan. The carve-out is the traditional method: subtract the share of volume that the IVR or bot will absorb, then plan the remainder with the usual queueing arithmetic. This page explains why that method was adequate for static deflection, why it fails once the automated share moves, and what a plan has to start from instead. It does not specify the replacement model; the Value-Based Planning Model owns the Erlang inversion and the four building blocks, and Value First, Then Route carries the five-stage chain and the routing mechanic. Nor does it argue what the carve-out habit does to value; that is The Value Destruction Risk in Service Automation. The scope here is the habit itself, the four ways the agentic boundary moves, and the trap the habit produces.

The old carve-out

Contact-center staffing descends from a single chain: forecast arrivals and handle time, convert them to offered load, and solve for the number of agents that meets a service-level target.[1] When self-service arrived, it entered that chain at the front. Total volume was drawn as a circle, a wedge labeled "IVR and bots" was cut out of it, and the plan was built for what remained. The automated share became a fixed reduction of arrivals: one parameter, set once per planning cycle, applied before anything else was calculated.

The method worked well enough, and it is worth being precise about why. Scripted deflection had four properties that made a fixed wedge a fair approximation:

  • It was deterministic. A menu tree or a rule-based bot handled exactly what was written into it, and its scope did not grow between planning cycles.
  • Its containment was stable. Last quarter's containment rate was a reasonable forecast for next quarter's.
  • Its failures were immediate and visible. A caller who pressed zero arrived in the human queue within seconds and was counted there.
  • The two populations were separable. The work the IVR took was work no agent would see again, so removing it from the forecast base did not distort what was left.

Under those conditions the carve-out is not a mistake. The operations literature has generally modeled self-service and channel choice upstream of the staffing calculation, with the arrivals that reach agents treated as exogenous.[2] The mistake is carrying the same approximation into a technology that has none of the four properties.

The agentic reality

Agentic systems reason over a case, call tools, and complete multi-step transactions rather than following a fixed script. Their scope expands as models, integrations and data improve. Their failures surface later, as partial completions, misrouted cases and repeat contacts that carry a history. The boundary between what the system handles and what people handle is therefore fluid, and it moves in four ways at once.

  1. Rebound demand. Freed capacity lowers queues and wait times. Demand that was previously suppressed by the cost of contact surfaces, and total volume rises. Volume responds to the service offered; it is not a fixed quantity to be divided between handlers.
  2. Escalation tax. Interactions the system fails to complete flow to human pools, carrying the cost of the attempt already made.
  3. Complexity concentration. Automation takes the most tractable work first. What remains in the human queue is harder per unit than the pre-automation average, so the remaining work gets harder even as it gets smaller. The effect was described for process control decades before agentic systems existed: the operator is left with the tasks the designer could not automate, and has less practice at them.[3]
  4. New capabilities. The system's scope grows, and interaction types that were human work last cycle are automated this cycle.

Two of these push human volume up, one makes each unit of human work heavier, and one pulls volume away. Their net effect on any given human pool cannot be read from a containment rate. The volume in a human pool can rise or fall between cycles, and the set of interaction types the system handles is itself a moving quantity. In the task-based view of automation, this is the expected outcome: machines substitute for the codifiable tasks and complement the rest, and the line between the two is redrawn as capability improves rather than fixed at deployment.[4]

The four second-order effects as working estimates

The conference presentation from which this page is drawn attaches a working figure to each of the four effects.[5] None of the figures is cited on the slides, and none should be read as a finding. Each effect has a wiki page that treats it from the literature, and where the sourced range differs from the deck's estimate, the sourced range governs.

Effect Deck's working estimate (uncited) Page that treats it
Rebound demand Freed capacity lowers queues, dormant demand surfaces; volume grows by around 20 percent Service Demand Rebound Model — sizes rebound from the energy-economics literature; its three rebound components total 30–70 percent of gross deflection
Escalation tax Failures flow to humans at five to ten times the original cost The Escalation Tax — the cascade-adjusted expected-cost formula and worked multipliers
Complexity concentration Each 10 percent of automation raises remaining handle time by about 6.5 percent The Hardening Residual — the sizing consequences; the Complexity Premium in the rebound model carries the handle-time form at 5–8 percent per ten points of containment
Savings erosion Year one realizes 50–70 percent of projected savings; by year three only 25–45 percent remains Service Demand Rebound Model — why projected savings systematically miss, and why rebound strengthens over time

The Interior Optimum (containment rate) treats the adjacent static question, why maximum containment is not the cost optimum. No page yet carries a multi-year decay curve for realized savings.

The presentation's summary line, that 50 percent containment does not mean 50 percent savings and that 25–35 percent is the figure to expect, is a rule of thumb in the same category. It is consistent in direction with the sourced treatments on the linked pages and should be quoted with the same caveat.

The trap

The presentation names the trap as assuming that automated volume is gone and never needs to be planned for again.[5] Stated plainly it sounds avoidable, but the carve-out produces it mechanically, through four omissions.

  • The forecast base is cut. Once the deflected wedge is removed, rebound arrives as an unexplained rise in a base that was never expected to grow. It is read as forecast error rather than as a predictable consequence of the automation.
  • Escalations have no budget line. Failures flow into human queues as volume with no origin, so the human pool is understaffed by the escalation share, and the shortfall is attributed to the humans.
  • Targets are not re-baselined. The cycle boundary passes without anyone resetting handle-time and quality targets on the new mix, so the targets are missed and the miss is read as deterioration. Why the miss is structural is the subject of The Hardening Residual; the omission here is that no step in the carve-out prompts the reset.
  • Year-one savings are booked as permanent. The business case is closed where realized savings are highest, and the decay is never planned for. The rebound model's rule is a measurement window of at least 24 months; anything shorter understates rebound systematically.

A fifth omission is organizational. Once the wedge is carved out, the automated volume tends to become the property of the technology function, and the planning function stops forecasting it. The boundary is then managed by nobody, at the moment it has begun to move.

A different way to plan

The presentation's alternative inverts the input: the value of every interaction, rather than the volume expected to be deflected, and the question of who should handle each type given its value and the system's capability for it.[5] The specification of that alternative belongs to other pages. What matters here is the one consequence that follows from the input alone: the automated share stops being a parameter set before planning and becomes an output of a decision taken inside it, re-taken each cycle as capability moves in either direction. The marketing-science treatment of service productivity supports the reversal: pushing productivity past its profit-maximizing level trades away more revenue than it saves in cost, so the level itself has to be chosen rather than maximized.[6] A carve-out that treats deflection as pure saving is, in that framing, maximizing productivity without asking where the optimum lies.

The alternative itself is documented on its own pages. The Value of an Hour supplies the time taxonomy that gives "value" an operational meaning. Tier 1 Is Not One Thing shows why productive time has to be split by value before it can be routed. Value First, Then Route gives the stage-by-stage mechanic, and the Value-Based Planning Model states the model that the mechanic serves.

Maturity Model Position

Under the WFM Labs Maturity Model™, the carve-out is the correct method at Level 2 and remains serviceable at Level 3, where deflection is scripted and containment is stable enough to treat as a parameter. The proposition becomes binding at Level 4, the first level at which agentic capability is part of the workforce being planned. At Level 5 the re-decision is continuous, and the four effects in the table are instrumented as live signals rather than planning assumptions.

See Also

References

  1. Gans, N., Koole, G., & Mandelbaum, A. (2003). "Telephone Call Centers: Tutorial, Review, and Research Prospects". Manufacturing & Service Operations Management 5 (2), 79–141. doi:10.1287/msom.5.2.79.16071.
  2. Akşin, Z., Armony, M., & Mehrotra, V. (2007). "The Modern Call Center: A Multi-Disciplinary Perspective on Operations Management Research". Production and Operations Management 16 (6), 665–688. doi:10.1111/j.1937-5956.2007.tb00288.x.
  3. Bainbridge, L. (1983). "Ironies of Automation". Automatica 19 (6), 775–779. doi:10.1016/0005-1098(83)90046-8.
  4. Autor, D. H. (2015). "Why Are There Still So Many Jobs? The History and Future of Workplace Automation". Journal of Economic Perspectives 29 (3), 3–30. doi:10.1257/jep.29.3.3.
  5. 5.0 5.1 5.2 Lango, T. (2026). Adaptive: Building Workforce Systems for an (Unpredictable) Future. Presentation, SWPP Annual Conference, slides 33–41. No sources are cited on the slides for the figures reproduced here; they are the presentation's working estimates.
  6. Rust, R. T., & Huang, M.-H. (2012). "Optimizing Service Productivity". Journal of Marketing 76 (2), 47–66. doi:10.1509/jm.10.0441.