Deferrable Work and the Idle Cushion

From WFM Labs
The proportional idle cushion falls steeply as a gate grows: at the small end it is where deferrable work fits; at the large end it is gone.

Deferrable work and the idle cushion describes the economics of absorbing work that can wait — email, case follow-up, documentation, back-office processing — into the idle capacity that small real-time service gates necessarily carry. The mechanism is the square-root safety-staffing law: the staff a gate needs to hold a service level is roughly its offered load plus a safety margin proportional to the square root of that load, so the proportional idle cushion shrinks as the gate grows and small gates carry several times the fractional cushion of large ones. That cushion is not waste — it is the price of the service level — but it is available, and deferrable work absorbed into it costs its marginal displacement, which is close to nothing. The staffing mathematics of blending inbound with deferrable work is owned by Blending and Deferred Workload, which this page assumes rather than restates; the concurrency mathematics across real-time channels by Multi-Channel and Blended Operations; the demand modeling of deferred work by Back Office and Knowledge Worker Workforce Management; the pooling penalty curve by Pooling Architecture in Service Workforces. What this page adds is how that work is costed and placed, and why the placement question has no standing answer.

The cushion

For a queue holding a service target, the required number of servers is approximately the offered load plus a safety term proportional to the square root of the load — the square-root staffing law, formalized in the many-server heavy-traffic regime by Halfin and Whitt[1] and developed for call-center dimensioning by Borst, Mandelbaum and Reiman.[2] Written with offered load R and a service-quality parameter β, staff is roughly R + β√R, and the proportional cushion — the safety staff as a fraction of the load — is β/√R. It falls as the gate grows. Under a constant β the ratio would go as the square root of the load ratio, but β is not constant under a service-level constraint. Computed exactly, a gate carrying four erlangs needs seven agents — a cushion of 75% of its load — against 267 agents at 256 erlangs, a cushion of 4.3%: roughly seventeen times the fractional cushion at the small end.[3]

The cushion is not waste but the reserve that holds the service level, and it is idle by design — Blending and Deferred Workload develops both points and the limits on how much of it can be filled. This page is about what the filling is worth and where the work should go.

What deferrable work does inside it

Work that tolerates delay — an email that can be answered within hours, a case note, a fulfillment step — can be absorbed into the cushion without displacing the real-time work the cushion exists to protect, provided it is dropped the moment a real-time contact arrives. Three consequences follow, and each contradicts the way deferrable work is usually costed.[4]

  • The loaded rate is the wrong denominator. The conventional comparison prices deferrable work at a loaded hourly rate in each candidate location and picks the cheaper. But deferrable work absorbed into already-paid-for idle time does not cost a loaded rate; it costs the marginal displacement of work that would otherwise not have been done, which is close to zero. Comparing two loaded rates compares two numbers neither of which describes the onshore case.
  • Stripping it out can destroy the utilization it was subsidizing. Remove all deferrable work from a small gate and its cushion becomes visible as pure idle cost with no productive outlet. The saving on the deferrable work can be smaller than the utilization lost behind it.
  • Gate expansion and offshoring deferrable work are substitutes. As gates widen, the proportional cushion shrinks and the capacity to absorb deferrable work disappears with it. Both moves may be worth making; their benefits cannot both be banked in full, and an estate that plans both without netting them against each other has counted the same hours twice (see Gate Expansion: What It Buys and What It Spends).

A worked illustration makes the denominator error concrete; the staffing figures are the exact ones above, and the six-minute item and the eight-hour day are illustrative. Take the four-erlang gate: seven agents staffed, four erlangs of load, so about three agent-equivalents of cushion — roughly twenty-four agent-hours over an eight-hour day that the service level requires and that sit idle between arrivals. If a deferrable item takes six minutes, the cushion can absorb of the order of two hundred items a day at a marginal cost that is close to zero, because the hours are already paid for and the alternative use of them is waiting. The conventional comparison prices those two hundred items at the gate's loaded hourly rate — twenty hours of loaded onshore cost — against a lower loaded rate elsewhere, and finds the offshore option cheaper by the rate difference. The comparison is between a number that describes the onshore case (near zero) and one that does not (twenty loaded hours), and it gets the sign wrong whenever the cushion is genuinely idle.

The formal home of the problem is call blending: a service-level-constrained real-time class and an infinitely backlogged deferrable class served by the same pool, with throughput on the deferrable class maximized subject to the constraint on the real-time one, typically through a threshold policy that reserves servers for the real-time class.[5] Chaining and Flexibility Design supplies the routing construct: deferrable work as a second skill on a real-time agent is a two-skill chain, viable on that page's own test only where the deferrable scope reaches competence in weeks rather than months.

The state-dependence

The counter-case does not win either, and that is the finding. On a day when the small gate is behind on its service level, the cushion is not sitting available; it is being consumed by the gate's own backlog, and the deferrable work absorbed into it is displacing real-time work it should not displace. The same email, in the same gate of the same size, is right onshore in the morning and right offshore by the afternoon. The right destination is a function of the system state at the moment the work arrives.

That has a direct implication for how the question is asked. "Should this class of deferrable work go offshore?" is a standing question, and the answer to a standing question is a rule. But the economics are not standing; they move with the state of every gate that could absorb the work. A preference — absorb deferrable work into small-gate cushions when they are available — is legitimate and can be encoded. A standing answer — this work type belongs in this location — is a category error, and it is the form in which the question is usually put. The general principle is the argument for a placement engine over a placement rule set: work types cannot be optimized one at a time, and not once, because moving one changes the economics of every other type sharing the same supply. Placement Engine Architecture describes the machinery that turns placement from a recurring argument into a produced answer — constraints in, placement out, re-run when conditions change — and the state-dependence here is the reason the re-run matters.

The measurement dependency

Idle capacity inside a gated group is real and, in most estates, invisible. An agent who picks up an email while waiting for a call is genuinely occupied but sits in an available state, so the work is recorded as idle; the cushion is understated where it is being used and overstated where it is not, depending on state conventions nobody set. Occupancy treats blending as raising measured occupancy; that holds only where the blended time is signalled, and in most estates it is not. Until the harvest is measured, the estate cannot know how much of the cushion is already absorbing deferrable work, and any before-and-after comparison of a transfer inherits the error. The measurement work precedes the placement decision in sequence, not beside it.

Failure modes

  • Pricing deferrable work at a loaded rate in both locations. The onshore cost is overstated by the whole cushion.
  • Banking the offshoring saving and the gate-expansion saving separately. The same hours are counted twice.
  • Answering the standing question. A rule is written for a quantity that moves hourly.
  • Absorbing deferrable work into a gate that is behind. The cushion is not there; the real-time target pays.
  • Deciding before measuring. The harvest is invisible, so the case is argued on assumed idle rather than observed.

Maturity Model Position

Blending and Deferred Workload places blending with inbound priority at Level 3 on the WFM Labs Maturity Model™ and dynamic blend ratios at Levels 4 and 5, and this page inherits that placement. What it adds is the placement question: costing deferrable work on marginal displacement rather than loaded rate is the Level 3 measurement discipline of Variance Harvesting applied to work placement, and deciding where the work goes on system state rather than by standing rule sits at Level 4, where the model replaces planning cycles with continuously re-run engines.

See Also

References

  1. Halfin, S., & Whitt, W. (1981). Heavy-traffic limits for queues with many exponential servers. Operations Research, 29(3), 567–588.
  2. Borst, S., Mandelbaum, A., & Reiman, M. I. (2004). Dimensioning large call centers. Operations Research, 52(1), 17–34.
  3. Figures from Pooling Architecture in Service Workforces, computed from Erlang C at a 300-second handle time and an 80%-in-20-seconds target.
  4. Practitioner observation from multi-node service estates in which the placement of deferrable work is a standing argument; a consistent pattern rather than a measured result.
  5. Gans, N., & Zhou, Y.-P. (2003). A call-routing problem with service-level constraints. Operations Research, 51(2), 255–271.