Chaining and Flexibility Design

From WFM Labs

Chaining is the design of limited, deliberately configured flexibility so that a small number of capabilities per resource delivers nearly the benefit of universal flexibility. The central result — established in manufacturing and reproduced in service operations — is that configuration matters more than quantity: resources able to perform only two tasks, connected into a single closed chain, capture close to the full throughput of a system in which every resource can perform every task. For a workforce, this reframes fungibility from a training-budget problem into a skill-map design problem, and it is the mechanism by which a fragmented estate can be made responsive without universal cross-training.

This page covers chain topology, the decomposition of work into chainable scopes, the placement of automated capacity inside the chain rather than outside it, and the treatment of sourcing and site decisions as chain design rather than as unit-cost decisions alone. For the data model behind a skill graph and the economics of the training decision, see Cross-Training and Skill Mix Strategy. For why fungibility matters to capacity at all, see Supply Elasticity in Workforce Planning.

The chaining result

Jordan and Graves established the finding in a manufacturing setting: limited flexibility, configured as a chain, captures approximately 98% of the throughput of a fully flexible system using resources capable of only two tasks.[1] The benefit arises from products and plants being connected into a single chain — not from any resource being universally capable.

Wallace and Whitt reproduced the result directly in the call centre. Where service time does not depend on call type or agent, two skills per agent in the right combinations performs almost as well as every agent holding every skill, and required staffing under limited cross-training is nearly the same as under full cross-training.[2]

Bassamboo, Randhawa and Van Mieghem extended it to parallel queueing systems, showing a tailored chaining configuration using dedicated and level-2 resources is asymptotically optimal — and that in some parameter regions the fully flexible resource is not worth using at all.[3]

Simchi-Levi and Wei went further, providing theoretical grounding for why the long chain performs as well as it does relative to sparse alternatives.[4]

Four independent results, across three decades and three settings, converge on the same conclusion. The practical statement: the barrier to fungibility is skill-map design, not training budget.

Topology: why configuration beats quantity

The property that generates the benefit is connectivity, not capability count. Three consequences follow, and each is routinely violated in practice.

The chain must close

A single long chain connecting all nodes materially outperforms several disconnected short chains using the same total training investment. Two separate five-node chains are not equivalent to one ten-node chain: capacity can circulate within each loop but cannot cross between them, so a surge in one loop cannot draw on slack in the other.

This is the single most consequential result for a fragmented or recently merged estate. Cross-training conducted within each business unit, segment or site — which is what happens by default, because that is where the training relationships already exist — produces exactly the disconnected-short-chain configuration. The investment is made, the flexibility is reported, and most of the available benefit is not realised because the loops never join.

The highest-value edges in a merged estate are the ones that cross the merger boundary, and they are the ones nobody proposes.

A broken chain loses most of its value

Because the benefit depends on the loop, a chain with one missing edge is not a chain with slightly less benefit — it is two shorter chains. Chains break silently through attrition, reorganisation, site closure and skill decay, and the break is invisible unless the graph is monitored as a graph.

Degree beyond two adds little

Once a chain is closed, adding a third and fourth skill per resource produces sharply diminishing returns. This is the result's most useful practical implication and the one most often ignored: an operation that has closed its chain should stop cross-training and spend the next increment elsewhere.

Designing a chain

  1. Define the nodes. A node is a unit of work that can be routed to independently. Nodes that cannot be routed to separately are one node regardless of how they are described organisationally.
  2. Measure node volumes and their correlation. Chaining pays where demand is variable and imperfectly correlated across nodes. Where two nodes surge together, an edge between them buys little.
  3. Lay out the loop. Order nodes so that each is connected to the next and the last connects back to the first. Prefer edges between nodes whose demands are negatively or weakly correlated, and between nodes of comparable skill depth.
  4. Cross the boundaries. Deliberately include edges crossing segment, site, brand and sourcing boundaries. These are the edges that close the loop and they will not be proposed spontaneously.
  5. Verify the routing engine can execute it. A chain that exists in the training records but not in the routing configuration does not exist. Realised flexibility is what the router will do at 10:47 on a Tuesday, not what the skills matrix says.
  6. Monitor the graph. Track edge currency and detect breaks. Treat a broken edge as an incident.

Decomposing work into chainable scopes

Chaining assumes a node can be learned. Where a role takes six to eight months to reach experienced-level performance, a second skill is not a training decision — it is a second career, and the chain will not be built regardless of how compelling the arithmetic is.

The resolution is to decompose the work before chaining it. A role is not an atom. Most complex servicing roles contain scopes that differ enormously in the time required to become competent: intake and verification, straightforward transactions, exception handling, and high-judgment resolution. Learning-curve effects apply to each separately.[5]

Decomposition changes what is possible in three ways:

  • It creates chainable nodes. A scope reaching competence in weeks can carry a second skill; a scope reaching competence in months cannot.
  • It widens the supply pool. A narrow, fast-to-proficiency scope can be filled by part-time, contingent and returning workers who could not fill the whole role. Chain design and staffing-model design are therefore complements rather than alternatives — see Supply Elasticity in Workforce Planning.
  • It concentrates the deep skill where it is needed. Removing routine scope from expert nodes raises the effective capacity of the scarcest resource without hiring.

The counter-consideration is real and should be stated: decomposition adds handoffs, and handoffs add transfer time, context loss and customer effort. A decomposition that halves training time and doubles transfers has not obviously improved anything. The test is whether the handoff can be made clean — which is the point at which automation becomes relevant.

Chaining automated capacity

Automation is usually positioned as deflection: a contact type is handled end to end by a system and leaves the human graph. Chaining offers a different placement, and the distinction is structural rather than semantic.

Deflection Chained automation
Topology Removes a node from the graph Adds a node inside the graph, with edges to human nodes
Effect on the residual Remaining work is more specialised, so the human chain is sparser Connectivity preserved or increased
Elasticity effect Reduces volume; may reduce the flexibility of what remains Adds a node with near-zero ramp time
Failure mode Contained work that should have escalated Mis-routing, and error propagation along the chain

The elasticity consequence is the important one. A deflection programme removes the most routine work — which is the most fungible work — leaving a residual that is harder, more specialised and less chainable than the original. An operation can therefore reduce its contact volume and simultaneously reduce its ability to respond to variability, which is a poor trade that is rarely priced.

Automation as an ideal chain node

Placed inside the graph, automated capacity has three properties no human node has. It can hold many skills simultaneously, so it is naturally high-degree. Its ramp time is near zero, so it is the only node that can be added inside a forecast horizon. And its capacity is elastic within the interval rather than fixed by a shift.

A high-degree node with elastic capacity is precisely what a sparse chain is missing.

Triage as the highest-value chained scope

Triage — classifying and routing an arriving contact to the correct node — is the highest-connectivity scope in any servicing graph, and therefore the highest-value place to put automated capability.

The argument is topological rather than about labour saved. In most operations, work enters the graph at the wrong node: the customer selects from an IVR menu or a web form, the selection is approximate, and the contact arrives somewhere plausible rather than correct. Every mis-entry is then resolved by a transfer, which consumes capacity at two nodes and destroys the benefit the chain was built to deliver.

The result is that realised flexibility is usually well below designed flexibility, and the gap is caused by entry error rather than by insufficient cross-training. Accurate triage closes that gap. It raises the effective connectivity of the existing graph without training anyone, which makes it the cheapest available flexibility intervention in an operation that has already invested in cross-training and not seen the return.

A second effect compounds it. Where automation performs the intake and verification portion of a contact and hands a partially-resolved case to a human node, the human scope becomes shorter and more uniform — which lowers its time to proficiency, which makes that node easier to chain. Chained automation increases chainability, rather than merely substituting for it.

Cautions

Three limits apply, and stating them is what separates the argument from a vendor claim.

Triage error propagates. A mis-routing node is worse than no routing node, because the error is made confidently and at scale. Triage accuracy must be measured against outcome — did the contact require a transfer — and not against a classification benchmark. A human re-route path must remain available and must be cheap to invoke.

The frontier is jagged. Dell'Acqua and colleagues document a boundary where some tasks fall easily within automated capability while others, superficially similar, fall entirely outside it, together with mis-calibrated trust in which workers over-rely on the system precisely where it is weakest.[6] Node boundaries must be drawn against measured capability, not against apparent task difficulty.

Emotional load routes differently from topic. In a randomised field experiment on agentic customer service, human intervention preserved quality in technical escalations but was markedly less effective in emotional escalations, where early intervention proved essential.[7] A triage node classifying only by topic will route emotionally-loaded contacts into scopes that handle them badly. Emotional load is a routing dimension in its own right, and detecting it late is close to useless.

Sourcing and site strategy as chain design

Location and sourcing decisions — moving work to a lower-cost site, a captive service centre or a contracted vendor — are usually taken as unit-cost decisions and evaluated on rate arbitrage. They are also chain topology decisions, and taking them without that lens is the most reliable way to buy a cost reduction and pay for it in responsiveness.

A sourcing decision does not change the price of a node. It moves the node, and in doing so it changes the graph.

Three failure modes

Whole-queue migration. The default unit of a sourcing decision is a queue, because that is the unit the routing system and the commercial construct both understand. But a queue contains scopes with very different proficiency curves. Moving it entire moves the complex work with the simple work, so the receiving site carries the longest ramp in the estate against the work least tolerant of error. The quality complaints that arrive some months later are read as supplier or site performance, when they were determined at the moment the migration unit was chosen.

Dedicated site pools. Work moved to a new location is customarily ring-fenced there — its own queues, its own agents, its own reporting. That is, precisely, a disconnected short chain. Capacity cannot circulate between the new site and the existing estate in either direction, so a surge at either end cannot draw on slack at the other. The cost saving is realised and the flexibility is silently forfeited.

Ramp omitted from the business case. A saving computed from rate differential and headcount assumes proficient capacity on day one. A new site is, by construction, entirely mid-ramp at the moment the operation begins depending on it. The saving is realised at proficiency, not at signature, and the interval between them carries dual-running cost, elevated handle time and elevated error. Where the saving is committed to a financial year, the ramp interval determines the latest date the move can start — a constraint that is usually discovered rather than planned.

Decompose before sourcing

The corrective is to make the decomposed scope, not the queue, the unit of the sourcing decision.

This is not merely a risk control; it improves the economics on three axes simultaneously. A fast-to-proficiency scope ramps sooner, so the saving arrives earlier and the dual-running interval shortens. It draws on a wider labour pool at the receiving location. And because it is a shallow scope, it can carry a second skill — so it remains chainable, where a deep migrated scope would not be.

The corollary is that chain design is a precondition for sourcing strategy rather than a parallel initiative. An operation that has decomposed its work can source precisely; one that has not can only move queues.

Chainability differs by tier

Receiving tier Chainability Why
Captive or in-house service centre Highest Common employer, systems, taxonomy and skill definitions; edges to the existing chain are configuration rather than negotiation
Contracted vendor, dedicated Moderate Edges are possible but must be specified commercially and are rarely written into the contract
Contracted vendor, shared or multi-client Lowest Capacity is not the buyer's to redirect; the chain terminates at the contract boundary

Where a captive tier exists and is under-used, it is the tier that preserves the most flexibility per unit of cost saved. That is a distinct argument from the usual continuity and control case for captive delivery, and it is generally the stronger one.

Cross-boundary edges are cheap insurance

The edges that connect a new site back into the existing chain are the ones nobody proposes, because they cross an organisational, commercial or geographic boundary that the sourcing decision has just drawn. They are also inexpensive — the chaining result requires two skills per resource, not universal capability, so connecting a site into the estate costs a second skill on a minority of its agents rather than a cross-training programme.

Specifying that connectivity at the point of sourcing costs little. Retrofitting it after a dedicated pool has been established costs a great deal more, because by then the pool has its own volumes, its own performance targets and a commercial construct written around its isolation.

Triage matters more once the estate is distributed

Entry error is more expensive in a multi-site model. A contact that enters at the wrong node may now cross a site boundary, a timezone and sometimes a contractual boundary before reaching the right one, and each of those raises the cost of the transfer above the single-site case.

Accurate triage therefore returns more in a distributed estate than in a consolidated one, and it does so without requiring cross-training across sites — which is the expensive and slow way to connect a distributed graph. For an operation pursuing a cost-driven sourcing programme, triage accuracy is among the few interventions that improve unit cost and responsiveness in the same direction.

What to require of a sourcing business case

  1. The migration unit — queue or decomposed scope — stated explicitly, with the proficiency curve of what is being moved
  2. Ramp cost and duration as an explicit offset, and the resulting latest start date for the saving to land in the target period
  3. The chain edges connecting the receiving location back to the existing estate, specified before the commercial construct is written
  4. The elasticity effect, stated alongside the cost effect. A case that reports only unit cost is optimising one term of a two-term problem
  5. For contracted delivery, whether the arrangement buys overflow to standing proficient capacity or is a hiring arrangement wearing an elasticity costume[8]

Failure modes

Failure What it looks like Correction
Disconnected short chains Cross-training reported as complete; flexibility never materialises across units Add the boundary-crossing edges that close the loop
Paper chain Skills matrix shows coverage; the router does not use it Audit realised routing against the designed graph
Silent breakage A chain degrades through attrition and nobody notices Monitor edge currency; treat a break as an incident
Skill decay Secondary skills exist but have not been used in months Deliberately route a minimum volume to secondary skills to keep them live
Over-training Third and fourth skills added at rising cost and negligible benefit Stop at a closed chain; spend the increment elsewhere
Chaining correlated nodes Edges connect nodes that surge together Prefer edges between weakly or negatively correlated demands
Deflection mistaken for flexibility Volume falls, responsiveness falls with it Place automation inside the graph and measure connectivity, not only containment
Sourcing by queue A migrated queue carries its complex scopes with it; quality falls months later and is attributed to the site Make the decomposed scope the migration unit
Ring-fenced new site Cost saving realised; capacity cannot circulate in either direction Specify cross-boundary edges at the point of sourcing, not afterwards

Measuring whether the chain works

Designed flexibility and realised flexibility are different quantities, and only the second matters.

  • Realised cross-skill utilisation — the proportion of contacts handled on a secondary skill. A designed chain with near-zero secondary utilisation is not operating.
  • Transfer rate by entry point — the direct measure of entry error, and the quantity accurate triage is supposed to move.
  • Chain connectivity — whether the graph is a single connected component. Compute it; do not assume it.
  • Response to imbalance — when one node runs hot and another cold, does work actually move, and within what time? This is the outcome the chain exists to produce.
  • Edge currency — the distribution of time since each secondary skill was last used.
  • Cross-boundary circulation — where the estate spans sites, tiers or suppliers, the volume actually moving across each boundary. A boundary with zero circulation is a chain break regardless of what the skill records show.

Maturity Model considerations

  • Levels 1–2. Cross-training is ad hoc and driven by individual capability rather than design. Chains, where they exist, are accidental and usually disconnected.
  • Level 3. A skill graph exists as data and routing can execute it. The typical finding at this level is that designed and realised flexibility diverge sharply.
  • Level 4. Chain topology is designed deliberately, monitored for breakage, and includes boundary-crossing edges. Decomposition of work into chainable scopes begins, and sourcing decisions are taken at scope level rather than queue level.
  • Level 5. Automated capacity is placed as nodes within the graph rather than outside it, triage raises effective connectivity, and human and automated capacity are chained together in one design.

See Also

References

  1. Jordan, W.C., Graves, S.C. (1995). Principles on the benefits of manufacturing process flexibility. Management Science.
  2. Wallace, R.B., Whitt, W. (2005). A staffing algorithm for call centers with skill-based routing. Manufacturing & Service Operations Management 7(4), 276–294.
  3. Bassamboo, A., Randhawa, R.S., Van Mieghem, J.A. (2012). A little flexibility is all you need: On the asymptotic value of flexible capacity in parallel queuing systems. Operations Research.
  4. Simchi-Levi, D., Wei, Y. (2012). Understanding the performance of the long chain and sparse designs in process flexibility. Operations Research 60(5), 1125–1141.
  5. Kim, Y., Krishnan, R., Argote, L. (2012). The learning curve of IT knowledge workers in a computing call center. Information Systems Research.
  6. Dell'Acqua, F., McFowland III, E., Mollick, E.R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., Lakhani, K.R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. SSRN 4573321. Published in Organization Science (2026).
  7. Wang, Y., Zhu, C., Feng, T., Lu, L.X., Jia, B. (2026). Agentic AI and human-in-the-loop interventions: Field experimental evidence from Alibaba's customer service operations. arXiv:2605.14830.
  8. Koçağa, Y.L., Armony, M., Ward, A.R. (2015). Staffing call centers with uncertain arrival rates and co-sourcing. Production and Operations Management. Preprint: arXiv:1404.2938.