Sourcing Strategy Under Cost Pressure
Sourcing strategy under cost pressure is the problem of reducing the cost of service delivery within a defined period, where the levers available are where work is performed, how it is divided, and what performs it. It differs from sourcing design in general — covered in Sourcing Design Axes: Node and Client Ownership and [[Service Chain Decomposition and Node Sourcing]] — by having a number and a date attached.
This page describes a two-horizon approach: a twelve-month programme that operates against constraints already visible, and a longer programme that builds the capability to see the rest. The end state of the second is not a better sourcing position. It is a placement function — where work is performed, computed continuously from stated objectives and recorded constraints, as the longest horizon of the same workforce model that ends in real-time reallocation.
The situation, as of August 2026
Automation has moved from conversational to transactional. Organisations planning their next financial year are, in many cases, being asked to remove a material share of service delivery cost — and to do so while the work itself is changing shape underneath them.
The instinct is to answer with a sourcing position: move a percentage offshore, consolidate onto a vendor, expand a captive footprint. Positions of that kind are the reason most such programmes under-deliver. They are decided once, against conditions that then change, and they cannot be re-derived when they do.
Why the decision resists a single answer
Three properties make the problem combinatorial rather than directional.
The option space is large. Every routing group, every work type and every delivery node combine into a number of candidate configurations far beyond what can be held in a discussion. Add the legitimate and differing opinions of the leaders responsible for each part of the estate, and the number of defensible answers is effectively unbounded.
The estate is not uniform. Where several business units have been combined, each arrives with its own sourcing mix, its own history and its own view of what should move. Each of those views was correct for the business that formed it. None were reconciled.
Quality cannot currently discriminate between options. Where quality is measured on different instruments in different parts of the estate, at sampling rates that cannot resolve small differences, the recurring argument — that a vendor or a low-cost centre performs less well than an in-country team — can be neither confirmed nor dismissed with evidence. It is therefore settled by seniority, which is to say not settled.
The consequence is that a sourcing debate can run indefinitely without converging. What resolves it is not a better position but a mechanism: constraints written down, objectives stated, and the distribution of work computed rather than argued.
The short horizon: the first twelve months
The short horizon operates against constraints that are already visible. It does not wait for the capability described below.
Set the internal target above the requirement
If the requirement is $XXM, set the internal target materially higher — a factor of
roughly one and a half is a reasonable starting point.
This is not optimism. Programmes of this kind lose candidate moves during execution: a constraint surfaces late, a receiving location cannot absorb the volume, a commercial commitment blocks a transfer. A target set at the required number guarantees a shortfall, because it sizes the search to exactly the answer. A higher target forces a wider search, and the surplus candidates are what absorb the losses.
State the denominator. A saving expressed without the base it is measured against cannot be checked, and estates commonly carry more than one plausible base.
The move that pays: deferrable work out, gates consolidated
The most reliable short-horizon move is to shift asynchronous or deferrable work — email, back-office fulfilment, follow-up tasks — out of routing groups and into lower-cost delivery. It is attractive because it carries no live customer interaction, and therefore the lowest structural risk to experience.
But the move alone does not produce a saving, and can produce a loss.
Removing deferrable work from a routing group reduces its offered load. Required staffing behaves approximately as , so the proportional safety cushion is and grows as load falls. A smaller group therefore carries a larger fractional cushion at the same service target: occupancy falls, and the work that remains costs more per unit than it did before.
Where the deferrable work was previously absorbing that cushion — being handled in time the organisation was paying for regardless — its removal converts productive time into idle time and the saving is offset or reversed.
The resolution is that the two actions are one action.
| Step | Effect |
|---|---|
| Remove deferrable work from a routing group | Offered load falls; proportional cushion grows; occupancy degrades |
| Consolidate that group with another | Load recovers; cushion shrinks; headcount falls out |
The saving comes from the consolidation, not from the transfer. The transfer is what makes the consolidation possible, by removing the work that differentiated the groups.
| Worked example — two pools of twenty | |||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
Two routing groups, twenty rostered each. Voice interaction 25 minutes, after-call work 5 minutes — so handle time is 30 minutes. Shrinkage 30%, giving fourteen agents on the phone. Service target 80% within 30 seconds. Each group carries 20 calls per hour. Email is not counted as shrinkage. It is productive work harvested from availability — the pockets of idle time that exist because the service level requires them. This matters, and is returned to below.
Reducing after-call work saves people. It removes four minutes from every interaction, so offered load falls from 10.00 to 8.67 Erlangs and one agent per pool comes off the phone. That is a genuine reduction, and it is available whether or not anything is consolidated. Moving the email out saves nobody. Offered load is unchanged, because the email was never part of it — it was being absorbed in idle time the voice service level requires regardless. The fourteen, now thirteen, agents are still needed to answer 80% of calls within 30 seconds. All that changes is that four agent-hours per hour of previously productive time becomes visibly idle. Note the occupancy. It falls from 71.4% to 66.7% after the after-call work is removed. The pool has become smaller in load terms, and a smaller pool must run at lower occupancy to hold the same service target. Each remaining interaction is now carried by a more expensive hour than before. The merge is where the headcount is. Combined offered load is 17.33 Erlangs, and the pooled group needs 22 agents where two separate groups needed 26. Occupancy recovers to 78.8%. Six of the eight heads removed across the whole exercise come from this step alone. The measurement problem underneath the example. The baseline shows four agent-hours per hour of idle capacity, and email was being harvested from it. How much was actually harvested is not observable in most estates. Availability is recorded as availability whether it was filled or not. So the productive work being lost when the email leaves is real, unmeasured, and will not appear in any before-and-after comparison — which is one reason such transfers frequently under-deliver against their business case. And the error runs in both directions. An agent who picks up an email or shifts channel while waiting for the next call has self-directed into productive work. They are genuinely occupied. But the state model does not see it: they are sitting in an available state, so the measurement records idle time. Occupancy is understated and the person is working. If that same agent instead moves into an auxiliary state defined as productive, the identical work is captured, and occupancy reflects it. The difference between those two cases is not the work. It is whether the agent signalled it. So the measured number depends on agent behaviour and state discipline rather than on what actually happened — and the same operation can appear more or less efficient depending on habits nobody has deliberately set. There is a third case that cuts the other way: an agent parked in a productive auxiliary state who is not producing. The state model can flatter as easily as it can undercount. This is discoverable, and it should be discovered. It does not change the argument — the consolidation is still where the headcount is — but it changes how much confidence any before-and-after comparison deserves, and it is a prerequisite for the next point. Where this leads. If idle capacity inside a gated group is real and currently invisible, the design response is not only to consolidate. It is to capture the idle deliberately. Idle-time automation platforms detect the pockets between interactions and push work into them — either frontline work that would otherwise queue elsewhere, or training, development and coaching that would otherwise require the agent to be taken off the floor entirely. Both convert availability the service level already requires into output. Neither is possible while the capacity is unmeasured, which is why the measurement work comes first. Figures are illustrative and computed with Erlang C at the stated assumptions. The pattern, not the arithmetic, is what transfers. |
This changes the selection rule. Candidates are not the groups with the most deferrable work, nor simply the largest groups. They are the groups that can be merged once the deferrable work is removed. Two mid-sized groups stripped and combined will frequently yield more than one large group stripped and left standing.
Search for capacity before scaling capacity
Where a receiving location already holds unused capacity, using it avoids the cost of expansion. Where it does not, expansion carries a ramp cost, and contraction later carries an exit cost. Programmes are commonly costed on the steady-state difference in unit rates and lose most of the benefit to the transitions at either end.
Establish existing headroom across the network before authorising any expansion.
Select on structural risk, not measured quality
Where quality instruments cannot resolve differences at the scale that would change a decision, quality cannot function as a selection criterion. Claiming that a programme will "minimise quality impact" is therefore unverifiable in both directions.
The defensible alternative is to select on structural risk. Work with no live interaction, no real-time judgement and no recovery moment — where an error is correctable before the customer encounters it — is structurally lower risk irrespective of how it is measured. That is the reason deferrable work moves first.
This produces a natural two-phase logic:
| Phase | Selection criterion | Because |
|---|---|---|
| First | Structural risk | The instrument cannot resolve measured differences |
| Second | Measured quality as an explicit lever | Once it can |
Which gives the measurement programme a purpose and a date rather than a claim to importance. The addressable set expands when the instrument improves — and not before.
Pair every move with a headcount exit
Moving work does not reduce cost. People leaving reduces cost.
For a period, both the origin and the receiving location are staffed for the same work, and the programme is running a cost increase. The saving is realised only when the origin headcount exits, and that exit is the slowest and most constrained part: notice periods, statutory protections, locations where transfer is not permitted at all, and funding that must come from somewhere.
Every candidate move should therefore carry an explicit exit path — which roles, from where, funded how, over what period. A move without one is cost duplication with a plan attached.
Where exits can be met by foregoing planned hiring or by not backfilling attrition, they are materially cheaper than severance. Those routes are finite and are usually claimed by other initiatives, which should be established rather than assumed.
Constraint-check candidates before ranking them
A routing group cannot be merged with another unless the work is mutually serviceable. Groups are held apart by language, by regulatory or contractual restriction, by a commercial commitment made at the point of sale, by platform capability, and — frequently — by nothing at all beyond the circumstances of their creation.
That last category is the least expensive material in the entire programme and is invisible until someone asks.
This does not require a complete constraint inventory. It requires one for the groups under consideration, which is a much smaller exercise.
Prefer moves that are cheap to reverse
Decisions taken under uncertainty should be weighted toward reversibility. Where a move proves wrong, the cost of undoing it — reintegration, renegotiation, rebuilding capability — determines whether it is undone at all, or merely endured.
A slightly less valuable move that can be reversed cheaply may be preferable to a better one that cannot.
The long horizon: beyond twelve months
The short horizon operates against what is visible. The long horizon builds the capability to see the rest, and to keep seeing it as conditions change.
The end state: a placement function
The long horizon does not end in a better sourcing position. It ends in a function.
A sourcing position is a fixed answer, and it begins decaying on the day it is written — which is why estates that hold one find themselves re-cutting it every eighteen months at considerable cost. A placement function differs in kind: it does not store the answer, it stores what produces the answer. Objectives, ordered and owned. Constraints, recorded with a cost to relax and a lead time. Node attributes, measured on one instrument. Where work is performed falls out of the middle, and it falls out again whenever an input moves.

Read the figure downward. Objectives set what is preferred; constraints set what is eligible; the optimisation scores the eligible set against the ordered objectives while carrying transition cost; placement is what comes out.
Profitability and durability sit below that line rather than inside it. Profitability is a function of revenue and expense — carrying it as a separate objective double-counts both and invites the optimisation of a derived quantity. Durability is what a function buys that a position cannot: the answer keeps being right while conditions change.
The vocabulary consequence is worth stating plainly. An onshore team, a low-cost service centre, a third-party vendor and an automated node are outputs of this function, not inputs to it. Each is where work lands under a given ordering of objectives and a given constraint set. Treating any of them as the strategy is what produces the percentage target — move thirty per cent offshore — decided before anyone has established what is being optimised or what is eligible.
Where this sits in the workforce model
Placement is not a separate machine to be built alongside the operation. It is the longest horizon of the workforce model the operation already runs.

The same estate is planned at five clock speeds. At twelve to thirty-six months, long-range planning and a simulation engine size capacity and design the network. At annual and quarterly, the capacity plan commits headcount and budget. At weekly and daily, forecasting and scheduling place people against interval demand. Intraday, automation and real-time reallocation move supply toward emerging demand. In real time, the contact platform routes one contact to one person.
These are the same operation at different clock speeds. Deciding which country serves a routing group and moving one person between queues at ten fifteen are both placement decisions. They differ in horizon, in which constraints are binding, and in how expensive the decision is to reverse. Intraday, nearly every constraint is hard; at three years, nearly none are. This is also why a single ordering of objectives has to serve all five horizons — where each horizon holds its own ordering, the estate optimises for cost at one and continuity at another, and the two quietly cancel.
Two components carry most of the weight, and they are usually the two that are missing.
The capability layer. An attribute model describing what each unit of supply is eligible, proficient and entitled to do — held above the contact platform rather than inside it, because a platform's skill model is built for stable skill sets and pools do not hold still. It is the join between what a person can do and where a piece of work may go. Without it, eligibility is encoded as static membership that nobody maintains, and the constraint set the placement function reads is fiction.
Advanced analytics. Elasticity, attribution and capability scoring are what convert an estate that describes what happened into one that estimates what would happen. This is commonly the acknowledged gap, and it is load-bearing here: a constraint monitor and a scoring model are analytics products, not reporting products.
Build the constraint inventory from the first day
The capability layer — an ontology describing every delivery node, its attributes, and every restriction bearing on what work may reach it — is a substantial build. The inventory of constraints is not.
It can begin immediately, in a spreadsheet, and it pays from the first month because the short horizon already needs it. For each routing group: what holds it apart, whether that restriction is contractual, regulatory, capability-based, platform-based or merely historical, what it would cost to relax, and how long relaxation would take.
Separating the two matters for durability. If the tooling is delayed, or priorities shift, or sponsorship changes, the inventory survives and continues to produce decisions. The engine makes the work faster; it does not make it possible.
Instrument the objectives before building the model
A placement model optimises whatever is measured about its options. Where cost is the only attribute recorded for a delivery node — a loaded rate and little else — the resulting model will produce a cost answer regardless of what its stated objectives claim.
Before the engine exists, therefore, establish at least one measured attribute per node for each objective the organisation intends to hold. Four are sufficient to describe most service estates.
| Objective | What it holds | Minimum measured attribute per node |
|---|---|---|
| Revenue and growth | contribution generated by the work, held as a property of the work rather than of the location | revenue or margin attributable to the work performed at the node |
| Expense | the fully loaded cost of performing it | cost per unit of work, inclusive of governance, coordination, ramp and exit |
| Quality and experience | what the customer receives | score on a single instrument, case-mix adjusted, at a resolution the sampling can actually support |
| Risk and flexibility | the speed and cost of adjusting capacity in either direction | notice period, cancellable share, time and cost to add capacity at proficiency, concentration on critical work |
Two quantities are frequently proposed as objectives and should not be carried as such. Profitability is a function of revenue and expense; holding it separately double-counts both. Durability — that a decision remains defensible as conditions change — is a consequence of building a function rather than a position, and is bought by the constraint monitor rather than traded against cost.
Transition cost belongs in the objective as a first-class term, not as a business-case footnote. A model that recommends placement without carrying the cost of arriving there will recommend moves that do not pay.
Why flexibility is an objective, not a constraint
Continuity is usually handled as a rule — critical work may not concentrate in one location — and flexibility is usually handled nowhere at all, or as a line in a risk register. Both belong inside the fourth objective, and the reason is economic rather than tidy.
A model that does not price flexibility will systematically buy the cheapest and most rigid capacity available. Rigidity is invisible in steady state. A node on a long committed term, with a lengthy notice period and a slow ramp, is indistinguishable from a flexible one on every attribute a cost model records — until demand moves, at which point the difference is the entire cost of the event. An objective set with no flexibility term in it will therefore recommend, correctly by its own arithmetic, precisely the estate that cannot respond.
Where demand can move sharply within a single quarter — travel, event-driven and weather-exposed demand, seasonal retail, and anything exposed to a systemic shock — the ability to add and remove cost quickly carries a value that can exceed the unit-rate difference it is usually traded against. It is bought deliberately, the way insurance is, or it is discovered to be absent at the worst possible moment.
Four attributes make it measurable rather than rhetorical.
- Notice period and cancellable share. What proportion of the node's capacity can be stood down, on what notice, at what penalty. A committed minimum volume is a floor under cost that no forecast can lower.
- Time and cost to add capacity at proficiency. Not time to hire. Time to productive contribution, which is a ramp question and belongs to the same body of work as Supply Elasticity in Workforce Planning.
- The asymmetry between adding and removing. These are rarely symmetric and should be held as two attributes, not one. A node that scales up in weeks and down in quarters behaves very differently from its mirror image, and a single "flexibility" score conceals which one is being bought.
- Concentration on critical work. Continuity expressed as a measured attribute rather than a separate rule — what share of a critical routing group a single node holds, and what the recovery path is.
The asymmetry is the one most often missed, because procurement negotiates the rate and the ramp while the exit terms are settled by whoever is least interested in them.
State the objective as an ordered list with an owner
The relative priority of the four objectives is a leadership choice, not a modelling one. It also legitimately changes — with the business cycle, with ownership, with strategic phase. A business under cost pressure and the same business after a demand shock will order them differently, and both orderings are correct at the time.
Publishing the objective as an ordered set of priorities, dated, owned, and reorderable by a named authority converts a change of priority into a governance act with a signature, rather than a crisis for the model. Without it, whoever holds the model chooses the objective by default, which is neither legitimate nor stable.
Treat automation as a node, and measure it accordingly
Where automation can perform a piece of work, the question of whether it should is the same class of decision as which location should — asked of a node with unusual attributes. It follows that the automated node requires attributes measured on the same instrument as every other: cost, capability, quality, and failure behaviour.
Containment rate is not such an attribute. It is an outcome, and it conflates three distinct states.
| State | What happened | Effect on workload |
|---|---|---|
| Deflected | The contact never arrived | Genuine reduction — and the only available signal about demand drivers |
| Contained | Arrived; automation completed it | Unchanged. Supply substituted, workload not removed |
| Escalated at depth | Automation completed part of it; a person completed the rest | Transformed — neither removed nor unchanged |
Most estates measure the first two as one figure and the third not at all. A binary containment measure cannot distinguish automation that completes most of an interaction and hands over cleanly from automation that completes little and hands over a customer who must begin again. Both are recorded as not contained, and the residual human work differs enormously.
The missing attribute is completion depth: the distribution of how much of an interaction automation finishes before handover, and the residual human effort per escalation. It serves three purposes at once — a capability measure of the automated node, comparable across products and versions; a workload input, since residual effort is real work; and the basis on which the cost of escalation becomes computable rather than asserted.
Without it, no comparison between automated and human delivery can be made, and every claim about how much work automation will absorb remains an assertion.
Watch what is binding, not only what is optimal
The durable output of the long horizon is not a placement but a constraint monitor: a standing answer to the question of what is currently holding each routing group apart, and what has changed.
Is the group held by a capability gap that a tooling programme is scheduled to close? By a commercial commitment that expires on a known date? By a platform limitation? By specialist knowledge that a shielding layer would make unnecessary? Each of those has a different remedy, a different cost and a different horizon, and each moves independently.
A model that computes placement is a calculator. One that reports what is binding and when it changes is an operating capability, and it is the difference between a strategy that decays and one that re-derives.
Anticipate a role that does not exist yet
As automation absorbs the routine portion of interactions, a class of work emerges that is neither fully automated nor conventionally handled: a person supervising several concurrent automated sessions, intervening on judgement, and approving escalation. The engagement is short — minutes rather than a full interaction — but it is not queue-based, and conventional staffing mathematics does not describe it.
Organisations typically create this function informally, project by project, as an audit or assurance activity. Designing it deliberately — including where it can be located, which differs from where delivery can be located — is a long-horizon requirement.
Prerequisites
Three, and they are not equally within one function's control.
Data definitions and standards. Occupancy, utilisation, handle time, volume, the treatment of non-voice work, and the accounting of every minute of paid time — defined identically across the estate, with the source system named for each. Without this, no comparison between delivery options means anything, and quality cannot become a lever.
The capability layer. The ontology, node attributes and eligibility model. Preceded, as above, by the inventory.
A minimal set of commercial service offerings. Every distinct tier sold is a constraint on where and how work may be delivered. Fewer tiers means a larger feasible set, which makes this the single largest expansion of what any placement model can compute.
Commercial offerings are a constraint class, not a strategy dependency. The business owns whether to simplify them. The model functions either way, and functions better with fewer — which is a stronger argument for simplification than consistency alone.
Failure modes
- Costing on unit rates alone. A vendor rate is a price containing a margin, and excludes ramp, exit, governance and coordination cost. Two options at the same apparent rate can carry very different real economics.
- Treating the move as the saving. Cost falls when people leave, not when work relocates.
- Stripping deferrable work without consolidating. Converts absorbed cushion into visible idle time and can produce a net loss.
- Selecting on measured quality that the instrument cannot resolve. Produces confident conclusions from noise.
- Optimising over poorly measured attributes. The option a model selects is disproportionately likely to be the one whose estimation error was most favourable, so realised performance will tend to disappoint the estimate. Shrink each node's most flattering attribute toward the fleet mean before placing work on it.
- Buying the cheapest capacity because flexibility was never priced. Where the objective set contains no flexibility term, the lowest-cost option wins by construction, and the estate discovers what it bought only when demand moves.
- Re-optimising every period without carrying transition cost. Under sunk costs the optimal policy includes a band of inaction. A model that recomputes an optimum each cycle and acts on it will generate churn whose transition cost exceeds its steady-state gain.
- Publishing a ranked league table of delivery locations. Sorting is on comparative advantage across several attributes, not overall quality; a single ranked ordering is both theoretically unsound and organisationally corrosive.
Maturity Model considerations
- Initial / Foundational — sourcing decided by position and seniority; quality compared across incompatible instruments; savings costed on unit rates.
- Progressive — constraints written down for the estate under active consideration; moves paired with exit paths; deferrable work and consolidation treated as one action.
- Advanced — objectives instrumented per node across all four, flexibility included; automation carried as a node with completion depth measured; placement computed rather than argued.
- Pioneering — a standing constraint monitor; placement re-derived continuously as constraints relax and objectives are reordered; one ordering of objectives serving every horizon, so that the same placement logic runs from the strategic horizon down to intraday.
See Also
- Sourcing Design Axes: Node and Client Ownership
- Service Chain Decomposition and Node Sourcing
- Business Process Outsourcing
- Supply Elasticity in Workforce Planning
- Chaining and Flexibility Design
- Occupancy
- Service Level
- Capacity Planning
- WFM Labs Maturity Model™
References
- Borst, S., Mandelbaum, A. & Reiman, M. (2004). "Dimensioning Large Call Centers." Operations Research 52(1):17–34.
- Whitt, W. (1999). "Partitioning Customers into Service Groups." Management Science 45(11):1579–1592.
- Mandelbaum, A. & Reiman, M. (1998). "On Pooling in Queueing Networks." Management Science 44(7):971–981.
- Smith, J.E. & Winkler, R.L. (2006). "The Optimizer's Curse." Management Science 52(3):311–322.
- Dixit, A. (1989). "Entry and Exit Decisions under Uncertainty." Journal of Political Economy 97(3):620–638.
- Williamson, O.E. (2008). "Outsourcing: Transaction Cost Economics and Supply Chain Management." Journal of Supply Chain Management 44(2):5–16.
- Larsen, M.M., Manning, S. & Pedersen, T. (2013). "Uncovering the Hidden Costs of Offshoring." Strategic Management Journal 34(5):533–552.
