Comparing Delivery Arrangements

From WFM Labs


Comparing delivery arrangements is the problem of deciding which of several ways of getting work done — in-house teams, lower-cost internal centres, contracted suppliers — is actually performing better, when they may be doing different kinds of work, on different channels, cut into different numbers of steps.

The problem is usually posed as a comparison between places or between suppliers. Posed that way it has no answer. The central claim of this page:

A labour source cannot be compared. Only a labour source doing a specified kind of work, on a specified channel, at a specified level of decomposition, on a specified mix of cases, can be compared.

Everything else follows from that. For the decision framework this measurement feeds, see Sourcing Design Axes: Node and Client Ownership. For how work is decomposed in the first place, see Service Chain Decomposition and Node Sourcing. For how many scored interactions any of it requires, see Sample Size and Detectable Difference in Quality Measurement.

The comparison key

What this means in practice. A monthly pack puts two delivery locations side by side. One scores 91%, the other 87%, and the four-point gap becomes the meeting.

What the pack does not say is that the first location handles simple date changes on voice, end to end, by people with four years' tenure — and the second handles disrupted-travel re-accommodation split across two teams, on email, where the first team researches and a second team completes it.

Those are not the same job. They are not even the same shape of job. The four points are measuring the work, the channel and the hand-off design all at once, and attributing the total to the location.

A comparison between delivery arrangements is valid only when four things are held constant. Together they form the comparison key:

Element Why it must match What varies if it does not
Node — the kind of work Intake, judgment and fulfilment demand different skills and carry different error costs You compare a research task against a decision task
Channel Voice, email and chat differ in expectation, pace and what the customer can see You compare a live conversation against a written reply
Decomposition depth End-to-end handling and chained handling fail in different ways and produce different metrics You compare one person's outcome against a chain's outcome
Case mix Harder work scores lower at identical capability You compare the difficulty of the work, not the quality of the delivery

Vary any element of the key and the comparison measures the key, not the arrangement. Most estate comparisons vary all four at once and report the result as a supplier or site difference.

The practical discipline is modest: state the key beside every comparison. Where the key cannot be matched, the honest output is that the comparison is unavailable — which is a finding, not a failure.

The only unit that survives different decompositions

What this means in practice. One arrangement resolves a request in a single twelve-minute call. Another splits it: a four-minute call to capture and research the request, then six minutes of offline work to complete it, by someone cheaper.

Compare them on handle time per interaction and the second looks dramatically better — its longest touch is four minutes. Compare them on cost per interaction and it looks better again, because it has two cheap touches instead of one expensive one.

Both readings are nonsense. The second arrangement used ten minutes of total effort across two people and produced two interactions where there was one. The only honest question is what it cost, end to end, to resolve the customer's request once — and whether the customer got the same outcome.

When arrangements cut the work differently, their node-level metrics cannot be compared, because they do not have the same nodes. Handle time, interactions per hour and per-interaction quality scores are all defined against a unit that has changed shape.

The resolved case is the only denominator that survives. Whatever path a request took, however many people touched it, the customer had one problem and now has one outcome.

Two measures follow, and both are comparable across any decomposition:

  • Chain-total effort per resolved case — every minute at every node, including the coordination time at each hand-off, and including rework and repeat contact.
  • Outcome quality per resolved case — whether the customer's problem was actually solved, judged end to end rather than per touch.

A decomposed arrangement that improves every node metric while raising chain-total effort per resolved case has made the operation worse and will report that it made it better. This is the single most common way a well-intentioned redesign is misjudged.

Decomposition changes how failure happens

What this means in practice. Think of a restaurant kitchen. One chef cooks your whole dish: there are few ways for it to go wrong, but every one of them depends on that chef's judgment, and a weak chef produces a weak dish with nothing to catch it.

Now split the job — one person preps, one cooks, one plates. Each task is simpler and easier to train, and a weak performer at any station is easier to spot and correct. But there are now three hand-offs where the order can be misread, an ingredient forgotten, or a special request lost.

Neither kitchen is better. They fail in different places, and an inspection regime designed for one is blind to the other. Watching the chef will not catch a lost note between stations.

The two arrangements have genuinely different risk profiles:

End-to-end delivery Chained delivery
Number of failure modes Few Many
Nature of each Complex, judgment-dependent Simple, procedural
Where failure concentrates In the interaction At the hand-offs
Easiest to prevent by Selecting and retaining scarce expertise Designing the hand-off and carrying context
What conventional interaction-level QA sees Almost all of it Only the part inside each touch

Decomposition trades a small number of complex failure modes for a larger number of simple ones. That is a real trade with a real upside — simple failures are more preventable, and simpler tasks reach competence far faster — but it relocates risk to exactly the place most quality programmes do not look.

The practical consequence: a chained arrangement can score worse on interaction-level quality while producing better outcomes per resolved case, or better on interactions while producing worse outcomes. Both happen. Only chain-level measurement distinguishes them.

Measuring a chained arrangement therefore requires seam instrumentation that end-to-end delivery never needed: repeat-contact rate by hand-off, time spent re-deriving what an earlier step already established, and completeness of the context passed forward.

What makes an arrangement fit for a node

What this means in practice. Asked which delivery option is better, most packs answer with a rate. But four different things decide whether an arrangement suits a piece of work, and price is the one that says least about whether it will succeed.

An arrangement can be the cheapest available and still be wrong for the work — because it has no depth of experience for the judgment involved, because it cannot flex when volume moves, or because once work is placed there it can never be moved back out without a negotiation.

Four properties, and a rate card exposes only the last:

Property The question it answers How it is evidenced
Proficiency depth Does this arrangement have people who are actually good at this node? Tenure distribution, ramp maturity, share of the team past the competence curve
Elasticity Can it absorb a surge, and give capacity back in a trough? What fraction of its capacity is genuinely variable at 30, 60 and 90 days
Chainability Can work move in and out of it without a negotiation? Whether connections to the wider estate are configuration or contract
Cost per resolved case at the required standard What does it actually cost to finish the job properly? Chain-total cost, including rework, escalation and management overhead

Note what is absent: who employs the staff, and where they sit. Neither predicts quality. Observed differences between an in-house centre and a supplier, or between two teams inside the same supplier, are generally explained by tenure stability and case mix. Employment relationship determines cost and chainability; it does not determine capability.

Measurement as an input to routing

What this means in practice. Most quality reporting exists to answer a backwards-facing question: was last month's delivery what was paid for? The answer arrives, someone is asked to explain a number, and nothing about tomorrow changes.

The same measurement, organised by the comparison key, answers a forward-facing question instead: given this piece of work, where should it go? One arrangement demonstrably handles disruption work better; another finishes routine fulfilment more cheaply at the same outcome; a third is the only one that can absorb a spike next Tuesday.

At that point the sourcing conversation stops being an annual negotiation about volumes and becomes a routing decision made continuously on evidence.

This is the destination the framework is for. Measurement organised as assurance produces a monthly argument; measurement organised by comparison key produces a fitness map — which arrangement is best at which node, on which channel, at what cost and with what flexibility.

Three properties make a fitness map usable:

  1. It is stated per key, not per supplier. An arrangement is not "good"; it is good at something specified.
  2. It carries its own precision. Every entry shows the observations behind it and the difference it could actually detect, so a reader knows which entries can bear weight. See Sample Size and Detectable Difference in Quality Measurement.
  3. It is expressed as attributes, not as a list of sites. Routing on computed eligibility rather than static assignment is what allows the map to change without an organisational change. See Next Generation Routing.

The map does not decide alone. Eligibility constraints and client-ownership commitments still bind, and both sit upstream of any routing preference — see Sourcing Design Axes: Node and Client Ownership. What the map removes is the need to settle placement by assertion.

Failure modes

Failure What it looks like Correction
Comparing labour sources "Site A versus Supplier B", with no key stated State node, channel, decomposition depth and case mix, or withdraw the comparison
Node metrics across decompositions Handle time or per-interaction quality compared between an end-to-end and a chained arrangement Compare chain-total effort and outcome per resolved case
Counting interactions as outcomes A chained design reports more, cheaper interactions and is judged an improvement Count resolved cases; two touches for one request are one case
Interaction-only QA on a chained design Every touch scores well; customers still report the process failed Instrument the hand-offs: repeat contact, re-derivation, context completeness
Fitness inferred from rate The cheapest arrangement is assumed suitable Assess proficiency depth, elasticity and chainability alongside cost
Fitness inferred from employer In-house assumed better, or supplier assumed cheaper-and-worse Look to tenure stability and case mix; employment relationship predicts cost, not quality
Measurement that only looks backwards A monthly argument about last month's number Organise by comparison key and use it to place work

Maturity Model considerations

  • Levels 1–2. Arrangements are compared as places or suppliers. No key is stated, and node metrics are compared across arrangements that cut the work differently.
  • Level 3. Channel and work type are separated in reporting, but decomposition depth and case mix still vary silently inside any comparison.
  • Level 4. Comparisons carry an explicit key and are made on resolved cases. Seam instrumentation exists where delivery is chained, and precision is reported alongside every figure.
  • Level 5. A fitness map by comparison key informs routing continuously, expressed as work attributes rather than as site assignments, so it survives both organisational and estate change.

See Also