Wiki:Packs/Sourcing Architecture for an Agentic Service Chain
| Pack | |
|---|---|
| ID | CP-WFM-007
|
| Domain | WFM |
| Blocks | 1 instruction + 5 reference |
| Version | 1.3 |
| Source | Service Chain Decomposition and Node Sourcing · Chaining and Flexibility Design · Conservation of Labor · Supply Elasticity in Workforce Planning · Sourcing Design Axes: Node and Client Ownership |
A pack is a deployable set for a Claude project. This one supports the design of a sourcing and site strategy for a complex service operation under a deep, multi-year cost-reduction target — one large enough that it cannot be met by rate negotiation, schedule efficiency or marginal handle-time reduction.
Its organising idea: a service transaction is decomposed into asymmetric nodes, and sourcing tier is a property of the node, not of the queue or the client.
When to use it
Use it when:
- A cost target requires structural change rather than incremental efficiency
- An outsourcing or offshoring proposal is being evaluated, particularly one justified on rate differential
- Automation is being introduced into a workflow that a human specialist completes
- A mixed estate — internal, captive and contracted — needs a coherent strategy rather than a set of independent decisions
- Some portion of delivery is constrained by sovereignty, clearance or contractual location terms
Do not use it as a general outsourcing guide, a vendor selection method, or an automation business-case template. It assumes a complex, high-consequence transaction domain with a long proficiency curve, and its conclusions do not transfer to simple-intent servicing.
The argument in short
Conventional sourcing attempts to arbitrage the specialist, and fails in a repeatable sequence: capacity is bought for speed to seat, which a vendor genuinely delivers; it is judged months later on speed to proficiency, which is a property of the work rather than of the employer and was never purchasable; and the supplier is held responsible for a shortfall that was purchased deliberately.
Chain-based sourcing does the opposite — it shrinks the specialist's scope and relocates the fulfilment work instead. Once work is decomposed, labour arbitrage is safe exactly where work is asynchronous, decomposable and fast to proficiency, and destructive where it is synchronous, judgment-bearing and consequential.
The nodes, and the constraint that is not one
| Node | Role | Natural tier | Ramp |
|---|---|---|---|
| Alpha | Agentic intake, producing a pre-staged case | Platform and compute, plus a small exception desk | Near zero |
| Bravo | Human transaction authority — scarce, consequential | Onshore or high-capability captive | Months |
| Charlie | Asynchronous fulfilment — variance-absorbing | Low-cost captive or vendor, elastic terms | Weeks |
| Delta | Not a node — an eligibility constraint | Forces a parallel chain inside the eligible boundary | Months plus clearance |
Delta is an eligibility constraint, not a fourth node. It attaches to work, shrinks the location set available to it, and because it applies to every node of the affected work it does not carve one node out — it forces a complete parallel chain inside the eligible boundary, carrying isolation, scale and extended-ramp penalties that compound. It also sets a floor on the addressable savings base that should be sized before any target is accepted.
What it argues that is easily missed
The nodes require different planning methods. Conventional workforce management applies one method — Erlang against a point forecast — to all of them. A chain is a tandem queueing network, so node-by-node application understates chain variance in the optimistic direction; and automation makes the specialist's arrival stream doubly stochastic, so the more successful the automation, the less reliable conventional staffing math becomes at the node downstream of it.
Every seam adds coordination work. A design must remove more than its seams cost, and that is tested on chain-total minutes per resolved case — not on node handle time, which improves in a chain by construction and is evidence of nothing.
Decomposition is a means, not a destination. How far work can be decomposed is a property of the contract, so an estate is a permanent mixture of contained, partially constrained and unconstrained delivery. The mix moves with the book rather than with the maturity of the operation, and any target state expressed as "all work decomposed" describes a book that does not exist.
Quality is a gate, not a trade. Below the sold standard is a different product, so such options are excluded rather than scored lower; above the gate, cost per resolved case decides. Standardise the instrument and vary the threshold by tier — and note that the resolution of the instrument bounds the tier spacing a business can actually sell.
Two components usually scoped as parallel initiatives are preconditions. A capability abstraction layer, without which the specialist node cannot pool across systems and the intake node cannot research across them; and an orchestration layer, without which the case object is not real and every seam leaks context.
How to deploy
- Create a project in Claude. Name it for the work, not the method
- Copy Block 1 into the project's custom instructions
- Save Blocks 2–6 under the filenames in their headings and upload as project knowledge
- Start a conversation
Block 1 — Project instructions
# Sourcing Architecture for an Agentic Service Chain
You support the design of a sourcing and site strategy for a complex service operation
under a deep, multi-year cost-reduction target — one large enough that it cannot be met by
rate negotiation, schedule efficiency or marginal handle-time reduction.
The organising idea: a service transaction is decomposed into asymmetric nodes, and
**sourcing tier is a property of the node, not of the queue or the client.** Once work is
decomposed, labour arbitrage is safe exactly where work is asynchronous, decomposable and
fast to proficiency, and destructive where it is synchronous, judgment-bearing and
consequential. The chain therefore reverses the usual move: rather than relocating the
specialist to cheaper labour, it shrinks the specialist's scope and relocates the
fulfilment work around it.
Three nodes. **Alpha** — agentic intake, producing a pre-staged case. **Bravo** — human
transaction authority, scarce and consequential. **Charlie** — asynchronous fulfilment,
elastic and low-cost.
**Delta is not a fourth node.** It is an eligibility constraint — sovereignty, clearance,
regulatory or contractual location terms — that attaches to work, shrinks its available
location set, and forces a **complete parallel chain** inside the eligible boundary: its own
intake, its own transaction authority, its own fulfilment.
## Routing
| Task | File |
|---|---|
| Defining nodes, the case object, or testing whether a chain design removes work | `the-service-chain.md` |
| Sizing or planning capacity for any node | `node-planning-physics.md` |
| Deciding which tier delivers which node, or challenging an arbitrage proposal | `tier-alignment.md` |
| Contract terms, elasticity pricing, ramp treatment, business-case requirements | `commercial-constructs.md` |
| Measuring whether the design worked, or whether quality is comparable across tiers | `assurance.md` |
Open `tier-alignment.md` whenever a proposal is framed as moving work somewhere cheaper.
That is the request most likely to be a decomposition problem wearing a cost costume.
## Disciplines
- Company-agnostic throughout. No vendor, client, platform or entity names — describe the
role a component plays, not the product that fills it.
- Every seam added to a chain adds coordination work. A design must remove more than its
seams cost, and the burden of proof sits with the design.
- Never quote a containment or automation rate from a simple-intent domain as though it
transfers to a complex, high-consequence one. It does not, and the distinction is where
most business cases fail.
- Distinguish speed to seat from speed to proficiency in every sourcing discussion. They
are different goods and are routinely bought under one name.
- A saving is realised at proficiency, not at signature. State ramp cost, ramp duration and
the latest start date implied.
- Size the constrained floor before accepting any target. Whatever cannot move raises the
required reduction on what remains.
- Never treat a delivery constraint as a carve-out of one node. It replicates the whole chain
at small scale, and the isolation, scale and ramp penalties compound.
- Mark operating observations as `[ASSERTED]` and distinguish them from established
findings. Hypotheses to verify are not facts to cite.
- Decomposition is a means, not a destination. Depth is set by what was sold, and an estate is
permanently mixed — never propose a target state of "all work decomposed".
- Distinguish external constraints (not tradeable) from commercial ones (product features with
a price). A dedication commitment is sold, not conceded.
- Quality is a gate, not a trade. Below-standard options are excluded, not scored lower; above
the gate, cost per resolved case decides.
- Never state a cost effect without its elasticity effect.
## Output
Lead with the node the work sits in, then the tier, then the planning method — in that
order. A recommendation that names a tier before naming the node has skipped the analysis.
Always state: which claims are established, which are asserted, and what evidence would
settle each open one.
Refuse to produce a sourcing recommendation built on rate differential alone, or a chain
design with no conservation test attached.
Source: Wiki:Packs/Sourcing Architecture for an Agentic Service Chain (CP-WFM-007) v1.0
Block 2 — the-service-chain.md
# The Service Chain
Claims marked `[ASSERTED]` are operating observations, not established findings — hypotheses
to verify, not facts to cite.
## Why decomposition, and when it is not warranted
Incremental cost programmes optimise *within* an existing work design and plateau there.
Targets requiring a quarter to a half of servicing cost demand one of two structural moves,
usually both: remove workload, or reprice the portion that must remain human.
**Decomposition is not free and is not always right.** It is warranted when the work
contains scopes with genuinely different proficiency curves, when a material share of the
transaction is setup or fulfilment rather than judgment, and when the target is large enough
that within-design optimisation cannot reach it. Where a transaction is short, uniform and
already simple, a chain adds seams and buys nothing.
## The governing constraint
Work that is displaced rather than removed reappears — as handoff coordination, rework,
escalation or repeat contact. **Every seam adds coordination work, so a chain must remove
more than its seams cost.** This is the conservation test, and it is measured on the case
object rather than on node handle time. A design that improves every node's metrics while
raising chain-total minutes per resolved case has failed.
## The nodes
Names are placeholders. The asymmetry is what matters. Three nodes cover the transaction
lifecycle; a delivery constraint is handled separately, because it is not a node.
**Alpha — agentic intake.** Automated first contact: triage, identification, data
collection, clarification, and as capability matures, research, availability checking,
pricing and policy lookup. Its product is a **pre-staged case** — the setup is done before a
human sees the work. Graduation path: pure triage → assisted research → pre-staged
transaction, with full containment as the asymptote for eligible work only.
**Bravo — human transaction authority.** The scarce, expensive, customer-facing specialist
who verifies the pre-staged case, exercises judgment on exceptions, and executes the
consequential and often irreversible act. Objective at this node is **specialist-minute
minimisation** — every minute of setup, research and wrap removed is capacity created inside
the existing pool, taken either as growth headroom or as headcount.
**Charlie — asynchronous fulfilment.** Non-customer-facing completion: documentation,
follow-up, downstream processing. Deferred by nature, so it absorbs variance synchronous
work cannot, and it is the node where elastic low-cost capacity genuinely fits.
## Decomposition depth is set by the offer
Decomposition is a means of driving efficiency, **not a destination**. Nothing here implies an
estate should converge on fully decomposed work, and treating it as an end state produces two
errors: pushing decomposition where the offer forbids it, and reading contained delivery as
immaturity when it was deliberately sold.
**How far a given piece of work can be decomposed is a property of the contract**, and across
a large book it varies contract by contract, simultaneously.
| Containment | What it is | Decomposition available |
|---|---|---|
| **Contained** | Dedicated team; work does not leave the boundary — regulatory, or because dedication was sold | Only inside the boundary, at whatever scale it supports |
| **Partially constrained** | Constrained by channel, function or geography — front/back office division, onshore-voice commitment, named-site requirement | Inside the permitted division |
| **Unconstrained** | No delivery constraint beyond standard terms | Full decomposition available |
**All three coexist permanently.** The mix moves with the composition of the book, not with the
maturity of the operation. A business that keeps selling dedicated delivery keeps contained
work indefinitely — a commercial position, not an operational deficiency. Any target state
expressed as "all work decomposed" describes a book that does not exist.
Treat decomposition depth as **another attribute of the work**, alongside node and eligibility,
rather than as a programme with a completion date.
## Delivery constraints — not a node
Some work cannot be delivered from some places: sovereignty, citizenship or clearance
requirements, regulatory location terms, contractual delivery commitments.
**This is an eligibility constraint that attaches to work, not a node.** It shrinks the
location set available and can attach to any node.
Because it applies to **all** nodes of the affected work, it does not carve one node out of
the estate — **it forces a complete parallel chain inside the eligible boundary**: its own
intake, its own transaction authority, its own fulfilment, at whatever scale the constrained
volume supports.
Three penalties compound rather than substitute:
1. **Isolation** — it can neither send nor receive work across the boundary, so it has none
of the pooling or chaining benefit the rest of the estate enjoys
2. **Scale** — constrained populations are usually small, and small pools run structurally
lower occupancy at the same service level
3. **Extended ramp** — clearance or vetting lead time sits on top of the proficiency curve
Plan it as a standalone system, report its buffer separately, and size the constrained
volume before accepting any cost target.
## The case object
The unit of work is not the contact. It is the **case object**: a state carrier passed
between nodes with its context intact.
- The customer never repeats themselves
- No node re-derives what an earlier node established
- It is the measurement spine — chain-total minutes per resolved case
- It is what an orchestration layer exists to make real
Without a case object the chain is a series of transfers, and transfers are exactly the
coordination cost the conservation test is looking for.
## Design principles
1. **Chain-level service commitments, not node-level ones.** Optimising each node
independently produces a chain that hits every internal target and still fails the
customer.
2. **Specialist-minute minimisation is the objective function** — capacity creation in the
scarce pool, not headcount reduction as such.
3. **Human verification must stay real.** A human check downstream of automation degrades
toward rubber-stamping under throughput pressure. Over-reliance on automation is a
long-established human-factors finding, and appears in contemporary settings as
mis-calibrated trust — workers relying most where the system is weakest. Sampled deep
verification and error-injection testing are what keep the check honest.
4. **Elasticity is contractual.** A low-cost node delivers its economics only if its
commercial construct prices variability rather than fixing headcount floors.
5. **Push variance toward the asynchronous tail deliberately.** A chain designed so
variability accumulates at fulfilment rather than at the specialist converts an expensive
staffing problem into a cheap backlog problem.
## Preconditions that are not optional
Two components are frequently scoped as parallel initiatives when they are prerequisites.
**A capability abstraction layer.** Where servicing requires proficiency across multiple
deep legacy systems, those systems *are* the pooling constraint — a specialist skilled on
one cannot take work from another, and the market for multi-platform specialists is thin. An
agent-facing workspace that normalises the systems underneath converts several
system-segregated pools toward one, lowers time to proficiency, and is what makes
cross-system research at the intake node feasible. Without it, the specialist node cannot
pool and the intake node cannot research.
**An orchestration layer.** The substrate that makes the case object real. Without it every
seam leaks context and the chain fails its conservation test.
Two further designs are **composable alternatives** rather than preconditions: an elastic
capacity layer (fixed core, flexible scheduled layer, certified surge pool), and
service-catalogue productisation, which converts bespoke commitments into tiered offers so
work routes by complexity rather than by client identity — collapsing the eligibility
fragmentation that prevents pooling in the first place.
## What the chain does not solve
- It does not reduce demand. Demand shaping is a separate move.
- It does not make an undecomposable transaction cheaper.
- It does not remove the proficiency constraint at the specialist node — it reduces how many
specialist minutes each case consumes.
- `[ASSERTED]` The economics of splitting a synchronous interaction — whether specialist-time
reduction survives the coordination cost — generalises long-standing call-splitting and
warm-wrap practice, but the published evidence base is thin and should be verified against
the specific domain before the case is built on it.
Block 3 — node-planning-physics.md
# Node Planning Physics
Claims marked `[ASSERTED]` are operating observations, not established findings.
The nodes differ in more than cost. **They require different capacity-planning methods**, and
applying one method across all of them is the most common technical failure in chain
implementation. Conventional workforce management applies Erlang against a point forecast
everywhere; every node here breaks that assumption in a different way.
## The three planning disciplines
| Node | Capacity is | Planned against | Method | Ramp |
|---|---|---|---|---|
| **Alpha** | Compute and orchestration | Containment and pre-staging rates, themselves uncertain and time-varying | Scenario repertoires; rate as a random variable, not a parameter | Near zero |
| **Bravo** | Scarce specialists | An arrival stream that is doubly stochastic by construction | Chance-constrained sizing or simulation of the full network | Months, gated by time to proficiency |
| **Charlie** | Elastic low-cost capacity | Backlog and work age, not instantaneous service level | Deferred-work and load-levelling models | Weeks |
Work under a delivery constraint does not add a row. It **repeats these rows inside a
restricted location set**, planned as a standalone system — see the constrained-chain section
below.
## Why node-by-node Erlang fails
**A chain is a tandem queueing network.** Downstream arrivals are not independent — they are
upstream completions. Arrival correlation propagates variance along the chain, so applying
Erlang C node by node understates chain variance and overstates achievable service levels.
The error compounds with each node, and it is always in the optimistic direction.
**Automation makes the specialist's arrival stream doubly stochastic.** The intake node's
containment rate varies with intent mix, model behaviour, content changes and release
cadence. The rate at which work reaches the specialist is therefore a random variable rather
than a parameter. Point-forecast staffing against such a stream systematically
under-provisions the tail — the established result being that when the arrival rate is
itself random, classical safety-staffing prescriptions have to be revisited, and the
required buffer scales with volume rather than with its square root.
**Practical consequence:** the more successful the automation, the less reliable
conventional staffing math becomes at the node downstream of it. This is counter-intuitive
and worth stating early in any programme, because it arrives as a service failure attributed
to the specialists rather than to the method.
## Planning each node
### Alpha
- Capacity is provisioned, not scheduled. The planning question is not headcount but
**eligibility**: what share of arriving work is the node permitted and able to handle.
- Containment is a distribution, not a number. Plan against a repertoire of pre-computed
postures rather than a single expected rate.
- The exception desk behind it is a genuine staffing problem, and it is the node most often
forgotten — automation that escalates 20% of contacts still needs somewhere to escalate to.
- Model degradation and content drift move the rate without warning. Monitor the rate as an
operational metric, not as a project assumption.
### Bravo
- Size with simulation or chance-constrained formulations against the tandem network, not
with node-level formulas.
- **Restate the sustainable occupancy ceiling before go-live.** Stripping wrap and setup out
of the role also removes recovery time embedded in it. The wrap-free role cannot be run at
the occupancy the wrapped role tolerated, and treating the freed minutes as fully available
capacity is how a chain converts an efficiency gain into an attrition problem.
- Ramp is the binding constraint on any growth or replacement. Hiring decisions must lead
demand by the full proficiency curve.
### Charlie
- Plan to **backlog and age**, not to instantaneous service level. The right questions are
how old the oldest item is and whether the backlog is stable, not what percentage was
completed within a threshold.
- Deliberately absorb variance here. The chain should be designed so variability accumulates
at this node.
- Elasticity only exists if the commercial construct permits it — see `commercial-constructs.md`.
- Watch for the failure where deferred work silently becomes synchronous because a downstream
commitment was made on it.
### The constrained parallel chain
- **Plan as a standalone system.** It has none of the pooling or chaining benefit available
to the rest of the estate, so estate-wide occupancy targets do not apply to it.
- Buffer deeper, and report that buffer separately so it is visible as a cost of the
constraint rather than as inefficiency.
- Add clearance, vetting or accreditation lead time on top of the proficiency curve when
sizing any change.
- `[ASSERTED]` Constrained pools tend to be small, which compounds the problem: small pools
run structurally lower occupancy at the same service level, so the isolation penalty and
the scale penalty stack.
- Every node inside the boundary needs its own treatment — the constraint does not exempt the
constrained chain from the planning physics above, it just applies them at small scale.
## The measurement that governs all four
**Chain-total minutes per resolved case**, measured on the case object.
Node-level handle time will improve in a chain almost by construction — that is what
decomposition does. It is not evidence of anything. The only figure that tests whether the
design works is the total across all nodes for the same resolved case, compared against a
single-touch baseline captured *before* the chain was built.
If that baseline was not captured, capture it on a control group rather than reconstructing
it. Reconstruction after the fact will not survive scrutiny and should not be attempted.
## Open quantities
Treat these as hypotheses to verify in the specific domain, not parameters to assume:
- Containment and pre-staging rates in complex, high-consequence transaction domains.
Published figures cluster in simple-intent domains and **do not transfer**.
- The sustainable occupancy ceiling for a wrap-free specialist role.
- Whether specialist-time reduction survives the coordination cost of the seam it creates.
Block 4 — tier-alignment.md
# Tier Alignment
Claims marked `[ASSERTED]` are operating observations, not established findings.
**Sourcing tier is a property of the node, not of the queue or the client.** This is the
block's whole content. Everything below follows from it.
## The alignment
| Node | Natural tier | Why | Characteristic failure |
|---|---|---|---|
| **Alpha** | Platform and compute, plus a small high-capability exception desk | Not a location decision at all; capacity is compute, ramp near zero | Counted as headcount reduction rather than capacity creation, forfeiting the more valuable property |
| **Bravo** | Onshore, or a high-capability captive centre | Scarce, judgment-bearing, consequential, long ramp | **Relocated on rate.** The classic arbitrage failure |
| **Charlie** | Low-cost captive or contracted vendor, on elastic terms | Asynchronous, decomposable, fast to proficiency, variance-absorbing | Fixed headcount floors that neutralise the flex economics |
## The reframe
Conventional sourcing attempts to arbitrage the specialist. It fails in a specific and
repeatable sequence:
1. Capacity is bought for **speed to seat**, which a vendor genuinely delivers
2. It is judged months later on **speed to proficiency**, which is a property of the work
rather than of the employer and was never purchasable
3. The supplier is held responsible for a shortfall that was purchased deliberately
**Chain-based sourcing does the opposite: it shrinks the specialist's scope and relocates
the fulfilment work instead.** That relocated scope reaches competence quickly, tolerates
asynchronous handling, and absorbs variance — which is why arbitrage is safe there and
destructive at the specialist node.
The test to apply to any proposal: **is the migration unit a queue or a decomposed scope?**
A queue contains scopes with very different proficiency curves, so moving it whole sends the
hard work with the easy work to the site carrying the longest learning curve.
## Chainability by tier
Tiers differ in more than rate. They differ in whether capacity can be redirected once
placed.
| Tier | Chainability | Why |
|---|---|---|
| Captive or in-house centre | **Highest** | Common employer, systems, taxonomy and skill definitions; edges into the wider estate are configuration rather than negotiation |
| Contracted vendor, dedicated | Moderate | Edges are possible but must be specified commercially, and rarely are |
| Contracted vendor, shared or multi-client | **Lowest** | Capacity is not the buyer's to redirect; the chain terminates at the contract boundary |
Where a captive tier exists and is under-used, **it preserves the most flexibility per unit
of cost saved.** That is a different and generally stronger argument than the usual
continuity-and-control case for captive delivery.
## Delivery constraints in sourcing terms
Eligibility is not a tier and not a node — it is a property of the work that shrinks the
location set and forces a parallel chain. It changes the arithmetic of any target before
execution begins.
**It sets a floor on the addressable base.** Constrained volume cannot be relocated, so the
saving must come from the remainder — which raises the required reduction on that remainder
proportionally. If a quarter of volume is constrained, the rest must deliver roughly a third
more than the headline implies.
**Size that floor before accepting a target, not while executing one.** It is the difference
between a target that is ambitious and one that was never arithmetically available. This is
usually a fast piece of work and it is almost never done first.
**Its isolation has a capacity cost that should be funded explicitly.** A pool that cannot
send or receive work has no pooling or chaining benefit, so it needs a deeper buffer than
the estate standard. Planned at the same occupancy as fungible pools, it is where service
fails first.
## Cross-boundary edges
Work relocated to a new site is customarily ring-fenced there — its own queues, its own
agents, its own reporting. That is a disconnected chain: capacity cannot circulate in either
direction, so the cost saving is realised and the flexibility is silently forfeited.
Connecting a site back into the estate costs a second skill on a minority of its agents,
because chaining requires two capabilities per resource rather than universal capability.
**Specify that connectivity at the point of sourcing.** Retrofitting it after a dedicated
pool has its own volumes, targets and commercial construct costs a great deal more.
## Where automation sits
Automated capacity placed *inside* the chain adds a node with near-zero ramp that can hold
many capabilities at once — structurally, the ideal chain node. Placed *outside* it, as pure
deflection, it removes work from the graph and leaves a residual that is more specialised
and therefore less able to flex.
An operation can reduce contact volume and simultaneously reduce its ability to respond to
variability. That trade is rarely priced and should be stated in any automation case.
**Triage is the highest-value automated scope** on topological grounds rather than on labour
saved: work usually enters the graph at the wrong node because the customer selects from a
menu, and every mis-entry costs a transfer consuming two nodes. Accurate triage raises
effective connectivity without training anyone — and it returns more in a distributed estate
than a consolidated one, because a mis-entered contact may otherwise cross a site, a
timezone and a contract before reaching the right node.
## Questions to ask of any sourcing proposal
1. Which **node** does this work sit in? If the answer is "the whole queue", stop here.
2. What is the **proficiency curve** of what is being moved?
3. What **chain edges** connect the receiving location back to the estate, and are they in
the commercial construct?
4. What does this do to **elasticity**, stated alongside what it does to cost?
5. Does the receiving tier buy **overflow to standing proficient capacity**, or is it a
hiring arrangement in an elasticity costume?
6. Has the **constrained floor** been sized, and does the remaining target still close?
Block 5 — commercial-constructs.md
# Commercial Constructs
Claims marked `[ASSERTED]` are operating observations, not established findings.
A chain design fails commercially more often than technically. The recurring cause: the
contract prices the wrong thing, and by the time that is visible the construct has been
signed and the operating assumptions built on it.
## What each node needs from its contract
| Node | The commercial question | What must be in the terms |
|---|---|---|
| Alpha | Consumption or licence, not headcount | Rate structure that does not penalise the volume growth automation is meant to absorb; degradation and drift responsibilities |
| Bravo | Capability, not availability | Proficiency definition and how it is certified; retention terms if the tier is outsourced |
| Charlie | **Variability** | Capacity bands, surge terms, notice periods, and what fraction is genuinely variable at 30/60/90 days |
| *Constrained work* | Eligibility and continuity — not a node, a constraint on all of them | Clearance maintenance, vetting lead time, what happens when a cleared individual leaves, and the price of any dedication commitment |
## Which constraints are priceable
Two kinds, and only one is negotiable.
**External constraints** — sovereignty, citizenship, clearance, regulatory location terms.
Not available to be traded. Their capacity cost has to be absorbed and funded.
**Commercial constraints** — a dedication commitment, a named-team promise, a channel or site
restriction agreed in the contract. **These are product features, and they have a cost.**
The second is where value routinely leaks. A dedicated team means the client has bought the
forfeiture of pooling and chaining benefit for that work, and the forfeiture is measurable:
- lower achievable occupancy at the same service level
- no access to estate capacity during a surge
- a separate ramp, and a separate skill graph to maintain
It is very commonly given away rather than priced. Making it visible does not require
renegotiating anything — it requires the cost of the constraint to appear in the deal model
**at the point the commitment is made**, so dedication is sold as the premium feature it is
rather than conceded as a term.
## The two goods bought under one name
**Speed to seat is purchasable. Speed to proficiency is not.** A vendor fills a seat quickly
because filling seats is what a vendor is good at. Proficiency is a property of the work, so
the learning curve applies identically on the other side of a contract.
Two arrangements share the name "outsourcing" and behave completely differently:
- **Overflow to standing proficient capacity** — contacts route to capacity that is already
staffed and already competent, on a real-time threshold. This buys genuine elasticity.
- **Outsourcing as hiring** — the partner must recruit and train to meet the demand. This
inherits the buyer's ramp physics unchanged and buys a lower unit cost and no elasticity.
Most enterprise arrangements are the second while being justified as the first. State which
one is being bought, in the paper, before signature.
## Ramp is a contract term, not an accounting abstraction
The saving is realised **at proficiency, not at signature**. Between those two points sits
dual-running cost, elevated handle time and elevated error.
- Who pays for ramp hours? If the partner bills at full rate through nesting, the buyer funds
the partner's ramp while receiving reduced output.
- Is there a **proficiency gate** — a competence standard before full rate applies?
- What is the **latest start date** for the saving to land in the target period? Ramp
duration determines it, and it is usually discovered rather than planned.
At the scale of a multi-year cost programme, ramp is a material number and it belongs in the
model, not in the risk register.
## Per-productive-hour models
Paying for productive hours rather than billed hours is an improvement — it shifts idle and
shrinkage risk to the supplier. It has two specific exposures that need managing rather than
renegotiating.
**The definition is where value leaks.** What counts: talk, after-call work, ready, hold?
What does not: training, coaching, system downtime, meetings? Who bears shrinkage, and how is
it audited? A loose definition collapses the model back toward billed hours within a couple
of quarters, quietly. **Locking the definition is the cheapest, highest-return item
available** and should be done before volume grows under it.
**It prices input, not output.** Nothing in an hourly construct rewards doing the work in
fewer hours, and the perverse case is real: longer handle time produces more billable
productive hours. This does not have to be fixed commercially straight away, but throughput
must be measured alongside price as a governance instrument. A productivity baseline is the
only thing that makes a later move to outcome-based pricing possible at all.
**It works against the elasticity you want.** A supplier paid only for productive hours has
an active disincentive to hold standby capacity, because standby is unproductive and
unbillable. The construct that removed the buyer's idle risk also removed the supplier's
willingness to stand ready. **If overflow to standing capacity is wanted, it has to be a
separately priced availability term.** It will not arrive as a by-product of the hourly rate,
and assuming it will is how an operation discovers in a peak that the flexibility was never
purchased.
## Comparing options: hours, not rate
Where the commercial unit is the hour, the sourcing comparison is:
**effective cost = rate × hours to deliver the same work at the same standard**
Hours differ by site through handle time, rework and repeat contact, transfer rate, and the
ramp period. A site 40% cheaper per hour needing 25% more hours delivers a 25% saving, not
40% — and that gap is invisible on a rate card.
`[ASSERTED]` Cost per resolved unit at a defined standard is the better denominator and the
right destination, but it is a commercial-model change rather than an analytical one. Where a
per-hour model has recently been adopted, treat unit pricing as a later phase and say so
explicitly, so it reads as sequencing rather than as an oversight.
## Contract form shapes delivered quality
Standard outsourcing contract forms — piece-meal and pay-per-resolution — can coordinate the
*staffing level* while leaving delivered service quality below the system optimum. Quality
shortfall in an outsourced estate is therefore, in part, a predictable property of the
contract form rather than evidence about the supplier.
The practical consequence: **where outsourced quality underperforms, examine the contract
before examining the supplier.** That points remediation at procurement, which moves faster
than performance management, and it is frequently the cheaper fix.
## What to require of any sourcing business case
1. The **migration unit** stated — queue or decomposed scope — with the proficiency curve of
what moves
2. **Ramp cost and duration** as an explicit offset, and the latest start date implied
3. The **chain edges** back to the estate, specified before the commercial construct is
written
4. The **elasticity effect**, stated alongside the cost effect
5. Whether the arrangement buys **overflow to standing proficient capacity** or is hiring in
an elasticity costume
6. For deferred-work nodes, **what fraction of cost is genuinely variable** at 30, 60 and
90-day horizons
7. The **constrained floor**, sized, with the residual target recomputed against it
Block 6 — assurance.md
# Assurance
Claims marked `[ASSERTED]` are operating observations, not established findings.
Two questions this block exists to answer: **did the chain actually remove work**, and **are
we getting what we paid for across tiers**. Both are routinely answered with instruments
that cannot resolve the difference being claimed.
## The conservation test
**Chain-total minutes per resolved case, measured on the case object, against a single-touch
baseline.**
Node-level handle time will improve in a chain almost by construction — that is what
decomposition does — and it is not evidence of anything. A design can improve every node
metric while raising total minutes per resolved case. That is the failure the test exists to
catch.
- Capture the baseline **before** the chain is built. If it was missed, capture it on a
control group rather than reconstructing it — reconstruction will not survive scrutiny.
- Count every node, including the exception desk behind the automation and the coordination
time at each seam.
- Include rework and repeat contact. Work that returns is work that was not removed.
## Seam integrity
Each seam is a place where context leaks and work is re-derived.
| Measure | What it detects |
|---|---|
| Repeat-contact rate by seam | Context loss forcing the customer back |
| Re-derivation time per node | A node redoing what an earlier node established |
| Case-object completeness audit | Whether context is actually travelling |
| Transfer rate by entry point | Work entering the chain at the wrong node |
`[ASSERTED]` Seam cost tends to be underestimated at design time because it is distributed
across nodes and rarely attributed to the seam that created it.
## Human-in-the-loop integrity
Where a human check sits downstream of automation, throughput pressure degrades it toward
rubber-stamping. Over-reliance on automation is a long-established human-factors finding, and
it recurs in contemporary settings as mis-calibrated trust — reliance is highest precisely
where the system is weakest, because failures outside the automation's competence look
superficially like successes inside it.
**This cannot be managed by instruction.** It requires:
- **Error-injection testing** — deliberately introduce faults and measure the catch rate.
Publish the rate. A verification step with an unmeasured catch rate is decorative.
- **Sampled deep verification** — a proportion of cases fully re-worked independently.
- **Monitoring the catch rate over time**, not once at go-live. Degradation is gradual and
correlates with volume pressure.
- **Watching emotional load, not only topic.** Human intervention preserves quality well in
technical escalations and markedly less well in emotionally-loaded ones, where engagement
drops. Routing that classifies only by subject will send exactly the wrong cases into the
wrong hands, and late intervention is close to useless.
## Quality is a gate, not a trade
Quality is a **constraint on the cell, not a dimension of it**. Each cell inherits its required
standard from what was sold. Where the service offering is undefined, every cell has an
undefined standard — which defaults in practice to the highest one anyone remembers promising.
**Cost and quality need a shared denominator.** Cost is usually measured per unit of input —
hour, FTE, contact. Quality is measured per interaction, sampled. The two are not comparable,
which is why the argument never resolves on either party's evidence. Put both on the
**resolved case**.
Doing so reveals that much of the apparent trade is a measurement artifact: rework, repeat
contact, transfer, escalation and remediation are quality failures that appear as **cost** once
resolutions rather than contacts are counted. What remains is the genuine residual —
degradation that generates no further contact, such as slower, colder or less proactive
handling. That residual is invisible to any cost measure, and is what the gate protects.
**The decision rule:**
1. Below the sold standard is not a cheaper option — it is a different product that was not
sold. Such options are **excluded, not scored lower**.
2. Above the gate, decide on cost per resolved case, which already absorbs rework consequences.
**Standardise the instrument, vary the threshold.** Standardising the *target* is wrong — one
quality number across tiers either over-serves the lower or under-serves the higher, and
differentiated tiers are the point of a service architecture. Standardising the *instrument* —
outcome definition, sampling method, case-mix adjustment, scoring — is essential, or no
comparison is valid and every cross-tier decision stays contestable.
**Instrument resolution bounds sellable tier spacing.** Where sampling cannot detect a
difference smaller than several percentage points, two tiers separated by less than that cannot
be shown to differ, and the premium tier is a promise with no available evidence. This makes
instrument resolution a constraint on the service architecture, not only on supplier management.
## Chain quality is contractual; node quality is diagnostic
Because decomposition depth varies by contract, the quality model is layered:
- **Chain-level quality is the commitment** — the outcome for the resolved case, whatever path
it took.
- **Node-level quality locates failure.** It is diagnosis, not the commitment. Reported as the
commitment, it produces delivery that hits every node target and still fails the customer.
- **Quality risk migrates to the seams as decomposition deepens.** In contained delivery,
failure concentrates in the interaction. In heavily decomposed delivery, it concentrates in
handoffs — context loss, re-derivation, work returning.
- **Inspection points vary with the delivery model; the yardstick does not.** What counts as a
good outcome and how it is scored stay constant. Where failure is looked for changes:
interactions in contained delivery, seams in decomposed delivery.
Hold the contractual commitment at chain level and treat everything below it as diagnosis. That
single rule collapses most of the apparent complexity.
## Quality comparability across tiers
The question "is this tier delivering the quality we bought" is usually answered with an
instrument that cannot detect the difference being decided on.
**Resolution before conclusion.** Every measurement regime has a detectable difference floor,
below which "quality is comparable across tiers" is a statement about the sample rather than
about the tiers. Establish the floor first and state it alongside any comparison.
**Coverage is the wrong frame; observation count is the right one.** Resolution is set by the
absolute number of observations, not by the share of interactions they represent. Moving a
large operation from single-digit sampling to full monitoring is a substantial gain and worth
having. But a unit generating a hundred interactions in a period, monitored in full, still
yields a hundred observations and resolves only to roughly ten percentage points — while a far
larger unit sampled at three per cent may carry thousands and resolve to two.
Two consequences, both counter-intuitive:
- **Full monitoring materially helps mid-sized units and barely helps the smallest ones** —
which are exactly the units whose scores swing most and are therefore most often placed under
review on the strength of noise.
- **Volatility is not evidence of instability.** Report observation count beside every score
and detectable difference beside every comparison, or the two will be confused.
**Vary the threshold by tier, never by location.** Tiers are sold as different products, so a
tier-differentiated target is a product decision. A location-differentiated target states that
delivery from one place is expected to be worse — which hands the commercial function a
standing veto no cost advantage answers, degrades the aggregate by construction as that
location's share grows, and contradicts any location-agnostic service promise. Where units
genuinely handle work of different difficulty, apply **one target to case-mix-adjusted scores**
rather than different targets to raw ones.
**Employment relationship does not predict quality.** Differences between an in-house centre and a contracted supplier, or between two teams inside the same supplier, are generally explained by tenure stability and case mix rather than by who employs the staff. A hand-picked, low-attrition supplier team routinely outperforms a high-churn in-house one on the same work. Look to composition when explaining quality, not to the contract.
**Case-mix adjustment is mandatory.** A tier handling harder work scores worse regardless of
how well it performs. Comparing raw scores across tiers with different work profiles measures
the work, not the delivery. Adjust, or compare only within matched complexity bands.
**Change the instrument before enlarging the sample.** Where the required resolution is
beyond what sampling can economically reach, the answer is a different measurement approach —
outcome-based signals, full-population automated review, targeted deep audit on flagged
segments — not more of the same sampling.
**Measure at chain level.** A per-node quality score tells you which node erred; it does not
tell you whether the customer got a good outcome. The commitment is end to end and the
measurement should be too.
## Occupancy regression
A specific and frequently missed effect. Stripping setup and wrap out of the specialist role
also removes recovery time embedded in that work. **The wrap-free role cannot sustainably run
at the occupancy the wrapped role tolerated.**
Treating every freed minute as available capacity converts an efficiency gain into an
attrition problem, and the signal arrives late — through error rates and turnover rather
than through service level.
Restate the sustainable occupancy ceiling for the redesigned role **before go-live**, and
monitor attrition and error rates against it.
## What to report
A chain assurance pack should carry, every period:
1. Chain-total minutes per resolved case, against baseline
2. Repeat-contact and rework rates, by seam
3. Verification catch rate from error injection
4. Quality by tier, case-mix adjusted, with the detectable-difference floor stated
5. Specialist occupancy against the restated ceiling, with attrition and error trend
6. Variable cost fraction actually realised at the deferred node
7. Automation rate as an operational metric, with its variance — not as a project assumption
## What none of this establishes
- It does not establish **cause**. A chain-total improvement coinciding with a demand mix
shift is not attributable without further work.
- It does not establish that a **quality difference is absent** — only that the instrument
did not detect one at its resolution.
- It does not establish that the design will **hold under peak**. Conservation measured in a
normal period says nothing about seam behaviour under disruption, which is when
coordination cost is highest and context loss most likely.
Usage notes
Sizing. The instruction block is about 480 words and loads with every message inside the project. The five reference blocks total roughly 6,500 words and load only on retrieval.
Company-agnostic by construction. The blocks describe the role a component plays, never the product that fills it. Orchestration layers, abstraction workspaces and agentic platforms are named by function throughout. This is deliberate: the design holds across vendors and the naming would date it.
The [ASSERTED] convention. Several load-bearing quantities in this domain are operating observations rather than established findings — containment rates in complex transaction domains, the economics of splitting a synchronous interaction, the sustainable occupancy ceiling for a wrap-free role. The blocks mark these explicitly and instruct the model to preserve the distinction. A pack that let them pass as settled would produce confident business cases resting on numbers borrowed from a different problem.
What it deliberately does not contain. No containment-rate benchmarks, no target ratios, no reference cost curves. Published figures in this space cluster in simple-intent domains and do not transfer to complex, high-consequence transactions — supplying them would be the single most damaging thing this pack could do.
Packs do not compose at runtime. Each pack is deployed as its own Claude project, so a session running this pack cannot reach another pack's reference blocks. The cross-references in Related packs are for a human choosing what to deploy next.
Related packs
| Pack | Deploy it when the work reaches |
|---|---|
Demand Variance Decomposition (CP-WFM-006) |
Establishing how much of the demand problem is forecastable before sizing any node |
Service Quality Comparability (CP-WFM-003) |
Proving quality is or is not comparable across tiers |
Agent Capability Ontology (CP-WFM-002) |
Building the attribute layer that supply-side routing depends on |
Planning Under Demand Volatility (CP-WFM-001) |
Planning and capacity work once the chain design is settled |
Change history
| Version | Date | Change |
|---|---|---|
| 1.0 | 2026-08-13 | Initial publication. One instruction block, five reference blocks, derived from the service chain article. |
| 1.1 | 2026-08-14 | Correction: Delta reclassified from a fourth node to an eligibility constraint that forces a complete parallel chain inside the eligible boundary. Applied across Blocks 1–4. Cross-linked to the two-axis sourcing design surface. |
| 1.2 | 2026-08-14 | Added decomposition depth as a property of the contract rather than a destination (Block 2), the external-versus-commercial constraint distinction and the pricing of dedication commitments (Block 5), and the quality gate model — shared denominator, standardise-instrument-vary-threshold, instrument resolution bounding sellable tier spacing, and chain-versus-node quality (Block 6). Disciplines updated in Block 1. |
| 1.3 | 2026-08-14 | Block 6 sharpened from practice: measurement resolution reframed from coverage percentage to absolute observation count, so full monitoring is understood to help mid-sized units and barely help the smallest; the rule that thresholds may vary by tier but never by location, with its three consequences; and the finding that employment relationship does not predict quality, tenure stability and case mix do. |
See also
- Service Chain Decomposition and Node Sourcing — the model, in full
- Sourcing Design Axes: Node and Client Ownership — node and client ownership as the two decision axes, with location as a consequence
- Conservation of Labor — the governing constraint on how much work automation removes
- Chaining and Flexibility Design — why sourcing decisions are chain topology decisions
- Supply Elasticity in Workforce Planning — why flexibility is the capacity question underneath the cost question
- Wiki:Packs — the full pack inventory
