Wiki:Packs/Service Quality Comparability
| Pack | |
|---|---|
| ID | CP-WFM-003
|
| Domain | WFM |
| Blocks | 1 instruction + 5 reference |
| Version | 1.1 |
| Source | Quality Management · Quality Assurance Platforms in Contact Centers · Speech Analytics · First Contact Resolution · Performance-Based Vendor Allocation Design · Business Process Outsourcing |
A pack is a deployable set for a Claude project. This one supports answering a question that sounds simple and is not: is our outsourced service worse than our in-house service, and by how much?
The question is hard because the delivery models are usually measured by different instruments at different coverage by different parties. One estate may auto-score 80% of interactions through speech analytics while another self-scores a 3% manual sample. Those two numbers look alike, sit in adjacent dashboard cells, and are not comparable. Reporting them side by side produces the sensation executives describe as bits and pieces — plausible figures that never quite add up to an answer.
This pack treats that as a measurement comparability problem rather than a reporting problem, and gives a route to defensible answers in weeks rather than quarters.
When to use it
Use it when:
- Service is delivered through more than one model — in-house, captive low-cost centre, third-party outsourcer — and their quality is being compared
- Multiple business segments have merged or been acquired, each carrying its own quality instrument and rubric
- A quality gap has been asserted and is about to drive a commercial, contractual or personnel decision
- Executives are asking why quality reporting never resolves into a clear picture
Do not use it to design a quality rubric or a sampling plan from scratch — Quality Management covers rubric design, sampling and calibration, and this pack assumes that material rather than repeating it.
The central finding
Sampling depth is not a matter of degree. It changes what a measurement can detect at all.
For a quality score near 85% on a delivery unit handling 20,000 interactions a month:
| Regime | Interactions scored | 95% confidence interval | Smallest difference it can detect between two units |
|---|---|---|---|
| Automated, 80% coverage | 16,000 | ±0.55 pp | ±0.8 pp |
| Manual sample, 3% coverage | 600 | ±2.86 pp | ±4.0 pp |
A three-point difference between two outsourcers measured at 3% sampling is not a finding — it is noise. And the gap does not close by sampling a little harder: detecting a 2 pp difference at conventional confidence and power needs roughly 5,000 scored interactions per unit — about 25% coverage, an eight-fold increase.
The conclusion is therefore not "sample more." It is "change instrument." That is a procurement and contracting decision, not a quality-team decision, and it is the single most useful thing this analysis produces.
The approach on one page

How to deploy
- Create a project in Claude. Name it for the work, not the method
- Copy Block 1 into the project's custom instructions
- Save Blocks 2–6 under the filenames in their headings and upload as project knowledge
- Start a conversation
Block 1 — Project instructions
# Service Quality Comparability
## Context
This project establishes whether service quality can legitimately be compared
across delivery models — in-house, captive low-cost centre, third-party
outsourcer — and across business segments that carry different measurement
instruments. The work spans inventorying what is measured, grading what may be
compared, sizing gaps that survive adjustment, and sequencing answers so an
executive gets something defensible in weeks.
The governing assumption: apparent quality differences between delivery models
are confounded until proven otherwise. Instrument, coverage, provenance, case
mix and survey response bias all differ systematically between in-house and
outsourced delivery, and each moves scores in a predictable direction.
## Routing
| Task | Open |
|---|---|
| Deciding which measures may be compared; what quality even means here | `quality-measure-families.md` |
| Coverage, sampling precision, who does the scoring, statistical power | `measurement-integrity.md` |
| Grading a number for comparability; drill structure; paired-metric rules | `comparability-grading.md` |
| Running the discovery interview with a delivery or vendor lead | `quality-discovery-protocol.md` |
| Sequencing deliverables; what to tell an executive and when | `rapid-answer-path.md` |
## Disciplines
- **A difference is not a finding until it survives adjustment.** State the
unadjusted gap, the adjustments applied, and the residual gap separately.
- **Never report a quality score without its coverage rate.** A score at 3%
coverage and a score at 80% coverage are different kinds of object.
- **Never report service level without a quality measure beside it.** Attainment
and quality move independently, and reporting one alone invites optimising it
at the other's expense.
- **Distinguish imprecision from bias.** Small samples are imprecise and that
shrinks with n. Self-selected or self-scored samples are biased and that does
not shrink with n.
- **Adjustment sizes a gap; it does not dissolve one.** Decomposition exists to
make a real difference defensible, not to explain it away. Say so explicitly
whenever adjustment reduces a headline number.
- **Name the unit.** Percent, percentage point and basis point are not
interchangeable, and mixing them is a common source of executive confusion.
- **Blanks are findings.** An instrument nobody can describe is a result.
- Do not assert a level of claim the evidence does not yet support. Use the
claim ladder in `rapid-answer-path.md` and say which rung you are on.
## Output
Deliver comparability verdicts, not league tables. Every comparative figure
carries a grade (A–D), a coverage rate, and the adjustments applied. Where a
comparison is not currently legitimate, say so plainly and state what would make
it legitimate. Show the arithmetic for any confidence interval or detectable
difference. Refuse to produce a ranked vendor scorecard from measures that have
not been graded.
Source: Wiki:Packs/Service Quality Comparability (CP-WFM-003) v1.1
Block 2 — quality-measure-families.md
# Three Families of Quality Measure
Figures marked [estimated] are judgment, not measurement.
"Quality" names at least three different things. They answer different
questions, fail for different reasons, and must never be averaged into a
composite. Most quality reporting confusion begins with a composite score that
mixes them.
## The families
| | Perceived | Assessed | Outcome |
|---|---|---|---|
| Question | How did it feel? | Did we follow the standard? | Did it actually work? |
| Examples | CSAT, post-interaction survey, NPS, effort score | Speech-analytics scoring, manual QA, compliance checks | Reopen and rework rate, escalation rate, first-contact resolution, error rate, complaint rate |
| Source | The customer | An observer or model | System of record events |
| Fails through | Response bias, instrument wording, trigger point | Coverage, rubric drift, scorer bias | Definition drift, event capture gaps |
| Comparable across differing instruments | Rarely | Rarely | **Usually** |
## Why outcome measures are the fast route
When segments have merged, each usually brings its own survey instrument and its
own QA rubric. Reconciling those is a multi-quarter programme: rubrics must be
rewritten, scorers recalibrated, survey instruments re-fielded and trended.
Outcome measures largely escape this. A reopened case is a reopened case
whatever rubric the site uses. An escalation is an escalation. A ticketing error
is an error. These are events in transactional systems, not judgments applied by
an observer, so they survive instrument heterogeneity.
**Therefore: lead cross-model comparison on outcome measures, and demote
perceived and assessed measures to supporting evidence carrying explicit
caveats, while instrument harmonisation runs in parallel.**
This is a sequencing recommendation, not a claim that outcome measures are
superior. They are narrower — they capture whether the work was done, not
whether it was done well or felt good. They are simply the family that is
comparable *today*.
## The minimum comparable set
Define a small set — typically five or six measures — that is defined
identically across every delivery model and segment, and to which all reporting
rolls up. Everything else is drill-down detail.
Candidate members, all outcome-family:
| Measure | Definition discipline required |
|---|---|
| Reopen / rework rate | What counts as a reopen, and the window |
| Escalation rate | Escalation to whom, and whether transfers count |
| First contact resolution | Contact window, and whether the customer or the system decides |
| Error rate | Which error classes, and who adjudicates |
| Complaint rate | Threshold for a complaint versus dissatisfaction |
| Repeat contact rate | Window and matching rule |
The definition discipline column is the real work. A measure defined loosely is
worse than no measure, because it will be compared anyway.
## The composite trap
A single blended quality score is attractive to executives and destroys
diagnosis. It mixes families, hides which component moved, and lets a decline in
one dimension be masked by a rise in another. Where a composite is demanded,
publish it only alongside the components, and never use it for cross-model
comparison.
## Limits
Outcome measures can be gamed where the operator controls event coding — an
escalation logged as a new contact, a reopen suppressed by leaving a case open.
Check that event definitions sit in systems the delivery party does not
administer. Where they do not, the measure drops a comparability grade.
Block 3 — measurement-integrity.md
# Measurement Integrity: Coverage, Provenance, Precision
Figures marked [estimated] are judgment, not measurement.
Three properties determine whether a quality number can carry weight. They fail
independently, and a number can be fatally compromised on any one of them while
looking healthy on the other two.
## 1. Coverage — how much of the work is measured
Coverage is the proportion of interactions actually scored. It varies enormously
between regimes: automated scoring routinely reaches 70–90% of voice
interactions, while manual QA typically samples 1–5%.
**Coverage asymmetry invalidates direct comparison.** If one delivery model is
scored at 80% and another at 3%, a difference between their scores confounds
delivery quality with measurement regime. Always publish coverage beside score.
## 2. Provenance — who does the scoring
| Provenance | Bias risk | Maximum grade |
|---|---|---|
| Our systems, our data, our scorers | Low | A |
| Our instrument, applied by the delivery party | Moderate | B |
| Delivery party's instrument, results reported to us | High | C |
| Delivery party selects the sample **and** scores it | Severe | D |
The bottom row deserves emphasis. When the party being measured chooses which
interactions are scored, the estimate is not merely imprecise — it is
**biased**, and *bias does not shrink as the sample grows*. A self-selected 3%
sample cannot be repaired by expanding it to 30%. Only changing who selects, or
who scores, repairs it.
## 3. Precision — what the sample can detect
Sampling depth is not a matter of degree; it changes what questions the
measurement can answer.
For a score near 85% on a unit handling 20,000 interactions per month:
```
Automated, 80% coverage n = 16,000 95% CI = +/- 0.55 pp
Manual sample, 3% n = 600 95% CI = +/- 2.86 pp
Smallest detectable difference between two units (95%):
both automated at 80% +/- 0.8 pp
both manual at 3% +/- 4.0 pp
one automated, one manual +/- 2.9 pp
```
Read the second row carefully. **Two outsourcers measured at 3% sampling cannot
be reliably distinguished unless they differ by more than about four points.**
Every monthly movement smaller than that is noise, and treating it as signal
produces exactly the churn of contradictory explanations that makes quality
reporting feel unreliable.
### Why sampling harder is not the answer
To detect a 2 pp difference at 95% confidence with 80% power requires roughly
**5,000 scored interactions per unit** — about **25% coverage** at 20,000
interactions a month. That is an eight-fold increase in manual scoring effort to
reach a resolution automated scoring already exceeds.
**The implication is a procurement decision, not a quality-team one:** either
move every delivery party onto the same automated instrument, or accept
permanently coarser measurement for the sampled estate and set contractual
thresholds wide enough to respect it. Thresholds tighter than the measurement
can resolve are unenforceable, and both parties will discover this during the
first dispute.
## Case mix — the fourth confound
Even at equal coverage, provenance and precision, delivery models rarely receive
the same work. Outsourced estates commonly take a different intent mix, a
different complexity band, different hours, and different language markets.
Adjust before comparing. At minimum, stratify by interaction type and complexity
band and compare within strata; where volumes permit, weight to a common
reference mix. An unadjusted cross-model comparison is a **Grade C** number
regardless of how good the underlying instrument is.
## The direction of the biases
Useful to state explicitly, because they mostly point the same way:
| Confound | Typical direction on outsourced score |
|---|---|
| Lower coverage | Wider interval, unstable month to month |
| Self-selection of sample | Upward |
| Self-scoring | Upward |
| Harder or less familiar case mix | Downward |
| Lower survey response rate | Unstable, usually upward |
Self-selection and self-scoring push the outsourced number **up**, while case
mix pushes it **down**. They do not cancel. Which dominates is an empirical
question, and answering it is the point of the adjustment work.
Block 4 — comparability-grading.md
# Grading and Drilling
Figures marked [estimated] are judgment, not measurement.
## The comparability grade
Every comparative figure carries a grade. The grade travels with the number into
every deck and dashboard cell it appears in.
| Grade | Conditions | Permitted use |
|---|---|---|
| **A** | Same instrument, same rubric version, coverage within 10 pp, case-mix adjusted, independent provenance | Compare freely; may drive contractual action |
| **B** | Differs on one dimension, and a stated adjustment has been applied | Compare with the adjustment quoted alongside |
| **C** | Differs on instrument or provenance, unadjusted | Directional only. **Never appears in a stated delta** |
| **D** | Self-selected or self-scored sample, or definitions not established | Report separately. Never in the same table as A or B |
The rule that does the work: **a Grade C number may not be subtracted from
another number.** Most disputed quality claims are Grade C figures presented as
Grade A deltas, and the grade alone resolves the argument without anyone having
to relitigate the underlying data.
## Two standing rules
**Coverage travels with score.** Every quality figure is published as a pair —
the score and the proportion of interactions it was computed from. A score
without coverage is not a number, it is an assertion.
**Service level never appears alone.** Attainment and quality move
independently, and an operation can hit every service-level target while
delivering the worst customer experience it has ever produced. Publishing
attainment without a quality measure beside it creates a standing incentive to
optimise the measured dimension at the expense of the unmeasured one. Present
them as a pair, always.
## The drill architecture
```
L0 Enterprise scorecard — minimum comparable set, by delivery model
L1 + segment (A / B / C / D)
L2 + channel (voice / chat / email)
L3 + provider and site
L4 + case-mix decomposition <-- the layer normally skipped
L5 interaction-level evidence
```
### Why L4 must precede L5
Most quality dashboards jump from an aggregate to listening to individual
interactions. That turns every performance review into anecdote-trading: each
party arrives with examples supporting its position, and the meeting resolves
nothing.
L4 asks a different question — *how much of this gap is explained by what work
each party received?* — and answers it before anyone reaches for a transcript.
Without it, performance management is a recurring negotiation about whose queue
was harder. With it, the residual gap is agreed before the conversation starts,
and the conversation can be about causes instead.
## Presenting a gap
State three numbers, never one:
```
Unadjusted difference e.g. 22 pp
Adjustments applied case mix -6, coverage -3, response bias -2
Residual difference 11 pp [Grade B]
```
This format is harder to attack than a single figure, and it survives a hostile
vendor review. It also disciplines the analyst: every adjustment must be named,
sized and justified, so the residual is defensible line by line.
## When the grade cannot be raised
Sometimes a comparison cannot be made legitimate with the instruments available.
Say so, and state the specific change that would fix it — usually a change of
instrument, of scoring party, or of sample selection. An honest "we cannot
currently answer this, and here is what it would take" is a better executive
deliverable than a confident number that collapses on inspection.
Block 5 — quality-discovery-protocol.md
# Discovery Protocol
The question set to run with whoever owns quality for a delivery estate. Run it
per segment, per channel, per delivery model — the answers routinely differ
across cells that were assumed identical.
Figures marked [estimated] are judgment, not measurement.
## How to run it
Ask for the artefact, not the assurance. "Yes we calibrate" and a calibration
report with variance figures are different answers. Where an artefact cannot be
produced, that is the finding — record it rather than accepting the description.
Expect the inventory itself to be the first real deliverable. In most estates
nobody has previously listed every instrument in one place, and the list alone
usually explains the reporting incoherence.
## A. Instrument inventory
1. What quality instruments exist for this cell? Automated scoring, manual QA,
customer survey, outcome metrics — list all.
2. What exactly does each score? Name the dimensions and their weights.
3. What scale is used, and what rubric version, dated?
4. Is the rubric identical to the one used for other delivery models? Where
does it differ?
## B. Coverage and selection
5. What proportion of interactions is scored, by channel?
6. How is the sample selected — random, systematic, or chosen by a person?
7. Who selects it?
8. Are any interaction types systematically excluded — short calls, transfers,
abandoned chats, non-primary languages?
## C. Provenance
9. Who applies the score — our staff, the delivery party's staff, a third party,
or a model?
10. Is the result computed in our systems from our data, or reported to us?
11. If reported, can the underlying interactions be independently re-scored?
12. When was the last cross-party calibration exercise, and what scorer variance
did it show?
## D. Survey measures
13. What is the survey response rate, by channel and by segment?
14. Is the instrument identical across segments — same wording, same scale, same
trigger point, same delay?
15. Is response rate correlated with outcome? (If satisfied customers respond at
a different rate, the mean is biased.)
## E. Outcome measures
16. Which outcome measures exist — reopen or rework, escalation, first contact
resolution, error, complaint, repeat contact?
17. Is each defined identically across delivery models? Ask for the definitions
in writing.
18. Which system holds the underlying events, and who administers it?
## F. Case mix
19. Can interaction volume be described by type and complexity band per cell?
20. Does the work mix differ materially between delivery models? By how much?
21. Has any comparison to date been case-mix adjusted?
## G. Governance
22. Who owns each definition, and who may change it?
23. Is there a change log? When did each definition last change?
24. What happens when a delivery party disputes a score, and how often does that
happen?
## Producing the comparability map
The output is a grid — segment x channel x delivery model — with each cell
carrying: instrument, coverage, provenance, rubric version, and the resulting
comparability grade.
Cells sharing an instrument, coverage band and provenance are comparable. Cells
that do not are not, and the map shows immediately which comparisons currently
being made in reporting are unsupported.
This grid is the artefact that converts "quality reporting feels like bits and
pieces" from a complaint into a diagnosis.
Block 6 — rapid-answer-path.md
# The Rapid Answer Path
Figures marked [estimated] are judgment, not measurement.
Executive sponsors need progress before the analysis is complete. The way to
give it without overclaiming is to be explicit about which rung of the claim
ladder the evidence currently supports.
## The claim ladder
| Rung | Claim | Requires | Typical elapsed |
|---|---|---|---|
| 1 | "Here is what we measure, where, and which comparisons are currently legitimate" | Instrument inventory | 1–2 weeks |
| 2 | "Here is the first like-for-like comparison" | Outcome-family baseline with coverage | 3–4 weeks |
| 3 | "The headline gap was X; the real gap is Y" | Case-mix adjustment | 6–8 weeks |
| 4 | "Here is the cause, and here is the contractual response" | Attribution and remedy design | One to two quarters |
**Never claim a rung you have not reached.** The credibility cost of retracting a
Rung 3 claim made at Rung 1 is far higher than the cost of saying "we are at
Rung 1, here is Rung 1's answer, and Rung 2 lands in three weeks."
## Rung 1 is always achievable, and it is more useful than it looks
The first deliverable is not a number. It is the comparability map: what is
measured in each cell, by whom, at what coverage, and which of the comparisons
currently circulating are supported.
This lands quickly because it requires no new data collection — only the
discovery protocol. And it directly answers the executive's actual complaint,
which is usually not "what is the number" but "why does this never resolve."
It also does something politically valuable: it converts a disagreement about
performance into an agreement about measurement, before anyone has to defend a
result.
## Sequencing the deliverables
**Weeks 1–2 — Comparability map.** Every cell graded. Every figure currently in
circulation regraded. Expect some widely quoted numbers to grade C or D; flag
these before someone else does.
**Weeks 3–4 — Outcome baseline.** Publish the minimum comparable set by segment,
channel and delivery model, with coverage beside every figure. This is the first
genuinely like-for-like view and usually the first time the estate has had one.
**Weeks 6–8 — Size the largest gap.** Take the biggest observed difference and
decompose it. Report unadjusted, adjustments, residual. Whether the residual is
large or small, the number is now defensible.
**Quarter 2 — Instrument convergence.** Decide whether to move every delivery
party onto a common automated instrument or to accept coarser measurement for
part of the estate and set thresholds accordingly. This is the long-term path,
and it is a contracting decision.
## Handling a gap that is already public
Where a large gap is already circulating and driving decisions, commission the
decomposition immediately, before it is used to make a commercial or personnel
judgment.
The reason is not to soften the finding. It is that an unadjusted gap will be
attacked on exactly these grounds during the first serious review, and an
adjusted gap that survives is far more forceful than a larger unadjusted one
that does not. Sizing the real difference protects a correct conclusion as much
as it prevents an incorrect one.
Say this explicitly when presenting, because "we are adjusting the number" is
otherwise easily heard as "we are explaining the problem away."
## What good looks like at the end
- Every quality figure in circulation carries a grade and a coverage rate
- No Grade C figure appears inside a stated difference
- Service level is never published without a quality measure beside it
- The residual gap between delivery models is agreed by all parties before the
performance conversation begins
- The instrument roadmap is a contractual commitment, not an aspiration
Usage notes
Sizing. The instruction block is roughly 470 words and loads with every message in the project. The five reference files total approximately 4,100 words and load only when retrieved.
Scope. The pack covers comparability, discovery and sequencing. It does not cover rubric design, sampling design or calibration mechanics — those are in Quality Management — nor the capability of the scoring platforms themselves, which is in Quality Assurance Platforms in Contact Centers and Speech Analytics.
Worked figures. The confidence intervals and detectable differences in Block 3 and in the summary table are computed for a score near 85% on 20,000 interactions per month, using the normal approximation to the binomial. They are illustrative of the relationship between coverage and resolution; recompute for the actual volumes and score levels before quoting.
Relationship to the capability ontology. Quality score is an attribute in the supply-side ontology described by Wiki:Packs/Agent Capability Ontology, and its coverage asymmetry is called out there as a trap. Where an organisation intends to route supply on performance data, comparability is a prerequisite: an allocation engine fed uncomparable quality scores will systematically misallocate work, at speed and invisibly.
Drift. Blocks 2 and 4 derive from the source articles named in the infobox. Blocks 3, 5 and 6 are original to this pack and have no article home yet — a maintenance risk, since there is no canonical version to regenerate from.
Change history
| Version | Date | Change |
|---|---|---|
| 1.0 | 2026-08-07 | Initial publication. One instruction block, five reference blocks. |
| 1.1 | 2026-08-07 | Added the approach diagram. |
See also
- Wiki:Packs — pack index and how packs work
- Wiki:Packs/Agent Capability Ontology — the supply-side attribute model this feeds
- Quality Management — rubric design, sampling design, calibration
- Quality Assurance Platforms in Contact Centers — the scoring platform estate
- Speech Analytics — automated interaction scoring
- First Contact Resolution — the most commonly used outcome measure
- Performance-Based Vendor Allocation Design — routing volume on measured performance
- Business Process Outsourcing — the outsourced delivery tier
