Quality Score and Customer Experience Index

From WFM Labs


A quality score and a customer experience index are two instruments that measure a service interaction from opposite ends. The quality score is produced by an evaluator inside the organization, applying a rubric to a sampled interaction, and it measures whether the work was done the way the organization intends. The customer experience index is produced from the customer's side, typically as a weighted composite of sentiment, resolution and ease, and it measures what the customer got. They answer different questions, they are controllable by different parties, they fail in different ways, and they belong side by side on one scorecard rather than blended into one number. This page sets out what each is for, where each breaks, and how to use them together.

The composite discussed here is a customer experience index. Some organizations label a similar composite an experience score without saying whose experience; where the components are customer sentiment, customer-confirmed resolution and customer effort, the index describes the customer's experience, and that is the reading used on this page. Employee experience is a separate construct with its own instruments.

What a quality score is

A quality score is a process measure. An evaluator listens to or reads a sampled interaction and scores it against a rubric: greeting and verification, adherence to procedure, accuracy of the information given, handling of the system, compliance items, and a set of behaviors the organization believes produce good outcomes. The score exists because the organization knows, or believes it knows, how the work should be done and wants to check that it was done that way. In control-theory terms it is behavior control, which is available only where the organization knows the transformation process well enough to specify the right behaviors (Ouchi 1979; see Specification and the Placement of Vendor Oversight).

Its properties follow from that.

  • It is controllable by the agent. Every item on the rubric is something the agent could have done differently. That makes it the right basis for coaching.
  • It is available immediately. An interaction can be scored the same day, so the score leads the outcome it is meant to predict.
  • It is diagnostic. A low score says which behavior was missing. A low customer rating does not.
  • It is sampled and judged. A typical program scores a small fraction of interactions, and the resolution of any comparison is set by the absolute number of observations rather than the sampling percentage (see Sample Size and Detectable Difference in Quality Measurement). Two evaluators scoring the same interaction disagree unless calibrated, and agreement is measured with a chance-corrected statistic such as Cohen's kappa, with values of 0.41 to 0.60 conventionally labeled moderate and anything below that fair or worse (Cohen 1960; Landis & Koch 1977). The benchmarks are convention, not empirically derived thresholds.
  • It measures conformance, not consequence. A study of operational determinants across call centers examined service level, speed of answer, abandonment, first-contact resolution, adherence, talk time, after-call work, turnover and others. Only the percentage of calls closed on first contact and average abandonment had a significant influence on caller satisfaction, and that influence was weak (Feinberg, Kim, Hokama, de Ruyter & Keen 2000). The study did not examine rubric-scored conformance to procedure. The inference drawn here is that a rubric measuring procedure is measuring something whose link to the customer's experience is untested, not something shown to be irrelevant.

Two further findings bound how a quality score should be used. Monitoring whose content is performance-related and whose purpose is developmental is associated with better agent well-being, while monitoring that agents perceive as intense is strongly associated with worse well-being and emotional exhaustion (Holman, Chissick & Totterdell 2002). And any measure that is used to reward or punish is subject to corruption pressure: the more a quantitative indicator is used for decision-making, the more it distorts the process it was meant to monitor (Campbell 1979; Ridgway 1956).

What a customer experience index is

A customer experience index is an outcome measure. It aggregates what the customer reports or what can be inferred from the customer's side of the interaction. The three components in the composite discussed here are typical.

Resolution asks whether the customer's problem was solved, from the customer's point of view, on that contact. Customer-confirmed resolution is the operational measure most consistently associated with satisfaction (Feinberg et al. 2000), and it is the component most directly tied to whether the customer has to contact again (see First Contact Resolution). Resolution as the customer reports it and resolution as the system infers it, from the absence of a repeat contact within a window, are different measures and should not be treated as one.

Ease asks how much effort the customer had to expend. In a study of more than 75,000 customers interacting with contact centers and self-service channels, a customer effort measure predicted loyalty better than satisfaction measures or net promoter measures, and reducing effort mattered more than attempts to delight (Dixon, Freeman & Toman 2010). Effort captures the things that sit outside a single agent's control: transfers, repeat contacts, channel switches, and policy that requires the customer to do the work.

Sentiment asks how the customer felt. It is collected by survey or inferred from speech or text by a classifier (see Speech Analytics). It is the noisiest of the three: survey sentiment is subject to non-response and to the mood of the moment, and inferred sentiment carries classifier error that varies by language and channel. Satisfaction is nonetheless the construct with the longest research lineage, modeled as a function of perceived quality, expectations and value, and predicting loyalty (Fornell, Johnson, Anderson, Cha & Bryant 1996).

A weighting that puts half the index on resolution and a quarter each on ease and sentiment is defensible against that evidence: resolution is among the few operational measures with any measured relationship to satisfaction, effort outpredicted satisfaction and net promoter measures for loyalty, and sentiment is the least reliable single reading. The weights are a hypothesis about what drives the outcome the organization cares about, and they should be validated against that outcome, typically repeat contact, retention or spend, rather than fixed by convention. Aggregation necessarily discards the information in the component spread. The standard guidance for building composites is that a composite should be built on an explicit theoretical framework, that weights should be tested for sensitivity, and that the components should always be reported alongside the composite (Nardo, Saisana, Saltelli, Tarantola, Hoffmann & Giovannini 2008).

The index's properties are the mirror image of the quality score's.

  • It is only partly controllable by the agent. Policy, product, prior contacts, wait time and the customer's own situation all move it.
  • It lags. Survey responses arrive after the interaction, response rates are low, and per-interaction readings are noisy, so the index is read at the level of a team, a period or a supplier rather than a contact.
  • It is not diagnostic. It says the outcome was poor, not why.
  • It is the measure the customer would recognize. It is the right basis for accountability, allocation and contract terms, because it is the closest available reading of what the organization is paying for.

The two instruments compared

Quality score Customer experience index
What it measures How the work was done, against the organization's intent What the customer got, in the customer's terms
Who produces it An evaluator inside the organization or the supplier The customer, by survey, or a classifier reading the customer's side
Control mode Behavior control; needs process knowledge Outcome control; needs a measurable result
Controllable by the agent Fully Partly
Timing Leading; same day Lagging; days to weeks, read over periods
Grain Per interaction, sampled Per interaction but noisy; meaningful per team, period or supplier
Diagnostic Yes; names the missing behavior No; names the outcome only
Main reliability threat Evaluator disagreement; small samples; scoring what is easy to score Non-response bias; classifier error; components hidden inside the composite
Corruption pressure when used for reward Teaching to the rubric; the measured party scoring itself Survey solicitation bias; gaming the resolution question
Right use Coaching, calibration, compliance, process conformance, diagnosing an outcome movement Accountability, allocation, supplier comparison, contract terms, validating the rubric

What each is for

The quality score is for the agent and the coach. It exists so that a supervisor can tell an agent which behavior to change, so that a compliance obligation can be evidenced, and so that a procedure can be checked for conformance before its outcome is visible. It is also the instrument through which a rubric is tested: if a behavior scored on the rubric does not move the customer index when it improves, the behavior is either mis-scored or does not matter.

The customer experience index is for the organization and its suppliers. It exists so that a team, a site or a supplier can be held to the outcome the customer experienced, on a reading the measured party did not produce. It is the right basis for a supplier at-risk fee, for allocation between suppliers, and for any comparison across delivery arrangements, because it does not depend on whose evaluators scored the interaction (see Performance-Based Vendor Allocation Design and Vendor Governance Placement).

The distinction maps onto the classic separation of leading and lagging measures in performance management: the process measure predicts and explains, the outcome measure confirms and holds to account, and a scorecard that carries only one of them is blind in one direction (Kaplan & Norton 1992, 1996; see Performance Management).

Should they be combined?

There are three ways to put the two instruments together. One does not work and two do, and where a single number is forced, a fourth pattern applies.

Blending them into one number does not work. A weighted sum of a quality score and a customer index mixes a measure the agent fully controls with one the agent partly controls, a measure produced by the organization with one produced by the customer, and a leading measure with a lagging one. The blend hides which component moved, which is the information a manager needs (Nardo et al. 2008). When several tasks are rewarded through one measure and they differ in how well they are measured, effort shifts toward the better-measured task (Holmström & Milgrom 1991); applied here, the evaluator-scored rubric is the better-measured component, so the expectation is that effort drifts toward the rubric and away from the customer. And where the operational line runs its own quality evaluation, a blend lets the measured party author part of its own outcome score, a condition under which reported and real performance have been shown to diverge (Bevan & Hood 2006).

Pairing them on one scorecard works. Both measures appear, for the same unit and the same period, with the components of the index reported beneath the composite. The customer index carries the accountability weight; the quality score carries the explanation. A movement in one that is not matched by the other is the signal to investigate: rising quality scores with a flat customer index mean the rubric is scoring things the customer does not experience; a falling customer index with steady quality scores means the cause is outside the agent's behavior, in policy, product, routing or upstream contact.

Linking them works, and is the step most often left out in the operator experience behind this page. The customer index is the criterion against which the rubric is validated. Behaviors on the rubric that predict resolution and ease are kept and weighted; behaviors that do not are removed, whatever their standing in the rubric's history. Run periodically, this keeps the quality score honest. It may also reduce the perceived intensity of monitoring that is associated with worse agent well-being (Holman et al. 2002), though no study cited here tests rubric length against perceived intensity. In the other direction, the quality score is the first place to look when the index moves, because it is the one that names a behavior.

Where a single number is unavoidable, for instance a supplier fee that must settle on one figure, the pattern this wiki's doctrine proposes is gate and score rather than a weighted blend: the compliance items on the rubric act as a pass-or-fail gate, and the customer index is the score. The gate protects the things that must never be missed; the score rewards the thing the customer paid for; and neither dilutes the other.

Designing the index so it can carry the weight

If the customer experience index is going to hold suppliers and teams to account, it has to be built to survive that use.

  • Measure resolution from the customer, not from the system, or report both and name which is which.
  • State the weights as a hypothesis and test them against a downstream outcome such as repeat contact or retention; revisit when the mix of work changes.
  • Report the components with the composite, every time. A composite reported alone is a number with no explanation.
  • Control the survey path. Who is invited, when, and by which channel determines the non-response pattern; a supplier that influences invitation has influence over the score.
  • Treat inferred sentiment as a separate series from surveyed sentiment, with its classifier accuracy measured per language and channel before it is weighted in.
  • Set the grain to the noise. Per-interaction readings are for coaching conversations, not for judgments; per-period readings at team or supplier level are where the index is decision-grade.

Where the evidence runs out

The weak relationship between operational conformance measures and caller satisfaction rests on one multi-center study from 2000; it is consistent with later practitioner experience but has not been replicated at scale in the peer-reviewed literature. The effort finding comes from a practitioner research organization's dataset reported in a management periodical rather than a peer-reviewed venue. The specific weighting of resolution, ease and sentiment is not tested by any study cited here; it is a defensible starting point, and the page says so. The relationship between a given organization's quality score and its customer index is an empirical fact about that organization and has to be measured there; nothing on this page substitutes for that measurement. The pairing, linking and gate-and-score patterns set out above are design doctrine derived from the cited control and incentive literature, not findings from it; no study cited here tests them directly. The maturity levels are a construction of this wiki.

Maturity Model considerations

At L1–L2, the quality score is the only instrument, it is used for both coaching and accountability, and the rubric is inherited rather than validated. At L3, a customer measure exists but is reported separately, on a different cadence, and the two are not reconciled; where a blend exists it was built by convention. At L4, both instruments sit on one scorecard with the components of the index visible, the customer index carries accountability and supplier terms, the quality score carries coaching, and the rubric is validated against the index on a stated cadence. At L5, the weights of the index are tested against downstream outcomes, inferred and surveyed sentiment are separate series with measured accuracy, and the gate-and-score pattern governs any place a single number is required.

See Also

References

  • Bevan, G. & Hood, C. (2006). What's measured is what matters: targets and gaming in the English public health care system. Public Administration, 84(3).
  • Campbell, D.T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1).
  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1).
  • Dixon, M., Freeman, K. & Toman, N. (2010). Stop trying to delight your customers. Harvard Business Review, 88(7–8).
  • Feinberg, R.A., Kim, I., Hokama, L., de Ruyter, K. & Keen, C. (2000). Operational determinants of caller satisfaction in the call center. International Journal of Service Industry Management, 11(2).
  • Fornell, C., Johnson, M.D., Anderson, E.W., Cha, J. & Bryant, B.E. (1996). The American Customer Satisfaction Index: nature, purpose, and findings. Journal of Marketing, 60(4).
  • Holman, D., Chissick, C. & Totterdell, P. (2002). The effects of performance monitoring on emotional labor and well-being in call centers. Motivation and Emotion, 26(1).
  • Holmström, B. & Milgrom, P. (1991). Multitask principal-agent analyses: incentive contracts, asset ownership, and job design. Journal of Law, Economics, and Organization, 7.
  • Kaplan, R.S. & Norton, D.P. (1992). The balanced scorecard: measures that drive performance. Harvard Business Review, 70(1).
  • Kaplan, R.S. & Norton, D.P. (1996). The Balanced Scorecard: Translating Strategy into Action. Harvard Business School Press.
  • Landis, J.R. & Koch, G.G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1).
  • Nardo, M., Saisana, M., Saltelli, A., Tarantola, S., Hoffmann, A. & Giovannini, E. (2008). Handbook on Constructing Composite Indicators: Methodology and User Guide. OECD and European Commission Joint Research Centre.
  • Ouchi, W.G. (1979). A conceptual framework for the design of organizational control mechanisms. Management Science, 25(9).
  • Ridgway, V.F. (1956). Dysfunctional consequences of performance measurements. Administrative Science Quarterly, 1(2).