Performance-Based Vendor Allocation Design

From WFM Labs

Performance-Based Vendor Allocation Design is the practice of distributing work among multiple outsourcing suppliers according to measured performance instead of fixed contractual shares. It appears wherever a contact center or back-office estate is delivered by several outsourcing vendors performing comparable work, and the buyer intends measurement to change supplier behavior. The design problem is harder than the scorecard it resembles. In service operations the difficulty of the work each vendor receives varies enough to dominate raw score differences — the condition that makes risk adjustment necessary when comparing hospitals or schools — so a mechanism built on unadjusted scores ranks the queues and not the suppliers.[1][2] Whether that condition holds in a particular estate is testable, and establishing it is the first step in any credible design.

A second constraint compounds the first: measurement produces signal continuously, while vendor capacity responds over months. A mechanism that moves work faster than suppliers can staff it oscillates, degrading service in both directions.

The four requirements

A mechanism that changes vendor behavior satisfies four conditions simultaneously. Failing any one produces activity without improvement.

Requirement Definition Consequence of failure
Comparable Differences in score reflect differences in performance, not differences in the work received Vendors optimize their intake instead of their service; the best-informed supplier wins
Consequential A better score produces a materially better commercial outcome Vendors acknowledge the scorecard and disregard it
Achievable A vendor can act on the signal within the time available before the consequence lands The consequence functions as a penalty for something the supplier could not have changed
Credible Vendors believe the measurement is honest Suppliers litigate the mechanism instead of responding to it

The first and third conditions absorb most of the design effort and are treated at length below. The second is largely a choice among commercial instruments. The fourth follows from publishing the method.

Comparability

Vendors do not receive identical work. Routing preference, language capability, client assignment, geography, time-zone coverage and historical accident all cause queues to differ, and the resulting difficulty spread is the dominant source of variation in raw scores. This is the problem addressed in health-care provider profiling and in educational value-added measurement, where outcomes must be compared across providers serving materially different populations.[1][2] The methods transfer directly to vendor comparison.

The underlying caution is older than either literature: ranking suppliers on outcomes dominated by variation in the system they operate within produces a ranking of the system, not of the suppliers.[3] See W. Edwards Deming and Statistical Process Control for WFM.

Comparability cannot be recovered by tuning a mechanism after launch. Once vendors have demonstrated that a comparison was unfair, the mechanism loses the credibility it depends on. Two constructions are used together.

Statistical adjustment

The expected outcome for each interaction is predicted from characteristics of the interaction itself, and the vendor is scored on the gap between observed and expected results instead of on the observed level.

The governing constraint is an admissibility criterion for covariates, not a list of them: a covariate must be knowable before the work is assigned, and outside the vendor's influence.

A covariate that fails the first test may have been affected by the treatment, a well-documented source of bias in which adjustment removes part of the effect being measured.[4][5] A covariate that fails the second lets a supplier reclassify performance failures as difficulty. Both errors bias in the vendor's favor, so neither is likely to be disputed by the party best placed to notice it.

Characteristics that generally qualify include what the customer is contacting about, the arrival channel, the product or segment, complexity of the underlying record, language and market, time of day and week, and whether an external disruption was active. Characteristics that generally do not qualify include anything describing how the interaction went, anything measured afterward, anything reflecting the vendor's own staffing or tenure, and anything the vendor reports about itself. Prior contact history is a subtle case: where repeat contact is itself the quality measure, adjusting for it partially conditions on the outcome.

Randomized assignment

Statistical adjustment corrects for differences that were measured. It cannot correct for differences that were not, and a fitted model contains no internal test of whether its covariate list is complete.

A fraction of work is therefore assigned without regard to routing preference, subject only to hard feasibility constraints such as language, licensing and security clearance. Because assignment is independent of every characteristic, measured or unmeasured, the resulting comparison is unbiased by construction. The size of the slice follows from the precision needed to detect a material disagreement between the two rankings — the same uncertainty calculation that gates share movement, described under guardrails below — and not from a fixed convention.

Its primary use is as a validity check on the adjustment model. Each cycle, vendors are ranked by adjusted score and separately by performance on the randomized slice. Persistent disagreement between the two rankings indicates that the covariate set is missing something material, and identifies the adjusted ranking as the one at fault. The design mirrors the quasi-experimental tests used to establish whether value-added estimates are biased by non-random assignment.[6]

Without randomized assignment, the claim that one vendor outperforms another and the claim that one vendor received easier work are observationally equivalent. Where a routing platform cannot randomize, this becomes a prerequisite to build, not a refinement to defer.

What the mechanism measures

Three components are measured separately, because they fail independently and a vendor may be strong on one while weak on another.

Quality. The load-bearing sub-measure is resolution, defined behaviorally: no further contact from the same customer about the same matter, through any channel, within a fixed window. A disposition-based definition, in which the handling agent marks the interaction resolved, is gamed within a single cycle and produces transfer-and-close behavior. A channel-scoped definition pays for pushing unresolved matters into a different queue, which is worse than no measure at all. See First Contact Resolution and Outcome-Based Work Measurement.

Efficiency. Cost per resolved unit, not cost per contact and not handle time. Handle-time efficiency pays vendors to end interactions, which suppresses resolution and returns as repeat volume the buyer pays for twice. Normalizing cost by resolution is the formulation in which efficiency and quality do not directly oppose one another — a specific instance of the multitask incentive problem, where rewarding an easily measured dimension diverts effort from a harder-to-measure one.[7]

This normalization places resolution inside two components, so its effective weight exceeds its nominal weight. The effective weight is published, since vendors will otherwise derive it privately and conclude the scorecard is not what it states.

Reliability. Delivery against committed staffing, service attainment, ramp attainment against agreed milestones, and attrition. Reliability carries independent weight because it protects the flexibility the buyer is paying for. A supplier scoring well on quality while missing every staffing commitment is worth less than a slightly weaker supplier that meets its ramp curve, and flexibility that is not scored will not be managed — which dissolves the rationale for multi-sourcing. Attrition is the leading indicator of the other two components and the figure most likely to be understated; it is normally required at individual level with a tenure distribution and reconciled against system-access records.

Two structural rules follow. Channels are scored separately. Voice, chat, asynchronous and back-office work carry different economics, ramp times and quality constructs, and a blended score conceals the distinctions a buyer would act on. It also forecloses the most valuable move available: concentrating each vendor into the work it is demonstrably best at. Measurement is taken from the buyer's systems. Vendor-supplied figures are inputs to reconciliation only.

Converting score into consequence

Two instruments are available, with materially different response times. Measurement produces signal continuously, but vendor capacity responds on the timescale of hiring, training and nesting — a matter of months for complex work (see Agent Onboarding and Nesting Period Management). That mismatch is the central design constraint, and it implies that the instrument which moves volume cannot also be the instrument that creates urgency.

Why volume allocation underperforms as an incentive

Where a share of volume is made contestable each cycle and awarded by score, the contestable pool bounds the entire prize. If twenty percent of volume is contestable across three vendors, the maximum a supplier can gain by moving from worst to best is a fraction of that twenty percent, and a realistic improvement moves a fraction of that again.

The committed base cannot be compressed far to widen the pool. Below a certain volume a vendor cannot sustain supervision, a training pipeline and site overhead; quality then degrades, and the mechanism manufactures the failure it penalizes. The volume lever is bounded by the same viability constraint that makes multi-sourcing workable. It is an effective strategic instrument, concentrating work over time where it performs best, and a weak incentive instrument.

Fee at risk

A portion of fee placed against the same index and settled monthly responds immediately, scales smoothly, requires no capacity change to act upon, and can be sized to matter without threatening supplier viability.

A single index drives both instruments. Two scorecards — one governing fee, another governing volume — allow a vendor to optimize the cheaper one and create an unresolvable dispute about which is authoritative.

Gainshare on volume reduction

Where vendor revenue scales with contact volume, the supplier is structurally opposed to the largest efficiency available to the buyer: removing contacts. Returning a share of verified volume reduction to the vendor realigns that interest. Absent such a term, a rational counterparty's reluctance is easily misread as a performance problem.

Share allocation and guardrails

Component scores are normalized to a 0–100 scale and weighted into a single index. A finer scale than a five-point one avoids compression, where all vendors cluster into a narrow band and the mechanism loses resolution. Weights are a governance decision, published to vendors and held for a stated period announced in advance, because changing weights mid-program is difficult to distinguish from changing the award after observing the result.

A softmax over deviations from the cohort mean is one allocation form:

share(v)=exp(λ(IvI¯))jexp(λ(IjI¯))

where Iv is vendor v's index, I¯ the cohort mean, and λ the sensitivity of share to a one-point index difference. This form is invariant to where the scale's zero sits. A ratio-of-powers formulation is not: under a power form, ratios compress as all vendors improve, and the mechanism weakens precisely as the program succeeds.

Worked illustration

Three vendors, a committed base of 70% of volume divided equally, 20% contestable, 10% held as surge capacity, and λ=0.09. Figures are illustrative.

Vendor Index Contestable share Total volume share
A 78 44.8% 32.3%
B 71 23.9% 28.1%
C 74 31.3% 29.6%

If vendor C improves from 74 to 79:

Vendor Index Contestable share Total volume share Change
A 78 38.1% 30.9% −1.4
B 71 20.3% 27.4% −0.7
C 79 41.7% 31.7% +2.1

A five-point index gain moves roughly two points of total volume, about a 7% increase in that vendor's work, arriving slowly. That arithmetic is the argument for pairing allocation with fee at risk. Raising λ does not substitute for it: a higher sensitivity parameter widens the spread between vendors at any moment but barely increases the reward for improving, because the contestable pool caps the prize regardless.

Guardrails

Guardrails are applied after the formula, not inside it.

Guardrail Basis
Floor at the committed base Vendor viability, derived from site economics
Ceiling on any single vendor Concentration risk, derived from recovery time and tolerable degradation
Cap on share movement per cycle Ramp feasibility
Minimum precision before any move Prevents reallocation on noise
Demonstrated capacity gate Binding, not advisory

Floors and ceilings asserted as round percentages, without a stated derivation, are difficult to defend under challenge. The capacity gate is the guardrail most often waived under schedule pressure: awarding volume beyond demonstrated capacity converts a quality reward into a service failure, which is then attributed to the vendor and not to the allocation. Volume that cannot be awarded to the leading performer cascades to the next supplier with capacity.

Sample thresholds require separate treatment. A residual from a model carrying several high-cardinality categorical controls has considerably wider uncertainty than a raw average computed on the same number of observations, so thresholds set for raw means are too permissive for adjusted scores. The gate is placed on the uncertainty of the estimated vendor effect: share moves only when a vendor's interval excludes the cohort mean. See Signal and Noise in WFM.

Stability under measurement–actuation delay

Fast measurement, slow actuation and a long transport delay describe a control system prone to oscillation. Uncontrolled, volume moves to the leading vendor, which becomes overloaded relative to its staffing, whose scores then fall, and volume returns. The dynamic is structurally the same as the amplification observed in supply chains when decisions are made on delayed information.[8]

Control Rationale
Long measurement window, weighted toward recent periods Single-cycle scores are too noisy to move volume on
Allocation cadence matched to ramp time, by channel Quarterly where hiring is slow, monthly where it is not
Deadband applied to cumulative drift since the last change Suppresses noise without penalizing a vendor improving steadily but slowly
Advance notice before an award takes effect An award a supplier cannot staff in time is worthless to both parties
Rate limit on share change per cycle Bounds the actuator

A related failure arises where the adjustment model is fitted once and left static while the composition of work shifts through automation, self-service or product change. The model progressively under-adjusts, and the drift is uneven, penalizing vendors serving the fastest-changing segments most. Refitting on the cadence at which the work actually changes, and publishing per-vendor mix composition as a control metric, addresses it.

Gaming exposures

Every measure in a consequential mechanism is optimized against. This is a general property of performance measurement, not a comment on supplier integrity.[9] See Goodhart's Law and Metric Gaming and Game Theory and Incentive Design in WFM.

Exposure Control
Closing interactions without resolving them Behavioral resolution definition; transfer-out and deflection rates scored under reliability. Rising transfer-out is the earliest indicator
Selecting favorable work Routing enforced in the buyer's platform; no vendor intake discretion; accept and defer patterns monitored
Influencing satisfaction measurement Survey solicitation triggered centrally by the buyer
Uneven automated-scoring coverage Coverage rate published per vendor per cycle. Systematic drop-out by route, codec, accent or language converts a measurement gap into a scoring difference; beyond a stated coverage gap the cycle's comparison is treated as void
Selecting audit samples All sampling designed and conducted by the buyer, stratified
Understating attrition Individual-level reporting reconciled against system-access records
Using disputes to delay Fixed dispute window; upheld disputes adjust the following cycle and never suspend the current one

Credibility operates as a control in its own right. Suppliers who believe the method is fair respond by improving; suppliers who believe it is arbitrary respond by arguing. Published methodology, published covariates and a genuine dispute path determine which behavior the mechanism produces, which makes transparency a design parameter and not a communications choice.

Governance

Four functions are required, however they are housed:

  • Approving allocations and any guardrail exception, with operations, sourcing, quality, finance and analytics represented
  • Owning the method — covariates, weights and sensitivity — and validating the adjusted ranking against the randomized slice. This function is deliberately separated from the body that approves awards, so the method cannot be adjusted to produce a preferred outcome
  • Reviewing with each vendor bilaterally and frequently
  • Calibrating across vendors through a shared methodology walkthrough, presenting aggregate distributions and not a leaderboard

Phasing and prerequisites

The mechanism is normally run in a live-but-inert state first, publishing scores and the share changes they would have caused while moving no volume. The dominant failure mode in these programs is going live on an unvalidated adjustment, having a vendor correctly demonstrate that its queue was harder, and losing credibility with the entire supplier community at once. A dry period surfaces that at no commercial cost. Widening the contestable share is deferred until the adjusted ranking and the randomized-slice ranking have agreed across consecutive cycles.

Two prerequisites are difficult to retrofit. Cross-channel identity resolution is required for the behavioral resolution definition, which is the load-bearing measure. Randomized assignment in the routing platform is required for the adjustment to be validated at all. Most other elements can be added incrementally.

A final consideration is whether the apparatus is worth its own cost. Randomized routing, identity resolution, interaction-level capture, periodic model refitting, independent audits and multiple governance bodies carry a standing overhead that is sized against the expected gain, using the methods in Applied Measurement and Estimation for WFM. Where the estate is small or the vendor set is short, the proportionate answer is often to begin with fee at risk on a simpler scorecard and defer allocation until the prerequisites exist for other reasons.

Maturity Model Position

Performance-based allocation maps unevenly onto the WFM Labs Maturity Model™, and the mapping is set by the comparability apparatus, not by whether allocation is performance-linked at all.

  • Level 1 — Initial (Emerging Operations) and Level 2 — Foundational (Traditional WFM Excellence): allocation is fixed by contract. Performance is governed through a monthly scorecard and a penalty or bonus tied to service level, with volume shares unchanged between contract cycles.
  • Level 3 — Progressive (Breaking the Monolith): allocation is linked to performance, but on unadjusted scores. Scorecards are produced faster and reviewed more often without the underlying comparison becoming defensible. Most programs that fail do so at this level, because the mechanism becomes consequential before it becomes comparable.
  • Level 4 — Advanced (The Ecosystem Emerges): case-mix adjustment validated against a randomized slice, a single index driving both fee and volume, guardrails derived from site economics and recovery time, gainshare on verified volume reduction, and allocation cadence matched to ramp. This is the level at which the mechanism becomes a planning instrument instead of a reporting artifact.
  • Level 5 — Pioneering (Enterprise-Wide Intelligence): vendor capacity and internal capacity are planned as a single supply pool, with allocation rebalanced continuously against a shared demand signal instead of on a governance calendar. Continuous rebalancing does not remove the delay problem described above; it relocates the controls, since the deadband and rate limit must then operate inside the routing layer rather than inside a quarterly committee decision. See Three-Pool Architecture.

The distinction that matters for assessment is between Level 3 and Level 4, and it is testable in a single question: what would falsify the current ranking? An operation with no answer to it is scoring queues.

Use this with Claude

A ready-to-deploy instruction set and reference files are at Wiki:Packs/Service Quality Comparability (CP-WFM-003).

See Also

References

  1. 1.0 1.1 Ash, A. S., Fienberg, S. E., Louis, T. A., Normand, S-L. T., Stukel, T. A. and Utts, J. (2012). Statistical Issues in Assessing Hospital Performance. Committee of Presidents of Statistical Societies white paper commissioned by the Centers for Medicare and Medicaid Services. PDF
  2. 2.0 2.1 Iezzoni, L. I. (ed.) (2013). Risk Adjustment for Measuring Health Care Outcomes, 4th edn. Chicago: Health Administration Press.
  3. Deming, W. E. (1986). Out of the Crisis. Cambridge, MA: MIT Center for Advanced Engineering Study.
  4. Rosenbaum, P. R. (1984). "The Consequences of Adjustment for a Concomitant Variable That Has Been Affected by the Treatment." Journal of the Royal Statistical Society, Series A, 147(5): 656–666. doi:10.2307/2981697
  5. Angrist, J. D. and Pischke, J-S. (2009). Mostly Harmless Econometrics: An Empiricist's Companion. Princeton: Princeton University Press, §3.2.3 ("Bad Control").
  6. Chetty, R., Friedman, J. N. and Rockoff, J. E. (2014). "Measuring the Impacts of Teachers I: Evaluating Bias in Teacher Value-Added Estimates." American Economic Review, 104(9): 2593–2632. doi:10.1257/aer.104.9.2593
  7. Holmström, B. and Milgrom, P. (1991). "Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design." Journal of Law, Economics, & Organization, 7 (Special Issue): 24–52. doi:10.1093/jleo/7.special_issue.24
  8. Lee, H. L., Padmanabhan, V. and Whang, S. (1997). "Information Distortion in a Supply Chain: The Bullwhip Effect." Management Science, 43(4): 546–558. doi:10.1287/mnsc.43.4.546
  9. Muller, J. Z. (2018). The Tyranny of Metrics. Princeton: Princeton University Press.