Vendor Operating Model Maturity

From WFM Labs

The vendor operating model maturity model grades the client-side capability to source, govern and get value from a portfolio of external and internal service delivery. It assesses the buying organisation, not the vendor — a distinction most published frameworks make and most practice does not.

It is designed for organisations running a multi-tier labour model: in-country employees, captive low-cost-country centres staffed by their own employees, and third-party BPO providers. Those three behave differently on cost, tenure, flexibility and relationship depth, and a model that treats them as three points on a cost curve will misplace work.

This page covers the maturity ladder. It does not restate the operational mechanics, which are covered at BPO and Vendor Management for WFM, Performance-Based Vendor Allocation Design, Build vs Buy and Vendor Governance and Offshoring and Nearshoring.

Why a client-side maturity model

Sourcing outcomes correlate more strongly with how intensively the relationship is managed than with which vendor was selected or what the contract said. Lacity and Willcocks, across three decades of empirical work on IT and business process sourcing, find consistently that selective outsourcing outperforms both total outsourcing and total insourcing, and that joint decisions taken by senior executives together with operational managers outperform either group deciding alone.[1]

The survey evidence is consistent with this. Only 46% of organisations that outsource report being satisfied with their provider. Organisations holding monthly formal business reviews are 1.3× more likely to be satisfied; those with a formal vendor management office are 1.2× more likely to be satisfied and 2× more likely to use a risk/reward commercial model.[2] Separately, organisations with a mature vendor management office report third-party spend reductions of 20% or more at 61%, against 22% for developing functions.[3]

Capability is the variable. The model below grades it.

The spine: unit of management

The level definitions derive from the eSourcing Capability Model for Client Organizations (eSCM-CL v1.1), developed by the IT Services Qualification Center, originally at Carnegie Mellon University, with client organisations including American Express, Boeing, BP, General Motors and L'Oréal.[4]

Its value is that levels are defined by what the organisation manages as a unit, not by adjectives:

Level Unit of management Practices at this level
1 0
2 The engagement 58
3 The portfolio 29
4 Business impact against external reality 8
5 Durability 0 new

The distribution is informative. The overwhelming majority of defined practice sits at Level 2 — running a single relationship properly. Only eight of ninety-five practices sit at Level 4, which is to say the value-management capabilities are, by construction, the ones almost no organisation has.

Staged, not capability-based

eSCM-CL states explicitly that it is a capability model rather than a maturity model: practices at a higher level may legitimately be adopted before all lower-level practices are complete.

This model departs from that deliberately. It is staged — an organisation occupies a level, and reaching it requires the levels beneath. The reason is audience: a staged ladder can place a function and describe the next plateau, which a capability profile cannot do as directly. The cost is real, and should be acknowledged wherever the model is used: staging conceals the common and legitimate pattern of an organisation strong on delivery and weak on exit.

The five levels

Level 1 — Transactional

Work is handed to a provider and becomes their problem. Sourcing is treated as a cost exercise rather than a management discipline. Capability is personal rather than institutional: it resides in whoever negotiated the arrangement, and leaves when they do. Performance is discussed when something breaks.

Level 2 — Managed engagement

Each relationship is run properly. Defined service levels, scheduled reviews, named owners on both sides, executive sponsorship. Measurement answers one question: is this provider meeting its service levels?

Every relationship is nonetheless managed in isolation. The estate has no shape, and nothing aggregates. This is where the majority of functions sit, and it is where the majority of published practice is directed.

Level 3 — Managed portfolio

The unit of management shifts from the deal to the organisation. Common process assets across providers, a single performance repository, measurement aggregated across engagements, risk managed across the estate rather than per contract, and a sourcing strategy that exists in writing with a named owner.

Work is placed deliberately across tiers by suitability rather than landing where history put it. This transition is the single most consistent inflection point across every framework surveyed, and it is the one most organisations have not made.

Level 4 — Value managed

Capability baselines sufficient to support statistically valid statements about expected performance and to compare processes against one another. External benchmarking used to set targets rather than merely to check price. The sourcing portfolio analysed for its impact on the business — retention, revenue, risk — rather than on service conformance.

Demand-side and supply-side accountability are explicitly assigned and separately measured. Commercial constructs carry a genuine outcome component with defined baselines and attribution rules.

Level 5 — Workforce orchestration

All labour — in-country, captive, contracted, contingent and digital — governed as one adaptive system and allocated dynamically as conditions change. Deloitte describes the emerging organisational form as an extended workforce management office, governing outsourcing providers, captive centres, freelancers and digital workers under a single operating model.[3]

Level 5 requires sustained demonstration, not achievement. Following eSCM-CL, it introduces no new capability: it requires that Level 4 capability be shown to hold across multiple cycles over a period of at least two years. A function cannot be assessed at Level 5 on a single assessment.

The seven dimensions

Dimension Weight Scope What it grades
Goals 10% estate Whether the footprint has a stated purpose beyond cost, and whether objectives are weighted
Roles 15% estate The retained organisation, and whether accountability tracks decision rights
Processes 15% engagement · FTE Routing, performance management, governance cadence, transition and exit
Relationships 10% engagement · FTE Provider partnership, and separately, continuity of the customer relationship
Technology 10% engagement · FTE Platform reach into provider operations, and measurement integrity
Commercial 20% engagement · spend Contracts, pricing constructs, flex mechanisms, and whether terms are operated
Portfolio 20% estate The multi-tier design itself, and how work is placed within it

Commercial and Portfolio carry the heaviest weight because they determine what is possible rather than how well it is done. An organisation cannot manage flexibly under a contract that prices per full-time equivalent, and cannot place work well without a portfolio design. The other five dimensions govern execution within those constraints.

The scope column is the second axis and matters as much as the weight. Three dimensions have exactly one answer for the organisation however many providers it runs; four vary between individual arrangements and are assessed differently. This is developed at #Assessing an estate with several providers.

Goals

Level Description
1 No stated rationale. Unit cost is the only figure quoted
2 Cost targets explicit per provider. Other objectives named but not scored
3 A written sourcing strategy with a rationale per tier. Flexibility, capability and coverage named alongside cost
4 Objectives weighted and traded explicitly. Flexibility is priced rather than assumed free. Value measured at business impact
5 The objective function is live, owned, and re-tested on a defined trigger

Roles

This dimension carries the model's central diagnostic.

Level Description
1 No retained function. Whoever signed the arrangement owns it
2 Named vendor managers per relationship, accountable for provider performance without holding the routing decision
3 A vendor management function with organisational standing. Demand-side decisions still made elsewhere, but visible and contestable
4 Demand shaping and provider performance re-integrated under one accountable structure — or, where they cannot be, the decision-owner is formally co-accountable. Demand-side capability has its own metrics
5 An intelligent client function shaping demand, sourcing and governing performance as one discipline across all labour types

The UK National Audit Office, examining government technology sourcing, documents both the mechanism and its consequence. Commercial teams lead procurement decisions "often without the benefit of digital expertise", with the result that requirements specialists consider essential "can be removed by commercial teams as 'savings' to the contract". The consequence is stated directly: "under-specifying may leave the buyer at risk of receiving a service that does not actually meet its needs yet still leave the buyer accountable for that failure."[5]

The same report records commercial directors identifying three necessary elements — demand planning and strategy; procurement, sourcing and contracting; and supplier performance and relationship management — and finding that practice "mainly addresses" the second.

The remedy in the literature is consistent and is worth stating plainly: accountability should track control.[6] A structure in which one function decides where work goes on cost grounds while another is measured on the outcome is a recognised failure mode, not a design. No published framework offers a responsibility matrix separating accountability for the sourcing decision from accountability for provider performance; every model treats the client as a single actor. That gap is why this dimension exists.

Processes

Level Description
1 Volume is transferred. Governance undefined
2 Defined service levels, scheduled reviews, a scorecard. Transition handled as a project when it arises
3 Common process assets across providers. Tiered governance. Case-mix adjustment, so reviews are not arguments about whose work was harder. Exit planned before it is needed. Governance intensity follows a stated rule — materiality, risk or work type — rather than history or account-team preference
4 Performance managed statistically — variance as well as level. Routing between tiers is a managed process with stated criteria. Where practice differs between arrangements the difference is documented and justified; unexplained variance is treated as a defect
5 Allocation adjusts dynamically to conditions within pre-agreed commercial bounds

Governance cadence is measurable and correlates with outcome: 54% of organisations hold monthly formal business reviews and 26% quarterly.[2] The three-tier structure documented by the University of Tennessee — daily operational, monthly joint operations, quarterly executive, with both parties represented at every tier — is the best-evidenced published design.[7] Governance itself is estimated to cost 3–8% of contract value, averaging approximately 4.2%.

Note what Level 3 does not ask for. It does not ask for uniform governance intensity. Governing a twelve-person arrangement as heavily as a five-hundred-person one is waste, and the reverse is exposure; proportionality is the test, not uniformity. What Level 3 asks is whether a rule exists and is applied. Level 4 then separates designed variance from drift: ask why one arrangement is governed differently from another, and if the honest answer refers to deal vintage, inherited teams or established habit, that is drift.

Relationships

This dimension covers two distinct things, and conflating them is a common error.

Level Provider relationship Customer relationship continuity
1 Transactional; escalation is the relationship Not considered
2 Professional, with named counterparts Acknowledged, not designed for
3 Partnership behaviours; transparency both directions Work requiring continuity is identified and deliberately placed
4 Survives a bad quarter without renegotiation Continuity measured as an outcome and priced into placement
5 Joint capability development Continuity designed across tiers; the customer experiences one service

Evidence status. The strongest available evidence that continuity of a named individual affects outcomes comes from healthcare, where a Norwegian registry study of 4,552,978 patients found a monotonic dose-response: longer relationship with the same general practitioner associated with lower mortality, fewer acute admissions and less out-of-hours use at every additional tenure band.[8] A separate English study of 230,472 patients found the effect strongest for the heaviest users of the service.[9] Meta-analytic work in marketing finds relationship effects are stronger when the relationship is held with an individual rather than with the selling firm.[10]

This is proxy evidence and must be labelled as such. No published study links service-agent continuity to client retention in commercial service industries. The mechanism plausibly transfers; the magnitudes do not.

Technology

Level Description
1 The provider reports its own performance; figures cannot be audited
2 Partial direct system access. Measurement partly ours, partly theirs
3 All measurement conducted in client systems from client data. Provider figures are inputs to reconciliation, never to scoring. Analytics coverage verified uniform across tiers
4 Instrumentation equal across all tiers, enabling like-for-like comparison. A coverage differential invalidates a comparison
5 Dynamic routing across tiers technically possible and routinely exercised

The case for grading this severely rests on a single striking finding. Asked what proportion of frontline staff leave annually, BPO executives self-reported a median of 25%. Benchmark data aggregated independently across the same population gives a median of 87%. At team-leader level the figures are 10% reported against 70% measured, with no respondent reporting above 60%.[2]

Whatever the true figure in any given operation, the implication for the model is unavoidable: decisions resting on provider-supplied performance data rest on sand. Level 3 is defined by the elimination of that dependency.

Commercial

Level Description
1 Per-FTE or per-hour pricing. Terms negotiated once and filed. No flexibility bands, no exit provision
2 Defined service levels and credits. Volume bands may exist but are not administered
3 Pricing construct matched to the purpose of the tier. Flexibility bands live and administered. Exit provisions written and tested. Cost measured per resolution rather than per contact. Cost expressed in one normalised unit across arrangements, so differently-priced contracts remain comparable
4 Hybrid pricing with a genuine outcome layer — defined indicators, normalised baseline, attribution rules, deadbands and audit rights. Benchmarking clauses with an expert-determination backstop. Where constructs differ, the difference traces to work type or tier rather than to deal vintage
5 Constructs differ deliberately by tier and work type, and are re-tested as the portfolio changes

Market calibration, which is what keeps the upper levels honestly rare:

Measure Value
Per-FTE pricing in use 53%[2]
Buyers open to outcome-based pricing 76%[11]
Buyers actually operating one 5%[11]
Organisations with a formal vendor management office 44%[2]

Note the tension in the first figure: per-FTE pricing dominates the market, and per-FTE is the construct that delivers the least flexibility — while flexibility is the second most cited reason for outsourcing, ranked identically by buyers and providers.[2]

Two traps the model grades against:

  • Service credits are a risk premium, not a behaviour lever. The larger the penalty a buyer demands, the higher the price a provider quotes; and a large credit payout reduces the provider's available cash to address root causes.
  • Cost per contact understates true cost. Where first-contact resolution runs at industry-typical levels, roughly 30% of contacts are repeats, and repeats cost materially more than an initial resolution. Published benchmarking puts average cost per contact at approximately $9.10 against cost per resolution of approximately $12.74 — an understatement of around 40%. A provider priced per contact has no incentive to reduce repeat contact, because repeats are revenue.

Portfolio

Level Description
1 Binary in-house or outsourced. Placement is historical. True footprint unknown
2 Footprint known. Placement decided per arrangement, on unit cost
3 Deliberate segmentation across all tiers by complexity, channel, value at stake and relationship dependence. The captive tier used purposefully rather than skipped
4 Placement carries a full cost of ownership including rework, escalation, attrition and relationship cost. Continuity risk quantified. Concentration quantified on every axis that carries risk — provider, geography, business line and work type — not only the one routinely reported
5 Portfolio re-optimised as conditions change; flexibility genuinely exercised rather than notional

Provider diversification and geographic diversification are different things, and they routinely diverge. An estate spread across four providers with no share above 35% reads as well diversified. If all four deliver from the same two countries, geographic concentration can exceed 80%, and four providers buys nothing against a country-level event. Reporting only provider share is the common case and scores zero at Level 4.

The three tiers

Tier Unit cost Tenure Flexibility Relationship depth
In-country employees Highest Long Low Highest
Captive low-cost-country centre Middle Long Low to moderate Can be high
Third-party provider Lowest Short by design High — this is the product Structurally hardest

The middle tier is the one most often absent from the conversation. A captive centre captures much of the unit-cost advantage while retaining tenure and relationship continuity, and explicitly does not provide flexibility. A model framed as in-house versus outsourced skips it, and that binary framing is itself a maturity marker.

There is one substantial piece of evidence bearing directly on the choice. A study of 150 North American firms over nine years, using national customer satisfaction index data cross-referenced against contemporaneous reporting to date offshoring events, found that outsourcing front-office functions is associated with lower customer satisfaction both offshore and onshore, at similar effect sizes — with the effect attributed to the organisational boundary rather than to geography. Back-office functions showed no significant satisfaction decline from offshoring.[12]

Read carefully, this supports the captive tier specifically: it is the boundary, not the border, that is associated with the satisfaction penalty.

What is not evidenced. No published study compares captive low-cost-country centres with third-party providers on attrition, tenure, domain knowledge or quality in a contact-centre context. The proposition that captive centres retain staff longer is a reasonable prior held consistently across the practitioner literature and demonstrated by none of it. This model does not rest on it.

Assessing an estate with several providers

The most common question put to this model is whether an organisation running several providers needs several assessments. It does not, and the reason matters more than the answer.

If a question cannot be answered once because providers genuinely differ, that is the result. Level 2 is the engagement; Level 3 is the estate. Needing one document per provider is the Level 2 finding. Splitting the assessment into separate documents does not solve that problem — it conceals it, by removing the only place the inconsistency would have appeared.

There is also no single correct way to split. A real estate varies by provider, business line, work type, geography and commercial construct simultaneously. Choosing one axis to organise separate assessments around silently suppresses the other four, and the suppressed axis is frequently the one carrying the risk.

Estate-level and engagement-level dimensions

Scope Dimensions Why
Estate-level Goals, Roles, Portfolio — 45% of weight One sourcing strategy, one retained function, one estate shape. These do not decompose. Assessing the shape of a portfolio from inside a single relationship is incoherent, so there is exactly one answer however many providers exist
Engagement-level Processes, Relationships, Technology, Commercial — 55% Contracts, system access and governance cadence are per-relationship by construction. Score the weakest material arrangement, then record where the item does and does not hold

Coverage, and the two readings it produces

For engagement-level dimensions, record coverage only where an item does not hold everywhere — typically a minority of items. This produces two readings of every dimension:

  • Floor — the weakest material arrangement. What can defensibly be asserted about the estate as a whole.
  • Coverage — the share-weighted reading. What is true of most of the population.

Both are honest; they answer different questions. The overall placement should be computed from the floor, because a staged model ought to report the figure it can defend rather than the flattering one.

The difference between them is the diagnostic, and it is best expressed two ways. Practice spread is the mean difference across level blocks on a continuous scale — how uneven practice actually is. Maturity gap is the difference in dimension score — what that unevenness costs in levels. The gap steps at the block-satisfaction threshold, so a very small spread sitting exactly on that threshold can produce a large gap; read the spread first and use the gap to size the prize.

The implication is the useful part. A wide spread means the capability exists somewhere and has not been generalised — which is the definition of managing engagements in isolation. Level 3 is not better practice; it is the same practice holding across the estate. Closing that difference therefore requires adopting nothing new, only applying what already works. That is a materially cheaper improvement path than the level number alone suggests.

Weighting basis

Processes, Relationships and Technology weight coverage by headcount. Commercial weights by spend. This is not a stylistic preference: the two measure different exposures and diverge in real estates, where an arrangement holding a modest share of headcount can hold a substantially larger share of spend. Weighting a commercial question by headcount would understate it.

Materiality

Coverage is assessed across material arrangements — those above a stated share of headcount or spend, on either measure rather than both. An arrangement can be small in headcount and large in spend, and that is precisely the one that must not be skipped.

Below the threshold, ask a single question instead: is the light touch deliberate and rule-based, or is it neglect? Proportionate governance is a Level 3 marker; uniform governance is not.

Tier and engagement are separate lenses

Arrangements nest inside a tier. An estate may have three tiers and ten arrangements, nine of which sit in one tier. Do not collapse the two decompositions into one.

Processes, Relationships and Technology are the three dimensions where practice diverges between tiers as well as between arrangements. For these, gather evidence per tier and score the weakest, on the same principle: the estate is only as governed as its least-governed part, and the common pattern is a well-run contracted estate sitting alongside an ungoverned captive one.

Scoring

Each evidence item is scored 1 where the practice is present and demonstrable, 0.5 where partial, and 0 where absent. Items not assessed are left blank and excluded from the mean rather than counted as zero.

A level block is satisfied at a mean of 0.75 or above. The level achieved is the highest level whose block and every block beneath it is satisfied. The dimension score is the level achieved plus the next block's mean, producing a continuous scale.

The staged rule is severe and deliberately so. Block means of 1, 0.5, 1, 1, 1 across the five levels yield Level 1, not Level 4. Strong advanced practice resting on an incomplete foundation earns no credit.

Commercial is scored on the contracted estate only. A captive centre has no commercial contract, and scoring it as though it does produces a meaningless figure; the internal service agreement, where one exists, belongs under Processes.

Level 5 cannot be awarded on a single assessment. It requires Level 4 capability demonstrated across at least two assessments spanning a minimum of two years.

Limitations

The staged form conceals real and legitimate unevenness; eSCM-CL's own position is that non-sequential capability adoption is valid, and a capability profile would be more actionable for an internal audience even though it is less usable for an executive one.

The market calibration figures are drawn from surveys that do not disclose sample size or methodology in their published form, and should be treated as directional. Several widely circulated outsourcing statistics were tested during the construction of this model and could not be traced to any source; they are excluded.

The maturity gap described above steps at the block-satisfaction threshold and can therefore overstate unevenness when block means sit close to 0.75 on either side. This is the staged form behaving as designed, but it makes the gap unsafe to read alone; the continuous spread measure exists for that reason.

The model grades capability, not outcome. A function can score well and still be sourcing badly if its strategy is wrong — capability is the ability to execute a strategy, not evidence that the strategy is correct.

See also

References

  1. Lacity, M. and Willcocks, L., "A Review of the IT Outsourcing Empirical Literature and Future Research Directions", Journal of Information Technology, 2010.
  2. 2.0 2.1 2.2 2.3 2.4 2.5 COPC Inc., Global Benchmarking Series 2022: Contact Center Outsourcing, March 2022. Surveys fielded September–December 2021, 900+ executives. COPC sells certification and vendor-management consulting; several findings nonetheless cut against provider interests.
  3. 3.0 3.1 Deloitte, Global Outsourcing Survey 2024. Sample size and methodology not disclosed in the published report; treat percentages as directional.
  4. ITSqc, eSourcing Capability Model for Client Organizations (eSCM-CL) v1.1, January 2010. 95 practices across 17 capability areas and five capability levels.
  5. National Audit Office, Government's approach to technology suppliers: addressing the challenges, HC 543, January 2025.
  6. Vitasek, K. et al., Unpacking Sourcing Business Models, University of Tennessee. States the principle that "a performance-based agreement should hold a supplier accountable only for what is under its control."
  7. Vitasek, K. et al., Unpacking Outsourcing Governance, 2nd edition, September 2022.
  8. Sandvik, H. and Hetlevik, Ø. et al., "Continuity in general practice as predictor of mortality, acute hospitalisation, and use of out-of-hours care", British Journal of General Practice 72(715), 2022.
  9. Barker, I., Steventon, A. and Deeny, S., "Association between continuity of care in general practice and hospital admissions", BMJ 356:j84, 2017.
  10. Palmatier, R. et al., "Factors Influencing the Effectiveness of Relationship Marketing: A Meta-Analysis", Journal of Marketing 70(4), 2006. 94 studies, approximately 38,000 relationships.
  11. 11.0 11.1 HFS Research, analysis of outcome-based pricing in customer experience services, 2026.
  12. Whitaker, J., Krishnan, M. S. and Fornell, C., "How Does Customer Service Offshoring Impact Customer Satisfaction?", Journal of Computer Information Systems, 2019. Earlier version presented at AMCIS 2006.