Interpreting WFM Maturity Assessments
Interpreting WFM Maturity Assessments is the discipline of converting completed maturity assessment results — dimension scores, interview responses, and evidence artifacts — into a defensible maturity position and an investment sequence. It is a distinct skill from administering an assessment: the instrument produces numbers, and interpretation determines what the numbers can honestly support. The discipline matters because the most common assessment failures occur after data collection: averaging away a gated model, accepting inflated self-reports, imputing scores the estate cannot produce, misreading scoring behavior that carries signal, drawing conclusions from multi-unit comparisons that confound assessor and capability, and converting an investment-prioritization instrument into a ranking exercise. This article covers the interpretation rules; the model itself is described at WFM Labs Maturity Model™, and level-to-level change management at Navigating WFM Maturity Transitions.
Gated ladders, not averages
Staged maturity models are defined by prerequisite structure: each level presupposes substantial satisfaction of the level below, a design principle established in the original Capability Maturity Model, whose staged representation exists precisely because capabilities at higher levels are unstable without lower-level foundations.[1] The interpretive consequence: an organization's level is bounded by its weakest gating dimension, and a flat average across dimensions manufactures a middle position nobody occupies. An estate with strong scheduling automation and no forecast discipline is not "a 2.5" — it is a Level 1 foundation carrying Level 3 tooling, which is a different diagnosis with a different remedy.
In practice, assessments are frequently presented as gated ladders and then computed as flat averages — the presentation and the arithmetic contradict each other, and the arithmetic wins unless interpretation reimposes the gates. The corrective is mechanical: compute dimension-level indications separately, state the binding gate explicitly, and report the level as "bounded by X" rather than as a decimal. A bounded-position statement reads, for example: "Level 2, bounded by Process — real-time and scheduling capability indicate Level 3, but the forecast cycle lacks the definitional consistency Level 3 presupposes."
Self-report inflation and the evidence rule
Self-assessment of capability is systematically miscalibrated, with the least capable performers overestimating most — the pattern documented by Kruger and Dunning for individual skill self-estimates.[2] Its transfer to organizational capability self-scoring is an inference rather than a documented result, but one field observation supports it consistently: in maturity assessment practice, self-reported capability scores have been observed to run on the order of two-thirds of a maturity point high in capability dimensions, collapsing on contact with artifacts — a claimed modeling capability falling from 3.0 to 1.0 when the model is asked for, a claimed scheduling practice falling from 5 to 2 when the artifact is requested [field-observed; practitioner experience across multi-unit assessments, not controlled measurement].
The correction is structural rather than arithmetic:
- The evidence rule. Any item scored at 4 or above requires a named evidence artifact — a file, a system screen, a produced report. Without one, the score is recorded as claimed rather than verified, and the two are carried as separate columns that are never merged.
- Ask for artifacts, not descriptions. Requests for an existing file tend to get answered; requests to describe how something works often return nothing. Where a description is genuinely needed, requesting it as a drawn workflow with names in the boxes converts an unanswerable question into a producible artifact.
- The credibility gap as a metric. Computing the position twice — once on verified and computed items only, once with claims included — yields a distance between the two positions that is itself a reportable finding about the estate's self-knowledge.
The not-assessable class
Some items cannot be scored because the estate cannot produce the number: a platform that does not compute the metric, a definition that differs across systems, a data feed that does not exist. Not-assessable is a different finding from scores-low, with a different remedy — instrumentation rather than practice change — and imputing a score for it, or averaging over it, destroys the assessment's most actionable output. The not-assessable inventory deserves its own section in any findings report, framed as the instrumentation agenda, because measurement-dependent ambitions at higher maturity levels are gated on precisely these items.
Scoring behavior as data
Measurement regimes shape the behavior of the measured — the dysfunctional consequences Ridgway catalogued apply directly to self-scored assessments.[3] Interpretation should therefore treat scoring behavior as signal rather than noise:
- A leader scoring their own area low is frequently signaling investment appetite, not merely position — record it as both.
- The failure mode to guard against is the high score given out of pride: consensus highs on capability dimensions are claims to test against artifacts, while consensus lows are findings that can be acted on directly.
Multi-unit estates and assessor effects
Where several units are assessed by several assessors, raw score differences confound three sources: genuine capability difference, platform or heritage difference, and assessor effect. Interpretation should identify which comparisons carry weight by construction:
- Bounds the assessor effect: units sharing platform and heritage but scored by different assessors — the spread among them is a free upper bound on assessor variation.
- Reads genuine difference: a single assessor scoring two different units — the cleanest available comparison, since the assessor is held constant.
- Uninterpretable: comparisons that vary platform, heritage, and assessor simultaneously — these cannot carry weight and should be labeled as such rather than silently included.
Where the same metric is computed differently across units or platforms, comparability itself is the finding; publishing a single blended score across incomparable instruments misstates what is known. A cheap calibration exercise — every assessor independently scoring one named unit on a small common item set — both quantifies disagreement and surfaces the definitional divergences that are usually the estate's deepest finding.
The purpose reframe
The single most consequential interpretive decision is what the assessment is for. Treated as a grade, it invites inflation, defensiveness, and ranking; treated as an investment-prioritization instrument, its output is hot spots and sequence rather than a number. The prohibition that protects the instrument follows directly: never publish a league table of units. The moment scores rank people, the assessment converts into a threat, and every future data collection inherits the distortion.
Two further disciplines follow. First, every identified gap is labeled by fix class, because the label is the action:
- Practice/configuration — cheap and fast; changeable by the teams themselves.
- Platform/instrumentation — must precede any measurement-dependent ambition; belongs on the technology roadmap.
- Structural — needs sponsorship and organizational decisions; the slowest class, so it is surfaced earliest.
Second, timeliness beats precision: a directional reading delivered before investment decisions are taken is worth more than a refined one delivered after, and interpretation cadence should be planned accordingly.
Maturity Model Position
The interpretation disciplines here apply at every level, but their yield differs. At Levels 1–2, the not-assessable inventory and the evidence rule dominate — most of the correction is separating claims from capability. At Level 3, multi-unit comparability and assessor-effect separation matter most, because integration work spans units. At Levels 4–5, the gates themselves become the interpretive focus: higher-level claims are frequent and the consistency checks (claimed evergreen planning against an annual budget cycle; claimed AI-inclusive planning with AI capacity absent from the plan) do the discriminating work.
Use this with Claude
A ready-to-deploy instruction set and reference files for analyzing assessment results are at Wiki:Packs/Maturity Assessment Analysis (CP-WFM-009). The vendor-estate sibling — for Vendor Operating Model Maturity results — is Wiki:Packs/Vendor Maturity Assessment Analysis (CP-WFM-010).
See Also
- WFM Labs Maturity Model™
- Navigating WFM Maturity Transitions
- Vendor Operating Model Maturity
- Executive Communication for WFM
References
- ↑ Paulk, Mark C.; Curtis, Bill; Chrissis, Mary Beth; Weber, Charles V. (1993). "Capability Maturity Model, Version 1.1". IEEE Software 10(4): 18–27. https://doi.org/10.1109/52.219617
- ↑ Kruger, Justin; Dunning, David (1999). "Unskilled and Unaware of It: How Difficulties in Recognizing One's Own Incompetence Lead to Inflated Self-Assessments". Journal of Personality and Social Psychology 77(6): 1121–1134. https://doi.org/10.1037/0022-3514.77.6.1121
- ↑ Ridgway, V. F. (1956). "Dysfunctional Consequences of Performance Measurements". Administrative Science Quarterly 1(2): 240–247. https://doi.org/10.2307/2390989
