WFM Labs Maturity Assessment Methodology

The WFM Labs Maturity Assessment Methodology is the procedure by which an organization's workforce management function is formally placed on the WFM Labs Maturity Model™: a sectioned questionnaire is scored against evidence, the scores are aggregated into a section profile, the profile is calibrated across the people who produced it, and the result is read back as an investment sequence rather than a grade. It is the operating manual for the instrument that WFM Assessment describes in outline — that page names the process domains, the four evidence sources, and the three deliverables; this page covers how the item bank is built, how evidence is attached to each score, how items roll up, how disagreement between raters is handled, and what the readback contains. It is distinct from The Maturity Diagnostic and Pathways, a five-pillar self-check answered in minutes; the methodology here is the multi-week engagement form that the self-check approximates. Interpreting WFM Maturity Assessments owns the rules for reading a completed result — the evidence rule, the not-assessable class, the league-table prohibition — and this page points to them where the procedure invokes them. The GRPI-T Framework supplies the pillars by which findings are sequenced; the Future WFM Operating Standard supplies the practice each pillar should exhibit at each level.
What the assessment measures

The instrument measures a workforce management function along two axes: a set of sections that partition the function into scorable areas, and the five levels of the maturity model — Initial, Foundational, Progressive, Advanced, Pioneering — against which each item's anchor text is written. The sections are the unit of measurement; the levels are the unit of interpretation.
The scored workbook has six sections. One is an explicit maturity block, with a short cluster of items under each of the five level headings so that a respondent's answers indicate which level descriptions the operation satisfies. The other five are capability areas: capacity planning (split into foundational planning, interval variance, occupancy and multi-skill calibration, volatility, and simulation), employee well-being and attrition, WFM talent, WFM process, and technology. Each section carries a fixed weight in the headline — a design choice that has changed between versions — with capacity planning the heaviest, and within it the assessor's rating of the capacity model and its flexibility the single largest input. The sections differ from the process domains WFM Assessment uses to describe scope: a domain describes what a function does, a section how the workbook scores it, and one domain may be scored across several sections.
Nor are the six sections the five GRPI-T pillars. The sections are how the instrument collects; the pillars are how findings are sequenced in the readback, because the pillar ordering rule — a symptom at one pillar usually has its cause one pillar higher — turns a list of low scores into an order of work. The mapping is approximate:
| GRPI-T pillar | Instrument sections that mostly feed it |
|---|---|
| Goals | Maturity block; capacity planning targets, occupancy bands, planning horizon |
| Roles | WFM talent: individual capability, leadership effectiveness, team structure and knowledge management |
| Processes | WFM process: forecasting, scheduling, real-time, cross-functional integration; capacity planning practice items |
| Interpersonal and Interconnected | Employee well-being and attrition; cross-functional integration items |
| Technology | Technology section; the source-system derivability check described below |
The instrument
The item bank is on the order of 140 scored items — the exact count varies by version — organized into the six sections and, within them, into blocks of three to twelve items. Every item is answered on a five-point scale in which 1 denotes high risk and 5 denotes low risk. The scale is a summated-rating design in the tradition of Likert,[1] but the anchors are behavioral rather than agree–disagree wherever possible. Four anchor families recur: frequency (Always to Never) for practice items; coverage (All to None) for items about definitions and data sources, such as whether shrinkage sources are defined and captured; a narrated rubric of five level paragraphs for the few items with recognizable stages; and a binary gate (Yes or No mapped to 1 or 5) in the maturity block, to establish whether a level description is met at all.
Two item types sit outside the summated scale. Assessor-completed ratings — capacity model coverage, capacity model flexibility, and the occupancy-limit judgment — are entered by the assessor after direct examination rather than answered by the client, on a lettered ladder from AAA to D− mapped to 5.0 down to 0.5. Derived inputs are computed from client data: the attrition score is a matrix lookup of attrition-versus-target against the direction and speed of change. Beside every scored item is an evidence-and-notes cell, which has more than once held the finding that most changed a readback.
Three administration rules apply. Score the run process, not the designed process — the respondent answers for what happened over the recent period, not for what the procedure document says. Blank means "don't know" and is excluded from the average, never scored as zero; early versions treated empty cells as zero and depressed one respondent's section score by roughly a third of a point, so completion rate per section is now reported as its own table — a section answered at sixty percent is itself a maturity signal. Reverse-scored and binary items are validated mechanically before aggregation — the workbook carries no data validation, so an integrity pass flags non-binary answers on binary items, half-points, out-of-range values, and reverse-worded items apparently answered forwards, and every adjustment is logged.
Evidence collection
WFM Assessment names the four evidence sources — document review, practitioner interviews, system review, and performance data analysis. The methodology assigns each a role and a weight of belief, because the four do not agree: document review supplies artifacts, interviews supply self-report and its corroboration, system review supplies direct observation of the platform, and performance data analysis supplies outcome history. Self-assessed capability is biased toward inflation at the individual level,[2] and self-report instruments carry common-method variance that inflates apparent consistency.[3] The engagement-level observation that shapes the methodology — capability scores collapsing on contact with artifacts, while items describing the operating environment tend to hold — is recorded at Interpreting WFM Maturity Assessments. The evidence classes are ordered accordingly:
| Class | What it is | Role in scoring |
|---|---|---|
| Self-report | The questionnaire, answered by a segment leader or a practitioner | The starting score; recorded as claimed until corroborated |
| Artifact | An existing file — a capacity model, a schedule, a shrinkage taxonomy, a forecast-error report | Converts a claimed score to verified; required for any item scored 4 or above |
| System observation | The assessor examining the WFM platform, routing configuration, or a report definition directly | Establishes what the platform can compute; the basis of the not-assessable class |
| Performance data | Interval-level service, forecast error, adherence, attrition, and occupancy history | Surfaces outcomes that contradict claimed process quality |
The engagement opens with a numbered data request built on the artifact-over-description rule at Interpreting WFM Maturity Assessments: organization charts, capacity models, forecast and schedule outputs, intraday logs, process documents, platform inventories, scorecards, prior performance reports, and routing configuration. Before the corresponding items are scored, source-system derivability is checked: for occupancy, utilization, shrinkage, handle time, and adherence, the assessor establishes whether the platform emits the components the metric needs, and a unit whose platform cannot compute a metric is recorded as not assessable rather than scored low.
The full engagement typically runs several weeks — problem-statement interview, short onsite, data request, interviews stratified from executive to practitioner level, interim readback, final report and deck — with the questionnaire answered before the interviews so that interview time goes to the items where self-report and artifact disagree.
Scoring and aggregation
Mechanics. An item block scores as the unweighted mean of its answered items; a section as the mean of its blocks, except where the workbook assigns an explicit split, as the capacity section does in favor of its assessor-completed model rating. The headline is the weighted mean of the six section scores, divided by the weight actually answered, so an unanswered section does not pull the total toward zero. Scoring is rebuilt independently from the raw item cells rather than read from the workbook's totals, because adapted workbooks accumulate hard-typed values over formulas and roll-ups that reference the wrong cells, and any difference between the rebuilt number and the workbook's is logged as a finding about the instrument.
Hazards. Three properties of the arithmetic must be corrected in interpretation.
- It is a flat average, not a gated ladder. The maturity block is presented as five levels and computed as a mean, so an operation can satisfy the Level 5 description while failing the Level 3 one, and the mean reports a comfortable middle. Staged maturity models define levels as prerequisite structures;[4] the methodology therefore never describes a unit as "at Level N" from the headline. Instead each section mean is banded to a level — 1.0 to 1.9 is Level 1, 2.0 to 2.9 is Level 2, and so on — and the diagnostic's rule is applied to the six banded section levels in place of its five pillar scores: the median is the routing key, any section two or more levels below the median is a drag indicator, and the position is stated as a pair, the level by median bounded by the lowest section. With six sections the median of the third and fourth banded levels is rounded down when they differ, so the routing key is never a level fewer than three sections have reached; it is deliberately unweighted, since the headline already carries the section weights.
- Item weighting inside a block does not exist, and block weighting across sections is uneven. A fifteen-item section carrying a fifth of the headline moves roughly four times as far per answer as a fifty-item section carrying a quarter. Equal weights within blocks are a defensible default — improper linear models with equal weights are robust[5] — and the report states the section weights it used.
- A material share of the headline is typed judgment. The assessor-completed ratings and the separately collected scores can together account for close to half of the headline in some versions, and the ratings use undefined terms — the coverage and flexibility of a capacity model — for which different assessors assume different denominators. Each is listed in its own table with the basis on which it was judged; where no common basis can be reconstructed across units, the sub-score is reported beside the headline rather than inside it.
Two scores are always carried: a claimed score from all answered items, and a verified score from items backed by an artifact, a system observation, or performance data. The distance between them is a reportable finding about the unit's self-knowledge, and has been observed to exceed the distance between units. Following Hubbard, a measurement is any observation that reduces uncertainty about a decision-relevant quantity;[6] the verified score reduces uncertainty, the claimed score records the prior, and the readback says which decisions the gap bears on.
Calibration across raters
One assessor scoring one unit has an unmeasured rater effect; several assessors scoring several units have a confounded one. The methodology treats calibration as a designed step.
Within a unit. Where a section is answered by more than one respondent — the maturity block deliberately is — the responses are not silently averaged; the spread is reported, and a leadership-versus-practitioner gap is read as a maturity indicator, the leader usually describing the intent of a process and the practitioner its execution.
Across units. Which cross-unit comparisons carry weight by construction is settled by the classes at Interpreting WFM Maturity Assessments. What the methodology adds is the specification of the calibration exercise that page names: every assessor independently scores one named unit on a common block of six to ten items, usually from the occupancy and multi-skill block where definitional divergence is most likely, before any cross-unit result is compared. Because the items are ordinal, agreement on the block is reported with a statistic that penalizes a one-versus-five disagreement more than a four-versus-five one: quadratic-weighted kappa for two raters,[7] and an intraclass correlation for more than two,[8] read against the magnitude benchmarks of Landis and Koch.[9] No numeric pass mark is set, because the exercise exists to surface definitional divergence rather than to certify raters, but agreement in the slight-to-fair range is grounds to report cross-unit comparisons as directional only. Each score also carries a confidence grade — measured with the method visible, measured but sampled or unmatched, reported without visible method, or modeled — and self-reported responses stay at "reported" until an artifact raises them.
Readback and pathway output
The readback is a written report and a deck built from it, in that order, because the report is what makes the deck defensible when a unit leader disputes a line.
The report opens with comparability, not scores: which metrics are defined the same way across units, which are comparable after a stated adjustment, and which measure different things under the same word. Opening with comparability stops a definitional problem being laundered into an apparently precise number. The profile comes next — six section scores per unit, read column by column, since two units with the same headline can have opposite shapes — then a variance map ranking blocks by their range across units, each wide range labelled capability, definitional, or assessor difference and graded. Findings by section follow, each with evidence, grade, and implication; then what the instrument could not see, framed as instrument gaps rather than respondent failures; then what follows, a pointer to discovery and standards work rather than a plan.
The position statement takes the pair form and never a single decimal alone. Every hot spot is labelled by fix class, using the three classes at Interpreting WFM Maturity Assessments. The pathway output is a sequenced roadmap in the pillar order the GRPI-T ordering rule implies — Goals before Roles before Processes before Interpersonal — with the drag section's move first and the rest in the following quarter; the minimum moves per level are on the diagnostic page and the target practice on the operating standard's pillar pages. Three rules protect the readback: no unit league table circulates; every score appears with its completion rate and its claimed-versus-verified split; and any unvalidated number carries a visible marker naming what must be checked, by whom, and by when.
Limitations
- Coverage gaps. The item bank has no item on utilization as distinct from occupancy, none on concurrency (every item assumes voice), one parenthetical mention of outsourced delivery, no capture of forecast error as a number, and nothing on cost or customer outcome. Vendor estates use the separate instrument at Vendor Operating Model Maturity.
- Scale defects. At least one anchor set is non-monotonic in its underlying variable — the occupancy-limit rubric peaks at "moderately below the theoretical maximum" rather than at the extreme — and self-assessors tend to read it as monotonic. The attrition matrix has cells out of order. Both are checked item by item before comparison.
- Version drift. Section weights, the split inside the capacity section, and the item count have changed between versions. Comparisons across engagements, or across units scored on different versions, are made at block level, not headline level.
- Small evidence base. Every engagement-derived statement on this page — the defects found, the typed share, the collapse of claimed scores on evidence, the claimed-versus-verified gap — comes from practitioner experience across a small number of engagements, not controlled measurement, and the transfer of individual-level self-assessment bias to organizational self-scoring is an inference. The rules are engineering responses to an observed failure, not measured effect sizes.
Maturity Model Position
Different parts of the method carry the load depending on where a unit sits; the cheap parts do the work at the bottom of the model and the expensive parts at the top. The maturity block's binary gates — the Yes-or-No items under the Level 1 and Level 2 headings — settle placement almost on their own, and the artifact request usually returns the first finding, that documented process is absent; the Level 1 Process Templates are the customary first output. The calibration block and the comparability section earn their cost only once an estate spans several units or platforms, characteriztically a Level 3 condition. The narrated level descriptions, the derivability check, and the banding-and-drag step discriminate at Levels 4 and 5, where claims of probabilistic planning, continuous re-baselining, and AI-inclusive capacity are tested against the plan artifacts themselves rather than accepted from a questionnaire.
See Also
- WFM Assessment — the assessment in outline: domains, evidence sources, deliverables, and when to conduct one
- The Maturity Diagnostic and Pathways — the five-pillar self-check this methodology extends, and the minimum moves per transition
- Interpreting WFM Maturity Assessments — the evidence rule, gated ladders, the not-assessable class, the comparison classes, and the league-table prohibition
- WFM Labs Maturity Model™ — the five level definitions the items are anchored to
- GRPI-T Framework — the five pillars and the ordering rule the readback sequences by
- Future WFM Operating Standard — the practice each pillar should exhibit at each level
- Navigating WFM Maturity Transitions — the change-management playbook once the pathway is chosen
- Vendor Operating Model Maturity — the sibling instrument for outsourced estates
- Wiki:Packs/Maturity Assessment Analysis — instruction set for analyzing a returned assessment with Claude
References
- ↑ Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22(140), 1–55.
- ↑ Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it: How difficulties in recognizing one's own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121–1134. https://doi.org/10.1037/0022-3514.77.6.1121
- ↑ Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879–903. https://doi.org/10.1037/0021-9010.88.5.879
- ↑ Paulk, M. C., Curtis, B., Chrissis, M. B., & Weber, C. V. (1993). Capability Maturity Model, version 1.1. IEEE Software, 10(4), 18–27. https://doi.org/10.1109/52.219617
- ↑ Dawes, R. M. (1979). The robust beauty of improper linear models in decision making. American Psychologist, 34(7), 571–582. https://doi.org/10.1037/0003-066X.34.7.571
- ↑ Hubbard, D. W. (2014). How to Measure Anything: Finding the Value of "Intangibles" in Business (3rd ed.). Wiley.
- ↑ Cohen, J. (1968). Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220. https://doi.org/10.1037/h0026256
- ↑ Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. https://doi.org/10.1037/0033-2909.86.2.420
- ↑ Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
