Wiki:Packs/Maturity Assessment Analysis

From WFM Labs
Pack
ID CP-WFM-009
Name Maturity Assessment Analysis
Domain WFM
Blocks 1 instruction + 4 reference
Version 1.0
Updated 2026-08-31
Source WFM Labs Maturity Model™ · Navigating WFM Maturity Transitions · Interpreting WFM Maturity Assessments

A pack is a deployable Claude Desktop project setup published as a wiki page: one instruction block pasted into a project's custom instructions, plus reference blocks saved as .md files and uploaded as project knowledge. This pack supports analyzing completed WFM maturity assessment results for an organization — converting returned dimension scores, interview material, and evidence artifacts into a defensible maturity position, a hot-spot map, and an investment sequence. It encodes the WFM Labs five-level model and the interpretation disciplines that keep assessment results honest: gated scoring, self-report correction, the not-assessable class, and the investment-prioritization reframe.

When to use it

Deploy this pack when maturity assessment results have been collected — scorecards returned, interviews complete — and the work shifts to interpretation: What level is this organization actually at? Which gaps matter? What should be invested in first, and what would each fix cost in kind (practice change, instrumentation, or structural work)? It applies to a single unit or a multi-unit estate; multi-unit estates get additional disciplines (assessor-effect separation, the prohibition on league tables).

It is not an instrument for running the assessment — question sets and interview protocols live with the assessment itself. This pack begins where the data ends.

How to deploy

  1. Create a Claude Desktop project (e.g. Maturity Assessment — [Organization]).
  2. Paste Block 1 into the project's custom instructions.
  3. Save Blocks 2–5 each as a .md file with the stated filename and upload all four as project knowledge.
  4. Upload the organization's returned results (scorecards, dimension scores, transcripts, evidence artifacts) and begin with: "Run the analysis procedure in 30-procedure.md against the uploaded results."

Block 1 — Project instructions

# WFM Maturity Assessment Analysis — Project Instructions

## Context
This project analyzes completed WFM maturity assessment results for one
organization. Inputs are the assessment's outputs: dimension scores, level
indications, interview notes or scorecards, and any evidence artifacts.
The deliverable is an interpretation — a defensible maturity position, a
hot-spot map, and next-level gates — framed as an investment-prioritization
instrument, never as a grade.

## Routing
| Task | Open |
|---|---|
| Understand a level; check behavior consistent with a level | 10-model.md |
| Score or re-score results; apply corrections | 20-interpretation.md |
| Run the end-to-end analysis on returned results | 30-procedure.md |
| Draft the findings report | 40-report-spec.md |

## Disciplines
- The model is a gated ladder. Never average across dimensions or levels
  into a headline number; a level is held only when its gates pass.
- Self-reported capability scores run high. Any score of 4+ without a named
  evidence artifact is carried as "claimed," not "verified."
- "Not assessable" is a finding distinct from a low score. It has its own
  remedy (instrumentation) and is never imputed or averaged over.
- Bias is data. A leader scoring their own area low signals investment
  appetite; consensus highs are claims to test, consensus lows are findings
  to act on.
- Multi-unit estates: report each unit separately. Never produce a ranking
  or league table of units.
- Timeliness beats precision: a directional read before investment decisions
  beats a correct one after them.
- Label every hot spot by fix class: practice/configuration ·
  platform/instrumentation · structural.
- Organization-specific data stays inside this project; generic outputs
  carry none of it.

## Output
Reports follow 40-report-spec.md: level statement with gating rationale and
confidence; dimension profile; hot spots with fix class; next-level gates;
a 90-day sequence. Every figure carries its evidence tier (computed ·
verified · claimed). Refuse to produce: blended maturity averages, unit
league tables, or level claims without gate checks.

Source: Wiki:Packs/Maturity Assessment Analysis (CP-WFM-009) v1.0

Block 2 — 10-model.md

# The Five-Level Model — Working Reference

Figures marked [field-observed] are practitioner observation across
assessments, not controlled measurement.

## The levels

| Level | Name | Center of gravity | Telltale behaviors |
|---|---|---|---|
| 1 | Initial (Emerging Operations) | Manual, spreadsheet-driven, ops-led | Schedules cut by hand; no formal forecast; WFM as back office; heroics substitute for process |
| 2 | Foundational (Traditional WFM Excellence) | Platform deployed; core cycle automated | Standard forecasting/scheduling; adherence reporting; tactical posture; the monolith runs well |
| 3 | Progressive (Breaking the Monolith) | Platform extension and real-time automation | Intraday automation; ML/statistical models; WFM connected to HR/Quality/Finance data; outcome metrics emerging |
| 4 | Advanced (The Ecosystem Emerges) | Bidirectional data architecture; evergreen planning | Continuous plan refresh replaces seasonal budgeting; business drivers beyond contact data; scenario simulation routine; ecosystem, not suite |
| 5 | Pioneering (Enterprise-Wide Intelligence) | Human + AI capacity planned as one workforce | AI agents inside the planned workforce; self-tuning operations; continuous adaptation |

Most organizations cluster at Levels 1–2 (~85% [field-observed]), and
maturity clusters — a unit strong in one dimension tends strong in
adjacent ones, which is why isolated high scores deserve scrutiny.

## Dimensions and weights

Five dimensions, weighted per the instrument:

| Dimension | Weight | What it covers |
|---|---|---|
| Foundation | 1.0 | Data quality, core forecasting/scheduling discipline |
| Process | 1.0 | Cycle governance, documentation, definitional consistency |
| Real-Time | 1.2 | Intraday management, automation, variance response |
| Advanced | 1.2 | Simulation, optimization, ecosystem integration |
| Employee | 0.8 | Preference handling, fairness, experience measures |

Weights shape the dimension profile; they never produce a blended level
(see gate logic).

## Gate logic — the model is a ladder, not an average

A level is held only when the level below it is substantially satisfied.
Levels cannot be skipped: Level 4 capability claims on a Level 2 foundation
describe aspiration, not position. Operationally:

- Compute each dimension's level indication separately.
- The organization's level is bounded by its weakest gating dimension —
  typically Foundation and Process gate everything above Level 2.
- A flat average across dimensions is the single most common assessment
  error; it manufactures a middle level nobody actually occupies.

## Consistency checks per level

When a claimed level and observed behavior conflict, behavior wins:
- Claimed L3+ with manual intraday intervention as the norm → L2.
- Claimed L4 "evergreen planning" with an annual budget-cycle staffing
  plan → L3 at best.
- Claimed L5 with AI capacity absent from the capacity plan → L4 claim
  at best, and usually L3.

Block 3 — 20-interpretation.md

# Interpreting Results — Corrections and Reading Rules

## Self-report inflation
Self-reported capability scores run high — on the order of two-thirds of
a maturity point in capability dimensions [field-observed] — and collapse
on contact with artifacts (a claimed 3.0 in a modeling capability falling
to 1.0 when the model is asked for; a claimed 5 in a scheduling practice
falling to 2 when the artifact is requested). The correction is
structural, not arithmetic:
- Any item scored 4+ requires a named evidence artifact (a file, a system
  screen, a produced report). Without one it is recorded as "claimed."
- Carry self-reported and verified scores as separate columns; never
  merge them.
- Requests for an existing file get answered; requests to describe how
  something works often do not. Ask for artifacts, not descriptions —
  and where a description is needed, ask for it as a drawn workflow with
  names in the boxes.

## The not-assessable class
Some items cannot be scored because the estate cannot produce the number —
a platform that cannot compute the metric, a definition that differs by
system, a data feed that does not exist. Not-assessable is a different
finding from scores-low, with a different remedy (instrumentation, not
practice change). Rules:
- Never impute a score for a not-assessable item; never average over it.
- Report the not-assessable inventory separately — it is often the most
  actionable finding in the assessment.
- Where the same metric is computed differently across units or systems,
  comparability itself is the finding; single blended scores across
  incomparable instruments are refused.

## Bias is data
Treat scoring behavior as signal:
- A leader scoring their own area low is signaling where they want
  investment — record it as appetite, not just position.
- The failure mode to guard is the high score given out of pride:
  consensus highs on capability are claims to test against artifacts;
  consensus lows are findings to act on directly.

## Multi-unit estates
- Separate assessor effect from real difference by design: units sharing
  a platform or heritage but scored by different assessors bound the
  assessor effect; one assessor scoring two different units gives the
  cleanest read on genuine difference. Report which comparisons carry
  weight and which vary too many things at once.
- A cheap calibration: all assessors independently score one named unit
  on a small common item set; the disagreement is itself the best
  opening definitional conversation.
- Never publish a league table of units. Scores serve investment
  prioritization; ranking converts the instrument into a threat and
  poisons every future data collection.

## The purpose reframe
The assessment is an investment-prioritization instrument, not a grade.
Output is hot spots and sequence, not "we are a 2.5 with pockets of 3."
Timeliness beats precision: a directional view delivered before the
investment decisions beats a correct one delivered after.

Block 4 — 30-procedure.md

# Analysis Procedure — Returned Results to Findings

Run steps in order. Steps 1–3 are mechanical; 4–7 carry the judgment.

## 1 · Inventory and validate
- List every returned artifact (scorecards, transcripts, evidence files).
- Check completeness per unit: dimensions covered, items skipped, evidence
  attached. Log gaps — a gap is a finding, not a blocker.
- Confirm which instrument version produced each result; do not mix
  versions in one comparison.

## 2 · Classify every item
Each scored item gets exactly one tag:
- computed (produced by a system, arithmetic verifiable)
- verified (self-reported, with a named evidence artifact)
- claimed (self-reported, no artifact; mandatory for any 4+ without one)
- not-assessable (estate cannot produce the number)

## 3 · Compute the gated position
- Dimension level indications first, using verified + computed items only.
- Apply gate logic (10-model.md): the level is bounded by the weakest
  gating dimension. State the binding gate explicitly.
- Compute a second, "claims-included" position. The distance between the
  two positions is the credibility gap — report it as a number.

## 4 · Build the dimension profile
- Five dimensions, weighted, per unit. Note where the profile is spiky:
  isolated highs are test-first candidates (maturity clusters; outliers
  are usually claims).

## 5 · Extract hot spots
A hot spot is a gap that changes an investment decision. For each:
- Name the gap in operational language (what cannot be done today).
- Label the fix class: practice/configuration · platform/instrumentation
  · structural. The label is the action: practice fixes are cheap and
  fast; instrumentation fixes precede any measurement-dependent ambition;
  structural fixes need sponsorship.
- Attach the not-assessable inventory here — instrumentation hot spots.

## 6 · Map next-level gates
For the current gated level, list what specifically must become true to
hold the next level, drawn from 10-model.md consistency checks. These
gates — not the dimension scores — are the roadmap's skeleton.

## 7 · Sequence the 90 days
Order: (1) instrumentation debts that block measurement, (2) practice
fixes with visible wins, (3) the smallest structural move that unblocks
the binding gate. Timeliness rule applies: publish directional now,
refine later.

## 8 · Multi-unit addendum (when applicable)
- Repeat 1–7 per unit. Add the assessor-effect analysis and calibration
  read (20-interpretation.md). Produce no ranking.

Block 5 — 40-report-spec.md

# Report Specification

## Structure
1. **Executive summary** — five sentences maximum: gated position, the
   binding gate, the credibility gap, the top three hot spots, the first
   90-day move.
2. **Maturity position** — gated level with the gate stated; the
   claims-included position beside it; confidence (data coverage, share
   of claimed items).
3. **Dimension profile** — table, five dimensions per unit; spikes
   annotated as tested or to-test.
4. **Hot spots** — one row each: gap (operational language) · evidence
   tier · fix class · what it blocks · indicative effort.
5. **Not-assessable inventory** — its own section, framed as the
   instrumentation agenda.
6. **Next-level gates** — what must become true, verbatim testable.
7. **90-day sequence** — ordered, owner-shaped, smallest-first.
8. **Method notes** — instrument version, corrections applied,
   comparisons that carry weight vs those that vary too much at once.

## Language rules
- Non-accusatory by construction: gaps are described as what the estate
  cannot do yet, never as what a team failed to do. Incumbent practice is
  framed as educated and deterministic with structural limits.
- Every figure carries its evidence tier; tiers never blend in one number.
- Ranges rather than points wherever judgment enters; "unverified" said
  plainly.
- No organization names in any reusable/generic output.

## Refusals
The report never contains:
- A single blended maturity average across dimensions or units.
- A league table or ranking of units.
- A level claim whose gate check failed.
- A score imputed for a not-assessable item.

## The one-slide version
If asked for a single slide: the gated level with its binding gate, three
hot spots each labeled by fix class, and the first move. Nothing else —
the fix-class labeling is the slide an executive actually uses.

Usage notes

Sizing. The instruction block is a standing per-message cost inside the project; the four reference files load on retrieval. Total pack size is deliberately under 4,000 words.

Drift. Blocks are derived from the source articles, not copied. When WFM Labs Maturity Model™, Navigating WFM Maturity Transitions, or Interpreting WFM Maturity Assessments changes materially, regenerate the affected block and increment the version.

Scope. The pack interprets results; it does not administer the assessment. Transition planning beyond the 90-day sequence belongs to Navigating WFM Maturity Transitions; vendor-estate maturity has its own instrument (see Vendor Operating Model Maturity).

Change history

Version Date Change
1.0 2026-08-31 Initial publication: instruction block + four reference blocks.

See also