Not Assessable Is Not Performing Badly

Not assessable is not performing badly is the rule that "this cannot be measured here" and "this is performing badly" are different findings, with different evidence, different remedies and different owners. The first calls for instrumentation; the second calls for a change in practice. Conflating them produces both false accusations and false reassurance. The rule generalizes a distinction that Interpreting WFM Maturity Assessments draws for maturity scoring, where a not-assessable item is recorded as its own class rather than imputed or averaged. This page applies the same distinction to any operational metric and to the conversation an operation has with the people who are measured by it. Three neighbors carry the arithmetic; this page uses only the figure the classification turns on and cites the rest. Instrument Effects During Measurement Rollout covers score movement caused by the instrument, Baselines Before Targets the sequence, and Sample Size and Detectable Difference in Quality Measurement what a given count of observations can distinguish.
Two findings that look alike
Both findings present as a number that is missing, low, or out of line with a peer's.
| Not assessable | Performing badly | |
|---|---|---|
| What the evidence shows | The instrument cannot produce a comparable figure for this unit — the metric is undefined here, defined differently, or measured on too few observations | A comparable figure exists, on a stable instrument, with enough observations, and it is low |
| What is wrong | The plumbing: a definition, a platform capability, a data feed, a sample | The practice: skill, process, staffing, supervision |
| Remedy | Instrument the unit; agree the definition; wire the feed; raise the observation count | Coach, restaff, redesign the process, change the target's precondition |
| Owner | The function that owns the instrument, usually workforce, quality or data | The function that runs the unit |
| Time to resolve | Weeks to months of engineering, then a baseline period | Whatever the practice change takes, measured against the baseline |
| What a score means in the meantime | Nothing comparable; a placeholder, recorded as such | A measurement, with its uncertainty stated |
The two findings are also different kinds of claim in the sense of Data Synthesis Before Decision. "Cannot be measured here" is usually an established finding about the instrument. "Performing badly" is an inference from a measurement, and it inherits whatever weakness the measurement carries.
Where the not-assessable finding comes from
The finding has three common sources, and naming the source is what keeps it from being misread as a judgment about people.
The instrument. One platform cannot compute a metric that another platform reports routinely, so the units on the first platform show no figure or a proxy. Two quality tools score against different rubrics, so their outputs share a name and nothing else. An occupancy figure depends on how agents signal state changes rather than on the work performed, so units with different desktop habits show different occupancy for the same workload. In each case the instrument failed to observe; the unit did not fail to perform.
The structure. Some work is shaped so that the standard instrument does not fit it. Concurrent chat breaks a time-accounting model that assumes one state at a time. After-hours work is disruption handling and does not resemble the daytime transaction it is measured against. A small unit produces too few observations for any score to be precise. The metric is well defined; the unit's structure places it outside the range in which the definition holds.
The inherited doctrine. Operations assembled from several predecessors carry several playbooks, each correct for the business that wrote it. One predecessor measured a quantity one way; another did not measure it at all. Neither was wrong. The absence of a comparable figure is the residue of a merger, not the residue of neglect.
In all three cases the finding is about the instrument, the structure or the doctrine. It is never about the person, and a report that lets it read as if it were has made an error of the same kind as an incorrect score. Deming argued that most of the variation attributed to individuals belongs to the system they work in.[1] An instrument that cannot see the system cannot fairly grade the individual.
Stating the correction before it happens
When a measuring stick is fixed — a definition agreed, a proxy replaced with the metric it stood in for, a rubric consolidated — some numbers will look worse than they did. Nothing in the operation changed; the instrument did, and units the old instrument flattered now read where they always were.
The discipline is to say this before the change, not after. A correction announced in advance is understood as an instrument event; the same correction discovered afterward in a dashboard is read as a performance collapse, and the people whose numbers moved defend themselves against an accusation nobody intended. The announcement has three parts: which metric changes, on what date, and in which direction the affected units should expect to move. Each figure produced after the change is stamped with the instrument version that produced it, so that a comparison across the boundary is visibly a comparison across two rulers. Instrument Effects During Measurement Rollout gives the diagnostic for confirming, after the fact, that a movement belongs to the instrument: every population moves together, and the least-attended population moves most.
A consequence for targets follows. A target set against the old ruler is unfounded against the new one, in the sense Baselines Before Targets uses the word. It is withdrawn, a baseline is measured on the corrected instrument, and only then is a target set.
The small-sample corollary
A unit whose score rests on very few observations is not assessable for comparison, whatever the dashboard shows. Thirty customer surveys in a month is a common count for a small team. A unit whose true satisfaction rate is 85 percent will, on thirty responses, plausibly read anywhere between about 68 and 94 percent in a given month; the 95 percent interval on a proportion at that sample size is roughly 25 points wide.[2] A month-to-month swing of ten points is inside that interval and carries no information about the team. A monthly ranking on that sample orders teams by sampling variation rather than by anything the teams did.
Tversky and Kahneman documented the tendency to treat small samples as representative of the population that produced them, and to read a pattern into a run that ordinary variation explains.[3][4] A ranking of teams on thirty surveys each will reorder itself every month, and the reordering will be narrated as improvement and decline. Sample Size and Detectable Difference in Quality Measurement gives the thresholds: below roughly 250 observations a score is an impression rather than a measurement, and two units need several hundred observations each before a gap of a few points can be distinguished from noise.
The corollary changes the classification. A small unit with a low score is not "performing badly." It is not assessable on this instrument at this observation count. The remedy is to raise the count — by pooling months, pooling units with the same work, or sampling more — before any inference is drawn. Publishing the sample size beside every score is the cheapest way to make the classification visible, because a reader who sees "n = 30" beside a figure will not mistake it for a measurement. Sample Size and Detectable Difference in Quality Measurement adds the operational form: a published suppression rule for units below a minimum count, applied openly rather than silently.
Consequences for how findings are communicated
Not-assessable units are reported as their own inventory, never imputed and never averaged into a group figure. Interpreting WFM Maturity Assessments applies two separate rules here. A score without a named artifact is recorded as claimed rather than verified, in columns that are never merged. The not-assessable inventory gets its own section of any findings report, framed as the instrumentation agenda. Both carry over to operational metrics.
Interpreting WFM Maturity Assessments prohibits league tables of units outright; where a unit in one is not assessable the table is additionally incoherent, because it ranks instrument coverage as if it were performance. The findings are instead sorted into three backlogs, matching the three fix classes that page names. A platform or instrumentation backlog holds the units whose metric the instrument cannot produce. A structural backlog holds the units whose work or size places them outside the instrument's range; it needs sponsorship and organizational decisions, and is surfaced earliest because it is slowest. A practice or configuration backlog holds the units that are measured and low. A finding whose source is inherited doctrine joins the platform backlog once the definition has been agreed, because what remains after the agreement is instrumentation work. The three have different owners and different budgets.
The order of remedies is fixed: instrument first, baseline second, judge third. Hubbard's observation that the decision to measure should follow from the decision the measurement serves applies here; the first decision a not-assessable unit needs is whether to instrument it, and no performance decision can precede that.[5]
Failure modes
| Failure mode | What it looks like | Countermeasure |
|---|---|---|
| Imputed score | A not-assessable unit is given a placeholder figure that is then compared | Own inventory; no imputation; sample size published beside every score |
| Personal attribution | An instrument gap is discussed as a supervisor's or team's shortfall | Name the source — instrument, structure or doctrine — in the finding itself |
| Silent recalibration | A definition is fixed and the resulting drop is read as a performance collapse | Announce the change, its date and its direction in advance; stamp the instrument version |
| Small-sample ranking | Teams on tens of observations are ranked monthly and narrated | Pool to the observation count the comparison needs before ranking |
| Merged backlogs | Instrumentation, structural and practice gaps are one list with one owner | Three backlogs matching the three fix classes, each with its own owner and budget |
Maturity Model Position
At Level 1 and Level 2 most of an estate is not assessable on at least one metric of consequence, and the rule is the main protection against misreading the dashboards that do exist. At Level 3 the instrument source of the finding appears to shrink as coverage rises, and the estate is most exposed to that appearance: Sample Size and Detectable Difference in Quality Measurement places at this level the belief that raising coverage resolves comparability. At Level 4 observation counts and intervals accompany every score, so the small-sample corollary is visible by construction. At Level 5 targets are derived rather than asserted and remediation is evaluated against a comparison group, so a not-assessable unit is a gap the loop names rather than a score it fills in.
See Also
- Interpreting WFM Maturity Assessments — the not-assessable class as applied to assessment scoring; the source of the distinction
- Instrument Effects During Measurement Rollout — the diagnostic for score movement caused by the instrument
- Baselines Before Targets — why a target set before a baseline is unfounded, and how to withdraw one
- Sample Size and Detectable Difference in Quality Measurement — the arithmetic behind the small-sample corollary
- Data Synthesis Before Decision — findings separated from inference, and the graded register
- Regression to the Mean in WFM — why the lowest-scoring units improve without intervention
- Statistical Process Control for WFM — separating common-cause from special-cause variation once a unit is measurable
- Goodhart's Law and Metric Gaming — what a target does to a measure once it is enforced
References
- ↑ Deming, W. E. (1986). Out of the Crisis. Cambridge, MA: MIT Center for Advanced Engineering Study. Chapters 3 and 11. ISBN 978-0-911379-01-9.
- ↑ Agresti, A., & Coull, B. A. (1998). "Approximate Is Better than 'Exact' for Interval Estimation of Binomial Proportions". The American Statistician 52 (2), 119–126. doi:10.1080/00031305.1998.10480550.
- ↑ Tversky, A., & Kahneman, D. (1971). "Belief in the Law of Small Numbers". Psychological Bulletin 76 (2), 105–110. doi:10.1037/h0031322.
- ↑ Kahneman, D. (2011). Thinking, Fast and Slow. New York: Farrar, Straus and Giroux. Chapter 10, "The Law of Small Numbers". ISBN 978-0-374-27563-1.
- ↑ Hubbard, D. W. (2014). How to Measure Anything: Finding the Value of "Intangibles" in Business (3rd ed.). Hoboken, NJ: Wiley. Chapter 7.
