Baselines Before Targets

Baselines before targets is the discipline of measuring the distribution of a metric on a stable instrument before setting a target on it, and of withdrawing targets that were set the other way round. A target set before a baseline exists is a claim about a distribution nobody has measured. It is not wrong in the way a miscalculation is wrong; it is unfounded, and an unfounded target does predictable damage — it makes every population look like a failure, it stops carrying information about which populations differ, and it invites the gaming that any unattainable target invites. This page sets out why targets so often precede baselines, what an unfounded target does, the sequence that corrects it, and how to withdraw a target without the withdrawal reading as retreat. The broader goal-setting discipline, including the practice of setting targets as a percentile of the measured distribution, is at Performance Management; the threshold arithmetic that makes stable operations look like failures is at Sample Size and Detectable Difference in Quality Measurement; the corruption of a measure once it becomes a target is at Goodhart's Law and Metric Gaming.
Why targets come first
Targets precede baselines for organizational reasons, none of them careless.
- Budgets need a number before the instrument exists. Planning cycles run on a calendar; instruments are procured, deployed and calibrated on a longer one. The target is written into the plan as a placeholder and is never revisited as one.
- Benchmarks are borrowed. A figure that is customary in another industry, another channel or another instrument is adopted as the target because it is available, without checking that it was produced on a comparable population by a comparable instrument.
- Aspiration is confused with expectation. A leadership team states where it wants the operation to be; the statement is recorded as the target; the target is then used as the expectation against which every period is judged.
- New instruments arrive with new metrics. A sentiment or quality platform introduces a metric the estate has never had, and a target is set on it in the first quarter, before the instrument has stabilized (see Instrument Effects During Measurement Rollout).[1]
The result is a target that sits above every measured population — sometimes far above the best of them — and that was never derived from any of them.
What an unfounded target does
Three things, each damaging on its own.
It converts the whole estate into a failure. If the best-performing population measured sits well below the target, every population is below target, and the review conversation is about a shortfall that describes no operational difference between them. The threshold arithmetic shows the milder version of the same problem: a target set at current average performance is missed by half of everything in any period, permanently and by construction.
It stops carrying information. A target's purpose is to discriminate — to separate populations that need attention from those that do not. A target no population reaches discriminates nothing. The metric continues to be reported, but the target attached to it is decoration, and reviewers learn to look past it.
It is ignored or gamed. Goal-setting research finds that difficult goals raise performance where goal commitment is high, and that commitment depends in part on the goal being seen as attainable; where a goal is judged unreachable, commitment falls.[2] An unreachable target is either ignored or gamed, and the estate ends up with the failure mode Goodhart's Law and Metric Gaming describes — a measure that has stopped indicating what it was meant to indicate — without ever having had a period in which the target was a fair test.
The sequence
The corrective is an ordering, and each step assumes the one before it.
- Stabilize the instrument. No baseline is taken while the instrument is being rolled out, recalibrated or redefined. The stability condition is stated in advance — a fixed number of periods on a fixed model version and coverage.
- Measure the distribution, per population. Not the average alone: the spread, the range across sites or arrangements, and the sample sizes that determine how much of the spread is noise. Populations that cannot yet produce the number are recorded as not-assessable rather than imputed (see Interpreting WFM Maturity Assessments).
- Set the target relative to the distribution. Performance Management gives the standard practice — a percentile of the measured distribution rather than a round number. Where the population is small enough that ordinary variation moves the score, the target is expressed as a band rather than a point, so that variation does not read as breach.
- Stamp it with the instrument and the as-of date. A target belongs to the instrument that produced its baseline. When the instrument changes, the target is re-derived, not carried forward.
- Recalibrate on a cadence. The recalibration discipline is at Performance Management; the point specific to this sequence is that it is scheduled in advance rather than triggered by a bad period.
Where the estate runs several delivery arrangements or locations, the sequence has one more rule: the instrument is standardized across all of them and the threshold varies only by what the customer bought, never by where the work is done. The arithmetic of why a location-varying target damages the estate's own headline number is at Mix Effects in Blended Quality Targets.
Withdrawing a target
A target already published without a baseline should be withdrawn, and the withdrawal is properly framed as a measurement finding.
- Publish the baseline as the finding. The news is not that the target was wrong; it is that the estate now knows its distribution for the first time, and the distribution is what the target should have been derived from.
- Replace the target with a band and a trajectory. A band around the measured distribution, with a stated direction and cadence of improvement, does the motivational work the target was meant to do and can actually be met.
- Name the original as a placeholder. A figure set before the instrument existed was, in fact, a placeholder pending measurement.
- Sequence it ahead of external challenge. A withdrawal initiated by the operation is recorded as a measurement finding; the same withdrawal made after an owner or client challenges the figure is recorded as a concession.[1]
Failure modes
- Adjusting the target to the best population. The target is lowered to whatever the top performer achieved, which is a baseline of one and inherits all of that population's noise.
- Keeping the target and re-explaining the miss. Each period's shortfall is attributed to a new cause, and the review cadence becomes a variance-explanation ritual.
- Setting targets during rollout. The first quarter's scores on a new instrument are treated as a baseline when they are the instrument settling.
- Borrowing without checking the instrument. A benchmark from another industry or channel is adopted without confirming it was measured on a comparable population.
Maturity Model Position
Setting baselines before targets is a Level 2 discipline on the WFM Labs Maturity Model™ in the Goals pillar — the level at which service targets are published and tracked — and a target set without a baseline is a Level 2 artifact wearing Level 1 foundations. It becomes decisive at Level 4, where targets are expressed as confidence bands and hit-rates: a band cannot be drawn around a distribution that has never been measured.
See Also
- Performance Management — the goal-setting framework and the percentile practice
- Sample Size and Detectable Difference in Quality Measurement — thresholds, and why stable operations get flagged
- Goodhart's Law and Metric Gaming — what happens to a measure once it is a target
- Instrument Effects During Measurement Rollout — why rollout-period data is not a baseline
- Mix Effects in Blended Quality Targets — why the threshold never varies by location
- Regression to the Mean in WFM — why an extreme period is not a baseline either
References
- ↑ 1.0 1.1 Practitioner observation from target-setting and quality-instrument work in multi-site service estates; a consistent pattern rather than a measured result.
- ↑ Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation: A 35-year odyssey. American Psychologist, 57(9), 705–717.
