Data Synthesis Before Decision

From WFM Labs
Many sources and several engines feeding one question. The register of graded claims is the step that turns data into an answer that carries its evidence.

Data synthesis before decision is the discipline of reducing the many reports an operation produces to one answer to one question, with the strength of that answer stated, before a decision is taken on it. It sits between the arrival of reports and the taking of a decision, a step most operations do not name and therefore do not staff. The subject is distinct from three neighbors. Signal and Noise in WFM concerns variation within a single metric on a single instrument, and Reporting and Analytics Framework concerns the architecture that produces and delivers reports. WFM Data Governance and Quality concerns the standing machinery — quality dimensions, lineage, ownership, master data — that keeps entities and definitions aligned across systems. This page concerns what a decision maker does when that machinery has not yet caught up and an answer is needed anyway: several systems, each correct by its own definitions, have been asked one question and have given several answers.

The condition

A consolidated or long-lived service operation typically runs several engines that each produce a number with the same name. A workforce platform, a telephony platform, a finance system and a human-resources system will each report a headcount, and the four figures will differ. Two or three quality tools with different scoring rules will each produce a number called "quality." An occupancy figure will depend on how agents signal state changes on a desktop, not only on the work performed. None of these engines is broken. Each answers a slightly different question with a slightly different population, and the difference is invisible on the page where the number appears.

A second form of the condition is structural rather than definitional. Status reporting can show a healthy summary sitting on top of overdue detail, because the template that carries the summary is not connected to the register that carries the detail. The summary is green; the line items beneath it are red. The cause is usually a wiring defect rather than misconduct: the summary artifact and the detail register were never connected, and each is maintained correctly by a different owner.

The combination produces an organization that holds a great deal of data and few answers. Any question of consequence — where work should sit, what automation can absorb, whether a unit is underperforming — can be supported by some report and contradicted by another. Discussion then turns on whose report is admitted rather than on what the operation is doing. The condition is common in operations formed by merger, where several playbooks were each correct for the business that wrote them and were never reconciled.

Three gaps between reports and an answer

The definition gap is the absence of an operational definition shared across engines. Deming argued that a measurement has no meaning without an operational definition — an agreed procedure that tells two people, working separately, how to produce the same figure.[1] Where definitions differ, comparisons mislead, commitments drift, and a well-run unit can be judged unfairly by a ruler that was never calibrated to it. Agreeing what words mean is usually the cheapest improvement available and the one most often deferred, because it produces no report of its own.

The register gap is the absence of any place where the answer to a question is held together with the evidence for it and the strength of that evidence. Without a register, each meeting reassembles the argument from whatever reports are to hand, and the assembly is never the same twice. Tetlock and Gardner argue that forecasting skill improves only where claims are stated precisely enough to be scored and the score is fed back. A vague, unscored claim cannot be revised, because there is nothing to revise it against.[2] An operation without a register cannot revise, because it cannot say what it previously believed.

The inference gap is the failure to separate what was measured from what was concluded. A finding — three systems report three headcounts — and an inference — the operation is over-hired — are different kinds of statement with different evidence requirements. When the two are written in the same register, the inference borrows the finding's authority. Kahneman's account of System 1 explains why this passes unnoticed: a coherent story feels true, and coherence is assessed on the information present, not on the information missing.[3]

Synthesis as a discipline

Synthesis is a discipline rather than a task because it has rules that hold regardless of the question. Four are load-bearing. The first is a precondition; the other three answer the three gaps in turn.

One question

Synthesis begins by stating the decision the answer will serve, not by collecting reports. Hubbard's method for measurement starts the same way: define the decision, identify what is uncertain about it, and only then ask what observation would reduce that uncertainty.[4] A question stated first excludes most of the reports available, which is the point. A report that cannot change the decision is not evidence for it, however accurate it is.

One definition, reconciled before aggregation

Numbers from different engines are mapped to one stated definition, with the differences explained, before they are summed, averaged or compared. Aggregating first and reconciling later produces totals that no engine can reproduce, and a total that cannot be reproduced cannot be defended when it is disputed. The reconciliation is done for the question at hand; the standing machinery that keeps definitions aligned across systems thereafter belongs to WFM Data Governance and Quality, which receives the mappings. The rule is also where the green-over-red condition is caught: a summary not derived from the detail beneath it is an asserted claim, and the fix is to compute it from the register.

One register of graded claims

Every claim that bears on the decision is held in one register, with a grade that states how strongly the evidence supports it. A four-grade scale is sufficient for most operational work.

Grade Meaning What it takes to hold the grade
Established Measured on a known instrument, or documented in an artifact that can be produced on request The artifact, its date, and the instrument or definition it used
Inferred Follows from established claims by stated reasoning, but has not itself been measured The chain of reasoning, written down, with the established claims it rests on
Asserted Stated by a person or a document without a producible artifact The name of the source and the date of the statement
Open Not yet known; a question rather than a claim The observation that would settle it, and who owns obtaining it

The grade is a property of the evidence, not of the claim's importance or of who made it. A senior leader's assertion is an asserted claim until an artifact is produced; a junior analyst's extract from a system of record is an established claim. Interpreting WFM Maturity Assessments applies the same boundary to assessment scores, where a score without a named artifact is recorded as claimed rather than verified and the two columns are never merged. The register is the general form of that rule, and it sets a norm: a line is overturned by producing the artifact that contradicts it, not by asserting more firmly. Two scales in Sourcing Strategy Under Imperfect Data grade different objects. Its two-tier rule, which keeps hard data and calibrated estimates apart, sorts evidence by whether it was measured or estimated, and its comparability study grades the comparison paths a placement decision runs on. The four grades here sort claims by whether a producible artifact stands behind them, for any decision rather than for placement.

Findings separated from inference

Every document that leaves the synthesis step carries its findings and its inferences in separate sections, or in separate columns. The separation is mechanical and is applied before the document is reviewed, not after. Its purpose is to prevent the inference gap from reopening every time a document is rewritten. A reader who disputes an inference can see which findings it rests on, and the reverse.

Why generative AI raises the stakes

A language model placed over the same stack of reports closes none of the three gaps. It reads every report, reconciles none of them, and returns a summary of unreconciled numbers in a form that no longer shows the mismatch. The cues that used to signal a problem — a footnote, a total that did not match, a person's visible hesitation — are absent from the output. The consequence and the countermeasures are the subject of AI Reads Everything and Thinks Nothing. For this page the point is that a tool which presents unsynthesized data as finished work removes the last natural prompt to perform the step.

Failure modes

Failure mode What it looks like Countermeasure
Report-first Reports are gathered before the question is stated; the decision is shaped by whatever was available State the decision first; admit only evidence that can change it
Authority grading Claims are graded by the seniority of their source rather than by their evidence Grade on artifacts; an assertion is asserted until an artifact is produced
Silent aggregation Figures from several definitions are summed into a total nobody can reproduce Reconcile to one definition before any aggregation; keep the mapping
Inference creep Conclusions are rewritten into the findings section over successive drafts Separate sections or columns, applied mechanically before each review
Decorative summary The summary is maintained by hand and drifts from the detail beneath it Derive the summary from the register; treat an underived summary as an asserted claim

Maturity Model Position

At Level 1 the definition gap exists without the engine multiplicity: one spreadsheet, unstated conventions, and no register. The full condition arrives at Level 2, where a platform, a telephony system and a reporting layer each carry their own figures. Level 2 names the definition gap and attacks it with explicit definitions and one versioned source of truth across those systems; the register and inference gaps remain untouched, and synthesis is a manual discipline carried by a few people. At Level 3 the rule registry and rollback discipline are the first governed register the operation keeps, but it governs automated actions rather than claims. At Level 4, where the plan is a distribution rather than a point, the discipline's contribution is that a claim's grade travels with the number into the model rather than being lost at the input boundary. At Level 5 the conflicts between systems of record are reconciled by stated policy and precedence rather than by a person at decision time, and synthesis becomes part of the operating loop rather than a review step.

See Also

References

  1. Deming, W. E. (1986). Out of the Crisis. Cambridge, MA: MIT Center for Advanced Engineering Study. Chapter 9, "Operational Definitions, Conformance, Performance". ISBN 978-0-911379-01-9.
  2. Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. New York: Crown. Chapters 3 and 7. ISBN 978-0-8041-3669-3.
  3. Kahneman, D. (2011). Thinking, Fast and Slow. New York: Farrar, Straus and Giroux. Chapters 7 and 19. ISBN 978-0-374-27563-1.
  4. Hubbard, D. W. (2014). How to Measure Anything: Finding the Value of "Intangibles" in Business (3rd ed.). Hoboken, NJ: Wiley. Chapters 4 and 7. ISBN 978-1-118-53927-9.