Post-Mortem and RCA for Workforce Operations
Part of the Planning Week chain · previous: Incident Management for Contact Centers · next: Call-Sharing Models and Cost Allocation
Post-Mortem and RCA for Workforce Operations is the learning half of a workforce function's incident section: the one condition under which a post-mortem is required, the record the review produces, root-cause analysis with the fishbone as its instrument, the problem record it opens and the specification revision it feeds. It matters because an incident process optimized for restoration does not find causes, and a function that closes tickets without a learning loop meets the same cause again with the same surprise. This page defers to Incident Management for Contact Centers for steps 10 to 12 of the process and the eight lifecycle facts, and to the fishbone page for the four arms; it adds the condition stated correctly, the records, and the loop into the Process Standardization Lifecycle. It produces the post-mortem record and the problem record for the incident section of a standard (the §3.x pattern of Anatomy of a ROC Standard) and is worked in a session with CP-OPS-003 ROC Standard Authoring once that pack is live.
When a post-mortem is required
The standard states one required condition, as a question with two parts, both of which must hold:
- Was the incident owned by the function, and was it rated severity 1?
Ownership decides who runs the review: an incident owned by another organization is reviewed there, with the function consulted, and the function's process routes it past its own post-mortem step. Severity decides whether a formal review is required at all. The two parts are written as one question on purpose. A gate written with one part in the step and the other in a parenthetical routes incidents into a step they do not qualify for; a severity 4 incident the function owned would reach a review meant for total loss of service, and the review would be skipped by hand and then by habit. The incident page carries the gate in this form at step 10, stating the question with both parts and separating what each decides; its routes are executable as written — yes to the post-mortem, else to monitoring root-cause actions, with an incident owned elsewhere reviewed by its owner.
A discretionary post-mortem may be called below the required condition by two routes the standard names: on recurrence, where the incident log shows the same arm and the same probable cause more than once inside a set-locally window, and on request, where the node owner or the book's planner asks because the miss matched no known event. Event Management carries the same two routes: a post-resolution analysis always at its first tier, and a report with an action plan on a recurring second tier. A discretionary review uses the same record and the same moderator rule; what differs is that nobody is obliged to hold it, so the request is recorded with its reason.
The post-mortem record
The review has three rules and one agenda. Incident Management for Contact Centers already fixes two of them and the agenda: an independent moderator drawn from outside the team that worked the incident, a balanced walk in which each item is examined for what went well and what did not, and the eight lifecycle facts recorded at closure reused deliberately as the review agenda — which is why the closure record and the review agenda are one list and not two. The standard adds the third rule and one reason. The rule: the goal is stated as a question at the top of the record, not as a verdict, so that the review can find that the process worked and the plan did not. The reason the moderator sits outside: the people who restored service are the least able to see what the process did to them and the most likely to be blamed by a review they run.
| Field | Content |
|---|---|
| ID · incident · date · moderator (seat) · attendees (seats) | PM- number; the IN- number it reviews; the review date; the outside moderator; who attended, as seats |
| The question | one sentence: what the review is trying to learn |
| The walk | the eight facts, each with what went well and what did not |
| Findings | what the review established, inferred, asserted or left open, one grade per finding |
| Actions | each with an owner as a seat and a date; recurring incidents entered on the collision calendar as events |
| Problem opened | the PB- number, or "none" with the reason |
| Specification revision requested | the L2 row or job aid to revise, routed to the process owner; or "none" |
Two literatures converge on the same rules: the blameless postmortem of reliability engineering, whose record names actions and conditions and never a person as the cause;[1] and the safety literature's finding that "human error" is where an investigation starts, not where it ends, because the error is a symptom of the conditions the process created.[2]
Root-cause analysis with the fishbone as the instrument
Root-cause analysis follows the review when a finding is left open or a problem is opened, and it runs in five phases: initiation (a request and an analyst named as a seat), information gathering, causal analysis, action planning and the repository. It is performed by the primary fix agent with the function monitoring completion, because the party that owns the cause owns the cure; the incident page carries that rule at step 12.
The instrument of the causal phase is the fishbone, and the standard is exact about what the instrument can and cannot say. The four arms name where a miss can come from; placing the miss on an arm is an association, made under time pressure at triage and re-made at leisure in the review. Cause is decided by people in the review, from the arm and the evidence, and the standard records both the triage assignment and the review's assignment so that the triage's own accuracy can be measured over time. The real-time agent team follows the same rule: its detector may write that service fell and staffing was short and may not write that the second caused the first. Correlation and Causation in WFM carries why the distinction holds; this page only applies it.
Two further tests belong to the causal phase. A miss that sits inside the process's control limits is common-cause variation and gets no root-cause analysis, however painful the interval; treating it as special cause produces actions that tamper with a stable process and make it worse.[3] Statistical Process Control for WFM carries the charts; the standard requires the control-chart state on the incident record, which the agent team's monitor already produces. And an analysis that stops at the first plausible cause is not finished: the causal phase asks why until the answer is a condition the function can change, which is the discipline the quality literature teaches as the difference between a proximate cause and a root cause.[4]
The problem record and the specification revision
A review that finds a recurring or underlying cause opens a problem record in the root-cause repository, outside the ticketing system, because a problem has no severity and outlives the tickets it produces (Event, Incident and Problem in Contact Centers fixes the object). The record carries: ID, the cause as a condition, the incidents it has produced, the arm they share, the owner of the cure as a seat, the action plan with dates, and its state (open, cure in progress, verified, closed). A problem is closed when the cure is verified by the absence of the incident class over a set-locally window, not when the action is done.
The problem record's most common consequence is a specification revision, and this is where the incident section meets the lifecycle. A recurring cause is usually a gap in a written process: an L2 row that did not cover the case, a job aid that named the wrong threshold, an event class the collision calendar had no rule for. The review routes the revision to the process owner as an odd case, and the Process Standardization Lifecycle's outcome review (step 3.1) absorbs it: the L2 row and the specification that reads from it are revised together, and the package's revision history records the incident that caused it. That is the loop by which incidents improve the standard rather than only the plan, and it is the loop a function without a lifecycle cannot close.
The section's flow
Filled as the §3.x method block of a standard, the post-mortem and root-cause flow is eight steps, decisions as questions, one page, every branch routed. It begins where the incident process's step 9 ends. Eight steps sits below the ten-to-fifteen range Process Decomposition (L0–L3) calls healthy, because this flow is the tail of another process rather than a whole one; it is drawn as a section method block, not as a standalone L1.
| # | Step | Type | Routes to |
|---|---|---|---|
| 1 | Read the closed incident record: nine fields, severity, owner, arm at triage, control-chart state | activity | 2 |
| 2 | Required: owned by the function and severity 1? | decision | yes: 4 · no: 3 |
| 3 | Discretionary: recurrence inside the window, or a request with a reason? | decision | yes: 4 · no: 8 |
| 4 | Appoint the moderator from outside the team; schedule; state the question | activity | 5 |
| 5 | Walk the eight facts, balanced; record findings with grades; enter recurring incidents on the collision calendar as events | activity | 6 |
| 6 | Cause established, or a problem to open? | decision | established, no problem: 8 · problem or open finding: 7 |
| 7 | Open the problem record; initiate root-cause analysis with the fishbone; route any specification revision to the process owner as an odd case | activity | 8 |
| 8 | Monitor any open root-cause actions from the incident's own closure and, where a problem was opened, the problem to verification; close | activity | End |
Worked example
Incident IN-001 of Wed 8 Apr 2026 closed at severity 2, owned by the function (Contact-Center Incident Severity Matrix carries the record). At step 2 the required condition did not hold, severity being 2; at step 3 the book's planner requested a review that afternoon with the reason that the miss matched no event on the collision calendar. PM-001 was moderated by the capacity-planning seat, outside the real-time team that worked it and outside scheduling, whose booking the review examined, with the question "why did a staffing gap of 11 percent [M] appear on a day with no event?" The walk found the process had worked: detected five minutes after start, engaged one minute later, repaired in twenty-nine [C]. It found the plan had not: a training pull had taken 14 agents off the floor on Wed 8 and Thu 9 Apr and was in nobody's calendar. The triage arm, arm 1 by the staffing gap, was confirmed at review with the cause named as an unplanned off-phone activity; arm 3 stayed ruled out. Findings: the pull, Established from the training roster; the absence of any calendar rule for pulls, Established; the size of the effect on the demand trend had the event not been found, Open.
Step 6 opened problem PB-001, "off-phone activities scheduled outside the collision calendar," owner the scheduling seat, with two actions: the pull entered as event EV-001 (confirmed by the planner, so the forecast loop could treat the two days as a supply-side break), and a specification revision routed to the scheduling section's process owner requiring any pull above a set-locally threshold to be entered as an event before it is booked. The revision entered the lifecycle's outcome review as an odd case; PB-001's state on Fri 8 May 2026 was "cure in progress," to be verified by the absence of an unscheduled pull incident over the following quarter [E].
The artifact this page produces
The post-mortem record, one per review, and the problem record, one per problem, for the incident section of a function's standard (the §3.x pattern of Anatomy of a ROC Standard), recorded against L0 card P-008 (T4). One filled example row of each:
| Record | ID | Reviews or arises from | Date | Owner or moderator (seat) | Finding or cause | Actions | State |
|---|---|---|---|---|---|---|---|
| Post-mortem | PM-001 | IN-001 (discretionary, on request) | Wed 8 Apr 2026 | the capacity-planning seat (moderator) | an unplanned training pull, in no calendar; process worked, plan did not | EV-001 entered; PB-001 opened; specification revision routed | closed |
| Problem | PB-001 | IN-001; arm 1 | Wed 8 Apr 2026 | the scheduling seat | off-phone activities scheduled outside the collision calendar | the pull-as-event rule at L2, owner the scheduling process owner, Fri 12 Jun 2026 | cure in progress |
Produced in a working session with CP-OPS-003 ROC Standard Authoring (the pack link is added when the pack is live); the filled set is part of blueprint v0.1.
What would change this
The page's central claim is that the required condition should have exactly two parts, ownership and severity 1, with everything below it discretionary. The observation that would overturn it is a function whose incident log, over a full cycle, showed that the causes removed by discretionary reviews of severity 2 and 3 incidents prevented more service loss than those removed by required reviews of severity 1, and that the discretionary route was called too rarely to capture them; that would argue for a required review on recurrence at severity 2, and the question would gain a third part.
How this connects
- Previous in the chain: Incident Management for Contact Centers — steps 10 to 12 of the process, where this page's flow begins
- Next in the chain: Call-Sharing Models and Cost Allocation — the first of the standard's routing sections
- Defers to: Incident Management for Contact Centers (the eight lifecycle facts and the closure loop) · Real-Time Cause and Effect Fishbone (the four arms) · Statistical Process Control for WFM (the control-chart state that separates common from special cause) · Correlation and Causation in WFM (why an arm is association and the review decides cause) · Event, Incident and Problem in Contact Centers (the problem as an object) · Process Standardization Lifecycle (the outcome review that absorbs the specification revision)
Maturity Model Position
A required post-mortem on a stated condition with an outside moderator is Level 2 discipline, the point at which incidents stop being handled by whoever notices them. Recording the triage arm and the review arm separately, requiring the control-chart state on the record, and closing problems on verified absence rather than completed action are Level 3 practices, and they are what let the real-time agent team's incident records start a review with the evidence attached rather than reconstructed. Four scales on this wiki use the word level; the launch page states which is which.
See Also
- Planning Week for a Workforce Function
- Incident Management for Contact Centers
- Real-Time Cause and Effect Fishbone
- Event, Incident and Problem in Contact Centers
- Contact-Center Incident Severity Matrix
- Process Standardization Lifecycle
- Statistical Process Control for WFM
- Correlation and Causation in WFM
- Real-Time Agents
References
- ↑ Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.) (2016). Site Reliability Engineering: How Google Runs Production Systems. Sebastopol, CA: O'Reilly. Chapter 15, Postmortem culture: learning from failure.
- ↑ Dekker, S. (2014). The Field Guide to Understanding 'Human Error' (3rd ed.). Farnham: Ashgate.
- ↑ Wheeler, D. J. (2000). Understanding Variation: The Key to Managing Chaos (2nd ed.). Knoxville, TN: SPC Press.
- ↑ Rooney, J. J., & Vanden Heuvel, L. N. (2004). "Root Cause Analysis for Beginners." Quality Progress, 37(7), 45–53.
