Event, Incident and Problem in Contact Centers

From WFM Labs

Part of the Planning Week chain · previous: Forecast Collision Calendar · next: Contact-Center Incident Severity Matrix Event, Incident and Problem in Contact Centers is the definitions block of the incident-management section of a workforce function's standard: three terms defined once, each with the ledger it lands in and the process it triggers, so that every other section, every notification and every post-mortem uses them the same way. It matters because the three words are used loosely on many floors [A] and in two different senses on this wiki, and a function that has not fixed them cannot tell a planning failure from an operational one or count either. The triad is inherited from IT service management, where it was separated for the same reason.[1] This page is narrow by design: Event Management carries the event taxonomy and a severity matrix, and Incident Management for Contact Centers carries the twelve-step process; this page defers to both and adds only the definitions, the ledgers, the triggers and the reconciliation. It produces the definitions block of a standard's incident section (the §3.x pattern of Anatomy of a ROC Standard) and is worked in a session with CP-OPS-003 ROC Standard Authoring once that pack is live.

The triad, defined once

Term Defined as Known when Rated by Lands in Triggers
Event A known activity or condition, with a date and an effect window, that is expected to move demand or supply: a go-live, a release, a campaign, a holiday at any node, a migration phase, a training pull Before it happens, or after the fact when a post-mortem finds it should have been known Its effect window and the layer it moves (volume, handle time, supply) The event ledger, whose human face is the collision calendar Forecasting (the three-step build's event adjustment) and scheduling; never the incident process by itself
Incident An unplanned occurrence, or a planned event whose effect exceeded its window, that is degrading service against target now, where the goal is restoration When a written entry criterion trips: a sustained service-level breach, a driver outside its band, an external party reporting impact Severity, from the severity matrix, set once and revised only with the reason recorded The incident record (the ticket), nine fields The incident process, steps 1 to 9; a post-mortem where the condition on Post-Mortem and RCA for Workforce Operations is met
Problem The recurring or underlying cause behind one or more incidents, whose removal prevents the next one When a post-mortem or root-cause analysis names it, or when the incident log shows the same cause on the same arm repeatedly Its recurrence and the class of incident it produces; never a severity of its own The problem record, in the root-cause repository outside the ticketing system Root-cause analysis, the specification revision it feeds, and, where the cause is a missing event, an entry on the collision calendar

The incident row narrows Incident Management for Contact Centers's "planned event or unplanned occurrence" to the case where the planned event exceeded its effect window, so that a degradation inside a known window stays with the forecast review; the narrowing is the standard's, and it is stated here. Three consequences follow from the table and are stated here so that no other page has to argue them. An event that was known and planned for is not an incident, even when service degrades inside its window; the degradation is the forecast's assumption being tested, and it belongs to the forecast review. An incident is not a problem, and closing the incident does not close the problem; an operation that counts closed tickets as resolved causes will meet the same cause again. And a problem is not rated by severity; it is rated by what it costs across the incidents it produces, which is why the problem record lives outside the ticketing system, where a severity field would be wrong.

Which ledger, which process

The standard fixes the home of each object because the three homes are read by different processes on different clocks. The event ledger is read daily by the forecast loop and weekly by scheduling; an event entered late is a forecast miss with a known cause. The incident record is read minute to minute during the incident and once at closure, when its nine fields become the instrument panel of the process itself, as Incident Management for Contact Centers describes for the eight lifecycle facts. The problem record is read at the standards committee's sitting and at the outcome review of the Process Standardization Lifecycle, because a recurring cause is usually a specification gap: a step the L2 did not cover, or an event class the calendar had no rule for.

The homes also fix who writes. The real-time desk writes incidents and proposes events; the planner confirms events; the post-mortem opens problems; the process owner closes them by revising the specification. The real-time agent team follows the same rule, issuing incidents with the evidence attached and proposing events it may not confirm.

Reconciling this wiki's two usages

Two pages on this wiki use "event" in different senses, and a function adopting the standard must know which it is using. Event Management defines an event as an unplanned condition that crosses a threshold requiring coordinated response, and rates it on a four-tier severity matrix; that is the service-management usage in which "event" is the generic name for anything the operation must respond to. Incident Management for Contact Centers and this page use the triad above, in which an event is known and dated and only an incident is severity-rated. The two are not in conflict once mapped: an event in the first sense, once it has crossed its threshold, is an incident in the second, and the first page's severity matrix is an incident severity matrix. The standard uses the triad, because a workforce function needs the planned-and-known object to have its own name; it is the object the collision calendar holds, and the object whose absence a post-mortem most often finds.

The mapping, stated once
Event Management says The standard says Why the standard's term
event (an unplanned condition above threshold) incident the planned, known object needs its own name
severity matrix (four tiers, illustrative values) incident severity matrix (row count set locally; six parameters fixed) see Contact-Center Incident Severity Matrix
post-resolution analysis (Sev 1; recurring Sev 2) post-mortem (required on one condition; discretionary otherwise) see Post-Mortem and RCA for Workforce Operations
the event taxonomy by type (capacity, demand, system, process, third-party) the incident's fault arm and lever the arm names where the miss came from; the type names what kind of thing broke; both are kept

The service-management standard on which the triad rests separates the three for a reason a workforce function shares: an incident process optimized for restoration will not find causes, and a problem process optimized for causes will not restore service in the interval.[2] The reliability-engineering literature draws the same line between the incident response and the postmortem, and treats the second as the place where the organization learns.[3]

The three health inputs

Every incident is a deviation in one or more of three inputs to contact-center health: demand (volume), supply (agent resources) and handle time. The incident page carries them as the three corrective levers; this page states them once as the frame a definitions block cites, because they are also how an event is classified (which layer it moves) and how a problem is classed (which input it keeps disturbing). The four arms of the Real-Time Cause and Effect Fishbone are the four ways a miss can arise across those three inputs, and the arm is recorded on the incident record as association, never as cause; the post-mortem decides cause.

Worked example

The chain's worked example carries one incident and its consequences through all three objects. On Wed 8 Apr 2026, the 38th day of the migration, the voice service level crossed below target at 09:15 with staffing 11 percent under schedule from 09:00 [M]; the entry criterion tripped, and the real-time team's issuer wrote incident IN-001 at 09:21, severity 2 from the function's matrix, arm 1 indicated and arm 3 ruled out, matched-events field empty, as Real-Time Agents records. Service recovered by 09:50 after a break shift for nine agents (proposal R-044). The incident closed at severity 2, so the required post-mortem condition was not met; the planner requested a discretionary review the same afternoon because the miss matched no event, and the review (PM-001) found a training pull that had taken 14 agents off the floor for two days and was in nobody's calendar.

That finding produced the other two objects. The training pull became event EV-001, "training pull, Wed 8 to Thu 9 Apr 2026, voice, in-house cohort," proposed by the issuer, confirmed by the planner and entered on the collision calendar after the fact, so that the daily forecast loop could keep the two days out of the demand trend as a supply-side break. And the review opened problem PB-001, "off-phone activities scheduled outside the collision calendar," in the root-cause repository, with the specification revision it feeds: an L2 row in the scheduling section requiring any pull above a set-locally threshold to be entered as an event before it is booked. PB-001 is not severity-rated; it is rated by the incidents it has produced (one, IN-001) and the ones it would produce unaddressed [E].

The artifact this page produces

The definitions block of the incident section of a function's standard (the §3.x pattern of Anatomy of a ROC Standard), recorded against L0 card P-008 (T4), one row per term. One filled example row:

Term Definition (as adopted) Ledger Written by Confirmed by Triggers Example (worked)
Event a known activity or condition with a date and an effect window, expected to move demand or supply the event ledger (the collision calendar) the real-time desk or the planner; the agent team proposes the book's planner the forecast's event adjustment; scheduling EV-001, training pull, Wed 8 to Thu 9 Apr 2026

Produced in a working session with CP-OPS-003 ROC Standard Authoring (the pack link is added when the pack is live); the filled set is part of blueprint v0.1.

What would change this

The page's central claim is that the planned, known object needs its own name and its own ledger, separate from the incident's. The observation that would overturn it is a function whose collision calendar and incident log, kept as one list with a planned/unplanned flag, produced forecast event adjustments and post-mortem findings at the same rate as a function keeping two ledgers, over a full cycle; the triad would then be a service-management convention rather than a workforce necessity, and this page would fold into Event Management as a mapping paragraph.

How this connects

Maturity Model Position

Fixing the triad is Level 2 work: it is the vocabulary a documented incident process needs before its steps can be written. Keeping the three ledgers on their three clocks and reading the problem record at the standards committee is Level 3 practice, and it is the precondition for the real-time agent team, whose issuer may write an incident and propose an event but may confirm neither. Four scales on this wiki use the word level; the launch page states which is which.

See Also

References

  1. AXELOS (2019). ITIL Foundation: ITIL 4 Edition. London: TSO. The event, incident and problem practices.
  2. International Organization for Standardization (2018). ISO/IEC 20000-1:2018 — Information technology — Service management — Part 1: Service management system requirements, clauses 8.6.1 and 8.6.3. Geneva: ISO.
  3. Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.) (2016). Site Reliability Engineering: How Google Runs Production Systems. Sebastopol, CA: O'Reilly. Chapters 14 and 15.