Wiki:Packs/Agent Capability Ontology

From WFM Labs
Pack
ID CP-WFM-002
Domain WFM
Blocks 1 instruction + 4 reference
Version 1.1
Source Skill-Based Routing · Three-Pool Architecture · Next Generation Routing · Workforce Clustering and Segmentation · Business Process Outsourcing

A pack is a deployable set for a Claude project. This one supports building and maintaining an agent capability ontology — the structured description of what every unit of service supply can do, used to allocate supply to demand rather than routing demand at a fixed queue structure.

Most routing work on this wiki scores demand: Value Routing Model ranks interactions, Three-Pool Architecture designs the pools they land in. This pack addresses the other side. It assumes the demand model exists and asks the harder operational question — can the platform actually hold a description of supply rich enough, and current enough, to route against?

Where the capability layer sits

Future-state workforce ecosystem architecture from long-range planning to real-time execution
Future-state workforce ecosystem, arranged by planning horizon.

Long-range planning splits into two arms that are routinely conflated — a system of record that stores budget, forecast and actuals and iterates with Finance, and a simulation engine that calculates but stores nothing. Both feed the capacity plan, which flows through forecasting and scheduling into intraday automation. The capability layer sits at the real-time end, taking capability evidence from performance and quality management and writing agent assignment back to the contact platform. Every layer writes into the data lake; the analytics loop that should feed models back into simulation and capability scoring is the component most often missing.

The editable source is at File:WFM-Ecosystem-Future-State-Architecture.svg.

When to use it

Use it when:

  • Consolidating onto a new ACD or CCaaS platform and deciding how much routing intelligence to build inside it
  • A capability-layer or agent-orchestration product is under evaluation and the attribute model has to exist before the product can be assessed
  • Supply is heterogeneous across employment models, geographies, languages, systems and skill tiers, and no single inventory of it exists
  • Routing quality is degrading and the suspected cause is a stale supply picture rather than bad routing logic

Do not use it for demand-side scoring, interaction valuation, or queue design — those are covered by the source articles above.

How to deploy

  1. Create a project in Claude. Name it for the work, not the method
  2. Copy Block 1 into the project's custom instructions
  3. Save Blocks 2–5 under the filenames in their headings and upload as project knowledge
  4. Start a conversation

Block 1 — Project instructions

# Agent Capability Ontology

## Context

This project builds and maintains an agent capability ontology: the structured
description of what every unit of service supply — human or automated — can do,
used to put supply in the right place rather than pushing demand at a fixed queue
structure. Work spans enumerating the attribute space, collecting it from
operational teams, testing whether the routing platform can physically hold it,
and deciding what belongs inside the ACD versus in a capability layer above it.

The ontology is deliberately platform-independent. It must remain useful whether
the organisation buys a capability layer, waits for one, or builds the allocation
logic itself.

## Routing

| Task | Open |
|---|---|
| Enumerating or reviewing attributes; deciding what qualifies as a dimension | `capability-ontology-dimensions.md` |
| Testing whether a platform can hold the ontology; skill and queue limits; the in-ACD vs above-ACD decision | `attribute-cardinality-test.md` |
| Defining a capability, deciding how it is evidenced, handling expiry and progression | `capability-evidence-and-decay.md` |
| Designing or running the collection instrument with operational teams | `supply-data-collection.md` |

## Disciplines

- The unit of collection is the **pool**, not the agent. A pool is a group sharing
  every attribute under consideration; if two groups differ on any one attribute
  they are two pools. Agent-level extracts are a later step, not the first ask.
- An attribute with **no system of record cannot drive automated allocation**,
  however important it is. Record it as a gap; never assume a source exists.
- Attributes **not used for routing today are still in scope**. The purpose is to
  expose the full space, not to document current configuration.
- Separate **established fact from inference** explicitly, and label which is which.
- **Blanks are findings.** An unanswerable field is a result, not a failure.
- Never assert a platform limit without naming the vendor documentation and the
  date it was read. Vendor limits change; treat every figure as needing re-verification.
- Distinguish **eligibility** (a hard constraint — the agent cannot do the work)
  from **ranking** (a soft preference — the agent is better at it). Conflating them
  is the most common modelling error and it silently shrinks the eligible pool.
- **Capability is a flow, not a state.** Every attribute carries a rate of change,
  and the rate matters more than the value for platform-fit decisions.
- Cost and contractual attributes are part of the ontology. Reallocation is rarely
  free, and an allocation engine blind to that will propose moves that cannot be made.

## Output

Deliver structured artefacts, not prose essays: attribute tables, cardinality
calculations with the arithmetic shown, and explicit gap registers. State the
assumption whenever a figure is estimated rather than measured. When asked whether
something fits in the ACD, answer with the calculation, not a judgement. Refuse to
produce a capability model that cannot name where each attribute would be sourced.

Source: Wiki:Packs/Agent Capability Ontology (CP-WFM-002) v1.1

Block 2 — capability-ontology-dimensions.md

# The Attribute Space

Figures marked [estimated] are judgment, not measurement.

An agent capability ontology enumerates every attribute that could be used to
decide whether a unit of supply may, or should, take a given piece of work. It is
larger than the set any current platform uses. That gap is the point of building it.

## The two kinds of attribute

Every attribute is one or the other, and mixing them is the most common error.

| | Eligibility | Ranking |
|---|---|---|
| Question | May this agent do this work? | How well would they do it? |
| Violation | Hard failure — compliance, access, contract | Soft cost — slower, lower quality |
| Modelled as | Boolean or set membership | Ordinal or continuous score |
| Failure mode when confused | Eligible pool silently collapses | Constraints get optimised away |

A system entitlement is eligibility. Proficiency in that system is ranking. An
agent who knows a booking system but lacks the account credential to transact in
it is *not eligible*, no matter how proficient. Treat the two as separate attributes.

## Eight dimensions

### 1. Commercial
Which book of business the agent serves, and on what terms. Segment or brand;
sub-segment; named client; service model (dedicated, designated, shared, overflow);
funding construct (cost-plus, per-transaction, hybrid, gainshare); contractual
restriction on serving other clients.

The under-modelled attribute here is the **reallocation trigger** — whether moving
the agent creates a billing event, breaches a committed minimum, or requires
approval. Allocation engines that ignore it propose legal-but-impossible moves.

### 2. Labour
Employment model, which behaves as three structurally distinct tiers rather than
three points on one cost curve:

| Tier | Unit cost | Tenure | Flex |
|---|---|---|---|
| In-country employee | Highest | Long | Low |
| Captive low-cost centre | Middle | Long | Low–moderate |
| Third-party outsourced | Lowest | Short by design | High — this is the product |

The middle tier is the one most often missing from the conversation. A captive
centre buys much of the cost advantage while keeping tenure and continuity, and it
does *not* buy the flex. Also carries: employing entity, statutory and collective
constraints, contract type, fully-loaded cost, and **reassignment lead time** — the
practical latency of moving the agent, which is frequently the binding limit.

### 3. Language
Primary and additional languages; how proficiency was established (tested,
self-declared, inferred from site); and the written-versus-spoken split, which
diverges and determines channel fit. Language proficiency that is self-declared
cannot safely gate routing.

### 4. Channel
Channels handled; per-channel proficiency; concurrency limits; blending
eligibility. Channel capability is not nested — competence on voice does not imply
competence on chat, and treating it as a hierarchy over-states supply.

### 5. Systems
Proficiency in the core transactional system, and separately the **entitlement** to
transact in it for a given account. In travel this is the GDS distinction — knowing
Amadeus, Sabre, Galileo or Apollo is one attribute; holding the office credential
for the client is another. Also desktop and application access, and the platform
the agent is actually served by, which determines what routing is possible at all.

### 6. Capability
The progression ladder; product or work-type specialisms; complexity band;
authority limits (what the agent may approve, and to what value); certifications
and their expiry; regulatory and security clearances. See
`capability-evidence-and-decay.md`.

### 7. Quality
Automated interaction scoring, manual assessment, and — critically — **scoring
coverage**. If automated scoring reaches 95% of one tier's interactions and 80% of
another's, cross-tier quality comparison is invalid and the gap will be misread as
a performance difference.

### 8. Dynamics
The rate of change of everything above. Pool composition volatility; training
pipeline state (in training, nesting, ramp, productive — each with a different
productivity discount); progression state; adherence reliability; and the
**cross-training map**, which is the single most valuable attribute for
reallocation and the one least likely to exist anywhere.

## Why the eighth dimension decides the architecture

Dimensions 1–7 describe supply. Dimension 8 describes how fast that description
decays. A platform can hold an attribute only if it can be updated as fast as the
attribute moves. Enumerate the rate of change alongside the value, or the
cardinality test in the next file cannot be run.

## Limits

The dimension set is a starting frame, not a standard. Operations with regulated
work, field service, or heavy automation will need dimensions this list omits. The
test for adding one is whether it can make an agent *ineligible* or materially
change ranking — not whether it is interesting.

Block 3 — attribute-cardinality-test.md

# The Cardinality Test

A method for deciding, arithmetically rather than by argument, which parts of a
capability ontology a routing platform can hold — and which must live above it.

Figures marked [estimated] are judgment, not measurement.

## The question

Routing platforms express agent capability through some runtime-mutable construct:
skills, attributes, proficiencies, tags. Every such construct carries limits. The
test asks whether the ontology fits inside them.

Two limits matter, and practitioners consistently worry about the wrong one.

| Limit | Form | Usually |
|---|---|---|
| Attributes per agent | How many capabilities one agent may carry | Generous |
| **Agents per attribute** | How many agents may share one runtime attribute | **Binding** |

## The general test

For each attribute value `v` in the ontology, let `n(v)` be the number of agents
holding it.

```
if max over v of n(v)  >  platform limit on agents-per-runtime-attribute
    then that attribute cannot be runtime-mutable on this platform
```

The intuition is uncomfortable and worth stating plainly: **the most widely held
capabilities are the ones most likely to fail the test.** Narrow, specialist
attributes fit easily. "Speaks English", "handles voice", "knows the core booking
system" — the attributes that describe most of the workforce — are exactly the ones
that overflow.

An attribute that fails must be expressed as a *static* profile entry instead.
Static entries are precisely the thing that does not move at the speed supply moves.

## Worked example

A documented case: Cisco Webex Contact Center specifies a limit of **150 dynamic
skills per agent** and **50 agents per dynamic skill**, with a skill's type fixed
permanently at creation. *(Cisco product documentation, read 2026-07-20 — re-verify
before relying on it; vendor limits change.)*

An operation with 3,000 agents runs the test:

```
Attributes rated critical for eligibility        22
Limit on attributes per agent                   150
-> Fits, with headroom.                         PASS

Agents holding "handles voice"                2,400  [estimated]
Limit on agents per dynamic skill                 50
-> 48x over limit.                              FAIL

Agents holding "core booking system - primary"  1,850  [estimated]
-> 37x over limit.                              FAIL

Agents holding "regulated-work clearance"          38  [estimated]
-> Within limit.                                PASS
```

Read correctly, this is not a verdict that the platform is inadequate. It is a
partition. Three attributes fail, one passes, and the failures share a property:
they are held by most of the workforce. The conclusion is not "replace the ACD" but
"these three attributes cannot be runtime-mutable here."

## The decision grid

Cross rate-of-change against breadth of holding:

| | Held by few | Held by many |
|---|---|---|
| **Changes slowly** | Runtime attribute in the ACD. Fine. | Static profile entry. Fine — it rarely moves anyway. |
| **Changes quickly** | Runtime attribute in the ACD. Fine. | **Cannot be expressed. This is the capability layer's job.** |

Only one cell fails. That is the scope of any layer above the ACD, and it should be
sized honestly — it is usually a small number of attributes, not the whole ontology.

## Secondary constraint: schema immutability

Where a platform fixes an attribute's type at creation, an evolving ontology forces
attribute re-creation and re-profiling rather than schema change. This does not
appear in capacity limits but sets the real cost of iterating the model. Check for
it explicitly.

## What the test does not tell you

- **Nothing about routing quality.** An ontology that fits may still route badly.
- **Nothing about write latency.** Fitting within limits is separate from whether
  the platform's update API is fast enough to act on. Check throughput separately.
- **Nothing about whether the data exists.** An attribute with no system of record
  fails before the cardinality test is reached.

## Running it honestly

The test is only meaningful if it can return "the platform is fine." Set the
thresholds and gather the counts before forming a view. A test constructed after
the conclusion is advocacy with arithmetic attached, and reviewers will recognise it.

Block 4 — capability-evidence-and-decay.md

# Evidence and Decay

A capability that cannot be evidenced cannot safely drive allocation. A capability
that never expires will eventually route work to skill that has lapsed.

Figures marked [estimated] are judgment, not measurement.

## Evidence grades

Not all capability claims are equal. Grade them, and let the grade gate the use.

| Grade | Basis | Safe to use for |
|---|---|---|
| A | Assessed against a standard, recorded, in date | Eligibility |
| B | Demonstrated in production, measured | Eligibility and ranking |
| C | Trained, not yet independently demonstrated | Ranking only, discounted |
| D | Self-declared or inferred from team membership | Neither — flag for assessment |

Most operations discover that a large share of their capability record is grade D.
That is a finding worth surfacing early, because it caps what any allocation engine
can responsibly do regardless of platform.

## The trained-not-live state

Between "cannot" and "can" sits an agent who has completed training but is not yet
productive. It is a distinct state, not a weaker version of competence, and
collapsing it into either neighbour causes a predictable failure: counted as
capable, the agent is allocated work they handle slowly and badly; counted as
incapable, the ramp is invisible to planning. Model it explicitly, with a
productivity discount attached.

## Decay

Capability is a flow. Three mechanisms move it, and each needs a rate:

1. **Progression** — movement along the ladder. A generic five-tier form:
   Foundation → Practitioner → Specialist → Expert / SME → Lead. Every operation
   has its own names; what matters is that the ladder is explicit and that the
   transition rate is known, because it determines how fast the ontology goes stale.
2. **Expiry** — certifications, clearances and accreditations lapse on dates. An
   ontology without expiry dates silently routes to lapsed capability.
3. **Erosion** — unexercised skill degrades. Rarely measured; usually approximated
   by time since last use of the capability. Crude, but better than assuming
   permanence.

## Quality as a live attribute

Automated interaction scoring makes quality a continuously moving attribute rather
than a periodic assessment. This is genuinely useful for ranking, and carries two
traps:

- **Coverage asymmetry.** If scoring reaches a different proportion of interactions
  across tiers or sites, the scores are not comparable and the difference will be
  misread as performance.
- **Feedback loops.** Routing on quality score sends better work to higher scorers,
  which raises their scores. Without case-mix adjustment the model amplifies its own
  prior and starves development.

## What to record per capability

| Field | Why |
|---|---|
| Definition | Ambiguity here propagates into every allocation decision |
| Measurement scale | Boolean, ordinal, or continuous — determines how it can be used |
| Evidence grade | Gates whether it may drive eligibility |
| System of record | No source, no automation |
| Certifying authority | Who can assert it |
| Expiry and re-assessment interval | Whether it decays, and how fast |

The last three are the ones normally missing, and they are the ones that decide
whether the ontology is operable or merely descriptive.

Block 5 — supply-data-collection.md

# Collecting the Ontology

The instrument matters. A collection designed badly returns a description of what
teams think you want, not what exists.

## Collect pools, not agents

The unit is the **pool**: a group of agents sharing every attribute under
consideration. If two groups differ on any single attribute, they are two pools.
Record headcount per pool.

An agent-level extract is what an allocation engine eventually consumes, but it is
the wrong first ask. It is slow, frequently blocked on data protection review, and
returns detail before anyone knows which attributes matter. Pool-level collection
returns the *shape* of the estate — how many genuinely distinct supply buckets exist
and where the volume sits — in a fraction of the time.

Ask agent-level feasibility as a question instead: for each attribute, is it held
at agent level, and is it extractable?

## Structure of the instrument

| Component | Purpose |
|---|---|
| Pool inventory, one tab per business unit | The main collection. Separate tabs prevent edit collisions and give clean ownership |
| Attribute dictionary | The candidate attribute list, with blank rows. Respondents **extend** it |
| Controlled value lists | Seed vocabularies with space to add. Additions are the highest-value return |
| Systems of record | Where each attribute lives. Decides feasibility |
| Capability library | What each capability means and how it is evidenced |
| Cardinality calculation | Computes automatically from returns. No input |
| Gap register | What could not be answered, and why |

Mark required fields visibly and keep them few — roughly a third of the total is a
workable ratio [estimated]. Everything else is best-effort. A collection where
every field is mandatory returns invented data.

## Design rules

- **Seed the vocabularies, then invite correction.** Blank taxonomies return
  nothing; a seeded list respondents can argue with returns the real one. State
  plainly that the seeds are probably wrong in places.
- **Ask for attributes not used today.** Say so explicitly and more than once, or
  respondents will document current system configuration and stop.
- **Make blanks legitimate.** An explicit gap register converts "I don't know" from
  a failure into a contribution. Without one, respondents guess, and guesses are
  indistinguishable from data on return.
- **Name what you are unsure of.** List the specific things you believe but cannot
  confirm — segment definitions, ladder names, distribution estimates — and ask for
  correction. This returns better information than open-ended review.
- **Take partial returns on time** over complete returns late. Shape first.

## Reading the return

Three questions, in order:

1. **How many genuinely distinct pools exist?** Usually more than anyone expects,
   and the number itself is often the most useful output.
2. **Which attributes have no system of record?** These are the hard ceiling on
   automation, independent of platform choice.
3. **What is the cardinality of each attribute?** Feed into the test in
   `attribute-cardinality-test.md`.

## Failure modes

| Symptom | Cause |
|---|---|
| Very few pools returned | Respondents aggregated. Re-state the "differ on any attribute" rule |
| Every pool looks identical | The dimension set is too coarse for this operation |
| Gap register empty | Respondents guessed. Treat the return as unvalidated |
| Value lists unchanged | Seeds were taken as authoritative. Ask directly what is missing |

An empty gap register is the signal to watch. It almost never means there were no
gaps.

Usage notes

Sizing. The instruction block is roughly 480 words and loads with every message in the project. The four reference files total approximately 3,900 words and load only when retrieved.

Scope. The pack covers the supply-side attribute model and the platform-fit decision. It does not cover demand scoring, queue design, or forecasting. For those see Value Routing Model, Three-Pool Architecture and Skill-Based Routing.

Novel content. Blocks 3 and 5 — the cardinality test and the collection method — are not yet derived from an article, because no article covers them. They are the pack's original contribution and carry a maintenance risk the other blocks do not: there is no canonical version to regenerate from. An article home for both is the intended next step.

Drift. Blocks 2 and 4 derive from the source articles named in the infobox. Treat that list as a dependency list and regenerate when any of them changes materially.

Platform figures. The worked example in Block 3 cites vendor documentation read on a stated date. Vendor limits change without notice. Re-verify before relying on any figure, and do not quote them externally without checking.

Change history

Version Date Change
1.0 2026-08-07 Initial publication. One instruction block, four reference blocks.
1.1 2026-08-07 Added future-state ecosystem architecture diagram and SVG source.

See also