Evaluating Agent-Assist Tools on Ramp Compression

From WFM Labs
Two learning curves from the same start; the gold area between them is months of proficiency bought.

Evaluating agent-assist tools on ramp compression is the argument that knowledge-assistance and copilot tools for service agents should be justified and measured on how much they shorten the time a new agent takes to reach proficiency, not on how much they shorten handle time — because the field evidence shows the gain concentrates in novices and barely touches experienced agents, because time to proficiency is the binding constraint in operations whose pool of already-proficient hires is running out, and because the handle-time case is the one that gets tools canceled when it fails to appear.

The technology itself — what assist tools do and how they work — is described at Agent Assist; this page is about how to justify and measure one. The capacity-modeling consequences of augmentation — productivity multipliers, adoption curves, role redesign — are owned by Workforce Planning for AI-Augmented Roles; the learning curve itself by Speed to proficiency curve; the argument that ramp compression is an elasticity intervention by Supply Elasticity in Workforce Planning. The parallel argument for complexity-reduction tooling — that accounts per person at held error and escalation rates beats handle time as the benefit metric — is owned by Pooling Architecture in Service Workforces; this page makes the equivalent case for assist tooling, where the denomination is ramp rather than width. This page covers the choice of evaluation metric and the design of the evaluation.

What the evidence shows

The best-identified field study of a generative assistant deployed to customer-support agents found a productivity gain of about fourteen percent on average, concentrated in the least experienced and lowest-skilled agents — roughly a third for novices — with little or no measurable effect on the most experienced; the authors' interpretation is that the tool disseminates the practices of top performers to those who have not yet acquired them.[1] A separate field experiment with knowledge workers found the same shape — the largest gains for those below the median performance before the tool — alongside a warning that the gain depends on whether the task falls inside the tool's competence.[2] Both studies measure productivity, not displacement, and neither settles whether the gain persists as novices become experts; but the shape is consistent and it is the shape of a ramp-compression effect, not a throughput effect.

The practical reading: the gain shows up as a novice performing closer to a tenured colleague than their tenure would predict, while the tenured colleague is largely unchanged. Translating that into months of ramp requires the operation's own curve — the percentage gain alone does not give it.

Why the handle-time case fails

Most assist-tool business cases are written on handle time — the tool will cut seconds from every contact, multiplied across the whole workforce. Agent Assist lists handle-time reduction among the technology's benefits, and it is real at the contact level; the argument here is about whether it survives aggregation to a whole-workforce business case. That aggregation fails for three reasons, and when it fails the tool is canceled for not doing something it was never going to do.[3]

  • The experienced majority does not move. If the gain is concentrated in novices and the workforce is mostly tenured, the whole-workforce handle-time effect is small by construction and is lost in ordinary variation.
  • Handle time is often not the binding constraint. Where an operation's problem is composition — too many small teams, too few agents able to serve enough of the work — a handle-time reduction does not relieve it; Supply Elasticity in Workforce Planning and Pooling Architecture in Service Workforces describe why composition binds before speed.
  • The pilot measures the wrong cohort. Pilots are staffed with experienced agents, because they are trusted to try new tools, and the cohort least likely to benefit produces the evidence on which the tool is judged.

Meanwhile the case the tool can actually make goes unmeasured: how much faster a new hire reaches the proficiency line, how many months of ramp are removed, and what that does to hiring where already-proficient people are scarce.

Ramp as the metric

The right denomination of benefit is months of proficiency bought — the area between the learning curve without the tool and the curve with it, up to the proficiency line — together with the downstream quantities that follow from it. Time from first contact to the tenured performance level is the practical proxy for that area and is the quantity the table below measures; two curves that cross the line at different points enclose different areas, and the proxy is read with that in mind.

Evaluation metrics for an assist tool, in order of relevance (rows 1–3 measurable in pilot; rows 4–5 at rollout)
Metric What it captures Cohort
Time to proficiency Weeks from first contact to the tenured-agent performance level, with and without the tool New hires, matched on start date and prior experience
Quality at held error and escalation rates Whether the ramp gain arrives without a rise in error or escalation — the discipline Pooling Architecture in Service Workforces applies to accounts per person New hires, tracked through ramp
Early-tenure retention Whether faster competence reduces the attrition that clusters in the first months New hires, at ninety days and six months
Hiring-profile width Whether the tool lets the operation hire from a wider or less experienced pool at the same quality Recruitment funnel, by prior-experience band
Handle time The conventional measure, reported last and by tenure band, so that a null result on the tenured band is expected rather than fatal New hires, by prior-experience band; tenured agents only where the rollout later covers them

Two design rules follow. The pilot cohort is new hires, not volunteers, because the effect lives there; and the comparison is against a contemporaneous cohort of new hires without the tool, because ramp curves shift with hiring quality, season and training changes, and a before-and-after comparison confounds all of them (the design is at A/B Testing for WFM Experiments). Handle time on the tenured population becomes measurable only at general rollout; in the pilot it is reported for the treated cohort alone. The design is not cheap to run: with ramps of three to six months, the comparison cannot be read before two hire classes have completed the curve, and the cohorts must be large enough for the difference in weeks to clear the spread of individual ramp times — the sizing arithmetic is at A/B Testing for WFM Experiments.[3]

Why it matters more than it used to

Supply Elasticity in Workforce Planning argues that where time to proficiency is the binding constraint on how fast staffed, proficient capacity can change, a technology that measurably shortens it functions as an elasticity intervention rather than as a productivity tool. The argument sharpens in operations whose external pool of already-proficient hires is exhausting — where, as Speed to proficiency curve describes, the constraint is a skill gate no employer funds people to walk through rather than a shrinking occupation, so that the operation must build proficiency it once bought. There, time to proficiency stops being an efficiency target and becomes the only way to staff the operation, and a tool that compresses it is worth more than any tool that shortens calls.

Failure modes

Beyond the three ways the handle-time case fails, two further errors recur.[3]

  • Ignoring quality at ramp. A faster ramp with higher error is not a ramp gain; it is a quality problem arriving early.
  • Canceling on the wrong metric. The handle-time case fails, the tool is withdrawn, and the ramp gain it was delivering is never measured.

Maturity Model Position

Cohort measurement and tenure banding are Level 3 capabilities on the WFM Labs Maturity Model™; evaluating a tool investment against curve-compression value is the Level 4 practice described at Speed to proficiency curve. Level 5: Adaptive Orchestration reaches the same reading of the evidence — amplification rather than replacement — from the operating-model side; this page supplies the evaluation design that would test it.

See Also

References

  1. Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942 (NBER Working Paper 31161, 2023).
  2. Dell'Acqua, F., McFowland, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School Working Paper 24-013.
  3. 3.0 3.1 3.2 Practitioner observation from agent-assist deployments and their business reviews in service operations; a consistent pattern rather than a measured result.