Back to Blog

I Watched an AI Rollout Win on Average and Lose on the Tail

An AI copilot cut average handle time by 25 percent, and the number went straight into the quarterly review. The slowest five percent of cases got 69 percent worse, and that number went nowhere. Netted in dollars, the same quarter was a $276,000 loss.

A mirror-polished navy slab reflecting one unbroken bar of light, tearing open at its right edge into a deep fissure lit by a single amber seam.
Direct answer

What does I Watched an AI Rollout Win on Average and Lose on the Tail mean in practice?

An AI rollout can win on average yet lose on the tail when efficiency gains on routine tasks disguise catastrophic delays on complex exceptions. Dr. Jonah Tebaa illustrates this with an AI copilot that reduced mean ticket resolution by 25.4 percent, saving $84,000, while driving 95th-percentile resolution times to 71 minutes and tripling SLA breaches into a $276,000 net quarterly loss.

Twenty-five percent faster, on average — that was the number that got an AI rollout approved for expansion before anyone in the room asked what happened to the cases it didn't help.

I see this shape often enough now that I no longer trust a single-number verdict on any operational AI deployment, no matter how good the number looks. An average compresses a distribution into one figure, and compression always hides something. The real question is never whether the mean moved. It is what moved with it, in the other direction, in the part of the distribution nobody put on a slide.

This is the conversation I have most often with COOs and CFOs who are deciding, this quarter, whether to renew or expand an AI system they already trust.

The case that changed how I read a quarterly review

Here is the version of this I keep coming back to, because the arithmetic is clean enough to show the entire mechanism in one pass.

A support operation runs 40,000 tickets a quarter. Before any AI involvement, mean resolution time is 14.2 minutes, and the 95th-percentile case — the slow, messy, multi-touch ticket — takes 42 minutes. There is a contractual SLA behind that number: any ticket that runs past 60 minutes costs $150 in penalties. About 3 percent of tickets breach that line, which is 1,200 tickets and $180,000 in penalties for the quarter.

Then an AI copilot goes live, drafting the agent's first response automatically. Mean resolution time drops to 10.6 minutes — a 25.4 percent improvement. That is the number that makes the quarterly review deck. It is the number that gets applause, and it is the number that gets the tool renewed and expanded into other teams.

Here is what does not make the deck. The 95th-percentile case does not get faster. It goes to 71 minutes — a 69 percent increase.

Why the same tool helps and hurts in the same quarter

The mechanism stops being mysterious the moment you segment the ticket volume instead of averaging across all of it. Roughly 80 percent of tickets are routine. The AI drafts a correct, ready-to-send response, and the agent approves it in seconds. That is where the entire mean improvement comes from, and it is genuinely real — no accounting trick, no rounding error.

The other 20 percent are the ambiguous, multi-issue, policy-exception cases — precisely the tickets that were already driving the 95th percentile before the AI arrived. For those, the draft is the wrong shape. It is built for the routine case, not for the exception sitting in front of the agent. The agent now has to read the draft, recognise it does not fit, correct or discard it, and then still do the manual work the case required in the first place. That is not a removed step. It is an added one, layered on top of work that was already the hardest and slowest 20 percent of the queue.

Run the dollars and the story reverses completely. The labour saving is real: 3.6 minutes saved per ticket, across 40,000 tickets, is 144,000 minutes — 2,400 hours — at a $35 loaded hourly cost. That is $84,000 saved for the quarter.

But the breach rate does not stay flat. It triples, from 3 percent to 9 percent. That is 3,600 breaching tickets instead of 1,200 — 2,400 more, at $150 each, which is $360,000 in new SLA penalties.

Net the two figures on the same period, in the same unit: $84,000 in savings minus $360,000 in new tail cost is a $276,000 net loss — in the exact quarter the dashboard reported a 25 percent win.

The variance audit

This is the check I now build into every AI value review before I will sign off on a renewal or an expansion decision. It is five steps, and I keep all five, because dropping any one of them lets the same blind spot back in.

  • Report the full distribution, not the mean alone. Pull p50, p90 and p95 — p99 wherever a contractual penalty attaches — before anyone presents an "average improved by X percent" line.
  • Segment before you average. Split routine cases from exception cases first. A blended mean will always flatter a tool that is good at the easy majority and bad at the hard minority — that is exactly what blending does.
  • Price the tail, not just the labour. Find the dollar figure attached to your worst-case threshold — SLA penalty, churn, regulatory fine, chargeback — and multiply it by the change in breach count, not the change in average time.
  • Net the two on the same period. Mean-driven savings minus tail-driven cost, same quarter, same unit. If you cannot produce that single net number, you do not have a value claim yet — you have a headline.
  • Re-run it quarterly. The tail moves independently of the mean as case mix and behaviour adapt. A system that nets positive in month one can net negative by month four with no code changing at all.

What I tell CFOs before they sign the renewal

None of this means the AI tool is bad, or that the deployment was a mistake. In the case above, the tool is doing exactly what it was built to do — it drafts fast, competent responses for the volume it was designed around, and it was never asked to handle exceptions. That is a scope gap, not a failure of the technology. The tool is a feature update away from handling more of that 20 percent well, but that is a roadmap conversation, not a renewal justification, and the two get conflated more often than they should. The actual failure is upstream of the tool entirely: nobody priced the tail before the renewal conversation started.

The average is a summary. It was never designed to be a verdict, and treating it as one is the mistake, not the AI system itself. A mean-based dashboard will report a win even when the underlying programme is a net financial loss, because the losses are sitting in a part of the distribution that nobody built a slide for.

So before you renew or expand any AI rollout on the strength of an average, ask the question that number cannot answer: what happened to the tail, and what does it cost in dollars, not minutes? If the team presenting the renewal cannot answer that in the same meeting, the average is not evidence yet. It is a claim waiting for the rest of the story. That gap — between what the average says and what the tail costs — is where AI renewal decisions are actually won or lost.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.

For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.

Frequently Asked Questions

How can an AI copilot improve average resolution times while increasing total operational costs?

An AI copilot can improve average metrics while causing net financial loss because averages compress full distributions. In Dr. Jonah Tebaa's example, routine cases constituting 80 percent of support volume sped up, driving a 25.4 percent drop in mean resolution time and saving $84,000 in labor. However, for the 20 percent exception cases, incorrect drafts added extra steps. This pushed 95th-percentile resolution times to 71 minutes, tripling expensive contractual SLA breaches and adding $360,000 in penalties, producing a net loss of $276,000.

Why do exception cases take longer to resolve after deploying an AI drafting tool?

Exception cases slow down because the generative tool is designed primarily for standard, routine inquiries. When ambiguous, multi-issue, or policy-exception tickets arrive, the AI copilot generates drafts that do not fit the problem. Rather than eliminating steps, this introduces an extra task for human agents. Agents must read the inaccurate draft, recognize that it fails to address the unique issue, discard or correct it, and subsequently complete the original manual work, increasing resolution time for the hardest 20 percent of support tickets.

What are the five steps of the variance audit recommended before renewing an AI system?

Dr. Jonah Tebaa outlines a five-step variance audit before signing renewal or expansion decisions. First, report the full distribution including p50, p90, p95, and p99 metrics rather than solely relying on the mean. Second, segment routine tickets from complex exceptions before averaging. Third, price the tail by multiplying worst-case penalties like SLA fines by the change in breach counts. Fourth, net labor savings against tail costs for the period. Fifth, rerun the audit quarterly as operational behaviors and ticket mixes shift over time.