Back to Blog

The Override Rate Rose in a Team I Hadn't Touched

An override rate went from 8 percent to 34 percent in six weeks, and nothing in the RFQ workflow, the pricing rules, or the AI system had changed. The cause was two AI jobs quietly sharing one uncapped queue — and the fix was a weekly ceiling, not a smarter model.

A single glass funnel on a brass stand is visibly jammed — steel bearings locked and blocking the throat, sand stalled above them, a brass disc tipped over the rim, only a thin trickle reaching a mostly empty dish — while in the soft-focus background three small separate funnels, each on its own brass stand, sit calmly with their dishes evenly and fully filled.
Direct answer

What does The Override Rate Rose in a Team I Hadn't Touched mean in practice?

When an untouched sales team's RFQ qualification override rate jumped from 8 to 34 percent, Dr. Jonah Tebaa found the failure stemmed from queue contention rather than model drift. The single AI system handled three uncapped workflows, causing quality degradation whenever follow-up email volume spiked. Resolving this issue required an apparatus of weekly volume ceilings with named owners: 40 emails for the sales ops manager, 25 RFQs for the sales director, and one weekly report for the VP, reducing overrides to 9 percent.

Thirty-four percent — that is where a sales director's override rate on RFQ qualification landed after six weeks, up from 8 percent. Nobody had changed the RFQ workflow. Nobody had touched the pricing rules the AI system checked against. Nobody had touched the system itself.

This is the moment in a diagnosis I watch people get wrong most often. A quality number moves in one part of the operation, and the instinct is to interrogate that exact part of the operation — the workflow, the rules, the person doing the overriding. That instinct is usually backwards. When nothing local to the number has changed, the cause is rarely local either.

Here is the account I keep coming back to, because the mechanism is visible enough, once you know where to look, that I now use it as a template for how I read every quality complaint about an AI system on a team.

A Number Moved in the Wrong Department

An 11-person sales team at a regional distributor had one AI system carrying a single open mandate across its sales operation. Over time it had absorbed three jobs that ran through it: drafting follow-up emails to quoted customers, roughly 40 a week; qualifying inbound RFQs against pricing rules, about 25 a week; and assembling the VP's pipeline report, one a week.

The number that mattered to the sales director was the override rate on RFQ qualification — how often he corrected the AI system's call before it went out. That rate had sat near 8 percent for months. Over six weeks it climbed to 34 percent, and by the time it landed on my desk, everyone closest to the RFQ side had already ruled themselves out. The pricing rules were unchanged. The RFQ format was unchanged. The person reviewing them was the same person, applying the same standard he had applied for a year. Only the number had moved.

Three Explanations That Were All Wrong

Three explanations get offered first, almost every time, and all three are worth naming because dismissing them properly is what clears the way to the real one.

Maybe the model drifted. The version history said otherwise — no update had shipped to the system in that six-week window. Nothing had changed under the hood to check.

Maybe the RFQs got harder. The mix said otherwise too. Average deal size, deal complexity, and client segment were flat against the same period a quarter earlier. Nothing about the input had shifted.

Maybe the sales director got stricter. This one looked more plausible until the overrides were mapped against the calendar. If a reviewer tightens his own standard, the extra scrutiny shows up evenly, spread across every day he works. These overrides were not spread evenly. They were bunched on specific days — and once I saw the shape of that clustering, the explanation stopped being about the RFQs at all.

What a Ceiling Actually Does

The days where RFQ overrides spiked were the same days follow-up email volume spiked. The two jobs — drafting follow-up emails and qualifying RFQs — were running through the same AI system with no weekly limit set on either one, which meant they were sharing a single queue with no rule for which job went first. On a day both arrived heavy, whichever job landed second got whatever attention was left over. The 2.1-day delay that had also crept into follow-up emails that quarter was not a separate problem. It was the same fault showing up in the other direction, on the days RFQ work happened to land first instead.

Once the cause was contention, not competence, the fix was not a better model or a retrained one. It was a weekly ceiling on each job and one named person watching each ceiling: a Follow-Up Drafter, capped at 40 emails a week and owned by the sales ops manager; an RFQ Qualifier, capped at 25 a week and owned by the sales director; and a Pipeline Reporter, one report a week, owned by the VP.

Within three weeks, the RFQ override rate was back to 9 percent, and follow-up emails were going out the same day again. Nothing about the AI system changed in that window — same model, same rules, same reviewers. What changed was that two jobs stopped sharing one uncapped queue.

This is the diagnostic sequence I now run before I let anyone touch the model itself:

  1. Did the drop begin when a different job was added to the same system? Read the log of tasks you assigned, not the model's version history.
  2. Is the degradation concentrated on peak-demand days rather than spread evenly? A model problem is flat. A contention problem has a shape.
  3. Does the failing job share a queue with a job whose volume grew?
  4. Set a weekly ceiling on each job and watch for recovery. If quality returns without touching the system, it was never the system.

The instinct to blame the model is understandable — it is the visible part, and it is the part everyone assumes owns the outcome. But in an AI-augmented team, a quality drop is more often a resourcing collision wearing a model problem's clothes. The tell is almost always the same: the number that falls belongs to the job nobody touched. Before anyone retrains, re-prompts, or replaces a system, I check what else was quietly added to its queue first. I go into this in more depth in my book on applied AI for organizations, and I keep the running version of my thinking on LinkedIn.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.

For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.

Frequently Asked Questions

Why did the sales director's RFQ qualification override rate jump from 8 percent to 34 percent?

The override rate increased because the single AI system was carrying an uncapped queue shared across multiple jobs. Specifically, qualifying inbound RFQs and drafting follow-up emails were running through the same system without volume limits or priority rules. On heavy days, whichever task arrived second received whatever capacity remained. The degradation was not caused by model drift, altered pricing rules, harder deal complexity, or stricter reviewer standards, but by queue contention occurring when follow-up email volume spiked alongside inbound RFQs.

How were model drift and input difficulty ruled out as causes for the AI quality drop?

Dr. Jonah Tebaa ruled out model drift by inspecting the version history, which confirmed that no updates had shipped to the AI system during the six-week window. Changing input difficulty was eliminated by examining deal characteristics: average deal size, deal complexity, and client segments remained flat compared to the same period a quarter earlier. Furthermore, the pricing rules and RFQ formats were completely unchanged, proving the quality drop was unrelated to model alterations or harder incoming data.

Why did calendar mapping disprove the theory that the sales director became stricter?

Calendar mapping showed that the overrides were bunched on specific days rather than distributed evenly across the schedule. If a human reviewer tightens their evaluation standard, the resulting scrutiny and corrections appear uniformly across every working day. In this case, the override spikes clustered specifically on the exact days that follow-up email volume also surged. This uneven calendar pattern revealed that the degradation had a distinct shape driven by workflow contention rather than a permanent personal shift in grading rigor.