Thirty-four percent — that is where a sales director's override rate on RFQ qualification landed after six weeks, up from 8 percent. Nobody had changed the RFQ workflow. Nobody had touched the pricing rules the AI system checked against. Nobody had touched the system itself.
This is the moment in a diagnosis I watch people get wrong most often. A quality number moves in one part of the operation, and the instinct is to interrogate that exact part of the operation — the workflow, the rules, the person doing the overriding. That instinct is usually backwards. When nothing local to the number has changed, the cause is rarely local either.
Here is the account I keep coming back to, because the mechanism is visible enough, once you know where to look, that I now use it as a template for how I read every quality complaint about an AI system on a team.
A Number Moved in the Wrong Department
An 11-person sales team at a regional distributor had one AI system carrying a single open mandate across its sales operation. Over time it had absorbed three jobs that ran through it: drafting follow-up emails to quoted customers, roughly 40 a week; qualifying inbound RFQs against pricing rules, about 25 a week; and assembling the VP's pipeline report, one a week.
The number that mattered to the sales director was the override rate on RFQ qualification — how often he corrected the AI system's call before it went out. That rate had sat near 8 percent for months. Over six weeks it climbed to 34 percent, and by the time it landed on my desk, everyone closest to the RFQ side had already ruled themselves out. The pricing rules were unchanged. The RFQ format was unchanged. The person reviewing them was the same person, applying the same standard he had applied for a year. Only the number had moved.
Three Explanations That Were All Wrong
Three explanations get offered first, almost every time, and all three are worth naming because dismissing them properly is what clears the way to the real one.
Maybe the model drifted. The version history said otherwise — no update had shipped to the system in that six-week window. Nothing had changed under the hood to check.
Maybe the RFQs got harder. The mix said otherwise too. Average deal size, deal complexity, and client segment were flat against the same period a quarter earlier. Nothing about the input had shifted.
Maybe the sales director got stricter. This one looked more plausible until the overrides were mapped against the calendar. If a reviewer tightens his own standard, the extra scrutiny shows up evenly, spread across every day he works. These overrides were not spread evenly. They were bunched on specific days — and once I saw the shape of that clustering, the explanation stopped being about the RFQs at all.
What a Ceiling Actually Does
The days where RFQ overrides spiked were the same days follow-up email volume spiked. The two jobs — drafting follow-up emails and qualifying RFQs — were running through the same AI system with no weekly limit set on either one, which meant they were sharing a single queue with no rule for which job went first. On a day both arrived heavy, whichever job landed second got whatever attention was left over. The 2.1-day delay that had also crept into follow-up emails that quarter was not a separate problem. It was the same fault showing up in the other direction, on the days RFQ work happened to land first instead.
Once the cause was contention, not competence, the fix was not a better model or a retrained one. It was a weekly ceiling on each job and one named person watching each ceiling: a Follow-Up Drafter, capped at 40 emails a week and owned by the sales ops manager; an RFQ Qualifier, capped at 25 a week and owned by the sales director; and a Pipeline Reporter, one report a week, owned by the VP.
Within three weeks, the RFQ override rate was back to 9 percent, and follow-up emails were going out the same day again. Nothing about the AI system changed in that window — same model, same rules, same reviewers. What changed was that two jobs stopped sharing one uncapped queue.
This is the diagnostic sequence I now run before I let anyone touch the model itself:
- Did the drop begin when a different job was added to the same system? Read the log of tasks you assigned, not the model's version history.
- Is the degradation concentrated on peak-demand days rather than spread evenly? A model problem is flat. A contention problem has a shape.
- Does the failing job share a queue with a job whose volume grew?
- Set a weekly ceiling on each job and watch for recovery. If quality returns without touching the system, it was never the system.
The instinct to blame the model is understandable — it is the visible part, and it is the part everyone assumes owns the outcome. But in an AI-augmented team, a quality drop is more often a resourcing collision wearing a model problem's clothes. The tell is almost always the same: the number that falls belongs to the job nobody touched. Before anyone retrains, re-prompts, or replaces a system, I check what else was quietly added to its queue first. I go into this in more depth in my book on applied AI for organizations, and I keep the running version of my thinking on LinkedIn.
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.
