
A controller leans over the ROI slide the week before it goes to the leadership offsite and asks one question that stops the meeting cold. "Where did the 38 minutes come from?"
Nobody in the room can point to a file. The number has lived in a planning deck for the better part of a year, cited in one steering update after another, footnoted as "baseline: 38 min/exception." Nobody remembers who first typed it. No one can produce the spreadsheet, the export, or the date range it was pulled from. It is not a bad number. It is not, strictly speaking, a number at all. It is a memory wearing the clothes of one.
What follows is a composite, illustrative case, built from a pattern I keep meeting across finance, operations, and shared-services teams, not one client's file, but arithmetically consistent with what recurs. A finance operations team at a mid-size B2B company runs AI-assisted triage on invoice exceptions, the mismatches, disputes, and flagged line items that used to sit in a queue until someone worked through them by hand.
The claimed baseline, recalled from that planning slide: 38 minutes per exception. The post-AI measurement, this time done properly: 21 minutes per exception, drawn from ticketing timestamps across 1,400 exceptions over 60 days. The claimed improvement is (38 minus 21) divided by 38, which comes to 44.7 percent, rounded up and presented as a 45 percent win.
Then finance does what finance does. Someone goes looking for the source of the 38 and finds, buried in an old shared drive, a contemporaneous dashboard export from the very month the baseline was supposedly measured. The actual average that month was 24.3 minutes per exception. The 38-minute figure turned out to be anchored to a single unrepresentative week, the one in which two staff members were out and the queue backed up. No one had done that on purpose, and no one had checked, either.
Recalculated against the real baseline, the improvement is (24.3 minus 21) divided by 24.3, or 13.6 percent, call it 14. Not nothing. Real, measurable, worth having. But less than a third of what the slide claimed.

The dollar version makes the gap harder to wave away. At a fully loaded cost of $42 an hour, or $0.70 a minute, and roughly 700 exceptions a month, the claimed savings were 17 minutes times $0.70 times 700, or $8,330 a month. The real savings, using the 3.3-minute true improvement, are $1,617 a month. The team had been reporting, in good faith, savings 5.15 times larger than what was actually happening. No one had lied, and no one had even been careless in the way people usually mean it. They had simply trusted a number that was never a measurement to begin with.
Why the Baseline Was Never a Number
I want to be precise about what went wrong here, because it is not the same failure I have written about before. It is not a denominator problem, where the wrong slice of the workload gets used as the base. It is not a translation problem, where a time saving gets mapped onto the wrong line of the P&L. Both of those assume the starting figure was at least real. Here, the starting figure was never real in the first place. It was recalled, not recorded.
Human memory is a bad instrument for measurement, and it fails in a specific, predictable direction: it anchors on whatever was most vivid, not whatever was most typical. The week two staff members were out was memorable. It generated complaints, a backlog, probably an escalation email. The eleven ordinary weeks around it generated nothing worth remembering. So when someone later asked, informally, "what was our average handle time before the AI rollout," the answer that surfaced was the number attached to the vivid week, not the representative one.
That number then did something numbers do once they appear on a slide: it calcified. No one re-derived it, because no one thought there was a reason to. It had a source once, technically, a bad week's worth of tickets glanced at by someone under deadline pressure, but that source was never archived, never dated, never distinguished from a real sampling exercise. By the time it mattered, it was indistinguishable, to everyone in the room, from a properly measured baseline. That is the actual failure mode: not a wrong number, but a number with no provenance, dressed convincingly enough that nobody thought to check.
The Baseline Capture Protocol
The fix is not complicated, and it does not require new tooling. It requires deciding, before the AI rollout begins, that the before number deserves the same rigor as the after number. I ask teams to run six steps, in order, every time:
- Freeze and timestamp a pre-change window of four to six weeks, never a single week. A single week measures whatever happened to be true that week. Four to six weeks measures the process.
- Archive the raw case-level data, not a summary average and never a slide. A slide is a claim about the data. The data is the evidence. Keep both, but only one of them counts as proof.
- Log anomalies during the window as they happen: staffing gaps, backlog spikes, policy changes, system outages. That way, a future reader can see whether the window was representative without having to reconstruct it from memory.
- Use an identical metric definition on both sides of the comparison. If "minutes per exception" excludes queue wait time before AI and includes it after, the comparison is already broken before the first number is calculated.
- Name one owner for the baseline file, with a fixed storage location and a stated retention date, so the file exists on purpose rather than by accident and can still be found in a year.
- Cross-check the baseline against an independent, contemporaneous source: a finance report, an ops dashboard, timesheets, before any ROI percentage built on it reaches a slide.
None of this is expensive. It is mostly the discipline of writing something down at the moment it is true, instead of trusting that someone will remember it accurately a year later. The finance team in this composite case did not need better AI. It needed a folder, a date, and one person responsible for what went in it.
When the Baseline Is Already Gone
Most organizations reading this are not standing at the start of a rollout. They are already past it, with an AI system live, a claimed win already reported upward, and no archived pre-change data to check it against. The honest move at that point is not to keep repeating the original percentage, and it is not to pretend the question can no longer be answered either.
Reconstruct from whatever proxy data survived: timesheets, staffing schedules, ticket backlogs, finance close reports, anything with a timestamp attached to the period before the change. These sources will not produce a clean point estimate, and that is fine. Report a range instead, with the uncertainty stated plainly, something closer to "handle time before the rollout was likely between 23 and 27 minutes, based on staffing records and backlog data from that quarter," rather than a confident single figure that implies precision the organization does not actually have. A wide, honest range survives a board question. A false point estimate does not survive the first person who asks where it came from.
The Question a Board Should Ask First
Before any AI ROI percentage reaches a board deck, I want one question asked and answered on the record: what was the before number, where is it stored, and who signed off on it. Not what is the after number, not what is the percentage improvement. The before number, first, because everything downstream of it inherits whatever it got wrong.
If the answer is a file with a date, a named owner, and an independent cross-check, the ROI figure that follows is worth taking seriously. If the answer is a shrug, or a slide, or "I think that's roughly what it was," the percentage attached to it is not a measurement. It is a guess dressed up as a measurement, and it deserves exactly the scrutiny that description implies.
I have written before about the denominator problem, how the wrong slice of workload can make a real result look larger than it is (see The Denominator Problem: Recalculating a 70 Percent AI Win), and about why hours saved on a screen rarely survive translation onto an actual P&L line (see why I stopped reporting hours saved). Both of those assume the starting number was real. This one does not get to make that assumption, and neither should you, until someone can show you where the before number actually lives.
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.
For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.
Related evidence: The ILO's assessment of global occupational exposure to generative AI finds that its main effect is likely to augment rather than automate jobs. (ILO working paper on generative AI and jobs)
Google's Rules of Machine Learning guidance tells teams to design and implement metrics before formalising what the system will do, and to track as much as possible in the existing system first. (Google's Rules of Machine Learning)