Back to Blog

The AI Baseline That Lived in Someone's Memory

A composite, illustrative case: a finance operations team reports a 45 percent efficiency win on AI-assisted invoice triage. The real number, once the true baseline surfaces, is 14 percent, and the gap traces back to a "before" figure no one ever recorded.

Direct answer

Where does an AI ROI baseline actually need to come from?

An unverified baseline distorted by human memory can drastically inflate perceived AI returns, as shown in a composite case where an unrecorded 38-minute baseline yielded a claimed 45 percent win instead of the true 14 percent calculated from an actual 24.3-minute record. To establish valid pre-change numbers, Dr. Jonah Tebaa prescribes the Baseline Capture Protocol, a six-step framework requiring teams to timestamp four to six weeks of raw case-level data, log anomalies, align definitions, assign an owner, and cross-check records.

A precision steel ruler with crisp engraved markings that stops short, continued across a dark surface by a smudged, hand-drawn pencil line leading back to a small pencil stub, symbolizing an AI ROI baseline that was recalled rather than recorded.

A controller leans over the ROI slide the week before it goes to the leadership offsite and asks one question that stops the meeting cold. "Where did the 38 minutes come from?"

Nobody in the room can point to a file. The number has lived in a planning deck for the better part of a year, cited in one steering update after another, footnoted as "baseline: 38 min/exception." Nobody remembers who first typed it. No one can produce the spreadsheet, the export, or the date range it was pulled from. It is not a bad number. It is not, strictly speaking, a number at all. It is a memory wearing the clothes of one.

What follows is a composite, illustrative case, built from a pattern I keep meeting across finance, operations, and shared-services teams, not one client's file, but arithmetically consistent with what recurs. A finance operations team at a mid-size B2B company runs AI-assisted triage on invoice exceptions, the mismatches, disputes, and flagged line items that used to sit in a queue until someone worked through them by hand.

The claimed baseline, recalled from that planning slide: 38 minutes per exception. The post-AI measurement, this time done properly: 21 minutes per exception, drawn from ticketing timestamps across 1,400 exceptions over 60 days. The claimed improvement is (38 minus 21) divided by 38, which comes to 44.7 percent, rounded up and presented as a 45 percent win.

Then finance does what finance does. Someone goes looking for the source of the 38 and finds, buried in an old shared drive, a contemporaneous dashboard export from the very month the baseline was supposedly measured. The actual average that month was 24.3 minutes per exception. The 38-minute figure turned out to be anchored to a single unrepresentative week, the one in which two staff members were out and the queue backed up. No one had done that on purpose, and no one had checked, either.

Recalculated against the real baseline, the improvement is (24.3 minus 21) divided by 24.3, or 13.6 percent, call it 14. Not nothing. Real, measurable, worth having. But less than a third of what the slide claimed.

A bar chart comparing three figures: a recalled baseline of 38 minutes per exception, a recorded contemporaneous baseline of 24.3 minutes per exception, and the actual post-AI result of 21 minutes per exception, showing the recalled number far overstating the true starting point, and the claimed 45 percent gain shrinking to a real 14 percent.
The recalled baseline, the recorded baseline, and the measured result. Two of the three numbers were never in dispute. The one that mattered most was the one nobody had actually kept.

The dollar version makes the gap harder to wave away. At a fully loaded cost of $42 an hour, or $0.70 a minute, and roughly 700 exceptions a month, the claimed savings were 17 minutes times $0.70 times 700, or $8,330 a month. The real savings, using the 3.3-minute true improvement, are $1,617 a month. The team had been reporting, in good faith, savings 5.15 times larger than what was actually happening. No one had lied, and no one had even been careless in the way people usually mean it. They had simply trusted a number that was never a measurement to begin with.

Why the Baseline Was Never a Number

I want to be precise about what went wrong here, because it is not the same failure I have written about before. It is not a denominator problem, where the wrong slice of the workload gets used as the base. It is not a translation problem, where a time saving gets mapped onto the wrong line of the P&L. Both of those assume the starting figure was at least real. Here, the starting figure was never real in the first place. It was recalled, not recorded.

Human memory is a bad instrument for measurement, and it fails in a specific, predictable direction: it anchors on whatever was most vivid, not whatever was most typical. The week two staff members were out was memorable. It generated complaints, a backlog, probably an escalation email. The eleven ordinary weeks around it generated nothing worth remembering. So when someone later asked, informally, "what was our average handle time before the AI rollout," the answer that surfaced was the number attached to the vivid week, not the representative one.

That number then did something numbers do once they appear on a slide: it calcified. No one re-derived it, because no one thought there was a reason to. It had a source once, technically, a bad week's worth of tickets glanced at by someone under deadline pressure, but that source was never archived, never dated, never distinguished from a real sampling exercise. By the time it mattered, it was indistinguishable, to everyone in the room, from a properly measured baseline. That is the actual failure mode: not a wrong number, but a number with no provenance, dressed convincingly enough that nobody thought to check.

The Baseline Capture Protocol

The fix is not complicated, and it does not require new tooling. It requires deciding, before the AI rollout begins, that the before number deserves the same rigor as the after number. I ask teams to run six steps, in order, every time:

  1. Freeze and timestamp a pre-change window of four to six weeks, never a single week. A single week measures whatever happened to be true that week. Four to six weeks measures the process.
  2. Archive the raw case-level data, not a summary average and never a slide. A slide is a claim about the data. The data is the evidence. Keep both, but only one of them counts as proof.
  3. Log anomalies during the window as they happen: staffing gaps, backlog spikes, policy changes, system outages. That way, a future reader can see whether the window was representative without having to reconstruct it from memory.
  4. Use an identical metric definition on both sides of the comparison. If "minutes per exception" excludes queue wait time before AI and includes it after, the comparison is already broken before the first number is calculated.
  5. Name one owner for the baseline file, with a fixed storage location and a stated retention date, so the file exists on purpose rather than by accident and can still be found in a year.
  6. Cross-check the baseline against an independent, contemporaneous source: a finance report, an ops dashboard, timesheets, before any ROI percentage built on it reaches a slide.

None of this is expensive. It is mostly the discipline of writing something down at the moment it is true, instead of trusting that someone will remember it accurately a year later. The finance team in this composite case did not need better AI. It needed a folder, a date, and one person responsible for what went in it.

When the Baseline Is Already Gone

Most organizations reading this are not standing at the start of a rollout. They are already past it, with an AI system live, a claimed win already reported upward, and no archived pre-change data to check it against. The honest move at that point is not to keep repeating the original percentage, and it is not to pretend the question can no longer be answered either.

Reconstruct from whatever proxy data survived: timesheets, staffing schedules, ticket backlogs, finance close reports, anything with a timestamp attached to the period before the change. These sources will not produce a clean point estimate, and that is fine. Report a range instead, with the uncertainty stated plainly, something closer to "handle time before the rollout was likely between 23 and 27 minutes, based on staffing records and backlog data from that quarter," rather than a confident single figure that implies precision the organization does not actually have. A wide, honest range survives a board question. A false point estimate does not survive the first person who asks where it came from.

The Question a Board Should Ask First

Before any AI ROI percentage reaches a board deck, I want one question asked and answered on the record: what was the before number, where is it stored, and who signed off on it. Not what is the after number, not what is the percentage improvement. The before number, first, because everything downstream of it inherits whatever it got wrong.

If the answer is a file with a date, a named owner, and an independent cross-check, the ROI figure that follows is worth taking seriously. If the answer is a shrug, or a slide, or "I think that's roughly what it was," the percentage attached to it is not a measurement. It is a guess dressed up as a measurement, and it deserves exactly the scrutiny that description implies.

I have written before about the denominator problem, how the wrong slice of workload can make a real result look larger than it is (see The Denominator Problem: Recalculating a 70 Percent AI Win), and about why hours saved on a screen rarely survive translation onto an actual P&L line (see why I stopped reporting hours saved). Both of those assume the starting number was real. This one does not get to make that assumption, and neither should you, until someone can show you where the before number actually lives.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.

For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.

Related evidence: The ILO's assessment of global occupational exposure to generative AI finds that its main effect is likely to augment rather than automate jobs. (ILO working paper on generative AI and jobs)

Google's Rules of Machine Learning guidance tells teams to design and implement metrics before formalising what the system will do, and to track as much as possible in the existing system first. (Google's Rules of Machine Learning)

Frequently Asked Questions

If we deployed AI without capturing a baseline, is our ROI number unusable?

Not unusable, but it needs to be relabeled. A recalled number should never be presented as a measured one. Reconstruct a range from the best proxy data available, timesheets, ticket backlogs, contemporaneous dashboard exports, state the uncertainty honestly, and stop reporting a single confident percentage until a real baseline exists for the next comparison period.

How long does a baseline capture window need to be?

Four to six weeks, never a single week. A single week absorbs whatever happened to be true that week, a staffing gap, a holiday, a system outage, and reports it as the normal state of the process. Four to six weeks is long enough to average out most one-off disruptions while still describing the process as it stood before the change.

Isn't this the same as the problem of hours-saved not mapping to real cost savings?

No, and the two should not be merged. The hours-saved problem is a translation error, converting a time reduction into a dollar figure on the wrong line of the P&L. This is a provenance error, further upstream: the before number itself was never timestamped, sampled, or archived. A perfectly translated ROI built on an unrecorded baseline is still wrong, because the number it started from was never real to begin with.

Who should own baseline capture before an AI project starts?

One named individual, not a committee and not whoever happens to be running the pilot. The owner is responsible for freezing the pre-change window, archiving the raw case-level data at a fixed location with a retention date, and producing that file on request, before any ROI percentage is allowed to reach a slide.