What does The Denominator Problem: Recalculating a 70 Percent AI Win mean in practice?
Dr. Jonah Tebaa reveals that recalculating a reported seventy percent AI win requires solving the denominator problem by factoring in downstream rework costs across the entire workload. In an invoice processing case study, clean straight-through automation reduced costs from $4.50 to $1.35, but three thousand flagged exceptions cost $6.20 each to correct. Dr. Jonah Tebaa demonstrates that applying a blended cost metric across all ten thousand invoices yields an actual 37.7 percent cost reduction.

A department head sent me a slide last week with one number circled: a 70 percent reduction in invoice processing cost after introducing AI-assisted coding. Attached to that number was a headcount plan, three roles cut from accounts payable by year end. Before we went further, I asked one question. Seventy percent of what, exactly?
The answer is the reason I am writing this. The 70 percent was real. It was also measuring the wrong population, and the gap between what it measured and what actually happened to the company's invoice spend is large enough to sink the plan built on top of it.
The report that said 70 percent
The company processes 10,000 invoices a month. Fully manual coding, a person opening each invoice, matching it to a purchase order, assigning a general ledger code, cost $4.50 per invoice. At 10,000 invoices, that is $45,000 a month, a number finance has tracked for years.
The AI system introduced this year handles a portion of those invoices straight through, with no human touching them beyond a periodic spot check. On those, the fully loaded cost, compute plus the sampled quality review, comes to $1.35 per invoice. Run the comparison on that population alone and the arithmetic is clean. $4.50 down to $1.35 is a 70 percent cost reduction. That is the number that made the slide, the deck, and the headcount plan.
It is also correct, as far as it goes. The problem is where it stops.
What got left out of the denominator
Of the 10,000 invoices, 7,000 went through cleanly. The other 3,000 did not. They were flagged for human review: ambiguous vendor formatting, a purchase order that did not match, a line item the model coded with low confidence. Someone on the accounts payable team had to open each of those, work out what the AI had done, decide whether to trust it, and correct it where it was wrong.
That correction work is not free, and it is not neutral. It costs $6.20 per invoice, higher than the $4.50 manual baseline, not lower. Reworking a partial, already-coded output takes longer than coding an invoice from a blank state. The reviewer has to reconstruct the AI's reasoning before finding the error in it, then override a system that is nominally supposed to be doing the work. That is a real, measurable cost, and it belongs in the same accounting as the 7,000 successes.
It was not in the report. The 70 percent figure was calculated only over the 7,000 straight-through invoices, the subset where the AI looked good. The 3,000 that pushed into rework were quietly dropped from the denominator before the percentage was calculated. Not through any intent to mislead. It is simply the natural thing to measure. You compare the new cost on the cases the new system actually processed cleanly, and you report that comparison as if it described the whole workload.
I see the same shape in customer support ticket deflection, in contract review, in claims triage, anywhere a system sorts work into an easy pile and a hard pile. The easy pile gets the case study. The hard pile gets the invoice for what full deployment actually costs.
The blended number finance should have used
The workload is 10,000 invoices a month, all of them, and that is the denominator that should govern any decision about headcount. Blend the two populations: 7,000 invoices at $1.35, plus 3,000 invoices at $6.20, divided by 10,000.
That comes to $2.805 per invoice, against a manual baseline of $4.50. The true reduction is 37.7 percent.
Thirty-seven point seven percent is a good result. It is a real, defensible improvement in a core operating cost, and on 10,000 invoices a month it works out to roughly $17,000 in monthly savings, worth the investment and worth expanding. It is also close to half of what the original report claimed. A headcount plan sized against a 70 percent reduction, when the true figure is 37.7 percent, overshoots the safe reduction by nearly two times. Cut three roles expecting a 70 percent efficiency gain, and the 3,000 invoices a month that still need a trained human to untangle them will not have anyone left to do it.
This is not a case against the technology. The blended number still justifies the project. It is a case for asking what the number is measuring before anyone builds a plan on top of it.
Four questions before you accept an AI ROI figure
In my work with executive teams evaluating AI performance claims, I have found the fastest way to catch this is a short, standing checklist, applied before any ROI figure is accepted into a budget or headcount decision.
- What is the denominator — the full workload, or only the cases that went through cleanly?
- Does the number include the cost of rework, escalation, and human correction downstream?
- Is the baseline it is compared against measuring the same population, over the same time window?
- Has the number been checked again after a full quarter, or is it still the pilot-week figure?
None of these questions require technical expertise. They require the discipline to ask what was left out before approving what was left in. A team reporting AI performance well will show the denominator without being asked. That is worth watching for, alongside the number itself.
The next AI performance report that lands on your desk will likely carry an equally impressive figure. Before it turns into a budget line or a headcount decision, ask what it was divided by.