Back to Blog

What Does 'A Human Reviewed It' Actually Have to Mean?

It is a claim a customer or regulator can test. Four tests and four honest levels, so you say only what your reviewers can prove. A composite worked example with numbers.

Direct answer

What Does 'A Human Reviewed It' Actually Have to Mean?

To legitimately claim an output was reviewed by a human, the process must satisfy Dr. Jonah Tebaa's four verifiable tests: SAW, TIME, POWER, and TRACE. In an illustrative composite case, a broker failed because reviewers lacked source policies and averaged thirty-eight seconds per item. Dr. Jonah Tebaa's Review Claim Ladder maps audit records to four tiers: Approved, Checked, Skimmed, or Sampled. Organizations must match public statements to the level they demonstrably pass.

Worn stone steps in darkness. A shaft of light falls on a low step, where a brass stopwatch and a single sheet of paper rest, while a closed, string-tied folder sits unopened in shadow on the top step. Headline: Say only what you can prove.

"Did a person check this?"

The email is short. A customer has been quoted an excess amount on a policy query, the figure is wrong, and the reply ends with a single line: "Did a person check this?"

The company's website says that every reply is reviewed by a human. So does its email footer. The honest answer, in the case I want to walk through, is that someone opened the message for about 38 seconds, without the policy document it cited in front of them. Both statements are, in a narrow sense, true. Only one of them would survive being read aloud to the customer.

The group below is a composite: a Gulf-based insurance broker using AI to draft replies to policy queries. Composite; numbers illustrative. They are not research and not a client's data. The structure is what matters, and you would replace the figures with your own.

The arithmetic behind the sentence

The broker sends about 1,200 AI-drafted replies a week, roughly 240 a day. Two reviewers each have about six hours a day for review, because the rest goes to calls, escalations and meetings. That is 720 minutes. Divide 720 by 240 and each reply can receive 3 minutes.

The timestamps tell a different story. The median time from opening a reply to sending it is 38 seconds. Of 1,200 replies in the week, 11 were edited, which is under 1 percent. And the reviewers see the draft but not the policy it cites.

None of this describes careless people. It describes a workload. Two people cannot give 240 items a day the attention the sentence promises, and the interface does not show them the one document that would reveal an error. I would redesign the process before I would criticise a single reviewer.

I have written separately about when a person is owed an explanation that AI was involved. This article is about something narrower: how to describe the review itself, once you have decided you will describe it. For how to carry out a review well, see the three-pass review checklist.

Four tests a review claim must pass

When a company tells me its output is reviewed, I ask four questions. Each can be answered from records, not opinions.

  • SAW. Did the reviewer see the full output and the source material it relies on, or only the AI draft? A reviewer who cannot see the policy cannot check the excess.
  • TIME. Do the reviewer hours available, divided by the items per day, give enough minutes per item to read it? Compare that figure with the measured open-to-send time.
  • POWER. Can the reviewer reject or change an item without penalty and without a second approval? If rejecting costs the reviewer a conversation with a manager, rejection will be rare.
  • TRACE. Does a record exist of who reviewed what, when, and what changed? It may cover every item or a defined sample, but it must exist before the dispute, not after.

Four levels, from strongest to weakest claim

The tests produce four levels of claim. I call the set the Review Claim Ladder. Each level has a sentence you may use and a condition you must meet.

  • 1. Approved by [named role]. All four tests, passed for every item. Use it only when the named role actually signs each one.
  • 2. Checked. SAW, POWER and TRACE for every item, against a defined checklist, with TIME documented. This is the honest word for a disciplined review that is not a sign-off.
  • 3. Skimmed. Every item is opened, but there is no checklist. It is a legitimate step in a process, but never call it "reviewed."
  • 4. Sampled. A stated fraction is fully reviewed, and the rest is machine-checked on named fields. It covers fewer items, but each claim it makes is one you can prove.

Sampled sits at the bottom because it promises the least coverage, not because it is the weakest practice. A true Sampled claim is stronger than a false Approved one. The rule is simple: publish the claim at the level you pass. To climb a rung, change the staffing or the volume, not the sentence.

The composite, re-run

Run the broker through the tests. SAW fails, because reviewers do not see the policy. TIME fails, because 38 seconds against 3 available minutes shows the minutes are not being used, and under 1 percent edited suggests reading is not what is happening. POWER and TRACE may well pass. With two failures, the honest level is Skimmed at best, and "every reply is reviewed by a human" overstates it.

There are two fixes, and I would do both.

  1. Put the relevant policy extract beside each draft, and check amounts and names in the draft against the policy record before the reviewer sees it. That repairs SAW and removes the commonest error class.
  2. Rewrite the claim: "One in five replies is fully checked by a named reviewer, and every reply's figures are verified against the policy record."

Check that the second claim is feasible. One in five of 240 is 48 full reviews a day. At 6 minutes each, that is 288 minutes, against 720 available. It fits, with room for escalations. The new claim is smaller than the old one. It is also one the broker could defend line by line.

A 20-minute self-test

You can run this on your own operation this week, with data you already hold.

  1. Write down the exact sentence you publish about review: website, contracts, email footers, tender responses.
  2. Pull last week's reviewer timestamps and compute the median open-to-send time per item.
  3. Divide reviewer hours spent on review by items reviewed, in minutes. Compare the result with step 2.
  4. Open three items. Ask whether the reviewer could see the source material the item relies on.
  5. Ask one reviewer, privately, what happens if they reject an item. Count the approvals required.
  6. Find your record. Can you show who reviewed one specific item, when, and what changed?
  7. Match your results to the ladder. Rewrite the sentence at the level you pass, and note what would need to change to reach the next rung.

The closing test is the one I use on every claim. Could you hand your reviewers' timestamps to the customer without flinching? If so, the claim is probably honest. If not, the sentence needs to change before the next complaint arrives. A smaller true claim is a stronger position than a larger one you hope nobody tests.

This is a method for describing a process accurately, not legal or regulatory advice. If a regulator or contract prescribes specific wording, that wording comes first.

Related evidence: The EU AI Act obliges providers of high-risk AI systems to report a serious incident to the market surveillance authorities immediately after establishing a causal link to the system, and in any event not later than 15 days after becoming aware of it — a disclosure deadline fixed in law rather than decided during the incident. (the EU AI Act's 15-day serious-incident reporting deadline)

NIST's AI RMF appendix on human-AI interaction notes that AI systems can autonomously make decisions, defer decision making to a human expert, or be used by a human decision maker as an additional opinion. (NIST AI RMF appendix on human-AI interaction)

Frequently asked questions

What does "every reply is reviewed by a human" have to mean to be a true claim?

It has to pass four tests for every item: the reviewer saw the full output and the source material it relies on, the reviewer had enough minutes to read it, the reviewer could reject or change it without penalty or a second approval, and a record exists of who reviewed it, when, and what changed. If any test fails, the word "reviewed" overstates what happened.

What is the difference between "checked", "skimmed" and "sampled" review?

Checked means every item is reviewed against a defined checklist, with the time documented. Skimmed means every item is opened but there is no checklist, so it should not be called reviewed. Sampled means a stated fraction is fully reviewed and the rest is machine-checked on named fields. Each is honest when it describes what actually happens.

How do I work out whether my reviewers have enough time?

Divide the reviewer hours actually available for review each day, in minutes, by the number of items per day. In my composite example, two reviewers with about six review hours each have 720 minutes for 240 replies, which is 3 minutes per reply. Then compare that with the measured median time from opening an item to sending it.

Is it acceptable to say only a sample of AI output is reviewed by a human?

Yes, provided the claim states the fraction, says who reviews it, and names what the rest is checked against. A claim such as one in five replies fully checked by a named reviewer, with every reply's figures verified against the policy record, is smaller than "every reply is reviewed", but it is true and it can be defended.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.