
"Did a person check this?"
The email is short. A customer has been quoted an excess amount on a policy query, the figure is wrong, and the reply ends with a single line: "Did a person check this?"
The company's website says that every reply is reviewed by a human. So does its email footer. The honest answer, in the case I want to walk through, is that someone opened the message for about 38 seconds, without the policy document it cited in front of them. Both statements are, in a narrow sense, true. Only one of them would survive being read aloud to the customer.
The group below is a composite: a Gulf-based insurance broker using AI to draft replies to policy queries. Composite; numbers illustrative. They are not research and not a client's data. The structure is what matters, and you would replace the figures with your own.
The arithmetic behind the sentence
The broker sends about 1,200 AI-drafted replies a week, roughly 240 a day. Two reviewers each have about six hours a day for review, because the rest goes to calls, escalations and meetings. That is 720 minutes. Divide 720 by 240 and each reply can receive 3 minutes.
The timestamps tell a different story. The median time from opening a reply to sending it is 38 seconds. Of 1,200 replies in the week, 11 were edited, which is under 1 percent. And the reviewers see the draft but not the policy it cites.
None of this describes careless people. It describes a workload. Two people cannot give 240 items a day the attention the sentence promises, and the interface does not show them the one document that would reveal an error. I would redesign the process before I would criticise a single reviewer.
I have written separately about when a person is owed an explanation that AI was involved. This article is about something narrower: how to describe the review itself, once you have decided you will describe it. For how to carry out a review well, see the three-pass review checklist.
Four tests a review claim must pass
When a company tells me its output is reviewed, I ask four questions. Each can be answered from records, not opinions.
- SAW. Did the reviewer see the full output and the source material it relies on, or only the AI draft? A reviewer who cannot see the policy cannot check the excess.
- TIME. Do the reviewer hours available, divided by the items per day, give enough minutes per item to read it? Compare that figure with the measured open-to-send time.
- POWER. Can the reviewer reject or change an item without penalty and without a second approval? If rejecting costs the reviewer a conversation with a manager, rejection will be rare.
- TRACE. Does a record exist of who reviewed what, when, and what changed? It may cover every item or a defined sample, but it must exist before the dispute, not after.
Four levels, from strongest to weakest claim
The tests produce four levels of claim. I call the set the Review Claim Ladder. Each level has a sentence you may use and a condition you must meet.
- 1. Approved by [named role]. All four tests, passed for every item. Use it only when the named role actually signs each one.
- 2. Checked. SAW, POWER and TRACE for every item, against a defined checklist, with TIME documented. This is the honest word for a disciplined review that is not a sign-off.
- 3. Skimmed. Every item is opened, but there is no checklist. It is a legitimate step in a process, but never call it "reviewed."
- 4. Sampled. A stated fraction is fully reviewed, and the rest is machine-checked on named fields. It covers fewer items, but each claim it makes is one you can prove.
Sampled sits at the bottom because it promises the least coverage, not because it is the weakest practice. A true Sampled claim is stronger than a false Approved one. The rule is simple: publish the claim at the level you pass. To climb a rung, change the staffing or the volume, not the sentence.
The composite, re-run
Run the broker through the tests. SAW fails, because reviewers do not see the policy. TIME fails, because 38 seconds against 3 available minutes shows the minutes are not being used, and under 1 percent edited suggests reading is not what is happening. POWER and TRACE may well pass. With two failures, the honest level is Skimmed at best, and "every reply is reviewed by a human" overstates it.
There are two fixes, and I would do both.
- Put the relevant policy extract beside each draft, and check amounts and names in the draft against the policy record before the reviewer sees it. That repairs SAW and removes the commonest error class.
- Rewrite the claim: "One in five replies is fully checked by a named reviewer, and every reply's figures are verified against the policy record."
Check that the second claim is feasible. One in five of 240 is 48 full reviews a day. At 6 minutes each, that is 288 minutes, against 720 available. It fits, with room for escalations. The new claim is smaller than the old one. It is also one the broker could defend line by line.
A 20-minute self-test
You can run this on your own operation this week, with data you already hold.
- Write down the exact sentence you publish about review: website, contracts, email footers, tender responses.
- Pull last week's reviewer timestamps and compute the median open-to-send time per item.
- Divide reviewer hours spent on review by items reviewed, in minutes. Compare the result with step 2.
- Open three items. Ask whether the reviewer could see the source material the item relies on.
- Ask one reviewer, privately, what happens if they reject an item. Count the approvals required.
- Find your record. Can you show who reviewed one specific item, when, and what changed?
- Match your results to the ladder. Rewrite the sentence at the level you pass, and note what would need to change to reach the next rung.
The closing test is the one I use on every claim. Could you hand your reviewers' timestamps to the customer without flinching? If so, the claim is probably honest. If not, the sentence needs to change before the next complaint arrives. A smaller true claim is a stronger position than a larger one you hope nobody tests.
This is a method for describing a process accurately, not legal or regulatory advice. If a regulator or contract prescribes specific wording, that wording comes first.
Related evidence: The EU AI Act obliges providers of high-risk AI systems to report a serious incident to the market surveillance authorities immediately after establishing a causal link to the system, and in any event not later than 15 days after becoming aware of it — a disclosure deadline fixed in law rather than decided during the incident. (the EU AI Act's 15-day serious-incident reporting deadline)
NIST's AI RMF appendix on human-AI interaction notes that AI systems can autonomously make decisions, defer decision making to a human expert, or be used by a human decision maker as an additional opinion. (NIST AI RMF appendix on human-AI interaction)