
Four hundred seventy out of five hundred. That was the number in the pilot deck for a proof-of-delivery photo checker: a logistics operation's AI tool built to look at a driver's delivery photo, confirm the package matched the order, and auto-close the ticket. Ninety-four percent accuracy, clean enough that operations signed off on a full rollout. Six weeks after go-live, on a matched sample of five hundred production photos, the same model scored 305. Sixty-one percent.
I want to be precise about what this example is: a composite, built from patterns I see repeatedly across logistics and field-service clients in the region, with numbers chosen to keep the arithmetic clean rather than lifted from one company's live report. But the shape of the gap is exactly what I watch happen, over and over, whenever a company scopes an AI pilot around photos, scans, or any input a human has to physically capture. The mistake is rarely in the model. It's in what the pilot quietly assumed about where its input would come from.
The Number That Looked Ready to Ship
The pilot corpus for this tool was 500 photos supplied by the operations team itself: well-lit, one package per frame, correct orientation, shot with the same two phones on a desk in the warehouse. Against that corpus, the model hit 94%. That number did real work. It justified the build cost, it set the rollout timeline, and it became the line in the board deck that made the case for closing tickets automatically instead of routing every delivery photo to a human reviewer.
None of that reasoning was unsound, given the input the pilot tested against. The problem is a category error that's easy to miss under deadline pressure: a pilot accuracy number describes the model's performance on the channel used to build it. It says nothing, by itself, about the channel the model will actually run against once real staff — not the ops team — are the ones pointing the camera. Ninety-four percent was a true fact about a photo set the company controlled. It was never a forecast about the photo set the company didn't yet have.
What Changed Between the Pilot and the Loading Dock
In production, the photos didn't come from the ops team. They came from drivers, forwarding delivery photos through a messaging app at the end of a shift: blurry from a moving hand, dark because the loading dock light had burned out, held at an angle because the phone was wedged against a clipboard, and occasionally the wrong photo entirely — a fuel receipt, a different stop, a blank frame from a pocket-dial. On a matched 500-photo production sample, the same unmodified model scored 61%: 165 more misses than the pilot number predicted.
The drivers were not wrong. Nobody trained them on photo composition, nobody told them a badly lit shot would silently fail a system they didn't know was scoring their submissions, and forwarding a photo through a messaging app is exactly how they'd been asked to report a delivery for years before any AI touched the workflow. The scoping was wrong. Nobody had defined, before the model was ever trained, what "the input" would actually look like once it left a controlled desk and entered a driver's pocket at the end of a twelve-hour shift. That gap between the channel used to build the system and the channel used to run it is close to the same failure I described in The Messages Your Test Set Never Saw, just showing up in a camera roll instead of a chat transcript.
The Step Nobody Scoped: Intake Normalization
The fix wasn't a bigger model, a retraining run, or a stricter policy memo to drivers. It was a pre-check inserted between "photo taken" and "photo scored": a lightweight screen for blur, crop and orientation, and basic lighting, running before the verification model ever saw the image. A photo that fails the gate is never scored at all. Instead, an automatic message goes straight back to the driver: retake the photo, here's what's wrong with it, resend. Three weeks of engineering time built that gate and its reject-and-resend loop.
On the same production sample, accuracy after the pre-check rose to 89%: 445 of 500. That's still short of the pilot's 94%, and it should be — a loading dock will never be lit like a warehouse desk. But the 11-point residual gap is now a known, monitored quantity that someone can own and improve, not a hidden 33-point miss discovered by customers calling in about tickets closed on bad data. This is the same lesson I keep returning to in why AI agents fail at the last mile: the failure almost never sits in the reasoning core. It sits in the handoff nobody assigned an owner to.
What makes this fix cheap, relative to the pilot itself, is that it never touches the model. Nobody relabeled data, retrained a network, or renegotiated the accuracy target with the board. The engineering effort went entirely into a gate that runs before the model is even called, and a message loop that already existed in a simpler form — drivers were already getting automated texts about their routes. The three weeks bought a known, bounded gap instead of an unbounded one, which is a different kind of engineering problem than "make the model smarter," and usually a cheaper one to solve on a deadline.
The Three-Question Intake Audit
I now run this audit before I trust any pilot accuracy number that involves a photo, scan, or document a field employee or customer has to capture themselves. Three questions, asked before the build is scoped, not after the pilot number is already in a board deck:
- Who captures the input in production, and with what device and behavior? Not who supplied the test set — who will actually be holding the phone, in what conditions, once the tool is live.
- What percentage of real inputs would fail a basic quality gate? Blur, crop, orientation, wrong file entirely. If nobody has measured this on a real production sample, the pilot number is a statement about the wrong population.
- What is the reject-and-resend path, and who owns it on day one? A quality gate with nobody watching the resend queue just adds a second silent failure on top of the first.
None of these questions require a data scientist. They require someone willing to ask what the real intake channel looks like before the accuracy number gets treated as a launch decision. A 94% pilot is not a lie. It's an answer to a narrower question than the one most rollout decisions actually need answered.