Back to Blog

The Pilot Photos Were Clean. The Real Ones Never Are.

A composite worked example of an AI pilot that scored 94%, a production channel that dropped it to 61%, and the three-question audit that would have caught the gap before launch.

Direct answer

Why does an AI pilot's accuracy drop sharply once it reaches production?

In a composite case, an AI delivery-photo checker dropped from 94 percent pilot accuracy to 61 percent in production because controlled test photos ignored real field conditions like blur, darkness, and angles. Dr. Jonah Tebaa resolves this intake scoping failure through intake normalization: a lightweight pre-check screening blur, crop, and lighting before scoring, paired with an automated reject-and-resend loop. Tebaa enforces this with a three-question intake audit evaluating production capture conditions, failure rates, and resend ownership.

A steel conveyor carries crisp photo prints of parcels on doorsteps through a softly lit inspection arch, while a gloved hand lifts aside one creased, rain-spotted, blurred print.

Four hundred seventy out of five hundred. That was the number in the pilot deck for a proof-of-delivery photo checker: a logistics operation's AI tool built to look at a driver's delivery photo, confirm the package matched the order, and auto-close the ticket. Ninety-four percent accuracy, clean enough that operations signed off on a full rollout. Six weeks after go-live, on a matched sample of five hundred production photos, the same model scored 305. Sixty-one percent.

I want to be precise about what this example is: a composite, built from patterns I see repeatedly across logistics and field-service clients in the region, with numbers chosen to keep the arithmetic clean rather than lifted from one company's live report. But the shape of the gap is exactly what I watch happen, over and over, whenever a company scopes an AI pilot around photos, scans, or any input a human has to physically capture. The mistake is rarely in the model. It's in what the pilot quietly assumed about where its input would come from.

The Number That Looked Ready to Ship

The pilot corpus for this tool was 500 photos supplied by the operations team itself: well-lit, one package per frame, correct orientation, shot with the same two phones on a desk in the warehouse. Against that corpus, the model hit 94%. That number did real work. It justified the build cost, it set the rollout timeline, and it became the line in the board deck that made the case for closing tickets automatically instead of routing every delivery photo to a human reviewer.

None of that reasoning was unsound, given the input the pilot tested against. The problem is a category error that's easy to miss under deadline pressure: a pilot accuracy number describes the model's performance on the channel used to build it. It says nothing, by itself, about the channel the model will actually run against once real staff — not the ops team — are the ones pointing the camera. Ninety-four percent was a true fact about a photo set the company controlled. It was never a forecast about the photo set the company didn't yet have.

What Changed Between the Pilot and the Loading Dock

In production, the photos didn't come from the ops team. They came from drivers, forwarding delivery photos through a messaging app at the end of a shift: blurry from a moving hand, dark because the loading dock light had burned out, held at an angle because the phone was wedged against a clipboard, and occasionally the wrong photo entirely — a fuel receipt, a different stop, a blank frame from a pocket-dial. On a matched 500-photo production sample, the same unmodified model scored 61%: 165 more misses than the pilot number predicted.

The drivers were not wrong. Nobody trained them on photo composition, nobody told them a badly lit shot would silently fail a system they didn't know was scoring their submissions, and forwarding a photo through a messaging app is exactly how they'd been asked to report a delivery for years before any AI touched the workflow. The scoping was wrong. Nobody had defined, before the model was ever trained, what "the input" would actually look like once it left a controlled desk and entered a driver's pocket at the end of a twelve-hour shift. That gap between the channel used to build the system and the channel used to run it is close to the same failure I described in The Messages Your Test Set Never Saw, just showing up in a camera roll instead of a chat transcript.

The Step Nobody Scoped: Intake Normalization

The fix wasn't a bigger model, a retraining run, or a stricter policy memo to drivers. It was a pre-check inserted between "photo taken" and "photo scored": a lightweight screen for blur, crop and orientation, and basic lighting, running before the verification model ever saw the image. A photo that fails the gate is never scored at all. Instead, an automatic message goes straight back to the driver: retake the photo, here's what's wrong with it, resend. Three weeks of engineering time built that gate and its reject-and-resend loop.

On the same production sample, accuracy after the pre-check rose to 89%: 445 of 500. That's still short of the pilot's 94%, and it should be — a loading dock will never be lit like a warehouse desk. But the 11-point residual gap is now a known, monitored quantity that someone can own and improve, not a hidden 33-point miss discovered by customers calling in about tickets closed on bad data. This is the same lesson I keep returning to in why AI agents fail at the last mile: the failure almost never sits in the reasoning core. It sits in the handoff nobody assigned an owner to.

What makes this fix cheap, relative to the pilot itself, is that it never touches the model. Nobody relabeled data, retrained a network, or renegotiated the accuracy target with the board. The engineering effort went entirely into a gate that runs before the model is even called, and a message loop that already existed in a simpler form — drivers were already getting automated texts about their routes. The three weeks bought a known, bounded gap instead of an unbounded one, which is a different kind of engineering problem than "make the model smarter," and usually a cheaper one to solve on a deadline.

The Three-Question Intake Audit

I now run this audit before I trust any pilot accuracy number that involves a photo, scan, or document a field employee or customer has to capture themselves. Three questions, asked before the build is scoped, not after the pilot number is already in a board deck:

  • Who captures the input in production, and with what device and behavior? Not who supplied the test set — who will actually be holding the phone, in what conditions, once the tool is live.
  • What percentage of real inputs would fail a basic quality gate? Blur, crop, orientation, wrong file entirely. If nobody has measured this on a real production sample, the pilot number is a statement about the wrong population.
  • What is the reject-and-resend path, and who owns it on day one? A quality gate with nobody watching the resend queue just adds a second silent failure on top of the first.

None of these questions require a data scientist. They require someone willing to ask what the real intake channel looks like before the accuracy number gets treated as a launch decision. A 94% pilot is not a lie. It's an answer to a narrower question than the one most rollout decisions actually need answered.

Frequently asked questions

If the model itself didn't change, why did accuracy drop so much?

Because accuracy is never a property of the model alone — it's a property of the model and the input distribution together. In this composite example, the same weights that scored 94% (470 of 500) against pilot photos scored 61% (305 of 500) against production photos. Nothing about the model degraded. The input it was being asked to read had simply changed shape: different lighting, different angles, different devices. Swap the input distribution and you swap the accuracy number, even with an unchanged model.

What does an intake normalization pre-check actually do?

It sits between the moment a photo is captured and the moment it reaches the scoring model. It checks basic quality signals — blur, crop and orientation, lighting — before anything is scored. A photo that fails the gate never reaches the model at all; instead, an automatic reject-and-resend prompt goes back to the driver asking for another shot. In the composite example, adding this single step brought production accuracy from 61% (305 of 500) to 89% (445 of 500), without retraining the underlying model at all.

Why didn't the pre-check close the entire gap back to 94%?

Because a loading dock at 6 a.m. is a permanently harder environment than a curated pilot set, and normalization can only screen out the worst inputs, not manufacture pilot-grade lighting on a moving dock. The residual 11-point gap in the composite example is not a failure of the fix; it's the honest cost of real conditions. The win isn't closing the gap to zero. It's converting a hidden 33-point miss into a known, monitored 11-point one that someone owns.

Who should own the reject-and-resend path once it exists?

Not the team that built the model. The reject-and-resend loop touches the driver relationship in real time, so it belongs with whoever already owns that relationship on day one — dispatch or field operations, not engineering. If nobody is named as the owner before launch, the resend prompts pile up unanswered, and the pre-check quietly turns into a second silent failure sitting on top of the first one.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.