Three weeks after launch, the system was still working. Nobody had turned it off. The dashboard still showed green. And underneath that green, a queue of forty-one cases the system could not resolve on its own was sitting untouched, growing by three or four a day, because no one had ever been told it was theirs to clear.
I have seen this pattern repeatedly, in different shapes, across different builds. The demo is flawless. The pilot goes well. Everyone in the room signs off. Then the system goes live against real volume, real edge cases, real customers who do not read the script — and it starts producing outputs a person has to catch, correct, or escalate. That queue is not a bug. It is the normal, expected residue of any system operating in the real world. The question that actually predicts whether an implementation survives is not "does it work in the demo." It is "what happens to the pile it leaves behind."
Most approval conversations never get there. They cover accuracy, cost, timeline, integration. They rarely cover who is standing at that queue in week three, whether anyone can see it filling up, and what happens the day it has to be turned off. Those gaps do not show up as risks on a build document. They show up as a quiet erosion three to six weeks after launch, at which point the fix is expensive and the trust in the system — deserved or not — is already gone.
Below are five questions I ask before I approve any AI build, in plain operator language. They are not about strategy or vendor selection. They are staffing, instrumentation, and process design decisions that have to be made before the build starts, because they are much harder to retrofit after.
Named capacity and a visible number
The first two questions are about whether the system has somewhere to send the work it cannot finish, and whether anyone can see how much of that work there is.
1. Who is the named person — not a team, not "the vendor" — responsible for the cases the system cannot resolve on its own? Not a shared inbox. Not a rotation nobody has actually staffed. A name, with the hours and the capacity to actually clear the volume, sized to what the system will realistically produce at real usage — not at demo usage, which is almost always an order of magnitude lower.
2. Can that person see the exception and override rate directly, without asking anyone for it? If getting that number requires a meeting, a report request, or a favor from engineering, it will not get checked in week three, when everyone is busy with the next thing. It needs to be a number on a screen the responsible person already looks at, updated on its own, with no dependency on someone remembering to pull it.
These are not abstract commitments made in a kickoff meeting. They are staffing lines and dashboard requirements that either exist on day one of real use or do not exist at all, because nobody circles back to add them after launch week ends.
How fast it can stop, and whether people are quietly avoiding it
The next two questions are less about the day-to-day operation and more about the system's relationship to the process around it — how reversible it is, and whether the team is actually using it as designed.
3. If this has to be switched off at 9am on a Tuesday, how long does that actually take, and who has the authority to make that call? Not in theory. In practice — with the specific dependencies this build has created, the data it now holds that nothing else holds, and the process steps that have started to assume it is there. If the honest answer is "we would need a week and three approvals," that is a fact worth knowing before launch, not after something has gone wrong and the fastest fix is blocked by its own rollback plan.
4. Is the team reshaping its process to use this well, or quietly building workarounds to avoid using it — and has anyone actually asked them? Workarounds are usually invisible from above. A team that does not like or trust a new step in its workflow does not file a complaint — it routes around the step, keeps a shadow spreadsheet, or does the task manually and updates the system after the fact to look compliant. The only way to catch this early is to ask, directly and specifically, before it hardens into habit.
The moment the build actually gets tested
Here is where most implementations die, and it is foreseeable well before launch: the first time the system makes a mistake that is visible to a customer, a regulator, or a senior person in the room — and consequential enough that it cannot be quietly corrected after the fact.
This moment is not a low-probability edge case. Given enough real volume, it is close to certain to happen, usually within the first few weeks. What varies enormously is what happens next. In some organizations, the moment is rehearsed: there is a known first move, a known second move, and a person who knows both. In others, it is improvised in real time by whoever happens to be in the room, which produces overreaction, underreaction, or a scramble that damages confidence in the system far more than the original mistake did.
5. What happens the first time it makes a visible, consequential mistake — has that moment been rehearsed, or will it be improvised? Rehearsed does not mean elaborate. It can be a single paragraph: here is who is told first, here is what we say, here is the threshold for pausing the system versus letting it continue while we investigate. Written down before launch, it takes five minutes to read. Improvised in the moment, it can take weeks to recover from.
These five questions, together, form a short checklist worth keeping on hand at approval stage:
- Who is the named person responsible for the cases the system cannot resolve on its own?
- Can that person see the exception and override rate directly, without asking anyone for it?
- How long would it actually take to switch this off at 9am on a Tuesday, and who has the authority to make that call?
- Is the team reshaping its process to use this well, or building workarounds — and has anyone asked them directly?
- Has the first visible, consequential mistake been rehearsed, or will it be improvised?
Use this before the build starts, not after it has already started producing a queue nobody owns. None of these five questions require technical expertise to ask. They require someone in the room willing to ask them before the enthusiasm of the demo carries the decision forward on its own.
