Back to Blog

Two AI Pilots Missed the Same Target. One Cost $310,000 More.

A composite six-initiative portfolio, two pilots, one identical missed target — and the single design choice that decided whether the miss cost two weeks or $310,000 more.

Direct answer

What makes an AI stop rule enforceable rather than just written down?

A stop rule only works if four things are fixed before funding moves: a numeric threshold, an independent reviewer who is not the sponsor, a review date that cannot slide, and a pre-named destination for the freed capital. In a composite case, the pilot missing one of these ran five months and $310,000 past its own stated target before being killed anyway; the pilot with all four in place was stopped within two weeks and freed $254,000 for reallocation.

Deep-navy engraving-style illustration of twin stone sluice gates. The left gate is closed, its channel dry below, with a hand in a white cuff and navy sleeve gripping the lowered lever. The right gate is open, its lever upright and untouched, with liquid-gold water pouring out of frame.

Two pilots sat in the same quarterly portfolio review this spring, part of the same six-initiative program, and missed the identical numeric target in the identical month. One was dead within two weeks. The other kept spending for five more months and cost $310,000 more before it was killed anyway - at a target it still hadn't hit.

Everything that follows is a composite, drawn from patterns I have watched repeat across portfolio reviews in my own work - not one client's file, not one company's ledger. I am labeling every dollar figure that way, up front, because I don't want a specific number mistaken for a case study with a name attached. The pattern is the part worth your time.

The portfolio: six AI initiatives, $2.4 million in annual capital, reviewed quarterly the way most boards I work with already run these things. Two of the six hit an unblock-rate criterion below 15 percent - the number both teams had written into their own launch plans as the point past which the pilot was supposed to stop, not a vague sense that things weren't working.

In the first, the sponsor who had championed the funding also held the review seat - the person who got to decide whether the missed number meant stop or one more quarter. At month four, the pilot registered 12 percent against its 15 percent floor. The sponsor asked for, and got, more time. Five months later it was still under its own target - 14 percent, an improvement that never crossed the line - and $310,000 in additional spend had gone into a pilot everyone eventually agreed should have stopped at month four.

In the second, the review seat belonged to the portfolio office - a function with no stake in whether this particular initiative succeeded. It missed the same criterion at month four. The review happened, the number was checked against the plan, and the pilot was killed within two weeks. Total spend: $86,000 against a $340,000 budget. The $254,000 that stayed unspent had already been named, before the pilot ever launched, as capital for two other initiatives in the same portfolio - both of which shipped by Q3.

The Number Was Never the Weak Point

Both pilots had a criterion that was specific, numeric, and written down before launch. Neither failure traces back to a vague stop rule - the kind that says "if it's not working out" and leaves everyone guessing what that means later. Both boards had done that part of the work correctly.

What differed was who sat in the review seat when the number came in short. In the expensive case, the person being asked to call time on the pilot was the same person who had built the funding case for it in the first place. That is not a criterion failure. It is a design failure in who decides - because it puts one person in the position of ruling against their own project, using a number they wrote themselves, on a schedule they control. Almost nobody does that well, and the ones who do are the exception you should not be designing your portfolio around.

Four Things Fixed Before the First Dollar Moves

None of this requires new software or a new committee structure. It requires four decisions made before funding is released, not after the pilot starts to wobble.

  1. Write the stop rule as a number, not a sentiment. "Unblock rate below 15 percent by month four" survives a review meeting. "If it's not working out" does not - it just moves the argument to whoever happens to be in the room.
  2. Name an independent reviewer for the review seat before launch. A portfolio office, a CFO's team, or a rotating peer sponsor with zero stake in this specific initiative's outcome. Never the sponsor, never the executive champion. The review seat is a role, and it has to sit outside the initiative it reviews.
  3. Set a hard calendar review date immune to "we're almost there." Fix it at approval, not at the review meeting itself, where the pilot's own momentum will always argue for one more quarter.
  4. Name the reallocation destination for freed capital in advance. When a pilot's budget already has a next home named before it's killed, ending it reads as a funding decision - not as a loss written off.
Line diagram of a closed sluice gate held shut by four numbered bolts: a numeric threshold, an independent reviewer, a fixed review date, and a named destination for freed capital.
A stop rule holds only when all four are fixed before funding moves.

Why Boards Resist Handing the Review Seat to Someone Else

The most common objection I hear is that naming an outside reviewer looks like a judgment on the sponsor - as though the board doesn't believe in the case they made. I think that reading has it backward. Leaving the review seat with the sponsor does not protect them; it puts them in an impossible position, asking them to rule against the project they staked their credibility on, using their own number, on their own schedule. The five-month, $310,000 drift in the costly pilot above is not a story about a bad sponsor. It is what happens to a good one, structurally, when nobody separates the case for funding from the decision to stop.

In markets across the Gulf and the Levant, where many of the boards I sit across from remain close to their founding families, the instinct is to keep review inside relationships that already work rather than hand it to a function that feels procedural. That instinct is understandable. It is also exactly the setup that produced the expensive pilot, not the cheap one.

The point of getting this right is not defense. A portfolio that kills its underperforming pilots on schedule, with the freed capital already named for its next stop, is a portfolio that funds its next wave on time. The $254,000 freed within two weeks in the composite case did not sit idle waiting for a committee to decide what to do with it - it went straight to two initiatives that shipped by Q3, on capital that would otherwise have stayed locked in a pilot already past its own target. Killing well is not the opposite of building. It is how a portfolio keeps building.

Before your next AI pilot goes to committee for funding, check whether it has a named reviewer sitting outside the sponsor's chain, a number instead of a sentiment, a fixed date, and a destination for the money if it stops. If any one of those four is missing, the stop rule on the page is a paragraph - not something that will actually happen.

For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.

Frequently asked questions

What is a composite case, and why label the dollar figures that way?

A composite case draws a pattern from multiple reviews I've watched rather than reporting one client's numbers. I label every dollar figure in the two-pilot example as composite and illustrative because the specific figures could otherwise be mistaken for one company's file, when the useful part is the pattern: an identical missed target producing two very different outcomes depending on who reviewed it.

Why does who reviews a missed target matter more than the target itself?

In the composite case, both pilots had a specific numeric stop rule written before launch, so the criterion itself was not the failure. What differed was who held the review seat when the number came up short. When the sponsor who built the funding case also decides whether to call the pilot a failure, the review has a structural bias toward one more quarter, regardless of how well the number was written.

Who should hold the review seat if not the sponsor?

Someone with no stake in that specific initiative's outcome: a portfolio office, a CFO's team, or a rotating peer sponsor reviewing an initiative they did not champion. The requirement is not seniority. It is having nothing to lose personally if the pilot is declared a failure.

How does naming a reallocation destination in advance change the outcome of killing a pilot?

In the composite case, the pilot killed within two weeks freed $254,000 that had already been named, before launch, as capital for two other initiatives - both of which shipped by Q3. Naming the destination in advance turns the kill decision into a funding decision the board can see, rather than a loss that has to be absorbed and re-argued after the fact.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.