
A Beirut logistics company had six AI use cases on a whiteboard and budget for one pilot. The value/effort matrix ranked route optimization first. The one that survived six weeks in production was ranked fourth.
I'm describing a composite here, drawn from a pattern across several engagements rather than any single client's story, because the pattern is what's worth passing on. The six candidates were the usual mix: route optimization for the delivery fleet, a WhatsApp bot to answer "where's my order" questions, an internal dispatch-scheduling assistant, an invoice-matching tool for the finance team, a demand-forecasting model for warehouse stock, and a driver-scoring system tied to bonuses. Run through the value/effort matrix every playbook ships with, expected return on one axis, build effort on the other, route optimization won by a wide margin. High value, moderate effort. The obvious first pilot.
It also depended on a continuous GPS feed from a fleet that lost signal outside Beirut's core more often than anyone had bothered to measure, feeding a routing engine whose worst suggestions needed a human to catch before a driver acted on them. Nobody had that human in place in week one. Six weeks in, the pilot was quietly shelved. The one that survived was the invoice-matching tool: ranked fourth, lower headline value, but self-contained, forgiving when it got something wrong, and running on two years of invoice data that was already sitting in the accounting system.
Why the Matrix You Brought From the Playbook Is Missing Two Axes
The standard value/effort matrix isn't wrong. It's incomplete for this operating environment, and the two axes it leaves out are exactly the ones that decide whether a pilot survives contact with production.
The first missing axis is infrastructure dependency. A global playbook is usually written by someone assuming a stable power grid, consistent broadband, and cloud services that rarely hiccup. That assumption quietly changes what "moderate effort" means. A use case that looks technically simple in a stable environment, like real-time GPS-based routing, turns operationally fragile the moment its live-data feed depends on a mobile network that drops out in specific neighborhoods at specific hours. The matrix scores the algorithm. It doesn't score the plumbing underneath it.
The second missing axis is the cost and speed of a wrong output reaching someone who has no way to tell it's wrong. A generic matrix folds risk into effort, if it considers it at all, as one soft variable. It doesn't distinguish between an AI draft a staffer reviews before anyone else sees it and an AI reply that goes straight to a customer over WhatsApp, a channel a lot of businesses in this region run as their primary support line, not a backup one. It doesn't ask whether the business runs on cash-heavy transactions where a pricing or inventory error can't simply be reversed with a chargeback. None of that makes AI unusable here. It means the order you run pilots in matters more than the playbook accounts for.
The Four-Question Triage
Before I rank use cases by value, I run each one through four questions. They don't replace the value/effort matrix; they sit in front of it. A candidate that scores badly on two or more of these should drop several places on the whiteboard, regardless of its projected return.
- Does it depend on infrastructure you don't fully control, running continuously? Scoring cue: batch or on-demand processing scores well; anything requiring an unbroken live feed (GPS, real-time inventory sync, a constant-connectivity assumption) scores poorly until you've checked the feed's actual uptime, not its spec sheet.
- What does a wrong output cost, and how fast does it reach someone outside your company? Scoring cue: an internal draft a staffer reviews scores well; an automated reply, price, or confirmation sent straight to a customer, especially over a channel like WhatsApp, scores poorly until you've built in a pause.
- Is a human already positioned to catch an error before it lands? Scoring cue: if someone already reviews this category of output as part of their existing job, score it well; if the pilot would be the first thing anyone reviews, score it poorly until that role is explicitly assigned.
- Can you start with data you already have today? Scoring cue: two or more years of clean, accessible records scores well; a use case that requires a new data-collection effort before the pilot can even start scores poorly, because you're now running two projects on the budget for one.
Scoring the Whiteboard
Back to the six candidates. Three of them make the shift clear enough that I'll walk each through all four questions, with scores that are illustrative rather than measured against any formal rubric.
Route optimization scored high on value, moderate on effort, and looked like the obvious pick. On the triage: infrastructure dependency, poor, it needed a continuous live GPS feed the fleet couldn't reliably provide. Cost of a wrong output, poor, a bad route reached a driver on the road within minutes, with no review step in between. Human backstop, none existed. Data in hand, actually fine, historical route and delivery-time data went back years. Three of four questions came back poor. The matrix never saw any of that.
The WhatsApp status bot scored respectably on value and looked simple to build. Infrastructure dependency, moderate, it needed reliable order-status data but not a continuous live feed. Cost of a wrong output, poor, a wrong delivery estimate went straight to a customer with no one reading it first. Human backstop, none by default, though one could be added cheaply. Data in hand, good, order records already existed. Two of four questions came back poor, one of them fixable. It ranked below the eventual winner but above route optimization.
Invoice matching scored lowest on headline value of the three, the least exciting pitch in the room. Infrastructure dependency, good, it ran in batches against records already sitting in the accounting system, nothing needed to be live. Cost of a wrong output, good, a mismatch got flagged for a person to review before anything was paid or booked, which was already standard practice. Human backstop, already in place, the finance team already checked invoices as part of their job. Data in hand, good, two years of clean records. Four of four came back good. It was ranked fourth on value. It was the only one of the six still running after six weeks.
What This Changes Before You Ever Talk to a Vendor
None of this replaces work I've written about elsewhere, and I'd rather point to it once than re-argue it here. Once you've picked a use case with this triage, you still have to work out what a fair AI price looks like in a market a lot of global pricing models weren't built for, which I covered in AI Pricing Was Never Built for Beirut. If the pilot touches a second country, and for a lot of companies in this region it will, data residency and compliance terms deserve their own pass before signature, which I walked through in One MENA Deal, Two Compliance Problems. Those are separate decisions from the one this piece is about. The triage here is about which use case earns the right to be the first one at all.
What changes, practically, is the conversation you walk into with a vendor. Instead of opening with "we want to automate X," you open already knowing why X is or isn't the right place to start, and you can ask a vendor to show you, specifically, how their tool handles a dropped connection, a wrong output, and a data gap, rather than taking their roadmap slide at face value. That's a shorter, cheaper conversation than finding out six weeks into a live pilot that the slide was wrong.
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.
For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.