
A vendor demo runs for forty minutes. The tool summarizes a curated document, drafts a clean follow-up email, and answers three questions before you have finished forming them. It is flawless. I still would not approve it on the strength of that.
Not because it failed. Because it succeeded on ground the vendor chose. A demo is staged inside the tool's designed use case: clean input, the exact task it was built for, run by someone who already knows where the seams are. My workflow isn't clean. It carries years of exceptions, half-finished context from three days earlier, formatting habits from clients who've never read a style guide, and constraints nobody wrote into a product spec. A demo tells you what a tool can do under ideal conditions. It tells you nothing about what it will cost you once it sits inside a real week.
Why the Demo Is the Wrong Evaluation Environment
The instinct most people follow when a new tool arrives is to hand it the hardest thing on the desk — the complex report, the tricky negotiation email, the edge case that's been sitting unresolved for a month. It feels like the fair test. It isn't. A hard task carries so much variance that almost any output looks impressive next to a blank page, and a single hard task tells you almost nothing about the four hundred easy ones you'll actually feed the tool that quarter.
What replaced the demo, in how I now evaluate tools before they join my workflow, is closer to an audition than a trial. You don't cast someone because they nailed the monologue they rehearsed for six months. You hand them a scene they've never seen and watch what happens once the polish runs out. I run five tests, in this order, before anything earns a permanent seat in my process.
- The Boring-Task Test. Give it your most repeated five-minute task, ten times in a row, and time the whole loop — including your own check.
- The Failure-Signature Test. Hand it a task just outside its stated scope and note whether it flags uncertainty or produces a confident wrong answer.
- The Edit-Distance Test. Count how much of its first output survives, untouched, into your final version.
- The Version-Drift Test. Find out how the tool gets updated, and whether your workflow has any way to notice if its behavior quietly changes.
- The Handback Test. Confirm you can still do the task yourself, cold, if the tool disappeared tomorrow.
Each one catches something the demo is built to hide.
The Boring-Task Test
Hero tasks flatter a tool. Routine tasks expose it. Say you screen inbound leads every morning — a five-minute read, a two-line judgment call, ten times before lunch. That's the task I'd hand a new tool first, not the annual strategy memo. Run it ten times in a row and time the full loop, not just the tool's output time. Include the seconds you spend rereading its answer, catching the one it got wrong, and fixing the phrasing that doesn't sound like you. Most tools that impress on a single hard task lose to a stopwatch on ten boring ones, because the check-and-correct step is where the real cost lives, and demos never include it.
If the tool doesn't beat your current baseline on the boring loop, no amount of brilliance on the hard task matters. The hard task happens once a quarter. The boring task happens every day, which means its price compounds and the hero task's does not.
The Failure-Signature Test
Every tool fails eventually. What matters is how. Hand it something just past the edge of what it's built for — a format it wasn't trained on, a request one step beyond its stated scope — and watch closely. A tool that stops, flags the gap, or says it isn't sure has given you something you can work with. A tool that produces a confident, fluent, wrong answer has given you a landmine with no warning label.
I've stopped ranking tools by accuracy rate alone. A tool that's right ninety percent of the time and honest about the other ten is safer in a live workflow than one that's right ninety-five percent of the time and silent about the rest, because the second one will eventually hand you a wrong answer dressed as a right one, at the exact moment you're least likely to double-check it. Loud failure is survivable. You catch it, you route around it, you keep going. Silent confident failure is disqualifying on its own, independent of how rarely it happens.
The Edit-Distance Test
This is the one number I trust more than any satisfaction score. After a tool produces its first output, I track how much of it survives, word for word, into what I actually ship or send. Not "did it help." Not "was I happy with it." How much of the actual text made it through untouched.
A tool can feel useful and still fail this test badly. It gave you a starting point instead of a blank page, so it felt like a win — but if you rewrote sixty percent of it anyway, you didn't save craft time, you moved it. You traded writing from scratch for editing someone else's draft, and editing someone else's draft is its own skill with its own time cost, one that doesn't always come in cheaper. The tools worth keeping are the ones where the survival ratio climbs the more you use them, because that means the tool is learning your standard, not just producing plausible text near it.
The Version-Drift and Handback Tests
The first three tests measure the tool as it is today. The last two measure whether you can trust it next month, and whether you're still safe if it's gone.
Ask, before you adopt anything, how it gets updated. Some tools change quietly, on a schedule you don't control, with no changelog and no way for you to notice that Tuesday's output quality isn't Monday's. If your workflow has no checkpoint built in — no periodic rerun of the boring-task and edit-distance tests against a fixed baseline — you won't catch drift until it's already cost you a bad week. A process that's brittle to a silent update isn't a stable instrument yet, no matter how well it performed on day one.
The test people skip most is the last one, and it's the one that tells you whether you've built leverage or dependency. Can you still do the task yourself, cold, without the tool, if it vanished tomorrow? Not "would you want to." Could you. If the honest answer is no — if the skill has atrophied past the point of recovery, or the workflow now assumes the tool's existence at a structural level — you haven't adopted a tool. You've outsourced a capability you can no longer reclaim, and you paid nothing for it explicitly, which is exactly why it's easy to miss until the day you need it and it isn't there.
None of this is an argument against adopting AI tools quickly. I adopt them often, and some earn a permanent place within a week. It's an argument against evaluating them the way vendors want you to — on their best day, on their chosen task, with the seams hidden. Run the audition instead. The demo tells you what a tool can do. The audition tells you what it will actually cost you, and that's the number that matters six months in.