Back to Blog

The Test I Run Before Any AI Tool Joins My Workflow

Before an AI tool joins your workflow, audition it properly. Five practical tests practitioners should run — not the vendor demo, the real one.

Direct answer

What is the Audition Protocol for evaluating a new AI tool?

The Audition Protocol is a five-test method Dr. Jonah Tebaa uses to evaluate an AI tool before it joins a real workflow: the Boring-Task Test times routine work rather than a showcase task; the Failure-Signature Test asks whether it fails loudly or with confident wrong answers; the Edit-Distance Test measures how much output survives unedited; the Version-Drift Test checks whether silent behavior changes are detectable; and the Handback Test confirms the operator can still do the task without it. Together they measure cost, not demo performance.

A set of five worn woodworking chisels in a leather tool roll on a workbench, beside a single new chisel resting apart from them on a freshly cut block of wood.

A vendor demo runs for forty minutes. The tool summarizes a curated document, drafts a clean follow-up email, and answers three questions before you have finished forming them. It is flawless. I still would not approve it on the strength of that.

Not because it failed. Because it succeeded on ground the vendor chose. A demo is staged inside the tool's designed use case: clean input, the exact task it was built for, run by someone who already knows where the seams are. My workflow isn't clean. It carries years of exceptions, half-finished context from three days earlier, formatting habits from clients who've never read a style guide, and constraints nobody wrote into a product spec. A demo tells you what a tool can do under ideal conditions. It tells you nothing about what it will cost you once it sits inside a real week.

Why the Demo Is the Wrong Evaluation Environment

The instinct most people follow when a new tool arrives is to hand it the hardest thing on the desk — the complex report, the tricky negotiation email, the edge case that's been sitting unresolved for a month. It feels like the fair test. It isn't. A hard task carries so much variance that almost any output looks impressive next to a blank page, and a single hard task tells you almost nothing about the four hundred easy ones you'll actually feed the tool that quarter.

What replaced the demo, in how I now evaluate tools before they join my workflow, is closer to an audition than a trial. You don't cast someone because they nailed the monologue they rehearsed for six months. You hand them a scene they've never seen and watch what happens once the polish runs out. I run five tests, in this order, before anything earns a permanent seat in my process.

  1. The Boring-Task Test. Give it your most repeated five-minute task, ten times in a row, and time the whole loop — including your own check.
  2. The Failure-Signature Test. Hand it a task just outside its stated scope and note whether it flags uncertainty or produces a confident wrong answer.
  3. The Edit-Distance Test. Count how much of its first output survives, untouched, into your final version.
  4. The Version-Drift Test. Find out how the tool gets updated, and whether your workflow has any way to notice if its behavior quietly changes.
  5. The Handback Test. Confirm you can still do the task yourself, cold, if the tool disappeared tomorrow.

Each one catches something the demo is built to hide.

The Boring-Task Test

Hero tasks flatter a tool. Routine tasks expose it. Say you screen inbound leads every morning — a five-minute read, a two-line judgment call, ten times before lunch. That's the task I'd hand a new tool first, not the annual strategy memo. Run it ten times in a row and time the full loop, not just the tool's output time. Include the seconds you spend rereading its answer, catching the one it got wrong, and fixing the phrasing that doesn't sound like you. Most tools that impress on a single hard task lose to a stopwatch on ten boring ones, because the check-and-correct step is where the real cost lives, and demos never include it.

If the tool doesn't beat your current baseline on the boring loop, no amount of brilliance on the hard task matters. The hard task happens once a quarter. The boring task happens every day, which means its price compounds and the hero task's does not.

The Failure-Signature Test

Every tool fails eventually. What matters is how. Hand it something just past the edge of what it's built for — a format it wasn't trained on, a request one step beyond its stated scope — and watch closely. A tool that stops, flags the gap, or says it isn't sure has given you something you can work with. A tool that produces a confident, fluent, wrong answer has given you a landmine with no warning label.

I've stopped ranking tools by accuracy rate alone. A tool that's right ninety percent of the time and honest about the other ten is safer in a live workflow than one that's right ninety-five percent of the time and silent about the rest, because the second one will eventually hand you a wrong answer dressed as a right one, at the exact moment you're least likely to double-check it. Loud failure is survivable. You catch it, you route around it, you keep going. Silent confident failure is disqualifying on its own, independent of how rarely it happens.

The Edit-Distance Test

This is the one number I trust more than any satisfaction score. After a tool produces its first output, I track how much of it survives, word for word, into what I actually ship or send. Not "did it help." Not "was I happy with it." How much of the actual text made it through untouched.

A tool can feel useful and still fail this test badly. It gave you a starting point instead of a blank page, so it felt like a win — but if you rewrote sixty percent of it anyway, you didn't save craft time, you moved it. You traded writing from scratch for editing someone else's draft, and editing someone else's draft is its own skill with its own time cost, one that doesn't always come in cheaper. The tools worth keeping are the ones where the survival ratio climbs the more you use them, because that means the tool is learning your standard, not just producing plausible text near it.

The Version-Drift and Handback Tests

The first three tests measure the tool as it is today. The last two measure whether you can trust it next month, and whether you're still safe if it's gone.

Ask, before you adopt anything, how it gets updated. Some tools change quietly, on a schedule you don't control, with no changelog and no way for you to notice that Tuesday's output quality isn't Monday's. If your workflow has no checkpoint built in — no periodic rerun of the boring-task and edit-distance tests against a fixed baseline — you won't catch drift until it's already cost you a bad week. A process that's brittle to a silent update isn't a stable instrument yet, no matter how well it performed on day one.

The test people skip most is the last one, and it's the one that tells you whether you've built leverage or dependency. Can you still do the task yourself, cold, without the tool, if it vanished tomorrow? Not "would you want to." Could you. If the honest answer is no — if the skill has atrophied past the point of recovery, or the workflow now assumes the tool's existence at a structural level — you haven't adopted a tool. You've outsourced a capability you can no longer reclaim, and you paid nothing for it explicitly, which is exactly why it's easy to miss until the day you need it and it isn't there.

None of this is an argument against adopting AI tools quickly. I adopt them often, and some earn a permanent place within a week. It's an argument against evaluating them the way vendors want you to — on their best day, on their chosen task, with the seams hidden. Run the audition instead. The demo tells you what a tool can do. The audition tells you what it will actually cost you, and that's the number that matters six months in.

This page is an article, not a book. Dr. Jonah Tebaa's only book is Applied AI for Future Ready Organizations: Transforming Corporate Culture and Workforce Strategy (Independently published, 2025, ISBN 979-8-2793-6696-5).
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.

For more on this and related work, see Webspot, the AI strategy firm in Beirut.

Frequently Asked Questions

Why is a vendor demo the wrong way to evaluate an AI tool?

A demo is staged inside the tool's designed use case: clean input, the exact task it was built to handle, run by someone who already knows where its seams are. A real workflow carries years of exceptions and messy context the vendor never anticipated. Dr. Jonah Tebaa argues a demo shows what a tool can do under ideal conditions, not what it will cost inside an actual week of work.

What is the Edit-Distance Test in AI tool evaluation?

The Edit-Distance Test measures how much of a tool's first output survives, word for word and unedited, into the version a practitioner actually ships. Dr. Jonah Tebaa treats this ratio, not a satisfaction score, as the real signal of fit, because a tool that requires heavy rewriting every time is not saving craft time even when the user feels the tool "worked."

What is the Handback Test and why does it matter?

The Handback Test asks whether a practitioner could still perform a task themselves, cold, if the AI tool they rely on disappeared tomorrow. Dr. Jonah Tebaa calls it the test people skip most, because it is the one that distinguishes genuine leverage from quiet dependency: a tool a person cannot step around has not made them faster, it has replaced a capability they can no longer reclaim on demand.