Back to Blog

Not Every AI Task Deserves Your Best Model

Three questions I ask before paying for the top tier, with a composite worked example of sorting eight recurring tasks in one 45-minute session.

Direct answer

How should a business decide which AI model tier each task needs?

Dr. Jonah Tebaa recommends deciding per task with three questions: what a wrong answer costs, whether a competent person can check the output in under a minute, and how often the task runs. Low cost and quick checking means the fast tier; real but checkable means standard; serious or hard to check means the top tier. Volume never promotes a task.

A craftsman's workbench in soft light with three tools of increasing precision, a broad chisel, a fine chisel and a jeweller's loupe, beside small paper tags, most of them placed by the simplest tool and one by the finest.

Here is a composite of a conversation I have had in several forms. A finance lead opens her company's monthly AI bill. In this illustration it comes to 1,900 dollars. She asks where it went. About 1,300 of it, roughly 70 percent, paid for summarising internal meeting notes: hundreds of them, each run on the most capable and most expensive model available. Meanwhile, the one task that needed careful reasoning, comparing a supplier contract against the company's standard terms, had been run once by a busy analyst on the quick, cheap setting. Nobody chose that. It was simply how the tools had been set up.

The bill was not the real finding. The real finding was that the spending and the risk were sitting in different places.

Why "Always Use the Best Model" Feels Safe and Isn't

The instinct is understandable. If the top model costs more, it must be safer, so use it everywhere and stop thinking about it. I hear this from operations leads, marketing heads and founders in firms of twenty to five hundred people, in Lebanon and across the region. Some pay for several seats and default every task to the top tier. Others do the reverse: they put everything on the cheapest option and wonder why quality is uneven.

Both defaults skip the same step, which is deciding by task. The top model costs more per run and usually responds more slowly. For routine work such as summarising notes, tagging tickets or drafting a standard reply, the quality ceiling rarely binds. A fast model already clears the bar, and the extra capability buys nothing you would notice.

The risk sits elsewhere. It sits in a handful of tasks where a subtle error is expensive and hard to spot: long documents, multi-step reasoning, numbers that flow into decisions. Spreading your budget evenly across every task means overspending on the many and, quite often, underprotecting the few.

The Three Questions I Ask

Before I choose a tier for any recurring task, I ask three questions, in this order.

  1. What does a wrong answer cost? Low means someone rereads a paragraph. Real means a customer gets a poor reply or a colleague loses an hour. Serious means a wrong price, clause or figure enters a decision.
  2. Can a competent person check the output in under a minute? A summary of a meeting they attended, yes. A comparison of two forty-page contracts, no.
  3. How many times a month does it run? Ten times, or ten thousand.

The answers map to a simple rule.

  • Fast tier: low cost of error and quickly checkable. This is where volume makes the savings real.
  • Standard tier: real cost of error but checkable, or you are not sure yet.
  • Top tier: serious cost of error, or not quickly checkable. Typical cases are multi-step reasoning, long documents and numbers that feed decisions.

One point is easy to miss. Volume never promotes a task to the top tier. A task that runs five thousand times a month at low stakes is a fast-tier task, and it is also the task where mis-tiering hurts your budget most. Volume raises the price of a wrong choice. It does not change what a wrong answer costs.

A Worked Example: Eight Tasks in 45 Minutes

The following is a composite with illustrative numbers, not a real client. Picture a 60-person distribution firm with eight recurring AI tasks. In one 45-minute session, the operations lead, the finance lead and a customer-service manager sorted them.

  • Summarising internal meeting notes, about 400 a month: fast. A wrong answer is cheap, and the person who attended can check it in seconds.
  • Categorising support tickets, about 1,200 a month: fast. Errors are visible in the queue and easy to correct.
  • Drafting replies to routine order-status emails, about 900 a month: fast, with a person sending each one.
  • Rewriting catalogue descriptions, about 150 a month: standard.
  • Translating customer notices between Arabic and English, about 60 a month: standard, because a bilingual colleague reads each one before it goes out.
  • Drafting weekly sales commentary, four a month: standard.
  • Comparing supplier contract clauses against standard terms, about six a month: top.
  • Reconciling supplier rebate calculations, about three a month: top.

Before the session, monthly spend was 1,900 dollars in this illustration, most of it on the top tier used by default, apart from the contract comparison, which sat on the quick setting. Afterwards, the three high-volume routine tasks moved down to the fast tier, the middle three to standard, and the two consequential tasks moved up. Monthly spend fell to about 700 dollars, a reduction of roughly 63 percent.

The number that mattered more was quality. The contract task now ran on the tier built for long documents and careful reasoning, and the analyst added a step: every flagged clause is read against the source before anyone relies on it. In the illustration, the first month's comparison caught a payment-terms difference that the earlier quick pass had missed. The savings paid for the improvement. Only two tasks, nine runs a month, sat on the most expensive tier, and those were the two where it earned its cost.

How to Run the Sort in One Sitting

  1. List every recurring AI task in the business, with a rough monthly run count. Ask the people doing the work, not just the invoice.
  2. Write next to each one what a wrong answer would cost: low, real or serious.
  3. Mark whether a competent person could verify the output in under a minute. Be honest about the ones that need a careful read.
  4. Assign a tier using the rule above. If you hesitate between two, choose the higher one for now.
  5. Write a one-line upgrade trigger per task, such as "move up if a reviewer finds two material errors in a month."
  6. Name one owner for the list and put a review a month out in the calendar.

Forty-five minutes is realistic for eight to twelve tasks. If it takes much longer, the tasks are probably not defined tightly enough, which is useful to learn in itself.

Where This Goes Wrong

Three mistakes come up repeatedly. The first is tiering by department instead of by task. "Marketing uses the fast tier, finance uses the top tier" sounds tidy, but marketing also drafts pricing pages that feed decisions, and finance also tags hundreds of routine invoices. The unit of decision is the task.

The second is saving pennies on a task whose output feeds a pricing or legal decision. A few dollars saved on a rebate calculation is a poor trade against one wrong figure in a negotiation.

The third is treating the sort as permanent. Models change, so a task that needed the top tier last year may not need it now, and the reverse can happen too. That is why step six has a date attached. I cover how to test a tool properly in The Test I Run Before Any AI Tool Joins My Workflow; here I only note that the sort is a living list, not a ruling.

The Discipline Is Matching Effort to Stakes

I think of this as ordinary management. A good manager does not give every task an hour of their most experienced person's time. They give the routine a quick glance and the contract a careful read. The discipline is the same with AI. Choose the instrument for the job, and keep your best one for the few jobs where a mistake would cost you.

Frequently asked questions

Do I need the most advanced AI model for everyday business tasks?

No. Routine work such as summarising notes, tagging tickets or drafting standard replies rarely reaches the ceiling of a fast model, and a person can check the result in seconds. Paying for the top tier there adds cost and delay without a visible quality gain. Keep the most capable model for the few tasks where a subtle error would be expensive.

What three questions decide which AI model tier a task should use?

Ask what a wrong answer would cost (low, real or serious), whether a competent person can check the output in under a minute, and how many times a month the task runs. Low cost plus quick checking points to the fast tier, real but checkable to standard, and serious or hard to check to the top tier.

When is paying for the top model actually worth it?

It is worth it when a wrong answer would be serious or when nobody can quickly verify the output. Typical cases are comparing long contracts, multi-step reasoning, and calculations whose results feed pricing, legal or financial decisions. In those cases the extra cost per run is small next to the cost of one undetected error.

How often should a team revisit which tasks use which model?

Review the list about a month after the first sort, then every quarter, and whenever a task's stakes change or a reviewer finds repeated errors. Give each task a one-line upgrade trigger so the decision is not left to memory. A list of eight to twelve tasks takes about half an hour to revisit.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.