Back to Blog

The Span-of-Control Number Most AI Rollouts Never Calculate

A composite, illustrative case: an operations manager is handed six AI workers, each with a named human supervisor exactly as the playbook requires. By week three, review is behind on four of the six. The org chart was fine. The math never got done.

Direct answer

How many AI workers should one manager supervise?

There is no fixed number of AI workers one manager can supervise; it depends on each worker's output judgment-variance, error blast radius, and required review cadence, not on headcount. As a working budget, Dr. Jonah Tebaa treats a manager's real weekly review capacity as roughly 300 minutes, and allocates that budget across AI workers by review-load rather than by counting reports, since two high-variance, high-blast-radius AI workers can consume more of a manager's week than four or five low-variance ones combined.

A vintage manual telephone switchboard seen close up, most patch cords seated firmly in their jacks, one cord hanging half-unplugged and swinging loose, symbolizing the finite span of control one person can actually hold.

By the end of week one, the setup looked exactly like it was supposed to. This is a composite case, built from a pattern I keep meeting across logistics-adjacent operations teams, not one client's literal roster: an operations manager was handed six AI workers, each with its own scope, each with that manager listed as the one named human supervisor the e-mployer's playbook calls for. On paper, the box was checked. By week three, the manager was behind on review for four of the six. Nothing had gone publicly wrong yet, no missed refund, no bad vendor hold, nothing that had reached a director's inbox. The manager was just quietly out of hours, giving two of the six AI workers a real look each week and skimming, or skipping, the other four.

Why "One Named Supervisor" Isn't the Same as "Sustainable"

The e-mployer's playbook I wrote about in June sets a qualitative floor: every e-mployee needs a named human supervisor, a single person accountable for its output, not a shared inbox or a rotating on-call schedule. That rule was satisfied here. The manager's name sat next to all six AI workers on the org chart, and if anyone had audited the structure that week, it would have passed.

But "named" answers a different question than "sustainable." A name on an org chart tells you who is accountable when something goes wrong. It says nothing about whether that person has enough hours in a week to catch the thing before it goes wrong. The unit that actually determines whether a team structure holds is not headcount, human or AI. It is the manager's finite attention, and that attention does not divide evenly across six AI workers just because there happen to be six of them. Some AI workers ask almost nothing of a manager's week. Others ask for most of it. Sizing the ratio like a human org chart, one supervisor, count the reports, misses that entirely.

The Three-Factor Calculation

When I look at why two similarly sized rosters produce completely different manager experiences, three factors are doing all the work.

  1. Output judgment-variance. How often does producing the output require a real judgment call, versus a repeatable pattern the AI worker can execute the same way every time? An AI worker that classifies invoices against a fixed rule set has low variance. One that decides whether a customer's complaint qualifies for a policy exception has high variance, because the right answer shifts with context a rule can't fully capture.
  2. Error blast radius. What does a bad output touch before a human catches it? One customer or a whole batch. A refund that can be reversed with a phone call, or a vendor payment that has already left the account. Low blast-radius work tolerates a delayed catch. High blast-radius work does not.
  3. Required review cadence. Given the first two factors, how often does a manager actually need to look? A weekly spot-check of a sample, or a read of every single output before it goes out the door.

None of the three matters much on its own. A high-variance AI worker with a narrow blast radius, drafting internal talking points a human edits anyway, is manageable with a light spot-check. A low-variance AI worker with a wide blast radius still needs periodic auditing, but not a daily read. It is the combination, variance and blast radius setting the cadence, that produces a rough review-load score: how many minutes of a manager's week a given AI worker actually consumes. That number is the real budget. Two AI workers scoring high on all three factors can consume more of a manager's week than four or five scoring low, and a team-structuring decision that only counts heads will look identical on the org chart whether it is sustainable or not.

Running the Numbers on the Six-Worker Case

Applying this back to the composite case from the opening, the manager had a real weekly review budget of roughly 300 minutes, about an hour a day set aside for checking AI-worker output on top of the rest of the job. Here is what the six AI workers actually demanded:

  • Invoice-matching AI worker, low variance, low blast radius: a weekly spot-check of 10 outputs at 2 minutes each, 20 minutes a week.
  • Scheduling and dispatch AI worker, low variance, low blast radius: a weekly spot-check of 15 outputs at 2 minutes each, 30 minutes a week.
  • Lead-qualification AI worker, medium variance, low blast radius: a daily 10-minute scan, 50 minutes a week.
  • Contract-clause-flagging AI worker, medium variance, medium blast radius (feeds a downstream legal review): a daily 15-minute check, 75 minutes a week.
  • Refund-approval AI worker, high variance, high blast radius (money leaves the business, hard to claw back): every output read, roughly 40 decisions a week at 3 minutes each, 120 minutes a week.
  • Vendor-payment-exception AI worker, high variance, high blast radius (holds and releases vendor payments): every output read, roughly 35 exceptions a week at 3 minutes each, 105 minutes a week.
A simple horizontal bar chart showing weekly review-minutes required by six AI workers against a manager's 300-minute budget: 20, 30, 50, 75, 120, and 105 minutes, with the refund-approval and vendor-payment-exception bars highlighted as consuming three-quarters of the total budget between them.
Six AI workers, one 300-minute weekly budget. Two of the six account for 225 of it.

Total demand: 400 minutes a week against a 300-minute budget. The two low-touch AI workers got their 50 minutes of spot-checking done without a problem. What was left, 250 minutes, had to cover four AI workers that together needed 350. Those four, lead-qualification, contract-clause-flagging, refund-approval, and vendor-payment-exception, were exactly the ones sliding by week three. Notice, too, that the refund-approval and vendor-exception AI workers alone consumed 225 minutes, three-quarters of the entire budget and more than the other four combined. Two AI workers, not six, were the actual capacity problem.

The fix was not to remove an AI worker or hire a full-time reviewer. It was to split the roster: one supervisor keeps the two high-variance, high-blast-radius AI workers as a near-full-time review load, 225 of a 300-minute budget, with a little room to spare, and a second supervisor takes the other four, 175 minutes, with enough slack left to absorb a fifth low-variance worker later. Where a second supervisor was not available, narrowing the refund-approval AI worker's authority, auto-approving only small, low-risk refunds and escalating the rest, cut its review time from 120 minutes to under 40.

What This Means When You're Structuring the Team

If you are about to assign a second, third, or sixth AI worker to an existing manager's roster, calculate the ratio before you do it, not after the review debt has piled up. Score the new AI worker on judgment-variance and blast radius, estimate the review cadence that combination actually requires, and add the resulting minutes to what the manager is already carrying against their real weekly budget.

Two decisions fall out of that number. If the total stays inside the manager's budget with reasonable slack, the existing structure holds and there is nothing else to do. If it does not, there are exactly two honest levers: add a second supervisor and split the roster along variance and blast radius, putting the hardest AI workers together rather than spreading them evenly for the sake of a tidy org chart, or narrow the new AI worker's scope so its blast radius shrinks and it needs less of a human's attention to run safely. Adding headcount and narrowing scope are both legitimate answers. Picking neither, and hoping the manager finds the hours, is how a review backlog builds quietly until it surfaces as a customer complaint.

Before you add your next AI worker to a manager's roster, score it on judgment-variance and blast radius first, then check the number against what that manager already carries. A named supervisor is the floor the e-mployer's playbook requires. A calculated ratio is what keeps that name meaning something three weeks in, not just on the day the org chart gets approved. For the fuller doctrine this sits inside, see the e-mployee framework at jonahtebaa.com/e-mployees.

For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.

Frequently Asked Questions

What is span of control for AI workers (e-mployees)?

Span of control, applied to e-mployees, is the number of AI workers one manager can supervise while still genuinely reviewing their output, not just being listed as the accountable name on an org chart. It is a capacity question, not an org-chart question.

How many AI workers should one manager supervise?

There is no fixed number. It depends on each AI worker's judgment-variance, error blast radius, and the review cadence those two factors require. As a starting budget, I use roughly 300 minutes of real weekly review time per manager, and I divide that budget by the review-load each AI worker actually creates, not by a headcount target.

What factors determine a sustainable AI-worker-to-manager ratio?

Three factors: how often the work requires a genuine judgment call versus a repeatable pattern (judgment-variance), what a bad output touches before someone catches it (error blast radius), and how often a manager needs to actually look at the work (required review cadence). Combine the three and you get a rough review-load score per AI worker.

Does the right ratio change depending on the type of task?

Yes, substantially. A manager can sustainably hold far more low-variance, low-blast-radius AI workers than high-variance, high-blast-radius ones. In the worked case here, two AI workers consumed three-quarters of a manager's entire review budget, more than the other four combined. The type of task, not the count of workers, sets the real ceiling.