
By the end of week one, the setup looked exactly like it was supposed to. This is a composite case, built from a pattern I keep meeting across logistics-adjacent operations teams, not one client's literal roster: an operations manager was handed six AI workers, each with its own scope, each with that manager listed as the one named human supervisor the e-mployer's playbook calls for. On paper, the box was checked. By week three, the manager was behind on review for four of the six. Nothing had gone publicly wrong yet, no missed refund, no bad vendor hold, nothing that had reached a director's inbox. The manager was just quietly out of hours, giving two of the six AI workers a real look each week and skimming, or skipping, the other four.
Why "One Named Supervisor" Isn't the Same as "Sustainable"
The e-mployer's playbook I wrote about in June sets a qualitative floor: every e-mployee needs a named human supervisor, a single person accountable for its output, not a shared inbox or a rotating on-call schedule. That rule was satisfied here. The manager's name sat next to all six AI workers on the org chart, and if anyone had audited the structure that week, it would have passed.
But "named" answers a different question than "sustainable." A name on an org chart tells you who is accountable when something goes wrong. It says nothing about whether that person has enough hours in a week to catch the thing before it goes wrong. The unit that actually determines whether a team structure holds is not headcount, human or AI. It is the manager's finite attention, and that attention does not divide evenly across six AI workers just because there happen to be six of them. Some AI workers ask almost nothing of a manager's week. Others ask for most of it. Sizing the ratio like a human org chart, one supervisor, count the reports, misses that entirely.
The Three-Factor Calculation
When I look at why two similarly sized rosters produce completely different manager experiences, three factors are doing all the work.
- Output judgment-variance. How often does producing the output require a real judgment call, versus a repeatable pattern the AI worker can execute the same way every time? An AI worker that classifies invoices against a fixed rule set has low variance. One that decides whether a customer's complaint qualifies for a policy exception has high variance, because the right answer shifts with context a rule can't fully capture.
- Error blast radius. What does a bad output touch before a human catches it? One customer or a whole batch. A refund that can be reversed with a phone call, or a vendor payment that has already left the account. Low blast-radius work tolerates a delayed catch. High blast-radius work does not.
- Required review cadence. Given the first two factors, how often does a manager actually need to look? A weekly spot-check of a sample, or a read of every single output before it goes out the door.
None of the three matters much on its own. A high-variance AI worker with a narrow blast radius, drafting internal talking points a human edits anyway, is manageable with a light spot-check. A low-variance AI worker with a wide blast radius still needs periodic auditing, but not a daily read. It is the combination, variance and blast radius setting the cadence, that produces a rough review-load score: how many minutes of a manager's week a given AI worker actually consumes. That number is the real budget. Two AI workers scoring high on all three factors can consume more of a manager's week than four or five scoring low, and a team-structuring decision that only counts heads will look identical on the org chart whether it is sustainable or not.
Running the Numbers on the Six-Worker Case
Applying this back to the composite case from the opening, the manager had a real weekly review budget of roughly 300 minutes, about an hour a day set aside for checking AI-worker output on top of the rest of the job. Here is what the six AI workers actually demanded:
- Invoice-matching AI worker, low variance, low blast radius: a weekly spot-check of 10 outputs at 2 minutes each, 20 minutes a week.
- Scheduling and dispatch AI worker, low variance, low blast radius: a weekly spot-check of 15 outputs at 2 minutes each, 30 minutes a week.
- Lead-qualification AI worker, medium variance, low blast radius: a daily 10-minute scan, 50 minutes a week.
- Contract-clause-flagging AI worker, medium variance, medium blast radius (feeds a downstream legal review): a daily 15-minute check, 75 minutes a week.
- Refund-approval AI worker, high variance, high blast radius (money leaves the business, hard to claw back): every output read, roughly 40 decisions a week at 3 minutes each, 120 minutes a week.
- Vendor-payment-exception AI worker, high variance, high blast radius (holds and releases vendor payments): every output read, roughly 35 exceptions a week at 3 minutes each, 105 minutes a week.

Total demand: 400 minutes a week against a 300-minute budget. The two low-touch AI workers got their 50 minutes of spot-checking done without a problem. What was left, 250 minutes, had to cover four AI workers that together needed 350. Those four, lead-qualification, contract-clause-flagging, refund-approval, and vendor-payment-exception, were exactly the ones sliding by week three. Notice, too, that the refund-approval and vendor-exception AI workers alone consumed 225 minutes, three-quarters of the entire budget and more than the other four combined. Two AI workers, not six, were the actual capacity problem.
The fix was not to remove an AI worker or hire a full-time reviewer. It was to split the roster: one supervisor keeps the two high-variance, high-blast-radius AI workers as a near-full-time review load, 225 of a 300-minute budget, with a little room to spare, and a second supervisor takes the other four, 175 minutes, with enough slack left to absorb a fifth low-variance worker later. Where a second supervisor was not available, narrowing the refund-approval AI worker's authority, auto-approving only small, low-risk refunds and escalating the rest, cut its review time from 120 minutes to under 40.
What This Means When You're Structuring the Team
If you are about to assign a second, third, or sixth AI worker to an existing manager's roster, calculate the ratio before you do it, not after the review debt has piled up. Score the new AI worker on judgment-variance and blast radius, estimate the review cadence that combination actually requires, and add the resulting minutes to what the manager is already carrying against their real weekly budget.
Two decisions fall out of that number. If the total stays inside the manager's budget with reasonable slack, the existing structure holds and there is nothing else to do. If it does not, there are exactly two honest levers: add a second supervisor and split the roster along variance and blast radius, putting the hardest AI workers together rather than spreading them evenly for the sake of a tidy org chart, or narrow the new AI worker's scope so its blast radius shrinks and it needs less of a human's attention to run safely. Adding headcount and narrowing scope are both legitimate answers. Picking neither, and hoping the manager finds the hours, is how a review backlog builds quietly until it surfaces as a customer complaint.
Before you add your next AI worker to a manager's roster, score it on judgment-variance and blast radius first, then check the number against what that manager already carries. A named supervisor is the floor the e-mployer's playbook requires. A calculated ratio is what keeps that name meaning something three weeks in, not just on the day the org chart gets approved. For the fuller doctrine this sits inside, see the e-mployee framework at jonahtebaa.com/e-mployees.
For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.