Back to Blog

The AI Accountability Trap: What Executives Must Never Delegate

Most AI failures are not technical failures. They are governance failures — moments where no human in authority had made an explicit decision about what the system was and was not allowed to do.

Direct answer

What can executives never delegate to an AI system?

Three judgments can never be delegated to an AI system or the team running it: which actions the system may take without human review, what counts as acceptable failure and at what error rate, and who owns the judgment that the system's objective is still the right one. Each is a values decision requiring executive judgment, not technical expertise. Visibility is not oversight — a usage dashboard reports what happened, not whether it was what you would have chosen. Dr. Jonah Tebaa calls this the AI accountability trap.

An executive hand pressing deliberately on a dark conference table — the gesture of a decision being made

The executives most exposed to AI risk are not the ones who said no.

They are the ones who said yes — approved the budget, endorsed the roadmap, celebrated the launch — and then handed it off.

I have seen this pattern more than once in my work with senior leaders across the region. A decision-maker does everything right on paper: hires capable people, brings in reputable vendors, allocates real resources. Then they step back, as good leaders do, and let their team execute.

The problem is not the stepping back. The problem is what stays behind when they do.

Delegation is a leadership skill. But there is a version of AI delegation that is not leadership — it is exposure. When you delegate AI execution without retaining judgment over its scope, its risk tolerance, and its failure modes, you have not handed off a task. You have handed off accountability. And accountability, unlike a project plan, does not transfer cleanly.

Decision One: What AI Is Permitted to Act On Without Human Review

Every AI system operates within a set of boundaries. The question is who defined them.

In most organizations, those boundaries were drawn by the team closest to the implementation — the product owners, the architects, the vendors who scoped the use cases. These are smart, capable people. But they are not you. They do not carry the same view of what a mistake costs.

When a system is permitted to act autonomously — to send a communication, make a recommendation that drives a business process, filter a decision — someone made a judgment call about where human review was necessary and where it was not. If you did not make that call explicitly, it was made without you.

This is not a technical question. It is a governance question. It belongs at your level.

Decision Two: What Constitutes Acceptable Failure

AI systems fail. This is not a defect — it is a property. The question that matters is not whether the system will make errors, but what kind of errors are acceptable and at what rate.

A system that occasionally surfaces the wrong content recommendation is a different problem from one that occasionally misclassifies a customer risk profile. Both are failures. The tolerance for each is radically different. And the tolerance threshold is a values decision, not a technical one.

I have sat in rooms where this conversation was avoided because it felt premature — "let's see how it performs first." That is a reasonable instinct in a pilot. It is not a governance posture for a deployed system. By the time you are watching it perform, the failure mode you did not define has already been implicitly accepted.

Someone has to draw this line. That someone is you.

Decision Three: Who Is Accountable When the System Optimizes for the Wrong Thing

This is the question most organizations have not answered, because it is uncomfortable.

AI systems optimize for what they are trained and instructed to optimize for. Sometimes the objective was specified correctly. Sometimes it was not. Sometimes it was correct at the time and is now misaligned with where the business has moved. When that misalignment surfaces — in a customer outcome, a regulatory conversation, a press inquiry — accountability does not flow to the algorithm.

It flows upward.

The question is not just who owns the system. It is who owns the judgment that the system's objective is still the right one. That review cannot live permanently at the technical layer. At some cadence, at some altitude, someone with decision authority has to look at what the system is optimizing for and confirm that it matches the organization's actual intent.

If that person is not you, you have delegated something that cannot be delegated.

The Law Now Names a Person, and It Has a Date

When I first wrote this, the argument stood on judgment alone. It no longer has to. For a large class of AI systems, European law now requires exactly what this article asks of you, and it requires it of a named human being rather than a department.

Article 26(2) of the EU AI Act is one sentence long: “Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support.” Read the nouns slowly. Not competence alone, which most organizations can claim. Authority — meaning the person assigned to oversee the system must be able to stop it. If you have named a reviewer who would need to escalate three levels before the system could be paused, you have satisfied the org chart and not the obligation.

The date moved, and it moved in a way that is easy to misread as a reprieve. Article 26 sits in Chapter III, Section 3 of the Act, and the Digital Omnibus on AI — Regulation (EU) 2026/1744 — rewrote the application clock for that whole chapter. Sections 1, 2 and 3 of Chapter III now apply from 2 December 2027 for systems classified as high-risk under Article 6(2) and Annex III, and from 2 August 2028 for the product-embedded systems under Article 6(1). The original date was 2 August 2026. The reason given in the recitals is administrative, not substantive: national competent authorities were not standing up fast enough, and the Commission judged that enforcing on the old date would raise implementation costs without a matching benefit.

Two things follow, and they point in opposite directions. The first is that nobody got a lighter obligation — the duty to name a competent, trained, empowered human is unchanged in substance, only later in application. The second is that the prohibitions never moved: Chapters I and II have applied since 2 February 2025, with a narrow set of new provisions landing 2 December 2026. A delay in the high-risk regime is not a delay in the parts of the Act that say certain uses are simply off the table.

If you are reading this from Beirut, Dubai or Riyadh, the natural reaction is that none of it binds you. Often it does not. It reaches you through customers: if your system touches EU users, or sits inside the supply chain of a company that deploys in the EU, the obligation arrives contractually long before any regulator does. In practice I have watched that clause show up in a procurement questionnaire well ahead of any legal analysis — which means the operational question is not are we in scope but can we answer the question when a client asks it. Three decisions, written down, with a name attached to each. That is the answer, and it is the same answer whether the pressure comes from a regulator, a customer or your own board.

Attaching a name is where most organizations stall, because the role Article 26(2) describes is one almost nobody has hired for. I wrote about that separately: why AI deployment creates a job nobody posts — the ownership gap between the team that builds the system and the person answerable for what it does.

The Reframe: Literacy, Not Micromanagement

None of this requires you to understand the model architecture or review the training data. That is not the ask.

The ask is that you be literate enough to make these three decisions deliberately — and to make them again when the context changes. What is in scope for autonomous action. What failure looks like and at what threshold it becomes unacceptable. Who answers for the objective.

These are strategic decisions. They require executive judgment. They do not require technical expertise. International guidance frames it the same way: UNESCO's Recommendation on the Ethics of Artificial Intelligence holds that the ethical deployment of AI systems "depends on their transparency and explainability" — and how much of either is enough for your organization is a judgment call that sits at your desk, not the vendor's.

The leaders I respect most in this space are not the ones who went deepest into the technology. They are the ones who stayed close to these three questions while letting their teams execute everything else. That is not micromanagement. That is leadership in a domain where the cost of judgment gaps is unusually high.

Visibility is not oversight. A dashboard showing AI usage metrics tells you what happened. It does not tell you whether what happened was what you would have chosen.

There is an honest question underneath all of this, and I think it is worth sitting with:

Have you made these three decisions explicitly — in writing, in a governance forum, in a conversation that produced a clear answer?

Or are you operating on the assumption that someone below you made them, and made them well?

I am not asking to create doubt about your team. I am asking because the gap between "I delegated the AI program" and "I delegated these three specific judgments" is where most of the exposure lives. And it is a gap that tends to be invisible until it is not.

If this maps to something you are navigating, I would be glad to think through it with you. You can reach me through jonahtebaa.com.

For a fast, direct answer on this, see the answer hub entry on who is accountable when an AI system gets it wrong.

Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf. This page is an article, not a book. Dr. Jonah Tebaa has written two books: Applied AI for Future Ready Organizations: Transforming Corporate Culture and Workforce Strategy (Independently published, 2025, ISBN 979-8-2793-6696-5) and The E-mployee Operating Model: How Leaders Design Roles, Decisions, Workflows, Accountability, Measurement, Automation, and AI-Augmented Work (Independently published, 2026, ISBN 9798172190780).

Frequently Asked Questions

Does the EU AI Act require a named human to oversee an AI system?

Yes, for high-risk systems. Article 26(2) requires that deployers “shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support.” Authority is the operative word: a named reviewer who cannot stop the system satisfies the org chart but not the obligation. Article 26 sits in Chapter III, Section 3, whose application date was moved by the Digital Omnibus on AI, Regulation (EU) 2026/1744, to 2 December 2027 for systems that are high-risk under Article 6(2) and Annex III, and to 2 August 2028 for product-embedded systems under Article 6(1). The original date was 2 August 2026. The delay is administrative and does not touch Chapters I and II, which have applied since 2 February 2025.

What are the three decisions an executive must never delegate on AI?

What the AI system is permitted to act on without human review; what constitutes acceptable failure and at what rate; and who is accountable when the system optimizes for the wrong thing. Delegating AI execution is ordinary leadership. Delegating those three judgments is not delegation at all, it is exposure, because accountability does not transfer cleanly the way a project plan does.

What is the AI accountability trap?

The executives most exposed to AI risk are not the ones who said no. They are the ones who approved the budget, endorsed the roadmap, celebrated the launch, and then handed it off. The trap is delegating AI execution without retaining judgment over the system's scope, its risk tolerance and its failure modes. The gap between delegating an AI program and delegating those three specific judgments is where most of the exposure lives, and it stays invisible until it is not.

Is an AI usage dashboard the same as executive oversight?

No. Visibility is not oversight. A dashboard showing AI usage metrics tells you what happened; it does not tell you whether what happened was what you would have chosen. Oversight means the decision was made explicitly, in writing, in a governance forum or in a conversation that produced a clear answer, about scope, acceptable failure and who owns the objective. Reporting after the fact cannot substitute for a judgment made in advance.

Who should decide what counts as acceptable AI failure?

The executive, not the technical team. AI systems fail; that is a property rather than a defect. The question that matters is which errors are acceptable and at what rate, and tolerance varies enormously: an occasional wrong content recommendation is a different problem from an occasionally misclassified customer risk profile. The threshold is a values decision, not a technical one. Waiting to see how the system performs quietly accepts a failure mode nobody defined.

Do executives need technical expertise to govern AI systems?

No. Governing AI does not require understanding model architecture or reviewing training data. It requires enough literacy to make three decisions deliberately, and to make them again when the context changes: what is in scope for autonomous action, what failure looks like and at what threshold it becomes unacceptable, and who answers for the objective. These are strategic decisions needing executive judgment. Staying close to them while teams execute everything else is leadership, not micromanagement.