A regional bank's credit team adds an AI risk-scoring layer that auto-flags loan applications below a threshold for "further review." In practice, further review means a junior analyst opens the file, glances at the flag, and confirms the rejection in under ninety seconds — roughly eight times out of ten, without re-examining the underlying numbers. The applicant receives a form letter. No one mentions that a model touched their case, because compliance decided that since a human "signed off," disclosure wasn't required.
I have watched versions of this exact pattern play out across banking, HR, and healthcare settings, and the reasoning is always the same: the deployment felt cautious internally, so no one asked whether it should be disclosed externally. That is the wrong test. Disclosure should not be triggered by how comfortable the deployer is with the model. It should be triggered by what the decision does to the person on the other side of it.
Why "It's Just a Recommendation Engine" Is Not a Disclosure Policy
Almost every organization I talk with about AI governance has arrived at some version of the same internal argument: the model doesn't decide anything, a human reviews the output, therefore there is nothing to disclose. I understand why this argument is appealing. It lets legal sign off quickly, it avoids an uncomfortable conversation with customers, and it is technically true in the narrowest sense — a human did, in fact, click a button.
But "a human reviewed it" is doing an enormous amount of unexamined work in that sentence. Review can mean genuine, independent re-analysis of a case. It can also mean rubber-stamping a recommendation in ninety seconds because the analyst is measured on throughput, trusts the model more than they trust their own judgment, or simply has no practical way to override it without triggering a lengthy escalation. Those are radically different situations, and treating them as equivalent because both technically involve "a human in the loop" is how organizations end up making decisions that are functionally automated while describing them, to regulators and customers alike, as human-made.
The comfort test — is this a tool we feel fine about internally — measures the wrong thing. It tells you whether your legal and technical teams are satisfied with the model's accuracy and your process controls. It tells you nothing about whether the person affected by the decision has any idea a model was involved, or would want to know if they did.
The Three-Part Test
In my work advising organizations on AI risk and deployment, I use a simpler standard, built around the person on the receiving end of the decision rather than the comfort level of the team that built it. I call it the Disclosure Test. It has three parts. If any two of the three apply, disclosure is warranted.
- The Stakes Test. Does the outcome materially affect the person's money, opportunity, health, or reputation? A loan rejection, a denied claim, a hiring decision, a diagnosis-adjacent recommendation — all clear yes. A product recommendation or a chatbot suggesting an article — no.
- The Substitution Test. Did the model make or substantially shape the outcome, rather than genuinely supporting a human who independently reviewed the case on its merits? The question is not whether a human touched the output. It is whether the human's judgment would plausibly have differed from the model's absent time pressure, incentive structure, or institutional deference to the score.
- The Contestability Test. If the person knew a model was involved, could they actually do something differently — appeal with new information, correct an error in the underlying record, negotiate terms, or walk away and go elsewhere? If disclosure would change nothing they could act on, it carries less weight. If it would let them contest a specific factor the model weighted, that changes the calculus entirely.
Run the bank example through it. Stakes: yes — a loan rejection affects the applicant's access to capital and, often, their credit profile going forward. Substitution: yes — the analyst's ninety-second confirmation is not independent review; it is deference to the score, dressed up as human oversight. Contestability: yes — if the applicant knew a model had flagged their file on a specific variable, they could correct an error in that variable, provide missing documentation, or request true human reconsideration. All three trigger. This was never a borderline case. It only looked like one because the internal test being applied was the wrong one.
What Disclosure Actually Looks Like (and What It Doesn't Require)
The reason organizations avoid this conversation is that they imagine disclosure means something disruptive — a warning banner, a legal disclaimer that spooks customers, an admission that erodes trust in the institution. In my experience, that fear is mostly unfounded, and it comes from confusing disclosure with confession.
Disclosure that passes the test I've described is usually a single, plain sentence, delivered at the point where it matters and paired with a real path forward. For the bank example, something as simple as: "This decision included an automated risk assessment. If you believe any information used in that assessment is incorrect, you can request a manual review here." That sentence does three things at once — it names the model's involvement, it respects the applicant's intelligence, and it gives the contestability test something to act on.
What disclosure does not require is a technical explanation of the model's architecture, a probability score attached to every decision, or a blanket statement on every page of a website that "AI may be used somewhere in this organization." Over-disclosure is its own failure mode — it either buries the meaningful instance in noise, or it triggers so much friction that teams quietly stop applying the policy where it actually matters. The test exists precisely to prevent both outcomes: it tells you when disclosure is owed, not to disclose everything, everywhere, by default.
For executives in banking, HR, healthcare, and lending, the pressure to get this right is no longer theoretical. Enforcement under the EU AI Act is ramping, data-protection regimes across the GCC are tightening their expectations around automated decision-making, and customers are starting to ask the question directly: did a machine decide this? Organizations that wait for a regulator or a lawsuit to force the answer will find themselves retrofitting disclosure under far worse conditions than the ones they face now, sitting down this quarter to write a customer notice or an HR policy.
The standard I'd offer is not complicated, and that is deliberate. Stop asking whether the model feels safe to your team. Start asking what it does to the person it affects, whether a human genuinely stood between the model and the outcome, and whether that person could act differently if they knew. Two out of three, and the answer is yes — you owe them an explanation.
Take your ten highest-volume customer questions. For each one, finish this sentence: "After this answer, the customer should be able to…" If the action is unclear, the interaction is not yet decision-ready.

