
At 11:40 p.m., a prospect in Riyadh types one line into a chat widget: can you guarantee delivery by Friday. It is an ordinary question, the kind a sales team fields dozens of times a week during business hours, with a manager two desks away to check the answer. At 11:40 p.m. there is no manager. There is only the AI, and whatever it has been told it is allowed to say.
There are three ways it can answer that question in the affirmative, and each one carries a different amount of weight. It can describe how quickly the company typically delivers. It can quote a date tied to this specific order. It can promise the delivery outright, in language a contract would use. Only the third of those actually matches the word "guarantee" the prospect used. If the AI reaches for that word anyway, because it sounds helpful and the model has no sense of legal weight, the company has made a commitment nobody in the building agreed to, hours before anyone reviews the transcript.
Why "We Trained It on Our Policies" Doesn't Solve This
Most leadership teams I talk to about customer-facing AI in sales have already done the obvious work. They fed the model their pricing sheet, their delivery windows, their refund policy. They tested it for accuracy and it passed. They conclude the risk is handled, because the AI isn't going to state a wrong price or invent a policy that doesn't exist.
That conclusion misses where the actual failure lives. The risk in a sales-facing AI is rarely factual. It is almost never a wrong number. It is a tone-of-certainty mismatch: the AI answers a quoting question with the same confident, complete-sentence certainty it uses to answer a general question, even though the two questions require completely different levels of authority to answer. A well-trained model does not sound less sure when it drifts from describing a typical timeline into promising a specific one. It sounds exactly as sure either way, because sounding sure is what fluent language does by default.
Customers read tone before they read fine print. A prospect who hears "we can guarantee that" does not mentally file it as "a chatbot's best estimate." They file it as the company's word. The gap between what the AI actually knows and what the business is willing to be held to is invisible in the transcript and very visible the first time someone tries to hold the company to it.
The Four-Rung Commitment Ladder
The fix is not a vaguer instruction to the model to "be careful" with commitments. Vague instructions do not survive contact with a fluent model that has no concept of legal exposure. The fix is a written taxonomy that sorts every question a customer might ask into one of four rungs, each with its own permission and its own test.
- Describe. General capability, timeline, or policy stated in the abstract, with no promise attached to this customer's order. Example: "Our typical turnaround is three to five business days." Reviewer's test: could this sentence be true on a slow week and a fast week both, without anyone being misled? One rung too high, this becomes a quote by accident, because a customer reasonably hears "typical" as "mine."
- Quote. A specific number or date tied to this customer's stated inputs, always time-boxed and explicitly reversible until confirmed by a human or a system of record. Example: "Based on your order details, the estimated delivery date is Friday, pending final confirmation." Reviewer's test: is there a visible condition attached, and would a reasonable person still read this as tentative? One rung too high, the word "pending" quietly disappears and a quote reads as a promise.
- Promise. A commitment the company will honor in writing if the customer holds it to that commitment tomorrow: guaranteed SLAs, contract terms, refund eligibility. Reviewer's test: if the customer forwards this exact sentence to a lawyer next week, does the company still stand behind it, unconditionally? One rung too high, an AI has bound the company to a promise that only a human with authority to make exceptions should ever issue.
- Escalate. Anything that requires human judgment, discretion, or exception-making, no matter how confidently the model could construct an answer. Reviewer's test: does answering this well require weighing something the AI has no visibility into, like account history, a hardship case, or a negotiated exception? One rung too high, the AI substitutes its own judgment for a decision that was never delegated to it in the first place.
The test question at each rung is deliberately simple enough for a non-technical reviewer, a sales manager or a compliance lead, to apply without reading a line of the underlying prompt. If a reviewer cannot answer the test question in a sentence, the question probably belongs one rung lower than wherever it currently sits.
Building the Permission Map Before Launch
Most teams write their AI's instructions around what the model knows. The permission map is built around what the business is willing to be held to, which is a different document entirely. Before any customer-facing sales AI goes live, or before an existing one is trusted with more traffic, I recommend a short protocol:
- Pull the sales team's actual FAQ list and the last two months of real chat and WhatsApp transcripts, not a hypothetical list someone writes from memory.
- Tag every question, individually, with the rung it belongs to. Do this in a spreadsheet, not inside the AI's configuration, so a non-technical reviewer can audit it later without opening the system.
- For every question tagged Promise or Escalate, write the exact fallback language the AI should use instead, so the model has somewhere safe to land rather than defaulting to whatever sounds most helpful.
- Write the rung directly into the AI's instructions as a hard boundary tied to the specific question pattern, not as a general tone note like "be cautious about guarantees." Tone notes get diluted by every other instruction competing for the model's attention. A boundary tied to a question pattern does not.
- Re-run the last two months of transcripts through the finished map before launch, and confirm out loud, with the sales lead in the room, that every answer the AI actually gave lands on the correct rung.
A Composite Example
The following is a composite drawn from a pattern across engagements, not a single identifiable client, offered because the pattern itself is the useful part. A mid-size Gulf services firm's chat AI began confirming delivery-date commitments during pre-sale conversations, because nobody had distinguished "confirm a date" from "describe our typical turnaround" in its instructions. The two phrases looked interchangeable to whoever wrote the original prompt. They are not interchangeable to a prospect deciding whether to sign, and they are not interchangeable to whoever eventually has to explain a missed date to that same prospect. The fix in that pattern was never a retrain. It was a one-page permission map, built the way I've described above, that took an afternoon to produce and closed the gap for every future conversation, not just the one that surfaced the problem.
If your sales AI has never been audited rung by rung, the audit is worth an afternoon before it costs you a signed conversation you cannot honor. Pull last month's transcripts, run them through the four questions above, and see how many answers are sitting one rung higher than anyone actually decided.
For more on this and related work, see BrianServes, the platform for deploying autonomous AI e-mployees and Webspot, the AI strategy firm in Beirut.
Related evidence: Article 26(2) of the EU AI Act requires deployers of high-risk AI systems to assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support — so the oversight duty lands on a named person who must actually be empowered to act, not on a department. Article 26 sits in Chapter III, Section 3, whose date of application was moved by Regulation (EU) 2026/1744 to 2 December 2027 for Annex III high-risk systems and 2 August 2028 for Article 6(1) product-embedded systems. (Article 26(2) of the EU AI Act)
NIST's AI RMF appendix on human-AI interaction notes that AI systems can autonomously make decisions, defer decision making to a human expert, or be used by a human decision maker as an additional opinion. (NIST AI RMF appendix on human-AI interaction)
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.