AInora
Voice AIRiskComplianceVendor Evaluation

What AI Should Not Handle on Business Calls (And Why It Matters)

JB
Justas ButkusFounder, Ainora
··12 min read

TL;DR

An AI phone agent should not give clinical, legal or financial advice; should not try to manage a caller in distress; should not complete an action that cannot be reversed without a person confirming it; should not decide an exception to a policy; should not present itself as a human being; and should not call people who never agreed to be called. Each of those refusals needs a destination behind it, or it is a dead end rather than a boundary. A supplier who cannot recite their own version of this list in specifics has not run the failure cases, which means the buyer runs them instead, live.

2 Aug 2026
AI Act Art. 50 applies
Source: AI Act Art. 113
$193,000
FTC order against DoNotPay
Source: FTC
$6m
FCC forfeiture, Kramer spoofed robocalls
Source: FCC
Art. 22
GDPR right to human intervention
Source: GDPR

What should AI not do in a business?

An AI system should not do the things whose failure mode is silent and expensive. On a business phone line that resolves to six specific refusals: advice that requires a licence, the handling of a caller in distress, any action that cannot be undone, the invention of an exception to a policy, pretending to be a human being, and calling people who never consented to be called. Everything else, and it is most of a working phone line, is a candidate for automation.

That list is not a confession of weak technology. It is the shape of a system somebody thought about. Almost every supplier in this category publishes the ceiling, meaning the longest possible list of what their agent can do. Very few publish the floor. The floor is what a competent operator actually has, because it is the part you can only write after you have sat with the failure cases and decided, in advance, what the system does when it does not know.

The rest of this piece takes each of the six in turn, with the public record underneath it rather than an assertion, and finishes with the question that lets you test any supplier on this, including us.

Why does a refusal list tell you more than a feature list?

Because a feature list is a claim about the good calls and a refusal list is a claim about the bad ones, and only one of those two categories generates complaints, refunds and regulatory correspondence.

There is a mechanical reason the bad calls are the ones that matter. When a language model is asked something it has no basis to answer, it does not stall. It produces an answer with the same fluency, the same speed and the same confident tone as everything else in the conversation. There is no audible tell. The caller has no way to distinguish a sentence grounded in your booking system from a sentence the model assembled because the shape of the question implied that shape of answer. This is the single most important property of the technology for a buyer to internalise, and it is the reason boundaries have to be structural rather than instructional. Telling an agent to be careful is not a control. Removing its ability to act, and giving every refusal a named destination, is.

The same reasoning runs through the difference between prompt instructions and encoded policy and through how an agent gets verified before it ever speaks to a customer. Boundaries without verification are a preference. Verification without boundaries has nothing to check against.

Should an AI phone agent give clinical, legal or financial advice?

No. The agent takes the request, records it accurately, and routes it to whoever is qualified to answer. It does not answer it. Not because a model could not produce something plausible in those domains, but precisely because it can, and a plausible wrong answer about a medication, a limitation period or a pension is the one that ends up in somebody's file.

Regulators have already been explicit about where the line sits. In February 2025 the US Federal Trade Commission finalised an order against DoNotPay, a service marketed as a robot lawyer, requiring $193,000 in monetary relief and notice to past subscribers. The FTC's complaint is worth reading carefully, because it is about process rather than output. According to the complaint, the company "did not test whether its 'AI lawyer' operated to the level of a human lawyer when generating legal documents and giving advice" and "did not hire or retain attorneys to test the quality and accuracy of its service's law-related features". The charge is not that a machine cannot draft. It is that nobody checked.

On the clinical side, the UK Medicines and Healthcare products Regulatory Agency published 2026 guidance on ambient voice technology products which draws the qualification line at exactly this point: a product that transcribes or summarises is one thing, and a product whose generated insights "provide suggested diagnoses or relevant follow-up and treatment options" is a medical device with everything that follows from that. The guidance also closes the obvious escape route, stating that "general disclaimers (for example 'this product is not for diagnosis') are not acceptable to demonstrate a product is not a medical device if medical claims are made or implied elsewhere". That sentence is about labelling and promotional material rather than about what an agent says out loud, but the reasoning travels: a disclaimer does not undo what the product is then presented as doing.

In financial services the framing is similar. ESMA's public statement of 30 May 2024 on the use of artificial intelligence in investment services expects firms using AI to comply with the existing MiFID II requirements, including their obligation to act in the best interest of the client, and lists overreliance on AI by both firms and clients among the risks. Read together, those three regulators are saying something consistent rather than something new: none of them treats the technology as creating a new category of adviser, and in each case the responsibility stays where it already sat.

The practical design consequence is small and specific. The agent is allowed to say what a service is, what it costs to attend, when the next opening is and who will be in the room. It is not allowed to say whether you should have it. This is the same boundary that governs prescription refill intake and front-desk work for financial advisers, where the intake is genuinely automatable and the advice is not.

What should an agent do when the caller is distressed?

Hand over, immediately, and without making the person explain themselves twice. Bereavement, an emergency, a person in evident crisis, a customer who has already been let down twice and is calling a third time: the design question is how quickly the agent notices, not how gracefully it copes.

The best-documented illustration of getting this wrong is not a phone system. In 2023 the US National Eating Disorders Association shut down a helpline it had run for more than twenty years, and a chatbot called Tessa was one of the resources it intended to promote in its place. Tessa had been built with preprogrammed responses written by eating disorder specialists, and one of its creators told NPR that everything was preprogrammed by design, because "AI isn't ready for this population". NPR reported that after the company operating Tessa added generative capability, letting it learn from new data and produce new responses, a consultant in the field who tested the chatbot was told she could lose one to two pounds per week and should have a calorie deficit of 500 to 1,000 calories per day. She sent the screenshots to the association, the chatbot was taken down within hours, and the company and the association both apologised. Nothing here was the model behaving strangely. A system built deliberately narrow for a high-risk population had been widened, and widening the capability is exactly the moment the question of what happens at the edges has to be asked again.

Financial services regulators have put the expectation in writing. The FCA's finalised guidance FG21/1 of 23 February 2021 on the fair treatment of vulnerable customers defines a vulnerable customer as "someone who, due to their personal circumstances, is especially susceptible to harm", a susceptibility the FCA ties directly to whether the firm is acting with appropriate levels of care, and asks firms to ensure frontline staff have the skills to recognise and respond to vulnerability. The guidance also asks firms, where possible, to offer multiple channels so vulnerable consumers have a choice, which makes a line with no fast exit to a person something to test against that sentence rather than something to assume satisfies it.

The measurable version of this boundary is not a sentiment score. It is two numbers: how many conversational turns pass before the handover begins, and whether the human who picks up already has the context. A transfer that dumps the caller back at the beginning is technically a handover and practically a second failure, which is the whole subject of warm transfer versus cold transfer. Sectors where this is the primary design constraint rather than an edge case include funeral homes, mental health practices and emergency dispatch.

Which actions should never happen without a person in the loop?

Money leaving an account. A cancellation. A refund. A change to a prescription. The release of a record to somebody who called and said they were entitled to it. Read operations are safe to automate long before write operations are, and irreversible writes are a category of their own. The agent may prepare one completely and propose it. A person commits it.

The clearest public illustration of what happens when that ordering is not enforced came from software development rather than telephony, which makes it easier to look at without arguing about voice. In July 2025, as The Register reported, the founder of a SaaS community business said an AI coding tool had deleted his database despite his instruction not to change any code without permission, and that it had been covering up bugs by creating fake data and fake reports. It then told him a rollback was impossible; in a post dated 19 July he wrote that the rollback in fact worked. Screenshots in the report have the agent acknowledging "a catastrophic error of judgement" and that it had "violated your explicit trust and instructions". The vendor had not addressed the posts when the report was published. Whatever the eventual account, the control in place was a written instruction, and a written instruction is not a control.

The regulatory expectation runs the same way. Article 14 of the EU AI Act, which governs human oversight of high-risk systems, requires such a system to be built so that the people to whom oversight is assigned are enabled, as appropriate and proportionate, to "decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output" and to interrupt it "through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state". Whether a given phone deployment falls inside that classification is a separate question, treated in the high-risk classification guide. The design principle stands either way: an override that exists only after the money has moved is not an override.

In practice this means sorting every action the agent can take into three lists before launch. Read: look up a booking, check an opening, confirm an address. Reversible write: create an appointment, log a note, add someone to a callback list. Irreversible write: take a payment, cancel, refund, release. The third list should be short, and everything on it needs a named human before it commits. What the agent is able to write at all is bounded by what it is connected to, which is the subject of the system of record page, and the controls around the data it touches are on security. Card payments carry their own separate constraints, covered on the PCI page and in taking payment over the phone.

Can an AI agent decide an exception to a policy?

It can apply a rule. It cannot write one. Exception handling is judgement about a particular person in a particular situation, and judgement is precisely what the architecture does not supply. When a caller asks for something the policy does not cover, the honest answer is that a person will decide, not an improvised yes.

In April 2025 a developer tools company found out what the alternative looks like. Users had been unexpectedly logged out when moving between machines. They wrote to support, and the reply that came back from the company's support address, generated by an AI agent rather than written by a person, told them this was expected behaviour under a policy limiting one subscription to one device. No such policy existed. The invented rule reached Reddit and Hacker News before the company corrected it.

Hey! We have no such policy. You're of course free to use Cursor on multiple machines.

The company described what had happened as "an incorrect response from a front-line AI support bot". That phrasing is accurate and slightly misleading at the same time. The response was incorrect, but nothing malfunctioned. The agent was asked why something was happening, it had no grounded answer, and it produced the most plausible-sounding one available. A policy that limits a subscription to one device is an entirely reasonable thing for a software company to have. That is exactly why the invention was credible, and why customers acted on it.

There is a legal echo of the same principle. Article 22 of the GDPR gives a person the right not to be subject to a decision "based solely on automated processing" which produces legal effects or similarly significantly affects them, and, where such processing is permitted, requires safeguards including "the right to obtain human intervention on the part of the controller, to express his or her point of view and to contest the decision". An architecture with no human anywhere in the decision path cannot produce that intervention after the fact, because there is nobody who made the decision to talk to. Broader data obligations on a call are covered in the GDPR guide for voice agents.

Should an AI voice agent ever pass as a person?

No, and the commercial argument against it is stronger than the legal one, which is already strong.

Article 50 of the EU AI Act requires that AI systems intended to interact directly with people are built so that those people "are informed that they are interacting with an AI system, unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect". Article 113 sets the general application date of the Regulation at 2 August 2026. The obligations differ by jurisdiction and are moving in one direction, which is mapped country by country in AI caller disclosure laws by country and at the telephony layer on the multi-country compliance page.

Outside the EU the enforcement has already been financial, and it is worth being exact about what was penalised. In September 2024 the FCC adopted a $6,000,000 forfeiture against a political consultant whose robocalls carried a deepfake, AI-generated clone of President Biden's voice, telling New Hampshire voters not to vote. The violation the Commission found was of the Truth in Caller ID Act, section 227(e) of the Communications Act, for the misleading caller ID those calls transmitted, rather than for the synthetic voice itself. That distinction is the useful part rather than a technicality. The cloned voice is what the regulator described at length and what the penalty was calibrated against, and the hook it hung on was an ordinary telephony rule that was already on the books. It is an election interference case rather than an ordinary business one, and it is the register in which regulators now discuss synthetic voice.

The commercial argument is simpler. A product engineered to pass as human is one rule change away from being unusable, and it trades a small short-term gain for the exact customer you least want to lose: the one who finds out. Callers who are told plainly, in one short sentence, that they are speaking to an automated assistant tend to keep talking, because what they wanted was the answer. We publish our own position on this on the disclosure page, and the mechanics of the call itself are on how it works.

Everyone's, which in practice means it has to be settled explicitly or it falls through. Consent lives upstream of the technology. Who is on the list, where the list came from and what those people were told are decided before a single call is placed, and a supplier who treats that as a problem for the buyer alone is telling you something about how they will behave when something else goes wrong.

The European baseline is old and specific. Article 13(1) of Directive 2002/58/EC, in the consolidated text as amended by Directive 2009/136/EC, provides that "the use of automated calling and communication systems without human intervention (automatic calling machines), facsimile machines (fax) or electronic mail for the purposes of direct marketing may be allowed only in respect of subscribers or users who have given their prior consent". That text predates modern voice AI by a decade and a half and describes it uncomfortably well, which is the point worth noticing rather than arguing about. National implementations vary, and the country by country position is on our Europe compliance guide.

The United States arrived at a similar place from a different direction. In a Declaratory Ruling adopted on 2 February 2024 and released on 8 February 2024, the FCC confirmed that calls using AI technologies that generate human voices fall inside the TCPA's restriction on an "artificial or prerecorded voice", on the reasoning that "they are 'artificial' voice messages because a person is not speaking them", with the consequence, in the ruling's own words, that callers must obtain the prior express consent of the called party to initiate such calls absent an emergency purpose or exemption. Where consent is captured as part of a funnel rather than bought in, that is a design problem with a clean solution, which is the subject of the registrant consent page.

One scoping note before the test. Nothing in this article is legal, medical or financial advice, and none of it substitutes for your own professional counsel. Every legal statement above is a description of what a named regulator or a named statute has published, with the source linked so you can read it in its own words.

How do you interrogate any supplier, including this one?

Ask them to name three things their agent is built to refuse, and to tell you what the caller actually hears in each case. Then ask which actions it can complete on its own and which require a person before they commit.

A supplier who has walked the failure cases answers in specifics and in seconds, without checking with anyone. A supplier who has not will answer with capabilities, because the ceiling is the only part of the system they have ever written down. The test costs nothing, it can be run before a contract exists, and it works on us as well as on anybody else. If we cannot answer it in the first conversation, that is information too.

A useful second question is where each refusal goes. "A human" is not an answer. A named queue, a named rota, a callback with a committed window, or an honest statement that nobody is available until Monday morning are all answers. Vague handovers are where callers get lost, and a boundary with nothing behind it is a dead end wearing a policy's clothes.

The refusalWhat the caller gets insteadThe question to ask a supplier
Clinical, legal or financial adviceAccurate intake of the question, routed to the qualified person, with a committed callback windowWhat does your agent say when someone asks it for advice it is not allowed to give?
Managing a distressed callerHandover inside the first turns, with the context already passed to the humanHow many conversational turns before the handover, and what does the human see when they pick up?
Completing an irreversible actionA prepared action, held for a named person to confirmWhich actions can the agent commit alone, and which are held?
Deciding a policy exceptionA clear statement that a person will decide, plus the request captured verbatimWhat stops the agent inventing a rule when it has no answer?
Passing as a personOne short, plain disclosure, and a straight answer whenever askedWhat exactly does the agent say when the caller asks whether it is a human?
Calling people who never agreedNo call at all until the source of the list is documentedWhat do you require from us about the list before you place the first call?

None of this is a claim that the boundaries make a system safe. They make it legible, which is the precondition for anything else being true about it. The buyer-facing summary of ours is on what we do not automate.

If it is useful to go through this properly, the format we run is a working session rather than a demo. Forty-five minutes on your actual call flow: we take the policies your front desk already follows, push on them until the edge cases fall out, and you keep the written version at the end whether or not anything else happens. If it does go further, the first piece of work is one workflow, missed calls and after hours, roughly two weeks, before anything else moves.

Frequently Asked Questions

It should not give clinical, legal or financial advice, try to manage a caller in distress, complete an action that cannot be reversed without a person confirming it, decide an exception to a policy, present itself as a human being, or call people who never agreed to be called. Each of those refusals should have a defined destination behind it: a named person, a queue, a callback with a committed window, or an honest statement that nobody is available yet.

The ones that actually happen are quiet. The agent invents a policy that does not exist and the caller believes it. It answers a question it should have routed, fluently and wrongly. It completes a cancellation or a refund nobody authorised. It keeps a distressed caller in a loop because nothing in its design tells it to give up early. None of these look like an outage, which is why they are usually found by a customer rather than by a dashboard.

Because instructions are not controls. When a model has no grounded answer it produces a plausible one at the same speed and in the same tone as everything else, with no audible tell. The July 2025 case in which the founder of a SaaS business reported that an AI coding agent deleted his database despite an explicit instruction not to change code without permission is the clean illustration: the only safeguard in place was a written instruction, and it held for exactly as long as the model chose to follow it.

It is a different question from whether it should. Card payments bring their own handling rules, and a payment is an irreversible action, which puts it in the category that needs a person or a purpose-built payment path rather than a conversational agent improvising. The safe pattern is that the agent prepares everything and hands the actual transaction to a mechanism designed for it.

Article 50 of the EU AI Act requires providers to design AI systems intended to interact directly with people so that those people are informed they are interacting with an AI system, unless that is obvious to a reasonably well-informed, observant and circumspect person, and Article 113 sets the general application date of the Regulation at 2 August 2026. This is a description of the Regulation rather than legal advice, and requirements vary by country. In practice a single short sentence at the top of the call, plus a straight answer whenever asked, covers it without costing much in answer rate.

The opposite. A supplier with no stated boundaries has the same failure modes as one with boundaries, they have just not written them down, which means the boundaries will be discovered on a live call instead of in a specification. The refusal list is the part of the design that can only be written after somebody has thought about the failure cases.

A refusal is the agent declining to produce an answer. An escalation is the agent moving the request somewhere it can be answered. In a working system every refusal is paired with an escalation, because a boundary with no destination behind it is a dead end for the caller rather than a policy.

Ask them to name three things their agent is built to refuse and what the caller hears in each case, then ask which actions it can complete alone and which are held for a person. A supplier who has run the failure cases answers in specifics and in seconds. One who has not will answer with capabilities.

JB
Justas Butkus

Founder & CEO, AInora

Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.

View all articles

Ready to try AI for your business?

Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.