Should an AI Agent Talk to Your Customers Directly?
Should an AI Agent Talk to Your Customers Directly?
Yes for questions where a wrong answer is cheap to correct, and no for anything that commits your company to something. The useful line is not simple versus complex. It is reversible versus binding. An agent that can quote a policy is fine. An agent that can create one is not.
This argument usually gets framed as a cost decision, which is why it keeps going badly. The real question is what you are willing to be held to when the system is confidently wrong, because you will be held to it.
Here is the case on both sides, and where we would draw the line for a B2B company.
What Is the Strongest Case for Letting an Agent Answer?
Speed and coverage, and the numbers behind that case are real. Klarna announced in February 2024 that its OpenAI powered assistant "has had 2.3 million conversations, two-thirds of Klarna's customer service chats" in its first month, doing "the equivalent work of 700 full-time agents."
The rest of that release is worth reading closely. Klarna said resolution went from 11 minutes to "less than 2 mins", that the assistant was "on par with human agents in regard to customer satisfaction score", and that it produced "a 25% drop in repeat inquiries". It was live in 23 markets and "more than 35 languages", with an estimated "$40 million USD in profit improvement" for 2024.
Vendors report similar shapes. The Fin product page claims "industry leading resolution rates, averaging 76% across 12,000+ customers, with many seeing over 85%" and "2 million weekly resolutions". Treat that as a marketing claim rather than an audited figure, but the direction is consistent.
What Do Those Numbers Not Tell You?
They do not tell you what happened next. Klarna's release described its first month. A first month is the period when the questions are the common ones, the content is fresh, and the team is watching closely. We would not treat any first month announcement, from any company, as evidence of what a system does in year two.
The language coverage number is the most interesting one and the least discussed. Thirty five languages is not something a support team of any size can staff. If your customers are spread across markets you cannot hire for, that is a genuinely different argument from cost reduction, and a better one.
What none of these figures address is the tail. A 76% resolution rate means roughly one in four conversations still needs a person, and those are the hard ones. Your plan has to be about that quarter, not the three quarters that were always going to be easy.
Where Does the Legal Line Sit?
In the European Union it is now explicit. Article 50 of the EU AI Act requires that "AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system", unless that is obvious to a reasonably well informed person.
The timing requirement is specific too. The information must be "provided to the natural persons concerned in a clear and distinguishable manner at the latest at the time of the first interaction or exposure". A disclosure buried in a terms page does not meet that. A line in the chat window before the first reply does.
Article 50 also covers generated content, requiring that outputs be "marked in a machine-readable format and detectable as artificially generated or manipulated" where the system produces synthetic audio, image, video or text. We covered the wider compliance picture in our piece on the EU AI Act and marketing automations.
What Is the Actual Risk You Are Taking?
That the agent says something your company then has to honour. A support answer is a statement from the business. If an agent invents a refund window, quotes a discount that does not exist, or describes a feature you do not ship, you are left choosing between honouring it and telling a customer your system lied to them.
Neither option is good, and the second is worse than people expect. "Our chatbot was wrong" is not a reassuring sentence for a buyer evaluating whether your product can be trusted with their data. In B2B, where the audience is small and talks to each other, one bad answer travels.
The mitigation is not better prompting. It is narrowing what the agent is allowed to assert. An agent that can only quote from published documentation and link to it has a much smaller blast radius than one asked to be helpful in general.
Where Would We Let an Agent Answer Unsupervised?
On the reversible, documented, and low stakes. Where is my invoice, how do I reset my password, what does this error code mean, where is the API reference for this endpoint. These have one correct answer, it is written down, and getting it wrong costs a follow up message.
| Question type | Agent alone | Why |
|---|---|---|
| Documented how to | Yes | One answer, already published, cheap to correct |
| Account status lookup | Yes, read only | Factual, from your own system |
| Pricing and contract terms | No | Commits the company |
| Refunds, credits, exceptions | No | Binding and hard to reverse |
| Security and compliance answers | No | Wrong answers become contractual |
| Anything a distressed customer asks | No | Escalation is the product |
Read only access is the mechanism that makes the top rows safe. If the agent cannot write, it cannot commit you. We set out how we think about that boundary in our piece on least privilege access for AI agents.
How Should the Handoff to a Human Work?
Immediately, visibly, and without making the customer repeat themselves. The most common failure we see is not a wrong answer. It is a loop where the customer asks for a person four times and the agent keeps trying to help. That converts a minor annoyance into a complaint.
Give people an explicit way out in the first message. A visible option to reach a person is not an admission of failure, and hiding it does not increase deflection so much as it increases frustration. It also reduces the number of people who go straight to your public channels instead.
When the handoff happens, the transcript must travel with it. If a person has to ask the customer what the agent already asked, you have spent the customer's patience twice and saved nothing.
What Should You Measure?
Measure resolution without recontact, not deflection. Deflection counts conversations that did not reach a person, which includes everyone who gave up. Resolution without recontact within seven days counts conversations that actually ended. Those two numbers can point in opposite directions.
Track escalation rate by topic as well as overall. A rising escalation rate on one topic is a documentation gap you can fix. A flat overall number hides it completely, which is why topic level reporting is worth the setup cost.
Then sample. Read fifty real transcripts a month, chosen at random rather than from the complaints queue. It is the only way to see the confidently wrong answers that nobody bothered to report.
Does This Change for Voice?
The stakes go up. Voice removes the pause where a person would have reread a sentence and noticed it was odd, and it makes the disclosure requirement harder to satisfy gracefully. A written notice at the top of a chat is easy. The spoken equivalent has to be said, every time, without sounding like a legal recording.
Voice also compresses the escalation window. In chat a customer can wait. On a call, a transfer that takes too long is an abandoned call and often an abandoned deal. If you cannot staff the transfer, the agent should not be answering the phone.
We went into the specifics of that in our piece on AI voice agents in B2B. The short version is that we would deploy voice later than chat, not earlier.
What Would We Actually Recommend?
Start with an agent that answers from your documentation, discloses what it is in its first message, has a visible route to a person, and cannot write to any system. That configuration captures most of the volume benefit and almost none of the liability.
Then expand one capability at a time, and only after reading transcripts for the current one. The teams that get into trouble are the ones that turn on account actions in month one because the demo made it look easy. The teams that do well add write access to a single narrow action and watch it for a month.
If you are weighing this up and want a second opinion on where your line should sit, we are happy to talk it through. Find us at phoenix.studio and bring a handful of real transcripts, not the demo.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.