Gartner has published two predictions that sit uncomfortably next to each other. In March 2025, the firm predicted that agentic AI would autonomously resolve 80% of common customer service issues by 2029, cutting operational costs by 30%. A year earlier, a Gartner survey of 5,728 customers found that 64% of them would rather companies not use AI for customer service at all, and 53% said they’d consider switching to a competitor if they found out a company planned to (Gartner, 2024 customer survey; 2025 resolution forecast).
Both numbers are real, from the same research firm. They’re not actually a contradiction once you separate what conversational AI agents are good at from what they’re bad at, and most vendor pages have no reason to tell you where that line sits. This is where it actually sits.

Conversational AI Agents vs. Chatbots: Not Every Bot Is an Agent
Here’s a cheap way to spot the difference: ask what happens after the first reply.
A scripted chatbot gives you an answer and waits for your next message, the same way it would for anyone, in the same order, off the same script.
A conversational AI agent keeps working. It pulls your order record, checks your account status, decides whether it can resolve this, and if it can’t, it hands you off with the whole conversation intact instead of making you start over with a person.
That’s not a marketing distinction. It’s the same reactive versus goal directed split that separates generative AI from agentic AI everywhere else.
A script responds to what’s in front of it. An agent is working toward an outcome, checking a return, closing a ticket, updating a record, and everything it does in between is in service of that outcome, not just the next reply.
Where AI Customer Service Agents Work: The Routine 70%
Gartner’s 80% by 2029 number isn’t optimism. It’s math on a specific kind of ticket: the ones where the answer already exists somewhere in a system and getting it right or wrong isn’t a judgment call.
- Where’s my order.
- Can I reschedule this.
- Why was I charged twice.
A conversational agent that can query the right database answers these correctly almost every time, because there’s nothing to interpret. The question has one right answer, and the agent’s job is retrieval, not judgment.
This is also the part every vendor page leads with, because it’s the part that makes the ROI slide look good. It’s real. A support team that stops manually answering “where’s my order” for the thousandth time gets real hours back.
It’s also the specific slice Pendoah’s AI for Customer Service work is built around: automating tier-one queries reliably, and routing anything more complex with full context rather than a cold handoff. But this 70% was also never the part that determined whether customers trusted the company.
That part comes next.
Where Customer Service AI Agents Fail: The Harder 30%
Gartner’s other number, the 64% of customers who’d rather a company skip AI entirely, isn’t people being anti-technology on principle. When Gartner asked why, the answers were specific: 60% worried it would be harder to reach an actual person, 42% worried it would confidently give them a wrong answer, and 46% worried the whole thing was really about cutting costs at their expense. Those aren’t hypothetical fears. They’re describing exactly what happens when a company points out a conversational agent at tickets it was never built to handle.
Klarna is the case that made this public.
In February 2024, the company’s CEO said its AI assistant was handling 75% of support chats, doing the work of roughly 700 human agents, and processing 2.3 million conversations in a single month. That part was true, and it was genuinely impressive. What wasn’t said out loud yet was what happened to the other 25%, the angry customer, the ambiguous dispute, the case that needed someone to actually use judgment.
By mid-2025, the same CEO admitted to Bloomberg that cost had been “too predominant” a factor in the original decision, that service quality had dropped, and that Klarna was rehiring humans and guaranteeing every customer a path to one (Silicon Canals, Klarna U-turns on AI).
The AI didn’t get worse at handling volume. It never had a volume problem. What broke was the assumption that the hard 25% would somehow shrink to fit the tool, instead of the tool being scoped to fit the 75% it was actually good at.

Klarna’s Mistake Wasn’t a Technology Problem
It’s tempting to read the Klarna story as “the model wasn’t good enough yet.” That’s not really what happened, and treating it that way is how the same mistake gets repeated with a better model next year.
A conversational agent that’s designed to sound confident by default will answer a question it has no business answering, because nothing in how it was built rewards it for saying “I don’t know, let me get you a person” instead of guessing well.
That’s a design decision, not a technology ceiling. The fix isn’t a smarter model, it’s an explicit, tested line for when the agent stops and hands off, decided before launch instead of discovered after a customer post about it.
That’s the same principle behind Human in the Loop: the goal was never to automate everything. It was to automate what’s actually safe to automate and be honest about the rest.
Questions to Ask Before Deploying an AI Customer Service Agent
Almost every conversational AI failure traces back to one of three questions that never got a real answer before go-live:
The three questions that separate a successful deployment from a public failure.
What actually triggers a handoff?
Not “the model decides it’s stuck.” A specific, testable threshold, repeated failed attempts, certain account states, detected frustration, that someone can point to and defend, not a vague hope that the system will know when it’s out of its depth.
What does the human on the other end receive?
If a customer has to re-explain their problem to a person after the bot already failed, you’ve built the exact experience that made 60% of Gartner’s respondents worry about reaching a human in the first place. The handoff needs the full transcript, not a cold transfer.
Is anyone tracking the escalations that didn’t happen?
Resolution rate is the easy number to report and the wrong one to optimize alone. The harder, more honest number is how often the agent should have handed off and didn’t. That’s usually where the actual damage was done, quietly, long before anyone noticed.
What 90/10 Actually Looks Like
This isn’t a customer service problem specifically. Pendoah has built the same shape of solution elsewhere, and it looks identical.
Worklighter’s document-automation engine auto-processes 90% of incoming invoices and routes the remaining 10 to 20% to a person, and that gap isn’t a defect anyone’s trying to close to zero. It’s where the ambiguous, judgment-heavy cases actually live, and forcing them through automation anyway is exactly the mistake Klarna made in a different industry.
The number that matters for a conversational agent isn’t how close you can get to 100% automated. It’s whether you actually know from real ticket data rather than a guess, what your version of that 90/10 split looks like, and whether the 10% has a real place to go.
Is Your Business Ready for Conversational AI Agents?
The honest version of “should we deploy this” isn’t a checklist; it’s one question with two very different answers depending on why you’re asking. If you’re looking at high-volume, well-defined requests where the answer already lives in a system you can query, and you’re doing this to free your team’s time for the harder cases, you’re describing the 70% Gartner’s forecast is actually about, and you’re probably ready.
If you’re looking at this because support headcount is the line item you want to shrink, or because you haven’t actually measured what share of your tickets are routine versus genuinely hard, or because there’s no tested, honest answer to what happens when the agent gets it wrong, you’re describing the setup that produced Klarna’s reversal.
The technology isn’t the variable that changes that outcome. The honesty of the scoping is, and it’s exactly what Pendoah’s Conversational AI Development Services start with before any bot gets built.
Key Questions to Ask Before You Deploy
- What percentage of our actual ticket volume is genuinely routine, based on real data rather than a guess?
- What specific signals trigger a handoff to a person, and has that threshold been tested against real conversations?
- Does a human agent receiving an escalated conversation get full context, or does the customer start over?
- Are we tracking missed escalations, cases the agent should have handed off and didn’t, not just resolution rate?
- If a customer explicitly asks for a human, how many steps does it take to reach one?
The Bottom Line
Conversational AI agents aren’t good or bad. They’re scoped well or scoped carelessly.
Gartner’s 80% and 64% aren’t a contradiction; they’re two ways of describing the same tool: excellent at the routine 70% and genuinely damaging when a company mistakes that for permission to point it at everything else.
Klarna found that line the expensive way. You don’t have to.
Want help figuring out which share of your support volume is actually ready for this, and which isn’t?
Book a consultation with Pendoah, or look through our case studies to see how this decision played out for other clients.