pendoah

  • Home
  • Insights
  • Blog
  • Conversational AI Agents for Customer Service: Where They Work and Fail

Conversational AI Agents for Customer Service: Where They Work and Fail

AI automation is not for every SMB. Here is how to know if you are ready

Content marketer with 3+ years of experience in AI and B2B growth, leading brand positioning and full-funnel execution across web, email, sales, and social channels.

Share
Table of Contents

Gartner has published two predictions that sit uncomfortably next to each other. In March 2025, the firm predicted that agentic AI would autonomously resolve 80% of common customer service issues by 2029, cutting operational costs by 30%. A year earlier, a Gartner survey of 5,728 customers found that 64% of them would rather companies not use AI for customer service at all, and 53% said they’d consider switching to a competitor if they found out a company planned to (Gartner, 2024 customer survey; 2025 resolution forecast).

Both numbers are real, from the same research firm. They’re not actually a contradiction once you separate what conversational AI agents are good at from what they’re bad at, and most vendor pages have no reason to tell you where that line sits. This is where it actually sits.

Gartner Stats Comparison-selection
Two real Gartner numbers, same firm: the case for conversational AI agents and the case against deploying them carelessly.

Conversational AI Agents vs. Chatbots: Not Every Bot Is an Agent

Here’s a cheap way to spot the difference: ask what happens after the first reply.

A scripted chatbot gives you an answer and waits for your next message, the same way it would for anyone, in the same order, off the same script.

A conversational AI agent keeps working. It pulls your order record, checks your account status, decides whether it can resolve this, and if it can’t, it hands you off with the whole conversation intact instead of making you start over with a person.

That’s not a marketing distinction. It’s the same reactive versus goal directed split that separates generative AI from agentic AI everywhere else.

A script responds to what’s in front of it. An agent is working toward an outcome, checking a return, closing a ticket, updating a record, and everything it does in between is in service of that outcome, not just the next reply.

Where AI Customer Service Agents Work: The Routine 70%

Gartner’s 80% by 2029 number isn’t optimism. It’s math on a specific kind of ticket: the ones where the answer already exists somewhere in a system and getting it right or wrong isn’t a judgment call.

  • Where’s my order.
  • Can I reschedule this.
  • Why was I charged twice.

A conversational agent that can query the right database answers these correctly almost every time, because there’s nothing to interpret. The question has one right answer, and the agent’s job is retrieval, not judgment.

This is also the part every vendor page leads with, because it’s the part that makes the ROI slide look good. It’s real. A support team that stops manually answering “where’s my order” for the thousandth time gets real hours back.

It’s also the specific slice Pendoah’s AI for Customer Service work is built around: automating tier-one queries reliably, and routing anything more complex with full context rather than a cold handoff. But this 70% was also never the part that determined whether customers trusted the company.

That part comes next.

Where Customer Service AI Agents Fail: The Harder 30%

Gartner’s other number, the 64% of customers who’d rather a company skip AI entirely, isn’t people being anti-technology on principle. When Gartner asked why, the answers were specific: 60% worried it would be harder to reach an actual person, 42% worried it would confidently give them a wrong answer, and 46% worried the whole thing was really about cutting costs at their expense. Those aren’t hypothetical fears. They’re describing exactly what happens when a company points out a conversational agent at tickets it was never built to handle.

Klarna is the case that made this public.

In February 2024, the company’s CEO said its AI assistant was handling 75% of support chats, doing the work of roughly 700 human agents, and processing 2.3 million conversations in a single month. That part was true, and it was genuinely impressive. What wasn’t said out loud yet was what happened to the other 25%, the angry customer, the ambiguous dispute, the case that needed someone to actually use judgment.

By mid-2025, the same CEO admitted to Bloomberg that cost had been “too predominant” a factor in the original decision, that service quality had dropped, and that Klarna was rehiring humans and guaranteeing every customer a path to one (Silicon Canals, Klarna U-turns on AI).

The AI didn’t get worse at handling volume. It never had a volume problem. What broke was the assumption that the hard 25% would somehow shrink to fit the tool, instead of the tool being scoped to fit the 75% it was actually good at.

Support Queue Breakdown-selection
Every support queue has the same two halves. Klarna’s failure was forcing the second half through a tool built for the first.

Klarna’s Mistake Wasn’t a Technology Problem

It’s tempting to read the Klarna story as “the model wasn’t good enough yet.” That’s not really what happened, and treating it that way is how the same mistake gets repeated with a better model next year.

A conversational agent that’s designed to sound confident by default will answer a question it has no business answering, because nothing in how it was built rewards it for saying “I don’t know, let me get you a person” instead of guessing well.

That’s a design decision, not a technology ceiling. The fix isn’t a smarter model, it’s an explicit, tested line for when the agent stops and hands off, decided before launch instead of discovered after a customer post about it.

That’s the same principle behind Human in the Loop: the goal was never to automate everything. It was to automate what’s actually safe to automate and be honest about the rest.

Questions to Ask Before Deploying an AI Customer Service Agent

Almost every conversational AI failure traces back to one of three questions that never got a real answer before go-live:

The three questions that separate a successful deployment from a public failure.

What actually triggers a handoff?

Not “the model decides it’s stuck.” A specific, testable threshold, repeated failed attempts, certain account states, detected frustration, that someone can point to and defend, not a vague hope that the system will know when it’s out of its depth.

What does the human on the other end receive?

If a customer has to re-explain their problem to a person after the bot already failed, you’ve built the exact experience that made 60% of Gartner’s respondents worry about reaching a human in the first place. The handoff needs the full transcript, not a cold transfer.

Is anyone tracking the escalations that didn’t happen?

Resolution rate is the easy number to report and the wrong one to optimize alone. The harder, more honest number is how often the agent should have handed off and didn’t. That’s usually where the actual damage was done, quietly, long before anyone noticed.

What 90/10 Actually Looks Like

This isn’t a customer service problem specifically. Pendoah has built the same shape of solution elsewhere, and it looks identical.

Worklighter’s document-automation engine auto-processes 90% of incoming invoices and routes the remaining 10 to 20% to a person, and that gap isn’t a defect anyone’s trying to close to zero. It’s where the ambiguous, judgment-heavy cases actually live, and forcing them through automation anyway is exactly the mistake Klarna made in a different industry.

The number that matters for a conversational agent isn’t how close you can get to 100% automated. It’s whether you actually know from real ticket data rather than a guess, what your version of that 90/10 split looks like, and whether the 10% has a real place to go.

Is Your Business Ready for Conversational AI Agents?

The honest version of “should we deploy this” isn’t a checklist; it’s one question with two very different answers depending on why you’re asking. If you’re looking at high-volume, well-defined requests where the answer already lives in a system you can query, and you’re doing this to free your team’s time for the harder cases, you’re describing the 70% Gartner’s forecast is actually about, and you’re probably ready.

If you’re looking at this because support headcount is the line item you want to shrink, or because you haven’t actually measured what share of your tickets are routine versus genuinely hard, or because there’s no tested, honest answer to what happens when the agent gets it wrong, you’re describing the setup that produced Klarna’s reversal.

The technology isn’t the variable that changes that outcome. The honesty of the scoping is, and it’s exactly what Pendoah’s Conversational AI Development Services start with before any bot gets built.

Key Questions to Ask Before You Deploy

  • What percentage of our actual ticket volume is genuinely routine, based on real data rather than a guess?
  • What specific signals trigger a handoff to a person, and has that threshold been tested against real conversations?
  • Does a human agent receiving an escalated conversation get full context, or does the customer start over?
  • Are we tracking missed escalations, cases the agent should have handed off and didn’t, not just resolution rate?
  • If a customer explicitly asks for a human, how many steps does it take to reach one?

The Bottom Line

Conversational AI agents aren’t good or bad. They’re scoped well or scoped carelessly.

Gartner’s 80% and 64% aren’t a contradiction; they’re two ways of describing the same tool: excellent at the routine 70% and genuinely damaging when a company mistakes that for permission to point it at everything else.

Klarna found that line the expensive way. You don’t have to.

Want help figuring out which share of your support volume is actually ready for this, and which isn’t?

Book a consultation with Pendoah, or look through our case studies to see how this decision played out for other clients.

Sources

  1. Gartner. “Gartner Predicts Agentic AI Will Autonomously Resolve 80% of Common Customer Service Issues Without Human Intervention by 2029.” March 5, 2025. Read the Gartner announcement
  2. Gartner. “Gartner Survey Finds 64% of Customers Would Prefer That Companies Didn’t Use AI for Customer Service.” July 9, 2024. Read the Gartner survey findings
  3. Silicon Canals. “Klarna U-turns on AI: Rehires Humans for Customer Support After Quality Complaints.” Read the Silicon Canals article
  4. Pendoah. “Worklighter Document Automation Engine.” View the Worklighter case study

Frequently Asked Questions About Conversational AI Agents

A chatbot answers one scripted question at a time. A conversational AI agent works toward an outcome, pulling data, taking action, and deciding whether to hand off, all inside one conversation instead of one reply.

Not for the range of things customer service handles. They’re strong on structured, high-volume, low-ambiguity requests. Gartner itself projects that no Fortune 500 company will have fully eliminated human customer service by 2028, and the emotionally charged or ambiguous cases are exactly why.

Its CEO said cost reduction had been too dominant a factor in the original rollout, and that service quality suffered as a result. Klarna is now rehiring human agents and guarantees every customer a path to one, while keeping AI on the higher-volume, routine work it was built for.

Not by resolution rate alone. Track how often it correctly hands off cases it can’t handle, not just how often it resolves what it can. An agent that never escalates can look successful right up until it produces a public failure.

Conversational AI agents work best for high-volume, repetitive customer requests where answers can be grounded in reliable company data and clear workflows. They can handle tasks like FAQs, order updates, account support, and basic troubleshooting, while more complex, sensitive, or high-stakes issues should be escalated to a human agent. 

Ready to See Your AI ROI?

Book a 30-minute regulatory assessment.

Subscribe

Get exclusive insights, curated resources and expert guidance.

Insights That Drive Decisions

Let's Turn Your AI Goals into Outcomes. Book a Strategy Call.