Generative AI in Customer Service: Promise Versus Real-World Performance

Generative AI in Customer Service: Promise Versus Real-World Performance

Epignosis Insights Research Desk — Compiled and analyzed from government, industry association, corporate, consulting, and news sources

Report ID: CB13 | Format: PDF, Excel | Publish Date: August 2026 | Pages: 120

Executive Summary

This report, produced by the Epignosis Insights Research Desk, synthesizes primary disclosures and third-party research to assess how generative AI is actually performing in customer service, as distinct from how it is being marketed. The gap is substantial. Gartner has told the market that agentic AI will autonomously resolve 80% of common customer service issues by 2029 while cutting operational costs by 30% — and, in the same research cycle, has warned that more than 40% of agentic AI projects will be canceled by the end of 2027 due to unclear business value and inadequate governance. 

Klarna, the most cited real-world case study in the industry, publicly claimed its AI assistant did the work of 700 customer service agents in 2024, then quietly began rehiring humans in 2025 after complex-case quality deteriorated, before stabilizing into a hybrid model that Klarna’s own Q3 2025 disclosures describe as automating the equivalent of 853 full-time positions and generating $60 million in savings. Reading these threads together, compiled from company filings, regulatory actions, consulting-firm forecasts, and direct news reporting, the clearest finding is that generative AI delivers real, measurable value in customer service — but almost never in the unqualified, fully autonomous form vendors describe in press releases.

The Promise: What the Projections Say

The most aggressive claims come from the analyst community itself. Gartner’s March 2025 forecast that agentic AI will autonomously resolve 80% of common customer service issues without human intervention by 2029 was widely circulated as evidence that full automation of service functions was imminent. McKinsey’s contact-center research offers a more measured version of the same optimism: the firm estimates generative AI could reduce human-serviced contacts by up to 50% in banking, telecommunications, and utilities, and documents a case study in which a 5,000-agent operation achieved a 14% increase in issue resolution per hour alongside a 9% reduction in handling time after deploying a generative AI assistant. Separately, McKinsey has quantified the cost impact directly: companies implementing AI in customer support have reduced average cost per interaction by 68%, from $4.60 to $1.45. These are not vendor marketing claims they are findings from one of the most conservative consulting firms in the market, which makes the subsequent gap between promise and delivery more notable, not less.

Case Study: Klarna’s Rise, Reversal, and Rebalancing

Klarna remains the industry’s defining case study precisely because the company was transparent about both the initial success and the subsequent correction. Launched in February 2024 on OpenAI’s models, Klarna’s AI assistant processed 2.3 million conversations in its first month, handled more than two-thirds of all customer chats across 23 markets and 35 languages, and cut average resolution time from 11 minutes to under 2 minutes. CEO Sebastian Siemiatkowski publicly framed the deployment as equivalent to replacing 700 full-time agents. By May 2025, in interviews with Bloomberg and Reuters, Siemiatkowski acknowledged the company had "focused too much on efficiency and cost" and that "the result was lower quality, and that’s not sustainable" and Klarna began rehiring human agents, initially through a flexible, remote contractor model. Trade publication CX Dive and other outlets characterized this directly as Klarna "turning back to people." Critically, subsequent reporting indicates this was less a full reversal than a scope correction: the AI assistant remained the front line for high-volume, routine queries, while humans were reintroduced specifically for premium accounts and complex, emotionally sensitive cases. By the third quarter of 2025, Klarna’s own earnings disclosures described the assistant as automating work equivalent to 853 full-time agents and generating $60 million in cumulative savings — a larger automated footprint than the original 700-agent claim, achieved after the company learned which categories of interaction AI could and could not safely own.

Klarna’s AI assistant performance metrics, pre- and post-deployment.
Figure 1: Klarna’s AI assistant performance metrics, pre- and post-deployment.

The Governance Problem: Agent Washing and the Promise-Reality Gap

Gartner’s own research provides the clearest documented evidence of a structural gap between marketing and deployed capability. The firm estimates that of the thousands of vendors currently marketing "agentic AI" customer service products, only approximately 130 offer genuine agentic capability systems that can reason across multi-step problems, coordinate across channels, and take autonomous action within enterprise systems. Gartner labels the remainder "agent washing": existing chatbots, robotic process automation tools, and virtual assistants rebranded with agentic terminology but no meaningful new capability. Industry analysis following up on Gartner’s research estimates that these relabeled systems typically handle only 20–30% of interactions in production, far short of the 80% autonomous-resolution figure Gartner itself projects for the technology category by 2029. Gartner has separately predicted that in 2026, one-third of companies will actively harm their customer experience by deploying AI prematurely through misread personalization, compliance violations, or poorly timed automated outreach eroding both brand trust and customer retention. Perhaps most tellingly, Gartner’s own research anticipates that roughly half of the companies that cut customer service staff in favor of AI will rehire for those roles by 2027, often under new job titles, effectively predicting the Klarna pattern will repeat industry-wide.

Gartner’s own forecasts, compared side by side.
Figure 2: Gartner’s own forecasts, compared side by side.

The Trust Gap: What Customers Actually Think

Independent survey research from Salesforce’s State of the AI Connected Customer study, based on responses from 16,585 consumers and business buyers worldwide, quantifies a consumer trust problem that sits underneath the deployment statistics. Seventy-two percent of consumers say they trust companies less than they did a year ago, and 60% believe that advances in AI make organizational trustworthiness even more important, not less. Transparency emerged as the central demand: 75% of customers say it is important to know whether they are communicating with an AI agent rather than a human, and 45% say they are more likely to use an AI agent at all if a clear escalation path to a human exists. These findings matter directly for the Klarna case: the company’s course correction toward a hybrid model, preserving human access for complex and premium interactions, aligns closely with what Salesforce’s survey data says customers actually want, rather than the fully autonomous model the original 2024 announcement implied.

Consumer trust and transparency preferences around AI agents.
Figure 3: Consumer trust and transparency preferences around AI agents.

Regulatory Scrutiny: The FTC and the Deception Question

U.S. regulators have moved from watching this space to actively enforcing against it. The Federal Trade Commission’s Operation AI Comply, launched in September 2024, has produced at least eight settled enforcement cases targeting companies that overstated AI capability, with the DoNotPay case an "AI lawyer" chatbot the FTC found was never adequately tested against the legal expertise it claimed resulting in a $193,000 settlement finalized in February 2025 and a permanent order barring unsubstantiated capability claims. In September 2025, the FTC broadened its focus with a Section 6(b) inquiry into seven major AI chatbot providers, formally requesting information on how these companies measure, test, and monitor chatbot safety and accuracy claims. A December 2025 executive order directed the FTC to issue a policy statement clarifying that while hallucinations themselves may not automatically violate Section 5 of the FTC Act, misrepresenting the likelihood or frequency of hallucinations to consumers can constitute deception — a standard directly relevant to any customer service deployment marketed as more reliable or more "human-equivalent" than its actual error rate supports.

Where Generative AI Actually Delivers Value

None of this evidence suggests generative AI is failing in customer service it suggests the value is concentrated in a narrower band than initial marketing implied. McKinsey’s documented case results 14% higher resolution per hour, 9% lower handling time, cost per interaction cut by roughly two-thirds are real, repeatable gains, concentrated specifically in high-volume, low-complexity, well-structured interactions: order status, password resets, refund tracking, appointment scheduling. Deutsche Telekom’s own operational leadership, cited in McKinsey’s contact-center research, frames the realistic efficiency gain at around 30% over two to three years — substantial, but far short of the 80% autonomous-resolution figure sometimes used in vendor marketing. The pattern that recurs across McKinsey’s research, Klarna’s disclosed metrics, and Salesforce’s consumer survey data is consistent: generative AI performs reliably on structured, low-emotional-stakes interactions, and underperforms — sometimes badly — on complex, ambiguous, or emotionally charged interactions where customers specifically want to reach a human.

Outlook: The Hybrid Model as Emerging Consensus

Taken together, the government enforcement record, the industry association and consulting-firm forecasts, and the corporate disclosures compiled in this report point toward a converging industry consensus: the winning model is hybrid, not fully autonomous. Gartner’s own rehiring prediction, Klarna’s realized operating model, Salesforce’s consumer preference data on escalation paths, and McKinsey’s realistic efficiency estimates all describe the same structure AI absorbing high-volume, low-complexity contact volume while human agents retain ownership of complexity, empathy, and judgment-dependent cases. For organizations evaluating generative AI customer service investments, the evidence compiled here, drawn from government, industry association, corporate, consulting, and news sources but assessed collectively by Epignosis Insights, suggests the more useful planning question is not "how much of customer service can AI replace," but "which specific interaction types can AI safely own, and what does the escalation path to a human need to look like for the rest." Companies that have asked the narrower question Klarna’s 2025 correction being the clearest public example appear to be converging on better outcomes than those that treated the 80%-autonomy narrative as an immediate deployment target.

Industry Response: Governance and Certification

Professional bodies within the customer experience field are beginning to formalize the governance frameworks this transition requires. The Customer Experience Professionals Association (CXPA), the field’s primary global non-profit body, has expanded its guidance and certification materials to address AI integration directly, publishing practical guidance for embedding AI into customer experience strategy, governance, and operations "while strengthening trust, leadership, and organizational readiness." This reflects a broader shift within the profession: CX leadership is increasingly treated as a governance discipline — defining which interaction types are appropriate for automation, setting escalation thresholds, and monitoring quality — rather than a purely technical deployment decision left to IT or vendor teams. That shift mirrors Gartner’s own diagnosis of why agentic AI projects fail: not because the underlying models lack capability, but because organizations deploy them "without a clear strategy, without understanding the complexity, and without the governance to manage what happens when something goes wrong," in the words of Gartner senior director analyst AnushreeVerma.

Sector Variation: Not All Customer Service Is Equal

The promise-versus-performance gap is not uniform across industries, and the underlying data compiled for this report shows meaningfully different automation ceilings by sector. McKinsey’s research on human-serviced contact reduction is explicitly scoped to banking, telecommunications, and utilities — sectors with high transaction volume and relatively standardized query types — where the firm sees up to 50% reduction potential. Deutsche Telekom’s own operational leadership, by contrast, points to telecommunications’ comparatively high service complexity as a reason to expect workforce efficiency gains closer to 30% over two to three years rather than the more aggressive figures cited for simpler transactional sectors. Financial services faces an additional layer of constraint: the same complexity that makes AI valuable for fraud checks and payment processing also invites the closest regulatory scrutiny, since misrepresenting AI accuracy or reliability in a financial context carries direct consumer-harm exposure of the kind the FTC has already signaled it will pursue. The practical implication for any organization benchmarking its own AI customer service investment against the statistics in this report is that industry context changes both the realistic automation ceiling and the acceptable risk tolerance for getting it wrong.

Data Table: Promise vs. Performance at a Glance

Metric Figure
Agentic AI autonomous resolution forecast (by 2029) 80%
Agentic AI projects predicted cancelled by 2027     40%+
Real agentic AI vendors (of thousands claiming it) ~130
Klarna AI assistant, FTE-equivalent automated (Q3 2025) 853
Consumers who trust companies less than a year ago 72%
Cost per interaction reduction with AI (contact centers) $4.60→$1.45
FTC AI enforcement cases settled since 2022 8+

Frequently Asked Questions

Did Klarna actually reverse its AI strategy?
Not fully; Klarna rebalanced toward a hybrid model, keeping AI as the front line for high-volume queries while rehiring humans for premium and complex-case support.
What share of "agentic AI" vendors offer real capability?
Gartner estimates only about 130 of the thousands of vendors marketing agentic AI products offer genuine agentic functionality.
How much do customers trust AI-driven customer service today?
Salesforce survey data shows 72% of consumers trust companies less than a year ago, and 75% want to know when they are talking to an AI agent.
Has the FTC taken action against overstated AI customer service claims?
Yes; the FTC has settled 8+ AI enforcement cases since 2022 under Section 5 of the FTC Act, including the $193,000 DoNotPay settlement.
Where does generative AI deliver the most reliable value?
McKinsey research shows the clearest gains in high-volume, low-complexity interactions, cutting cost per interaction from $4.60 to $1.45 on average.

Small Analyst Support Card

Need Help Choosing the Right Report

Talk to Our Analyst