New Epignosis Insights Report Exposes the Gap Between Generative AI's Customer Service Promise and Real-World Performance
Epignosis Insights, a Global Market Research and publishing firm, today released a new Customer Experience Report, “Generative AI in Customer Service: Promise Versus Real-World Performance,” assessing how generative AI is actually performing in customer service operations, as distinct from how it is being marketed.
A Widening Gap Between Forecast and Deployment
The report finds the gap between projection and reality substantial. Gartner has told the market that agentic AI will autonomously resolve 80% of common customer service issues by 2029 while cutting operational costs by 30% — and, in the same research cycle, has warned that more than 40% of agentic AI projects will be canceled by the end of 2027 due to unclear business value and inadequate governance. Klarna, the most cited real-world case study in the industry, publicly claimed its AI assistant did the work of 700 customer service agents in 2024, then quietly began rehiring humans in 2025 after complex-case quality deteriorated, before stabilizing into a hybrid model that its own Q3 2025 disclosures describe as automating the equivalent of 853 full-time positions and generating $60 million in savings.
What the Projections Say
McKinsey's contact-center research offers a more measured version of industry optimism, the report notes, estimating generative AI could reduce human-serviced contacts by up to 50% in banking, telecommunications, and utilities, and documenting a case study in which a 5,000-agent operation achieved a 14% increase in issue resolution per hour alongside a 9% reduction in handling time after deploying a generative AI assistant. Separately, McKinsey has quantified the cost impact directly: companies implementing AI in customer support have reduced average cost per interaction by 68%, from $4.60 to $1.45.
Klarna's Rise, Reversal, and Rebalancing
Klarna remains the industry's defining case study precisely because the company was transparent about both the initial success and the subsequent correction, according to the report. Launched in February 2024, Klarna's AI assistant processed 2.3 million conversations in its first month, handled more than two-thirds of all customer chats across 23 markets and 35 languages, and cut average resolution time from 11 minutes to under 2 minutes. By May 2025, however, Klarna's CEO publicly acknowledged the company had focused too much on efficiency and cost, resulting in lower quality, and the company began rehiring human agents. Subsequent reporting indicates this was less a full reversal than a scope correction: the AI assistant remained the front line for high-volume, routine queries, while humans were reintroduced specifically for premium accounts and complex, emotionally sensitive cases.
Agent Washing and the Governance Problem
Gartner's own research provides the clearest documented evidence of a structural gap between marketing and deployed capability, the report finds. The firm estimates that of the thousands of vendors currently marketing “agentic AI” customer service products, only approximately 130 offer genuine agentic capability — systems that can reason across multi-step problems, coordinate across channels, and take autonomous action within enterprise systems. Gartner labels the remainder “agent washing”: existing chatbots and automation tools rebranded with agentic terminology but no meaningful new capability, typically handling only 20–30% of interactions in production. Gartner has separately predicted that roughly half of the companies that cut customer service staff in favor of AI will rehire for those roles by 2027, often under new job titles.
“The evidence shows generative AI delivers real, measurable value in customer service — but almost never in the unqualified, fully autonomous form vendors describe in press releases. The winning model is hybrid, not fully autonomous”
The Consumer Trust Gap
Independent survey research from Salesforce, cited in the report and based on responses from 16,585 consumers and business buyers worldwide, quantifies a consumer trust problem sitting underneath the deployment statistics. Seventy-two percent of consumers say they trust companies less than they did a year ago, and 60% believe that advances in AI make organizational trustworthiness even more important, not less. Seventy-five percent of customers say it is important to know whether they are communicating with an AI agent rather than a human, and 45% say they are more likely to use an AI agent at all if a clear escalation path to a human exists — findings that align closely with Klarna's course correction toward a hybrid model.
Regulators Are Watching Closely
U.S. regulators have moved from watching this space to actively enforcing against it, the report details. The Federal Trade Commission's Operation AI Comply, launched in September 2024, has produced at least eight settled enforcement cases targeting companies that overstated AI capability, including a $193,000 settlement over an “AI lawyer” chatbot that was never adequately tested against the legal expertise it claimed. In September 2025, the FTC broadened its focus with a formal inquiry into seven major AI chatbot providers, and a December 2025 executive order directed the FTC to clarify that misrepresenting the likelihood of AI hallucinations to consumers can constitute deception under the FTC Act.
Where Generative AI Actually Delivers Value
None of this evidence suggests generative AI is failing in customer service, the report emphasizes — it suggests the value is concentrated in a narrower band than initial marketing implied. McKinsey's documented case results — 14% higher resolution per hour, 9% lower handling time, cost per interaction cut by roughly two-thirds — are real, repeatable gains, concentrated specifically in high-volume, low-complexity, well-structured interactions such as order status, password resets, and appointment scheduling. Deutsche Telekom's own operational leadership frames a realistic efficiency gain at around 30% over two to three years, substantially short of the 80% autonomous-resolution figure sometimes used in vendor marketing.
Outlook: Hybrid Models Emerging as Industry Consensus
Taken together, the government enforcement record, industry association and consulting-firm forecasts, and corporate disclosures compiled in the report point toward a converging industry consensus: the winning model is hybrid, not fully autonomous. The report argues the more useful planning question for organizations is not “how much of customer service can AI replace,” but “which specific interaction types can AI safely own, and what does the escalation path to a human need to look like for the rest” — with companies asking that narrower question, Klarna's 2025 correction being the clearest public example, appearing to converge on better outcomes than those chasing an immediate full-autonomy target.