ScruTool
Technology

Why AI Support Bots Fail on Real Customer Questions (And How to Diagnose Yours)

AI support bot failing on real customer questions? Discover the 5 failure layers, how to audit your chatbot, spot hidden problems, and improve customer support.

Sep 8, 2026 17 min read

In April 2025, developers using the Cursor code editor started getting logged out whenever they switched between machines. Several emailed support and received a reply from someone called Sam, explaining that this was expected behaviour under a policy limiting each subscription to a single device. No such policy existed. Sam was an AI support bot, and it had invented the rule. Users posted the reply to Reddit and Hacker News, cancellations followed, and Cursor co-founder Michael Truell publicly apologised, tracing the logouts to a session management bug.

One detail from that incident matters more than the rest. Sam did not tell everyone the same thing. Some users were given the fabricated restriction, others were not, so customers comparing notes could not work out whether the policy was real.

That inconsistency is a diagnostic signal, and it points at something far more specific than "AI makes mistakes." Support bots fail at one of five identifiable layers, and the layer determines the fix. This article shows you how to find yours. There is no product at the end of it.

The Metric That Hides the Problem

Most teams discover their bot is failing months after it started failing, because the dashboard has been reporting success the entire time.

Deflection and containment are not resolution

Deflection counts conversations that never became tickets. Containment counts conversations that ended inside the bot. Neither tells you the customer got what they came for. A customer who gives up in frustration and never contacts you again registers as a win on both.

Resolution is the only measure tied to an outcome, and the gap between the three is wide. A Gartner survey of 5,728 customers found that just 14% of customer service issues are fully resolved in self-service. Even for problems the customers themselves described as very simple, the figure only reached 36%, while 73% of customers use self-service at some point in their journey. The most common reason for failure was unglamorous: in 43% of cases, people could not find content relevant to their issue. A further 45% who started in self-service said the company did not understand what they were trying to do.

The cost case for automation almost always rests on the assumption that a contained conversation is a resolved one. Gartner benchmarks self-service at roughly $1.84 per contact against $13.50 for agent-assisted support. That arithmetic only works if the cheap contact actually ended the problem. When it does not, you have spent $1.84 irritating someone before spending the $13.50 anyway, and the agent now starts from a worse position than if the bot had never intervened.

Why the demo score never survives production

Vendor demonstrations routinely show 90% or better automation. Production figures across real implementations land closer to 55% to 70%.

A large part of that gap is variance rather than capability. Tau-bench, a benchmark built by Sierra to test agents on realistic customer service tasks with domain policies and working tools, introduced a metric called pass^k: the probability that all k attempts at the same task succeed, rather than at least one. In the original 2024 results, GPT-4o scored 61% on a single attempt at retail tasks and under 25% across eight. Same task, same model, wildly different outcomes. The paper is worth reading in full if you are evaluating vendors: arxiv.org/abs/2406.12045.

Support is a repetition business. A bot that answers correctly once during a scripted demo tells you almost nothing about what happens when 400 people ask the same question on a Monday morning.

What a Real Customer Question Actually Looks Like

The second reason bots pass internal testing and fail live is that the test set is wrong. Teams build their test questions from knowledge base article titles, which are, by definition, the questions the company already knows how to answer.

Real questions arriving at support are the residue of everything self-service already handled. They are the hard ones by construction.

Six kinds of question, and how far automation gets

Question typeExampleRealistic automation ceiling
Factual lookup"What are your delivery times to Manchester?"High. This is what bots are good at.
Account-specific"Where is order 88214?"High, but only if the bot can read your order system.
Compound"I was charged twice and I also need to change my delivery address."Low. Two intents, and most bots answer one then drop the other.
Policy exception"I know it is past 30 days, but the box was never opened."Low. Requires discretion the bot does not have.
Emotionally loaded"This is the fourth time I have contacted you about this."Should not be automated at all.
Underspecified"It is not working."Depends entirely on whether the bot asks before answering.

Run your last 200 tickets through that table before you run them through anything else. If the bottom four rows account for most of your volume, your automation ceiling is structural, and no amount of prompt tuning moves it.

Your customers do not use your words

A customer who encounters a documentation heading like "Subscription Lifecycle Management" will write "how do I stop paying you." That mismatch is not a wording preference. It is a retrieval failure waiting to happen, which is where the anatomy of the problem properly begins.

The Five Layers Where AI Support Bots Break

Vendor blog posts tend to describe failures in terms of what the bot lacked: empathy, context, training. That framing is not actionable. A support conversation moves through five distinct stages, and a failure at each one looks different, is measured differently and is fixed differently.

The Five-Layer Failure Stack. Original framework.

Layer 1: Retrieval, where the answer was never found

The bot searches your knowledge base and pulls back the wrong documents, or nothing useful at all. Research on production retrieval systems has consistently identified this stage as the dominant failure point, with the canonical taxonomy naming missing content, correct documents that ranked too low to be returned, and retrieved passages that never made it into the model's context.

Retrieval failures are also the most common thing teams misdiagnose as a model problem. Swapping to a more capable model does nothing if the required document does not exist, is three versions out of date, or is written in language no customer would ever type.

How to spot it:

Log what was retrieved, not just what was said. If the answer document was never in the top results, the model was never given a chance.

Layer 2: Grounding, where the bot invents a policy

The retrieval step returns something thin or ambiguous, and the model fills the gap with a fluent, confident, entirely fictional answer. This is what happened to Cursor, and the reason it caused cancellations rather than confusion is that the invented policy sounded exactly like a real one.

Fluency and accuracy are separate properties of a language model, and customers have no way to tell them apart. The fix is to make the bot willing to say it does not know. A support bot that refuses 15% of questions and is right about the rest is worth considerably more than one that answers everything and is wrong 15% of the time, because the first failure mode routes to a human and the second one reaches the customer.

Layer 3: State, where the bot loses the thread

The customer explains their situation across several messages. By turn five, detail from turn two has vanished, and the bot asks for information it was already given.

Microsoft Research and Salesforce Research tested this directly, simulating over 200,000 conversations across leading open and closed models. Performance dropped by an average of 39% in multi-turn conversations where instructions were revealed gradually rather than stated up front. The researchers found the cause was not lost capability but a sharp rise in unreliability, and the degradation appeared in conversations as short as two turns.

This is worth sitting with, because real support conversations are almost never fully specified in the first message. Underspecification is the normal condition, not an edge case. Larger context windows do not solve it, since the problem is that the model commits to an interpretation early and then defends it.

Layer 4: Action, where the bot can answer but cannot do

The customer asks for a refund. The bot explains the refund policy accurately, then tells them to contact support. Nothing was wrong with the answer, and nothing was resolved either.

Read-only bots hit a hard ceiling on anything requiring a write: refunds, cancellations, address changes, plan swaps, replacement orders. This ceiling has nothing to do with model quality and everything to do with integration work nobody scoped. It is also the layer where the cheapest-looking deployments run out of road fastest.

Layer 5: Handoff, where escalation makes things worse

The bot gives up and transfers to a human, and the human receives a ticket with no history attached. The customer starts again from the beginning, having already spent eight minutes with a machine.

A correct handoff carries the full conversation, what the bot attempted, how it classified the issue, the customer's account history and a suggested next step. Anything less converts a failed automation into a worse-than-baseline human interaction, because the agent is now handling both the original problem and the frustration the bot created.

LayerSymptom in the transcriptWhat to inspectWhere the fix lives
Retrieval"That doesn't answer my question"Retrieved document list per turnKnowledge base coverage and structure
GroundingConfident answer contradicting real policyWhether the claim traces to a sourceRefusal thresholds and citation requirements
StateBot re-asks for known informationTurn number where detail was droppedExplicit state extraction, early clarification
ActionCorrect answer, no resolutionWhich write actions existSystem integration and permissions
HandoffCustomer repeats everything to the agentContents of the escalation payloadHandoff package design

The Real Question Audit: Diagnosing Your Own Bot in an Afternoon

The next step is turning the framework above into a number you can act on. This takes roughly 90 minutes and needs nothing beyond your ticket export and a spreadsheet.

1.     Build the test set from tickets, not imagination. Pull 100 real tickets, weighted toward escalations, complaints, long threads and repeat contacts. Do not use knowledge base titles. Using questions you already documented guarantees a false pass, which is exactly how most bots earn their green dashboard.

2.     Replay every question five times. Ask each question five separate times in fresh sessions. Variance is the finding, not noise. If the same question produces three different answers, you have a reliability problem that no single-run accuracy score will ever surface.

3.     Score by layer, not by correct or incorrect. For every failure, assign it to one of the five layers instead of marking it right or wrong. Binary scoring tells you the bot is bad. Layer scoring tells you what to do on Monday.

4.     Read the distribution. A retrieval-heavy profile means your content is the problem and your model is fine. A grounding-heavy profile means your refusal behaviour is too permissive. A handoff-heavy profile means the bot is roughly working and your escalation design is not.

5.     Set a reliability gate. Decide the consistency threshold you need before you widen scope, then hold to it. The chart below shows why single-run accuracy flatters you.

The practical translation is uncomfortable. At a thousand conversations a day, a bot that is right 80% of the time produces 200 wrong answers daily and roughly 73,000 a year. The percentage reads like a pass. The absolute number reads like a staffing plan.

Fixing It, In Order of Impact

Once you know your distribution, the sequence matters as much as the interventions.

•     Fix the knowledge layer before you touch the model. Rewriting articles in customer language and closing coverage gaps beats a model upgrade in almost every retrieval-dominant profile, and costs less. Gartner's own recommendation points the same way: make content creation part of the resolution workflow so agents write knowledge as they close tickets, rather than treating documentation as a separate project nobody has time for.

•     Narrow the scope until it holds, then widen. Restrict the bot to the question types where it demonstrably clears your reliability gate, and route everything else immediately. A bot handling 30% of volume well is worth more than one handling 70% unreliably, because the second one damages trust on the easy questions too.

•     Design escalation as a feature, not a fallback. Trigger on sentiment and on repeated failure, escalate after two or three unsuccessful attempts or the moment a customer asks for a person, and never bury the human option. Gartner found that the single biggest customer concern about AI in service, cited by 60% of respondents, is that it makes reaching a human harder.

•     Close the loop monthly. Every answer a customer flags as wrong should trigger a content review and land in your eval set. Without this, the audit becomes a one-off snapshot rather than a control system.

The Liability Nobody Budgeted For

There is a layer beneath all five that most deployment plans skip entirely, and a Canadian tribunal has already ruled on it.

In Moffatt v. Air Canada (2024 BCCRT 149), a customer relied on an airline chatbot's description of how to claim a bereavement fare. The advice was wrong. Air Canada argued, in the tribunal's summary, that the chatbot was a separate legal entity responsible for its own actions. The tribunal rejected this, finding that the airline was responsible for all information on its website whether it appeared on a static page or came from a chatbot, and awarded the customer $812.02 in damages. It also dismissed the suggestion that the customer should have cross-checked the bot against another page of the same site.

The damages were small. The principle is not. Your bot's answers are your company's statements, which makes logging and retention of automated support decisions an operational requirement rather than a nice-to-have. If you cannot reconstruct what your bot told a specific customer on a specific date and what source it drew on, you cannot defend the interaction.

Where the Honest Boundary Sits

The best-documented correction in this category belongs to Klarna. In February 2024 the company announced that its OpenAI-built assistant had handled 2.3 million conversations in its first month, work it equated to roughly 700 agents, with projected annual savings around $40 million.

By 2025, chief executive Sebastian Siemiatkowski was describing the outcome differently, telling Bloomberg that cost had become too dominant an evaluation factor and that the result was lower quality. He added that customers must always be able to reach a human. Klarna has publicly disputed the "reversal" framing, stating that it continues to invest heavily in AI and that the new arrangement is a dual-track approach combining automation with human support rather than a retreat from it. Both things are true at once, and the honest reading is that the company mismeasured, then corrected.

The pattern is broad enough that Gartner now forecasts that by 2027, half of companies that attributed headcount reductions to AI will rehire staff for similar functions under different job titles. Emily Potosky, a senior director of research in Gartner's customer service practice, put the reasoning plainly in the accompanying release: "AI simply isn’t mature enough to fully replace the expertise, empathy, and judgment that human agents provide." The same body of research found that 64% of customers would prefer companies did not use AI for service at all, and 53% would consider switching to a competitor over it.

None of this argues against automating support. It argues against automating it on a number that does not measure whether customers got what they needed.

Run the audit and one of two things happens. Your failures cluster in retrieval, state or handoff, which means you have an engineering problem with a known fix and a defined cost. Or they cluster in grounding and action, which usually means the bot has been pointed at questions that require discretion or system access it will not have this year, and the answer is to shrink its remit rather than improve its prompts.

The second result is far more common than the first, and it is the one teams resist, because narrowing scope looks like admitting the project failed. It is the opposite. A support bot that handles a third of your volume and is trusted on all of it compounds in value, because every reliable answer teaches customers that the bot is worth trying. A bot stretched across everything teaches them the reverse, and that lesson is much harder to unteach than it was to trigger.

Where This Leaves You

Everything above reduces to a single operating decision, and it is not which vendor to buy.

Before your next renewal conversation, run the audit on a hundred real tickets and put the layer distribution next to your reported deflection rate in the same document. Those two numbers rarely agree. The disagreement between them is the most useful thing your support function will produce this quarter, because one of them describes what you have been reporting upward and the other describes what your customers actually experienced.

Then set a consistency threshold you are willing to enforce. Decide what a question type has to clear across repeated runs before the bot handles it without supervision, write the number down, and hold every scope expansion to it. Teams that skip this do not fail loudly. They drift, adding coverage a percentage point at a time until the bot is fielding questions nobody ever tested it on, and the first person to notice is a customer.

The uncomfortable part is that a good result often looks like a smaller deployment. Pulling your bot back to the question types it clears reliably will show up as a lower automation rate on the dashboard and a higher one in the only place that matters, which is whether the customer's problem ended.

In two years, the teams running support automation that customers trust will be the ones that measured early and were willing to hand work back to humans before their customers forced the decision.

Community

Discussion

Join the discussion and share your perspective.

Related Articles