Deflection that does not cost you the customer
An agent that never hands over will show you an excellent deflection rate. It will also be quietly generating the repeat contacts you see next week.
AI & AgentforceThere is a straightforward way to get your deflection rate to ninety percent. Stop the agent from escalating.
Nobody sets out to do this. It happens because deflection is the number on the slide, escalation is not, and every incremental decision that makes the agent try a little harder before handing over moves the number in the right direction. Six weeks later the deflection chart looks excellent and the phone queue has not moved, because the same customers are coming back through a different door.
Why deflection alone is a vanity metric
Deflection measures conversations the agent handled without a human. It says nothing about whether the customer's problem was solved, and those two things come apart in a specific and predictable way.
A customer asks a question. The agent gives a plausible, incorrect answer. The customer accepts it, leaves, discovers it was wrong, and calls three days later. That conversation counted as deflected. It also generated a phone contact, a longer one than the original would have been, with a customer who now trusts you less.
The metric is not wrong, it is incomplete. It needs three companions.
| Metric | What it catches | Warning sign |
|---|---|---|
| Deflection rate | Volume handled without a human | Rising while others deteriorate |
| Repeat contact within 7 days | Answers that did not actually resolve | Rising alongside deflection |
| Escalation rate | Whether the agent knows its limits | Falling toward zero |
| Sampled accuracy review | Whether answers were correct | Anything below your published bar |
The combination that should alarm you is deflection up, escalation down, repeat contact up. That is not an agent getting better. That is an agent getting worse at knowing when to stop.
The escalation rate paradox
Most organisations treat a falling escalation rate as improvement. Early in a deployment it usually is not.
A new agent should escalate a lot, because it is meeting intents outside its grounding for the first time and correctly declining them. As grounding improves, escalation falls for the right reason. But escalation also falls when triggers are tuned too tightly, when confidence thresholds drift, or when someone adds an instruction telling the agent to try harder before handing over.
The two are indistinguishable from the rate alone. What separates them is sampled review: read escalations that did not happen. Take a sample of conversations the agent completed, and ask whether any of them should have gone to a person. That is a tedious hour a week and it is the only reliable way to tell improvement from suppression.
Escalation design: three parts
Escalation is usually built as one thing, a route. It is three, and the two that get skipped are the ones customers notice.
Triggers
Four categories, and you need all four.
Explicit. The customer asks for a person. This should be honoured immediately and without negotiation, every time. An agent that asks a customer who has requested a human to first describe their problem again is generating a complaint.
Behavioural. Repeated rephrasing, rising message length, or the same question asked twice in different words. These are the signals that a conversation is going wrong before the customer says so.
Confidence. Low retrieval confidence, or no authoritative source found. Necessary, and insufficient on its own.
Business rule. This is the category that gets left out, and it is the one that matters most. Any mention of a regulator, an ombudsman, a complaint, financial hardship, bereavement, or a vulnerable-customer signal should escalate regardless of how confident the agent is. A confident wrong answer on a hardship enquiry is precisely the failure a confidence threshold cannot catch, because the agent is not uncertain, it is wrong about what kind of conversation it is in.
Routing
Escalated conversations are not average conversations. The straightforward questions have already been handled, so what reaches a human is skewed toward the difficult, the sensitive and the unusual.
Route accordingly. If agent escalations enter the same queue with the same skill requirements as everything else, you have concentrated your hardest contacts into a channel that was staffed for a normal mix. Several organisations we have worked with ended up defining a distinct skill for agent escalations, staffed by more experienced people, precisely because the composition changed.
Carried context
The human should be able to see, without opening anything, what the customer said, what the agent tried, and what it could not resolve.
This is the part that decides how the escalation feels. A handover with context reads as being passed to someone who can help. A handover without it reads as starting again, and the customer has now spent five minutes with an agent for nothing. That experience is worse than no agent at all, which is a genuine risk rather than a rhetorical one.
What to instrument before launch
Instrument these while you build, because retrofitting them onto a live deployment means a gap in the baseline exactly when you need it.
Conversation-level: intent, actions taken, whether escalated, trigger that fired, and outcome. Customer-level: contacts in the following seven days, across all channels, which is the join most organisations cannot currently make and the one that reveals whether deflection was real.
That cross-channel join is worth the effort. Deflection measured within a single channel will always look better than deflection measured across the estate, because the customer who was not helped simply used a different channel, and the single-channel view is structurally incapable of seeing it.
The first ninety days of measurement
The measurement cadence should be heaviest immediately after launch and settle as the picture stabilises. Front-loading it is what lets you distinguish a real problem from normal early noise.
Weeks one and two. Read every escalated conversation and a sample of completed ones. Not a metric review, actual reading. Expect to change escalation triggers in this period, because the triggers written before launch are hypotheses about behaviour, and behaviour is always slightly different.
Weeks three to six. Move to sampling. Establish the repeat-contact baseline, which needs the cross-channel join and is the number most likely to be unavailable when you first ask for it. Start reporting the four metrics together rather than deflection alone, and set that expectation with stakeholders now, while nobody is defending a figure.
Weeks seven to twelve. Look for drift. Compare accuracy against the launch baseline, watch escalation rate for suppression rather than improvement, and check actions per conversation, which is both a cost signal and an early warning that retrieval is degrading.
Reading a bad set of numbers
The four metrics move together in recognisable patterns, and the pattern usually names the cause faster than investigation does.
| Pattern | Likely cause | Where to look |
|---|---|---|
| Deflection up, repeat contact up | Confidently wrong answers | Grounding: contradictions or stale sources |
| Deflection up, escalation near zero | Escalation suppressed | Triggers, and any instruction encouraging persistence |
| Escalation up, accuracy fine | Scope wider than grounding | Intent list against what is actually being asked |
| Accuracy down, nothing else moved | Something changed upstream | Recent knowledge edits, permission changes, data refresh |
| Actions per conversation rising | Retrieval degrading | Source structure, retitling, duplicates reappearing |
The last row is worth watching hardest, because it moves before customers notice. An agent working harder to reach the same answer is the earliest available signal that grounding is decaying, and it shows up on the consumption bill weeks before it shows up in a complaint.
Rebuilding trust after a bad launch
Some readers will arrive here after the agent has already gone out and gone badly. That is recoverable, and the recovery has a shape worth knowing, because the instinctive response makes it worse.
The instinct is to improve the agent. Tune the prompts, add knowledge, widen the triggers, and wait for the numbers to recover. What that misses is that the damage is rarely evenly distributed. It sits in a small number of intents where the agent was confidently wrong, and those intents are also where the complaints came from.
Narrow first, then rebuild. Take the agent down to the intents where sampled accuracy was genuinely high, which is usually a smaller set than anybody expects, and let it run there while you remediate the rest. A narrow agent that is right is recoverable. A broad agent that is sometimes wrong trains your customers to bypass it, and once they have learned that, later improvements do not get noticed because nobody is using it.
The internal version of the same problem is worse. If service staff decided in week two that escalations arrive without useful context, they will stop reading the context permanently, and fixing the payload three months later does not undo that. Front-line trust is spent once.
Reporting it upward
The gap between how this gets measured and how it gets reported causes more programme damage than any technical failure, because a deployment judged on the wrong number gets steered toward the wrong behaviour.
If the executive metric is deflection, the deployment will optimise for deflection, and the mechanism available for that is escalating less. Nobody will decide to do this. It will emerge from a series of individually reasonable tuning decisions, and the numbers will look excellent throughout.
The fix is to agree the reporting set before launch, while the figures are hypothetical and nobody is defending them. Report the four metrics together, always, with deflection presented as efficiency achieved within accuracy and repeat-contact constraints rather than as the headline. That framing is harder to sell in week one and it is the only version that survives the first bad month.
What automation should not do
Two categories are worth ruling out explicitly, because both look automatable and neither is.
Judgement your business has never written down. Goodwill decisions, eligibility calls at the margin, and complaint outcomes. If three experienced staff would answer differently, the agent is being asked to decide rather than retrieve, and no amount of grounding fixes that. The narrow version, where the agent gathers information and prepares the case for a human decision, captures most of the handling time with none of the exposure.
Conversations where the emotional register is the point. A bereavement notification, a serious complaint, a customer in financial difficulty. The agent may well produce a technically correct answer, and producing it will make things worse. These belong in the business-rule trigger category, escalating on detection regardless of confidence.
The scoring framework for deciding which candidates fall into these categories is in Agentforce use cases that survive a business case.
Setting the bar
Decide in advance what performance would cause you to narrow scope or switch the agent off, and write it down before launch while nobody is invested in the answer.
In practice the useful bar is not a deflection target. It is a pair: accuracy on the sampled review must stay above a stated level, and repeat contact must not rise against the pre-agent baseline. Deflection is then the efficiency you achieve within those constraints, which is the correct relationship between the three and the opposite of how most programmes report them.
The architecture underneath all of this, including how escalation reaches a human through Omni-Channel, is set out in Agentforce and Service Cloud: what actually connects to what.



