Transforming Customer Service with Smarter Case Management and Routing
A support operation where the case taxonomy had grown one value per escalation, and routing sent work to whoever was nearest rather than whoever could resolve it.
Consumer Goods & RetailA case enters from one of four channels and is classified once, at arrival, into a two-level taxonomy that records why the customer made contact. Routing then reads the skill the work requires and the capacity actually measured for that work type, rather than headcount or proximity, so a combination nobody is staffing becomes visible instead of silent. The case is then either resolved with a knowledge article attached to it or escalated on a named reason. Closure records the resolution as a second fact rather than over the top of the contact reason, and the reopen path reattaches a repeat contact to the original case so a deflection cannot be counted twice. The cycle closes at review, where the escape hatch free text and the reason to resolution cross-tabulation feed the next revision of the taxonomy.
Arrival
- Web form
- Email-to-Case
- Retail partner portal
- Inbound call
Four channels, one Case object. Nothing may open work outside the case record, which is what makes volume countable at all.
Classification
- Case Type
- Case Reason
- Escape hatch with required free text
- Record type
Decides what the customer said, in the customer language, once. The resolution field stays empty here by design.
Routing
- Omni-Channel
- Skill requirement
- Capacity weight by work type
- Queue
- Overflow path
Guarantees the case reaches someone permitted and able to resolve it, and that an unstaffed skill combination raises rather than waits.
Work
- Knowledge article
- Article attached to case
- Escalation reason
- Flow for context assembly
Decides resolve or escalate. An escalation has to name why, so the escalation path is itself reportable.
Closure
- Resolution field
- Linked article
- Closure evidence by reason
- Customer confirmation
Guarantees the contact reason survives closure. The resolution is recorded as a second fact, never over the first.
Reopen and review
- Repeat contact window
- Reopen on the original case
- Escape hatch review
- Reason to resolution report
Decides whether a deflection was real, and feeds the free text agents had to type back into the value list.
A worked scenario showing how we approach this problem. The architecture and decisions are our real practice; it is not an account of one named customer.
The problem, as it actually presented
The brief in this situation is almost never "our case taxonomy is broken". It is "cases go to the wrong people".
That is a real complaint and it is not the root of anything. Routing reads classification, so if the classification is noise, faster routing distributes the noise more efficiently and everybody agrees afterwards that the new tool did not help. Trace it back and the same two failures sit underneath.
The first is a case taxonomy that stopped carrying information some years earlier. Case Type and Case Reason were configured once, with the sample values, before the first user logged anything. After that the list grew by request: every escalation that no existing value described produced a new one, added the same afternoon by whoever held the permission, because the alternative was blocking a colleague during an incident. Several years of that produces a list holding Delivery, Delivery Issue, Late Delivery and Order Not Received, all in active use, none owned, and no report able to say what delivery volume is.
The second is that routing distributes by proximity. Whoever is idle, whoever sits in the right queue, whoever the supervisor reassigns. Nothing in that model asks whether the person can resolve the case, and in a business running direct-to-consumer support alongside a retail partner channel that question has a real answer. Partner cases carry commercial terms, credit rules and a different escalation path, and a consumer agent working one produces a plausible answer somebody has to undo later.
What we designed, and why
Rebuilding the taxonomy rather than tidying it
Editing a grown taxonomy value by value does not work, because the values are not individually wrong. They are collectively incoherent, and each one has somebody who asked for it.
So the list is rebuilt from evidence rather than negotiated. Read the opening message on a sample of recent cases, not the subject the agent typed and not the resolution notes, and cluster what the customer actually said. Those clusters become the values. The result looks less tidy on a slide than an org chart does and works far better in a case feed, because an agent recognises a value rather than translating a conversation into an internal structure the customer knows nothing about.
Two levels, not three. A third level is chosen last, when the agent has already spent attention on the first two, and it is where selection accuracy collapses. It is also the level reporting aggregates away, so the detail is captured at cost and discarded at analysis.
We keep exactly one escape hatch and it requires free text. That is what stops the list rotting again: the text is read on a schedule, recurring patterns are promoted into real values, and the escape hatch becomes the mechanism by which the taxonomy learns rather than the hole it leaks through.
Classifying once, not twice
The objection to answer before any of this is agreed is agent workload. Nobody will accept a design that asks a person to classify the same case twice.
It does not, because the contact reason and the resolution are different facts recorded at different moments. The reason is set at creation, in the customer's words, and is hard to change after first save. The resolution is set at closure by the person closing it, and describes what turned out to be true.
Recorded in one field the resolution always wins, because closure is where the required fields live and the agent touches the record last. The declined payment disappears, and what remains is useful to the team fixing address validation and useless to anyone asking what generates demand. Two fields, each written once, at the moment its author knows the answer.
The cross-tabulation of the two is the most useful report in service and almost nobody has it. Reasons resolving to a single consistent outcome are the automation candidates. Reasons scattering across many resolutions are where diagnosis is happening, and those need knowledge investment rather than a bot.
A queue and a skill are different statements
A queue is a statement about what kind of work this is. A skill is a statement about who is able to do it. The two get collapsed constantly, and when queues are named after teams the consequence is that work is routed by reporting line: a reorganisation silently rewrites the routing model, and a cross-skilled agent is reachable through only one door.
The rule we apply is that a skill is added only where a genuine capability gap exists and the gap is stable enough to maintain. Partner channel authorisation qualifies, because a consumer agent is not permitted to agree partner commercial terms. Language qualifies. Seniority and account familiarity usually do not, because every skill added multiplies the combinations somebody has to keep staffed, and an unstaffed combination is a case that waits with nothing visibly wrong.
Capacity that was measured
Routing then reads skill and capacity, and capacity is the number almost always guessed.
The usual pattern is one capacity value per agent, typed into a required field during a workshop, then applied identically to a chat and to a partner credit dispute. Those are not equivalent units of work, so the model is wrong from the first day and everything downstream inherits it.
Instead, measure how long each work type occupies an agent and how many of that type one agent can hold before quality degrades, then derive the weights from that. The absolute ceiling is arbitrary. The ratios between the weights are the entire model.
The operational change follows directly. The routing rule now reads skill and measured capacity, so a queue nobody is staffing cannot quietly receive work. It raises on the overflow path instead, and that path is designed before go-live rather than discovered during the first peak.
Knowledge agents actually use
Knowledge fails in a specific way, and it is not that articles are missing.
It is that two articles describe the same process differently, both are published, neither is dated in a way an agent can act on, and the agent cannot tell which is current. The rational response is to stop consulting knowledge and ask a colleague, which is what agents do. A knowledge base that contradicts itself is worse than none, because none is at least honest about the gap, while a contradictory one costs the search and then misleads.
So the build starts with ownership rather than volume. Every article has a named owner and a review date. Articles are attached to the cases they resolve, so usage is visible rather than assumed. Contradictions are settled by retiring one article, never by publishing a third that reconciles them. And the reason to resolution report says where articles are worth writing, which is a better commissioning input than asking teams what they think is missing.
Deflection, counted honestly
Self-service deflection is the number most often reported and most often wrong, because the naive count treats every session that did not create a case as a success. Abandonment produces exactly that signal. So does a customer who read the article, did not get an answer, and phoned instead. Counted naively, the worst outcomes look like the best ones.
The honest count checks the repeat contact. A session counts as deflected only if the same customer did not arrive through another channel within a window the business agrees in advance. When they do arrive, the contact reattaches to the original case rather than opening a new one, so the second contact is visible as a repeat instead of adding to volume as a fresh success.
The decisions that were contested
Retiring values teams had asked for. Every value on the old list had an owner who requested it, several of them still in the business. The rule applied was that a value stays only if a decision changes depending on whether it is selected: different routing, different article, different report line, different process fix. Values failing that test were made inactive rather than deleted, so historical records keep their meaning. The conversations were not comfortable.
Not deploying skills-based routing everywhere. There was appetite to model every specialism as a skill on the grounds that precision is good. We argued against it, because each additional skill multiplies the combinations that must be covered on every roster, and the failure mode is silent. A case waits for a combination nobody is working that day, and nothing alerts, because nothing is technically wrong.
Keeping the partner channel on its own queue. The tidier design routes everything into one pool and lets skills sort it out. The business reason not to is that partner cases carry different commercial terms, a different escalation path and a different service expectation, and flattening them into consumer volume would have degraded both.
Refusing to automate the escape hatch. There was a proposal to have a rule read the free text and assign a value automatically. We declined for the first phase. A confident automatic classification on ambiguous text produces data that looks clean and is not, and the escape hatch is the one place the taxonomy is supposed to be learning from a human who could not find the right answer.
What changes
The outcomes worth claiming here are operational, and they follow from the decisions rather than from effort.
| Before | After | What made the difference |
|---|---|---|
| Taxonomy grown one value per escalation | Two levels, drawn from what customers said | Rebuilt from case openings rather than edited value by value |
| Reason overwritten by the resolution at closure | Two fields, each written once by whoever knows | The reason locked after first save, the resolution a separate closure fact |
| Work routed to whoever is nearest | Routed on skill and measured capacity | Capacity weights derived per work type instead of one number per agent |
| Unstaffed queues absorb work silently | Unstaffed combinations raise on the overflow path | The overflow path designed before go-live, not after the first peak |
| Articles contradict each other, agents ask a colleague | Every article owned, dated and attached to the cases it resolves | Ownership before volume, and retirement instead of reconciliation |
| Deflection counted as sessions without a case | Deflection counted net of repeat contact | Repeat contacts reattach to the original case |
The second-order effect is the one service leaders notice. When classification is trustworthy, the weekly meeting stops arguing about whether the numbers are real and starts arguing about what to do, which is the argument worth having.
What should stay human on purpose
Some of this could be automated further and should not be.
The escape hatch review stays manual, because it is the sensing mechanism for a taxonomy that will otherwise drift again. Complaint handling and anything carrying a compensation decision stays with a person, because the cost of a confidently wrong automated answer there is a customer relationship rather than a rework task. The first partner contact after an incident stays human, because that conversation is about trust rather than resolution.
Automation is best spent assembling context for the person who decides, not taking the decision away from them.
What we would tell you before starting
If your service reports are filtered before anyone reads them, do not start with routing. Start with the classification those reports depend on, get the value list agreed and owned by one named person, and be prepared for that to take longer than the configuration. Every implementation of this shape that goes badly goes badly there.
And be honest about scale. If one team of a handful of people handles everything and they are genuinely interchangeable, this design is over-engineering. A shared queue, a short taxonomy and a supervisor who can see the backlog is the right answer at that size, and adding skills, capacity weights and overflow paths buys coordination you do not need while costing maintenance you do. The design above earns its cost when the channels genuinely differ, when the work requires capabilities not everyone has, or when volume has passed the point where one person can hold the queue in their head.
