The metric that matters
Not deflection — containment with satisfaction. A customer who gave up and left counts as deflected in most dashboards. The honest measure is conversations genuinely resolved, where the customer was satisfied with the resolution.
AI support agents are the most commonly proposed and most commonly disappointing AI deployment. The gap is almost always in how the business case was modelled.
Model the return before building
Work through these with your own numbers. If the arithmetic does not hold here, it will not hold after you have spent the budget.
| Input | Where to get it |
|---|---|
| Monthly ticket volume | Helpdesk reporting |
| Share that are repeat questions | Sample 200 tickets and categorise |
| Average handling time | Helpdesk reporting |
| Fully-loaded cost per support hour | Finance |
| Realistic containment rate | Start at 30% of the repeat-question share |
| Build cost | Vendor quote, itemised |
| Monthly running cost | Model API plus infrastructure |
The step most businesses skip is the second row. Sample two hundred real tickets and categorise them honestly. Most teams believe their volume is more repetitive than it is — and that share is the ceiling on what any agent can contain. Everything else in the model depends on this number.
Where the value actually comes from
Headcount reduction is the least reliable source of return. These are more dependable:
- Absorbing growth. Handling 40% more volume without adding staff is easier to achieve and easier to defend than cutting existing staff.
- Out-of-hours coverage. Enquiries arriving at 11pm answered immediately rather than going cold overnight.
- Faster first response on everything, including tickets that ultimately escalate.
- Better escalations. An agent that gathers context before handing over saves the human several minutes per ticket.
- Staff retention. Repetitive query handling is the part of support work people leave over.
The strongest business cases we see are framed as capacity, not cost reduction. "Handle next year's growth without hiring three more people" is both more achievable and easier to get approved than "replace two people".
What determines whether it works
Your documentation quality
This is the single largest determinant and it has nothing to do with the AI. An agent grounded in vague, outdated or contradictory documentation produces vague, outdated, contradictory answers with more confidence than a human would.
Do this before commissioning anything: take your twenty most common questions and try to answer each purely from your existing documentation. Whatever you cannot answer, the agent will not be able to either. Fixing that documentation frequently reduces ticket volume on its own — sometimes enough that the AI project becomes unnecessary.
Escalation design
The moment that decides customer perception is the handoff. Done well — context preserved, no repetition required — customers barely notice. Done badly, the customer explains everything twice and concludes the AI wasted their time.
Honesty about limits
An agent that confidently answers questions outside its knowledge damages trust faster than one that says "I'll get someone who can help with that". Refusal behaviour is a feature.
Where AI support genuinely should not go
- Complaints and emotionally charged issues. An upset customer routed to a bot escalates.
- Anything involving money movement without human confirmation.
- Account security and access recovery.
- High-value B2B relationships, where the relationship is the product.
- Regulated advice — financial, medical, legal.
Measuring honestly after launch
| Metric | What it tells you | Trap |
|---|---|---|
| Containment rate | Resolved without a human | Counts abandonment as success |
| CSAT on contained conversations | Whether resolution was good | Only surveying escalations hides the problem |
| Escalation rate | How often it hands off | Low is not automatically good |
| Repeat contact within 48h | Whether it truly resolved | Most honest single metric |
| Cost per contained conversation | Actual unit economics | Ignores the escalations it also cost you |
The repeat-contact metric is worth prioritising. A customer who comes back within two days was not helped, regardless of what the containment dashboard says.
A deployment that protects your reputation
- Fix the documentation first. Measure ticket volume after; sometimes that alone changes the business case.
- Deploy internally — let your own support team use it as a lookup tool before customers see it.
- Launch on a narrow topic set where documentation is strongest.
- Make human escalation obvious and easy from the first message.
- Widen scope on evidence, using the logs to find what people actually ask.
Modelling a support agent business case? Send us a sample of your ticket categories — we will tell you honestly what containment is realistic. See our AI agent service and internal knowledge base guide.