The metric that matters

Not deflection — containment with satisfaction. A customer who gave up and left counts as deflected in most dashboards. The honest measure is conversations genuinely resolved, where the customer was satisfied with the resolution.

AI support agents are the most commonly proposed and most commonly disappointing AI deployment. The gap is almost always in how the business case was modelled.

Model the return before building

Work through these with your own numbers. If the arithmetic does not hold here, it will not hold after you have spent the budget.

InputWhere to get it
Monthly ticket volumeHelpdesk reporting
Share that are repeat questionsSample 200 tickets and categorise
Average handling timeHelpdesk reporting
Fully-loaded cost per support hourFinance
Realistic containment rateStart at 30% of the repeat-question share
Build costVendor quote, itemised
Monthly running costModel API plus infrastructure

The step most businesses skip is the second row. Sample two hundred real tickets and categorise them honestly. Most teams believe their volume is more repetitive than it is — and that share is the ceiling on what any agent can contain. Everything else in the model depends on this number.

Where the value actually comes from

Headcount reduction is the least reliable source of return. These are more dependable:

  • Absorbing growth. Handling 40% more volume without adding staff is easier to achieve and easier to defend than cutting existing staff.
  • Out-of-hours coverage. Enquiries arriving at 11pm answered immediately rather than going cold overnight.
  • Faster first response on everything, including tickets that ultimately escalate.
  • Better escalations. An agent that gathers context before handing over saves the human several minutes per ticket.
  • Staff retention. Repetitive query handling is the part of support work people leave over.

The strongest business cases we see are framed as capacity, not cost reduction. "Handle next year's growth without hiring three more people" is both more achievable and easier to get approved than "replace two people".

What determines whether it works

Your documentation quality

This is the single largest determinant and it has nothing to do with the AI. An agent grounded in vague, outdated or contradictory documentation produces vague, outdated, contradictory answers with more confidence than a human would.

Do this before commissioning anything: take your twenty most common questions and try to answer each purely from your existing documentation. Whatever you cannot answer, the agent will not be able to either. Fixing that documentation frequently reduces ticket volume on its own — sometimes enough that the AI project becomes unnecessary.

Escalation design

The moment that decides customer perception is the handoff. Done well — context preserved, no repetition required — customers barely notice. Done badly, the customer explains everything twice and concludes the AI wasted their time.

Honesty about limits

An agent that confidently answers questions outside its knowledge damages trust faster than one that says "I'll get someone who can help with that". Refusal behaviour is a feature.

Where AI support genuinely should not go

  • Complaints and emotionally charged issues. An upset customer routed to a bot escalates.
  • Anything involving money movement without human confirmation.
  • Account security and access recovery.
  • High-value B2B relationships, where the relationship is the product.
  • Regulated advice — financial, medical, legal.

Measuring honestly after launch

MetricWhat it tells youTrap
Containment rateResolved without a humanCounts abandonment as success
CSAT on contained conversationsWhether resolution was goodOnly surveying escalations hides the problem
Escalation rateHow often it hands offLow is not automatically good
Repeat contact within 48hWhether it truly resolvedMost honest single metric
Cost per contained conversationActual unit economicsIgnores the escalations it also cost you

The repeat-contact metric is worth prioritising. A customer who comes back within two days was not helped, regardless of what the containment dashboard says.

A deployment that protects your reputation

  1. Fix the documentation first. Measure ticket volume after; sometimes that alone changes the business case.
  2. Deploy internally — let your own support team use it as a lookup tool before customers see it.
  3. Launch on a narrow topic set where documentation is strongest.
  4. Make human escalation obvious and easy from the first message.
  5. Widen scope on evidence, using the logs to find what people actually ask.

Modelling a support agent business case? Send us a sample of your ticket categories — we will tell you honestly what containment is realistic. See our AI agent service and internal knowledge base guide.

Frequently asked questions

For a well-built agent grounded in good documentation, 25–45% of incoming volume is a realistic range for most businesses. Claims above 70% usually count any conversation the agent touched, including ones the customer abandoned in frustration.
Usually not immediately, and framing it that way tends to backfire. What it reliably does is absorb volume growth without adding headcount, and free existing staff from repetitive queries. Businesses that deploy expecting immediate cuts often end up with worse service and no saving.
Containment rate (resolved without human involvement), customer satisfaction on contained conversations specifically, and escalation quality. Deflection alone is a vanity metric — a customer who gave up counts as deflected.