Your support AI dashboard says the automated resolution rate went up again. Your agents say the queue got harder. Both can be true. And the gap between them is what this guide is about.

An automated resolution rate is a fraction. The number on the dashboard depends on what the vendor puts in the numerator, what it leaves out of the denominator and when it counts the conversation as finished. Change any of those and the rate moves without the service changing.

So before you fund the next phase of automation, define the number you’re paying for. Then add the four measures that show whether the work went away or moved: verified resolution, safe escalation, repeat contact and cost per resolved case.

What an automated resolution rate counts

Every AI support platform reports a headline rate. The definitions differ, and the differences are large enough to change a business case.

Zendesk groups outcomes into automated resolution tiers. An assisted escalation means the bot collected details and a person finished the job. A contained resolution means the customer stopped asking. A verified resolution means a language model read the ended conversation and judged that the request was resolved. And only the verified tier draws on the paid allowance.

Intercom counts a Fin resolution when the customer confirms the answer helped, or leaves without asking for more help. It calls the second case an assumed resolution. But both are billed at the same price.

Outcome the platform reports What happened What it proves about the customer’s problem
Bot replied The bot sent an answer Nothing yet
Contained (assumed) The customer did not ask for a person and the conversation timed out The customer stopped talking. They may have solved it, or given up
Verified A model read the transcript and judged the request resolved The transcript reads as resolved. Nothing outside the chat was checked
Confirmed The customer said the answer helped Strong signal, and rare, because few customers reply to say thanks
Assisted escalation The bot gathered details, then a person resolved it Agent time was saved. The resolution belongs to the agent

Note the timing. Zendesk ends an email conversation 72 hours after the last message, and a messaging conversation two hours after it by default. A resolution can’t be verified until the conversation has ended, so the newest days on any dashboard undercount until they settle.

From all contacts to billed resolutionsA ledger of four bars. All contacts, 10,000. The bot took part in 6,000. 3,000 were contained, meaning the customer stopped asking. 1,500 were verified as resolved. A bracket shows that the blended rate counts contained and verified together, 50 percent of the bot's conversations, while billing counts verified only, 25 percent. Conceptual illustration with assumed figures.One month, assumed figuresAll contacts10,000Bot took part6,000Contained3,000Verified1,500Blended rate50%Billed25%Both rates divide by the 6,000 conversations the bot handled
Two rates from the same month

The blended rate counts contained conversations, where the customer stopped asking, together with verified ones. The invoice counts verified only, so the dashboard can show double the billed rate.

Why the dashboard and the invoice disagree

A support team we worked with pays per verified resolution. Their vendor’s admin dashboard showed a headline rate near 50 percent. The billed verified rate for the same weeks was near 25 percent. But nobody had made an error. The two numbers counted different things.

Three causes explained the whole gap:

  • Numerator. The headline blended contained and verified conversations. Billing counted verified only. That alone roughly doubled the displayed rate.
  • Date basis. The vendor’s export grouped conversations by end time. The dashboard grouped them by start time. Conversations that ran overnight moved between days.
  • Time zone. The export used UTC. The account was set to a European time zone. Around 60 conversations changed day on a single date from that alone.

The verified rate was stable under every grouping we tried. So that became the defensible number for the finance team. The blended rate was useful for the service team, as a ceiling on what verification could reach.

Report both numbers, label each with its definition, and don’t let one stand in for the other in a board pack.

Choosing the denominator for an automated resolution rate

The numerator gets the attention. The denominator decides whether the rate describes the whole service or a flattering slice of it.

Denominator What it includes Who tends to use it What it hides
All inbound contacts Every conversation and ticket in the period, whether the bot saw it or not Finance, the COO Nothing. It needs a count the bot platform doesn’t hold
AI-handled conversations Conversations the bot took part in Vendors, by default Contacts that went straight to email or phone
In-scope conversations AI-handled minus request types the bot is not allowed to handle The service team, for tuning Whether scope is growing or shrinking

A worked example shows the effect. Assume 10,000 contacts in a month, of which 6,000 reach the bot and 1,500 end as verified resolutions.

  • Verified rate of AI-handled conversations: 1,500 / 6,000 = 25 percent.
  • Verified rate of all contacts: 1,500 / 10,000 = 15 percent.
  • Involvement rate: 6,000 / 10,000 = 60 percent.

All three are correct. The first is a quality number. The second is the business number, because it tells you what share of the whole workload went away. The third tells you where to look for growth. And a strong first number on a small second number is the usual shape of an early deployment.

Reviewed transcripts add a fourth base. If 800 of 1,000 eligible conversations had transcripts available and reviewers verified 400, the verified share is 40 percent of eligible conversations and 50 percent of reviewed ones. So show the 200 unavailable transcripts as unknown. An export gap isn’t a quiet day.

Verifying a resolution

A verified resolution is a claim about the customer’s problem. The vendor’s own check reads the transcript with a model. That is a reasonable first check, and it’s also the check that decides your invoice. So we add an independent one.

The method we use grades a sample of billed resolutions with a second model, against a rubric the service owner approved. Every verdict stores the transcript, the rubric version, the model and the reason. Humans review the disagreements and every sensitive case.

Escalated, reopened orrated negativeNo hard signalConfident and notsensitiveDoubt or sensitive topicBilled resolutionDeterministic checksNot resolvedSecond model gradestranscriptVerdict with evidenceHuman reviewerDisputed resolution listWeekly service review
Grading a billed resolution with rules, a second model and a reviewer

Three details make this work in practice:

  1. Deterministic checks go first. A conversation that was escalated, reopened or rated negative isn’t resolved. No model is needed to say so, and the result is reproducible.
  2. The grader knows the policy. If phase one says the bot must escalate every account change, an escalation is correct behaviour. Without the policy in front of it, a grader marks designed behaviour as failure.
  3. Humans set the bar. A small gold set of human-graded conversations shows how often the model grader agrees with people. Publish that agreement figure next to the grade.

The output is a disputed-resolution rate: billed resolutions that a second reading didn’t support. Expect single digits once the bot is tuned. And a rising figure is a knowledge or policy problem. It’s also money you should discuss with the vendor.

Testing the handoff from the receiving side

In a first phase the bot escalates anything that touches an account. So most of the value comes from the handoff, and most of the hidden cost hides there too. A bot that says “I have passed this to the team” has made a promise, and the customer treats it as one.

Testing a handoff from the receiving sideLeft, a chat bubble in which the bot tells the customer it has passed the case to the team. An arrow leads to the right, where the receiving agent's view lists four checks: the right queue, an owner assigned, a summary with the problem and steps tried, and what is still missing. A counter at the bottom reads questions repeated: zero. Conceptual illustration.What the bot saidI have passed thisto the team.What the agent gotRight queueOwner assignedProblem and steps triedWhat is still missingQuestions repeated0
The promise and the evidence

The bot's message is what the customer heard. The four checks and the repeat counter are what the test scores, because they decide how long the case takes from here.

We test the handoff from the receiving side:

  1. Confirm the case met an agreed escalation rule, such as a missing identity check or an exception outside policy.
  2. Check the bot made no promise it couldn’t keep, and disclosed nothing it shouldn’t.
  3. Find the routing event and confirm the case reached the right queue with an owner.
  4. Read the summary the agent received. It should hold the problem, the steps tried and what is still missing.
  5. Follow the case to its end and count the questions the customer had to answer twice.

Measure agent handling time on escalated cases separately from ordinary cases. A bot that collects details well cuts that time. But a bot that sends a thin summary raises it. An average across both hides which one you have.

Include the closed-queue case in the test. When the receiving team is offline, the bot should state the approved next step without inventing a response time.

Repeat contact and what it proves

A customer who comes back within a week is a signal, and only a signal. They may have a new question, or the same question on a different channel, or a real unresolved problem.

To turn the signal into a measure, define four things:

  • The window. Seven days is common. A window that hasn’t closed yet gives an immature result, so report it as such.
  • Identity coverage. If only authenticated customers can be linked, say what share of contacts that is.
  • Issue matching. A same-issue repeat needs the subject and outcome to match, with human review on the ambiguous ones.
  • The two measures. Keep broad repeat contact and confirmed same-issue recurrence as separate lines. They answer different questions.

Then read the repeats by theme before touching a prompt. Repeated withdrawal questions at a group running several consumer brands turned out to be a hold-period policy the customers couldn’t find. So that was a content change, and it took a morning.

Automations that pay before the bot resolves anything

Full resolution is the last step to earn. The earlier steps are native to the support tools you already run, and each one has a measure of its own. We built exactly this for a support team whose agents were sorting tickets before they could answer them.

Automation What it does What to measure Typical risk
Triage and routing Reads the ticket and sends it to the right queue with a priority Time to correct owner, avoidable reassignments Wrong queue on an urgent case
Tagging Applies consistent categories for reporting Share of tickets with a usable category, reporting effort Tag drift that breaks trend lines
Data collection Asks for the order number, account email or screenshots before an agent looks Questions the agent no longer asks, handling time on escalated cases Collecting data nobody uses
Draft replies for review Prepares an answer from approved help content, and an agent sends it Edit rate, time per reply, wrong-answer rate on sampled drafts An agent approving drafts without reading them
Summaries Condenses a long thread for the next person Time to first useful action after handoff A summary that drops the missing information

The draft-reply pattern deserves a note. An agent decides what reaches the customer, so the quality bar is lower than for unaided resolution and the savings arrive earlier. The measure is the edit rate on a sample. And if agents rewrite half the drafts, the help content needs work before the bot does.

Observability across the support stack

The measures above come from four systems, and each one keeps a different clock. Verification settles days after the conversation. The ticketing system records escalations and reopens in its own time zone. Agent handling time lives in a workforce tool. Customer identity lives in your own systems, if anywhere.

To report one honest number a week, we found these habits necessary:

  • Pull the vendor export daily and keep every version. The files are immutable per day, and the latest days change as verification settles.
  • Store every graded verdict with its inputs. The transcript, the rubric version, the model and the reviewer. A rerun should give the same result.
  • Use one time zone and one date basis everywhere. Pick the conversation end time, in the account’s time zone, and state it on the dashboard.
  • Show missing data as missing. An export that failed is a gap on the chart. It’s never a zero.
  • Keep the billed number visible. The usage figure in the vendor’s admin area is the invoice. Reconcile your operational number against it each cycle.

Sensitive categories sit in a separate lane. For a regulated consumer business, a conversation about self-exclusion or a vulnerable customer is a compliance event. Count those incidents as an absolute number with a target of zero.

Never average them into a quality score, and make them ineligible for a verified resolution by policy. Our AI governance service sets these boundaries before a bot goes live.

Cost per resolved case

The model bill is the line people ask about, and it’s the smallest one. A cost per resolved case includes everything the service spends to produce a resolution it can stand behind.

The example below uses assumptions, labelled as such. It is not a customer result.

Line Monthly assumption Note
Verified resolutions 1,500 25 percent of 6,000 AI-handled conversations
Platform fee 1,500 EUR A per-resolution price near 1 EUR, in line with Intercom’s published outcome price
Model, hosting and monitoring 800 EUR Grading models, dashboards, exports
Grading and review time 700 EUR 20 hours of reviewer time at 35 EUR
Rework 900 EUR 10 percent of verified cases return within 7 days and cost 6 EUR each to handle
Total 3,900 EUR
Net resolved cases 1,350 Verified minus the cases that came back
Cost per resolved case 2.89 EUR 3,900 / 1,350

Set that against the fully loaded cost of an agent-handled contact for the same request types. We treat it as an assumption of 6 EUR here, and you should replace it with your own figure.

The saving is real at that spread. But it disappears if rework climbs to 30 percent, or if the review time to keep the bot honest doubles.

Cost per resolved case as a ledgerLeft, a ledger in euros: platform fee 1,500, model and monitoring 800, grading and review 700, rework 900, total 3,900, divided by 1,350 net resolved cases gives 2.89 per case. Right, two bars compare 2.89 for an AI-resolved case with an assumed 6.00 for an agent-handled contact. Conceptual illustration with assumed figures.Monthly cost, EURPlatform fee1,500Model and monitoring800Grading and review700Rework900Total3,900Divided by 1,350 net resolved casesPer case2.89AI resolved6.00Agent handledAssumptions, never results
The lines that go missing

Grading, review and rework are the cost of a resolution you can defend. Leave them out and the model bill alone makes any bot look cheap.

Two rules keep the comparison honest. Use the same request population and the same period for the before and after. And count capacity as a saving only when a real cost changes: fewer contractor hours, less overtime, a hiring plan that moved. Released capacity that absorbs growth is valuable, and it’s a different line in the business case.

The weekly review that assigns work

The review we run turns each measure into an owner and a task, so a bad week produces a change rather than a chart.

Measure Definition Owner of a bad week
Verified resolution Verified resolutions / AI-handled conversations in a settled period Service owner
Disputed resolution Billed resolutions a second grade rejected / billed resolutions Knowledge owner, then the vendor
Safe escalation Handoffs that reached the right queue with a usable summary / sampled handoffs Integration owner
Same-issue repeat Confirmed same-issue returns within 7 days / identity-linked resolutions Content owner or product owner
Compliance incidents Sensitive conversations mishandled, as a count Compliance owner, target zero
Cost per resolved case Total service cost / net resolved cases Finance, with the service owner

Expand one request type at a time. Agree a baseline, a settled review period and the conditions that pause automation, such as a compliance incident or a broken handoff. Then decide per request type: expand, repair or hold.

A mixed result supports a narrow rollout. It doesn’t support a service-wide one, and a rising headline rate doesn’t change that.

For the support teams we’ve worked with, the sequence has been the same. Triage and drafting first, because they pay early and carry little risk. Then guided resolution with escalation on every action. Then, once verified resolution and repeat contact hold steady for a quarter, automated actions on the narrow request types that earned it.

Our process automation service runs that sequence, with the measures in this guide in place from the first week.

Frequently asked questions

What is a good automated resolution rate?

It depends on what the rate counts. A verified rate of 25 to 40 percent of AI-handled conversations is common for a first phase that escalates every account action. A blended rate that includes unverified conversations can show twice that on the same traffic. Compare like with like, and track the trend for one request type at a time.

Is containment rate the same as resolution rate?

No. Containment means the customer did not ask for a person and the conversation ended. Resolution means the problem was solved. A customer who gave up is contained. A verified resolution needs a check on the transcript, or a confirmed change in another system.

Can AI grade its own support conversations?

A separate model can grade transcripts against a written rubric, and it scales to every conversation. It still needs calibration against human reviewers, an option to say the evidence is insufficient, and a human on sensitive cases. We treat the model grade as a first pass, never as the only check.

Does a returning customer prove the first answer failed?

Not on its own. The customer may have a new question or may use another channel. A same-issue repeat needs identity matching, an issue match and a complete follow-up window. Broad repeat contact is still useful as a signal for investigation.

How do we work out the cost per resolved case?

Add the platform fees, model and monitoring costs, grading and review time and the cost of rework for the period. Divide by the verified resolutions in the same scope. Compare that with the fully loaded cost of an agent-handled contact for the same request types.

How we automate processes Talk to us about your project