AI Resolution Rate Benchmarks: What Counts as Resolved? | Nousu
All articles
8 min

AI resolution rates: honest benchmarks for ecommerce support

Vendors advertise 70 to 85% resolution rates. Here is what resolved should actually mean, the denominator games to watch for, and a realistic 90-day ramp.

Resolution rateBenchmarksAI support
Ezra Klijnsma, founder of Nousu

Ezra Klijnsma

Founder of Nousu, building AI support agents for online stores

An AI resolution rate is the share of customer conversations an AI closes without a human touching them. Vendors commonly advertise 70 to 85%. Buyers usually measure less, because the advertised number depends entirely on what the vendor counts as resolved and what they quietly leave out of the denominator.

This post is the definition, the games, and a checklist. Nousu sells an AI support agent, so we have skin in this game too. That is exactly why we would rather define the metric honestly than win a bidding war of fictional percentages.

Why the advertised number and your number diverge

A vendor demo shows 80%. Ninety days into your contract, your own count says 45%. Neither number is necessarily a lie. They are answers to different questions.

"Resolution rate" has no standard definition. Every vendor picks their own numerator and their own denominator, and both choices move the result by tens of points. Until you know both choices, the percentage is decoration. A vendor who will not define the metric on the spot is telling you something about the metric.

What counts as a resolved conversation?

A conversation is resolved when all three of these hold:

  • No human touch. Nobody on your team read it, edited a draft, or approved a reply. If a human approved it, that is assistance. Useful, but a different metric.
  • No reopen. The customer did not come back about the same issue within a reasonable window, on any channel. A reply that produces a second, angrier email tomorrow resolved nothing.
  • The customer got the outcome. They asked where the order was and got a delivery window. They wanted a return and the label is in their inbox. The thing they came for happened.

Hold that against the two impostors:

  • Deflected. The AI pointed the customer at a help article or a tracking link and the conversation ended. Maybe the link answered it. Often the customer had the link already, which is why they wrote in. Deflections have a habit of coming back as new tickets, sometimes on a channel you monitor less.
  • Abandoned. The customer gave up and closed the tab. Some dashboards count this as a success because no human touched it and nobody reopened it. By that logic, an unplugged phone line resolves 100% of calls.

The difference between resolving and deflecting is the same difference as between an agent and a chatbot, which we cover in chatbot vs AI agent. A chatbot, as a category, deflects well. Resolving requires reading the actual order and doing the actual work.

How vendors inflate a resolution rate: the denominator games

The numerator games above are half the story. The denominator is where the real creativity lives. Watch for these:

  • Excluding handoffs. "Of the conversations the AI handled, it resolved 85%." And the conversations it routed to your team in the first ten seconds? Not handled, so not counted. The AI gets to pick its own exam questions.
  • Counting deflections as resolutions. Every link click, every "was this helpful?" left unanswered, every conversation that simply stopped. If the definition of resolved is "the conversation ended", the rate will be spectacular.
  • Counting the closed tab. The purest form of the previous game. Silence scored as satisfaction.
  • Measuring the honeymoon topics. The rate is computed over FAQ-shaped questions only, or over a pilot period where the AI was scoped to the easiest three topics. Real queues are messier than pilots.

None of this requires bad faith. It requires only that marketing picks the most flattering defensible definition, which is marketing's job. Your job as a buyer is to ask for the definition before you compare any two numbers.

The realistic ramp: week one versus day ninety

The other honest thing nobody puts on a pricing page: resolution rates ramp. Week one is low, and it should be. The agent does not yet know your edge cases, your team has autonomy dialed down while it earns trust, and half the interesting topics are still routed to humans on purpose.

Then it climbs. You review the handoffs, tune the topics, flip the routine ones to autopilot one by one. Merchants typically automate 50 to 70% of routine tickets within the first 90 days. Note both qualifiers in that sentence. Routine tickets, not all tickets. Within 90 days, not at launch. Any vendor quoting one number for day one and day ninety is describing at most one of those days accurately.

A vendor who tells you a ramp story is being straight with you. A vendor who promises the plateau on launch day is quoting the end of someone else's ramp as your starting point.

Queue mix sets the ceiling

Here is the part that makes every cross-store benchmark suspect: your maximum achievable rate is a property of your queue, not of the software.

Order status, returns, and cancellations typically make up 30 to 50% of an ecommerce queue (we broke down the biggest slice in what is WISMO). Those tickets are answerable from order and carrier data with no judgment call, which makes them almost fully automatable. A store where they dominate has a high ceiling.

Now take a custom-goods store. The queue is design approvals, artwork problems, and complaints about a specific misprinted product. Those need eyes and judgment. The ceiling is structurally lower, and no vendor can change that with a better model. Same software, different store, twenty points of difference. Which is why a benchmark that does not mention queue mix is not a benchmark. It is an average of stores that are not yours.

What to ask a vendor about their resolution rate

Ask these in the demo. Write the answers down, they will matter at renewal.

  • What exactly counts as resolved? No human touch, no reopen, and the customer got the outcome, or something softer?
  • Are handoffs to my team in the denominator, or excluded from it?
  • Is a deflection to a help article counted as a resolution?
  • What happens to a conversation where the customer just stops replying? Resolved, abandoned, or excluded?
  • What is the reopen window, and does it span channels?
  • What rate should I expect in week one, and what at day ninety, for a queue with my ticket mix?
  • Can I see the per-topic breakdown in the product, not a PDF?

The point of the checklist is not to catch a vendor lying. It is to force every vendor onto the same definition so the numbers become comparable. Once the definitions match, most of the gap between competing claims quietly disappears.

What resolution rate Nousu claims

Nousu publishes ramp expectations, not a launch-day promise. We tell merchants the same thing this post says: routine tickets first, 50 to 70% of them automated within the first 90 days, ceiling set by your queue mix, autonomy expanded topic by topic as the agent earns it. How that works in practice is on the AI agent page.

And we will publish our own production resolution numbers, with the definition from this post attached, when we have cohorts worth publishing. Not before. A benchmark post that ends with an invented statistic would be a strange way to argue for honest benchmarks.