Guide

Auto QA: how AI call scoring actually works,
and when not to trust it

What happens between a recording and a score, which criteria AI can judge reliably, where it gets things wrong, and what our own human-versus-AI data shows.

We are a vendor, so treat the enthusiasm accordingly. Every number below comes from our own production corpus with sample sizes and time windows attached, including the ones that make our product look worse than the brochure.

What auto-QA is

Auto-QA is the automatic evaluation of every customer conversation against your own quality scorecard, producing a per-call score in which each criterion is marked pass or fail and tied back to the exact moment in the transcript that justifies it.

It replaces the sampling step in a traditional QA program, not the program itself. The rubric is still yours, the standards are still yours, and the coaching is still done by people. What changes is that the scorecard gets applied to 100% of calls instead of the 2 to 4% a human team can physically review.

The definition is deliberately narrow, because the loose version ("AI that listens to your calls") describes at least three different products; we drew the boundaries in how auto-QA, conversation intelligence and speech analytics differ. This page is about the evaluation layer: how it works, what it gets wrong, and what a QA program looks like on the other side of the switch.

Three things auto-QA is not

It is not recording

Voxjar does not record calls and most auto-QA tools do not either, because your phone system, dialer, or CRM already produces the recordings. Recording, transcription, analysis, and evaluation are four separate jobs, and vendors occupy different combinations of them.

It is not real-time agent assist

Live prompting during a call is a different product with a different failure mode. Auto-QA is evaluation after the conversation, against a standard you wrote in advance.

It is not a replacement for QA analysts

It does the listening, establishing what happened on every call instead of the 2-3% a human team can reach, and hands analysts a complete evidence base for the judgment work: rubrics, calibration, disputes, and coaching.

Recording, transcription, analysis, and evaluation are four separate jobs. Knowing which one a vendor actually does is most of the buying decision.

How AI scoring actually works, end to end

Most explanations stop at "AI analyzes the call." Here is the actual pipeline, seven stages, and where each one goes wrong.

  1. 1

    Ingestion

    Recordings arrive through an integration, upload, SFTP, or API, with metadata: agent, queue, direction, disposition, campaign. That metadata decides which scorecard applies to which call, and a sales scorecard applied to a support call generates garbage that looks exactly like real data.

  2. 2

    Transcription

    Audio becomes speaker-separated, timestamped text, and everything downstream inherits its quality. Diarization errors, where the agent's words are attributed to the customer, are the most damaging failure here, because "did the agent state the disclosure" is answered by looking at who said what.

  3. 3

    Criterion evaluation

    The core step, and the design decision that matters is that each criterion is evaluated separately. The system does not read the call, form an impression, then decompose it into scores. It asks one question of the transcript, answers it, and moves on. That is what makes results traceable and stops a strong opening from inflating the compliance section.

  4. 4

    Evidence extraction

    For each answered criterion, the system pulls the quote and timestamp behind the answer. A criterion marked "no" should point at the absence in a specific place ("no verification language appears before account details are discussed at 0:42").

  5. 5

    Score assembly

    Results roll up through your weights, with auto-fail gates outside the weighted math, zeroing the call regardless of everything else, exactly as on a paper evaluation form.

  6. 6

    Routing and flagging

    Low scores, auto-fails, and unusual patterns go to a review queue. Most implementations quietly fail here: a system that flags 600 calls a month at a team with capacity for 40 is functionally a system that flags nothing.

  7. 7

    Human review

    People calibrate, adjudicate disputes, override, and feed overrides back into rubric changes.

Stage 5 rolls up exactly as on a paper evaluation form. The mechanics of call scoring are unchanged from the manual era: automation changes coverage and cost, not the logic, which is why call center quality assurance software is only ever as good as the rubric you give it. If your scorecard was bad on paper, it is now bad on 100% of calls.

What an AI-scorable criterion looks like, and what one does not

This is the single biggest predictor of whether an auto-QA deployment works, and almost nobody writes about it. Teams import their existing scorecard, get inconsistent results on a third of the criteria, and conclude AI scoring is unreliable. Usually the criterion was unscorable by a human too, and manual review was hiding that behind reviewer confidence.

A criterion is AI-scorable when a competent stranger, given only the transcript, would answer it the way you would: it refers to something observable in the conversation, has a single decision point, and states the pass condition rather than implying it.

Not scorable Why it fails Scorable rewrite
Was the agent professional? No definition. Every reviewer supplies their own. Did the agent avoid interrupting the customer, and refrain from slang, sarcasm, or criticism of the company or colleagues?
Did the agent show empathy? Names an internal state, not a behavior. Did the agent verbally acknowledge the customer's stated problem or frustration before moving to a solution?
Did the agent handle the objection well? "Well" is undefined, and the criterion bundles detection with quality. Split in two: (a) Did the customer raise an objection? (b) If yes, did the agent respond to the specific objection raised before restating the offer?
Did the agent follow the process? Which process. All of it. One criterion per step, each naming the step: "Did the agent confirm the account number before discussing the balance?"
Was the call resolved? Depends on facts outside the conversation. Did the agent state a resolution, or set a specific next action with an owner and a timeframe, before ending the call?
Did the agent build rapport? Unfalsifiable. Any friendly exchange qualifies, or none does. Did the agent use the customer's name at least once after the greeting?

Every rewrite converts a judgment into an observation, and several are narrower than the original intent. "Used the customer's name" is not rapport. It is a measurable proxy for rapport, and it is honest about being a proxy. A rubric of honest proxies beats a rubric of noble abstractions that score randomly.

One question per criterion

"Did the agent verify identity and explain the recording notice" is unanswerable when one happened and the other did not.

Say what "not applicable" looks like

A large share of scoring noise is criteria applied to calls where they make no sense. "Did the agent overcome the objection" on a call with no objection has to return N/A rather than a fail.

Our scorecard templates are written in this style, and the AI QA scorecard builder applies these constraints as you write criteria rather than after you discover them.

What changes economically

The usual pitch is that automation is cheaper than analysts. True, and the least interesting part.

Manual QA reviews 2 to 4% of calls. A sample that size is a legitimate instrument for estimating a team average and a poor one for nearly everything else people ask of it. We worked the arithmetic in why sampling 2% of calls answers the wrong question: at eight calls per agent per month the 95% confidence interval on an agent's score is roughly plus or minus ten points, and a random sample detects your coverage rate's share of your compliance incidents and no more. The sample size calculator sizes your own. The bigger argument is what full coverage finds that sampling structurally cannot, and the clearest example in our corpus is this.

From our production corpus

An energy retailer, 63,752 scored sales calls over 556 days

Two criteria were scored on every call: whether the agent mentioned the product, and whether a sale was made. Cross-tabulated, the result is not a weak correlation, it is a floor.

5.09%
Conversion when the agent mentioned the product (35,285 calls)
0.00%
Conversion when it never came up. Zero sales across 28,467 calls
44.7%
Share of calls where nobody attempted the offer
Call segment Calls Sales Conversion
Agent mentioned the product 35,285 1,795 5.09%
Agent never mentioned it 28,467 0 0.00%
Total 63,752 1,795 2.82%

You cannot sell something you never mention. So this team's reported 2.82% is not a measure of closing ability. It is closing ability, 5.09%, diluted by the 44.7% of calls where nobody attempted.

And the fix is arithmetic rather than talent. Hold closing skill exactly constant and lift the mention rate from 55% to 80%: 0.80 x 63,752 = 51,002 calls with an offer, at 5.09%, is roughly 2,596 sales against 1,795 actual. A 45% increase with nobody getting better at selling.

The same absence, in three other operations

A weight-loss clinic BPO

4,021 calls over 307 days, booking at 55.1% overall. Agents who offered the first available appointment booked 86.4% (n=2,437) against 30.0% who did not (n=160). Agents who attempted to overcome the objection booked 65.5% (n=1,924) against 9.5% (n=231).

A B2B software support line

Across 6,933 conversations over 309 days with free or unregistered accounts, 90.7% never mentioned the paid upgrade path. On a separate criterion, across 19,943 callers over 338 days, 77.3% were never told to expect a feedback survey, so that company's CSAT rests on the primed 23%.

A home services operation

62.6% of 52,646 calls over 458 days ended with no appointment set. The ask is somebody's secondary job, and it is the first thing to disappear under volume.

Read the BPO figures as correlation, not causation. The negative cells are small and the causality plausibly runs both ways, since an agent on a call that is clearly going nowhere may not bother offering a slot. It rhymes with the energy finding; it does not prove it.

None of this is findable by sampling. At 3% coverage the missing offers look like a handful of weak calls rather than half the operation. Full coverage is what turns "some agents forget to ask" into a number with a revenue figure attached, and that has almost nothing to do with the cost of analysts.

Price the coverage argument for your own operation

The cost side is the smaller argument, but it is the one a budget owner asks about first. Put in your call volume, your reviewer pay rate, and how long a review actually takes, and compare what hand-scoring costs you now against scoring every call.

Your Metrics

Enter your current call volume and manual QA costs to see your potential savings.

$

Fully loaded cost including benefits if applicable.

3x

It typically takes 3x the duration of a call to listen, review, and fill out a scorecard.

Projected Annual Savings

$21,852

+1,839% ROI

Manual QA (100%) $1,920/mo
Voxjar AI (100%) $99/mo
Monthly Savings $1,821

Ready to realize these savings?

Get started with a free AI call evaluation.

Start a Free Evaluation

Where AI scoring is unreliable

Now the part the category does not publish. We have a corpus of 2,424 calls scored by both a human reviewer and the AI on the same scorecard, across 122 different scorecards. Same call, same rubric, two scorers.

Human vs AI, matched pairs Value
Matched pairs 2,424
Distinct scorecards 122
Human average score 78.2
AI average score 68.5
Calls where the human scored higher 36.0%

Read the last two rows together, because they point in different directions. The averages say the AI is nearly ten points harsher. The win rate says humans scored higher on only about a third of calls. Both cannot describe a uniform difference in strictness. What they describe is two scorers landing in a similar place on most calls and diverging hard on a minority, with the AI far below the human when they split.

Our working hypothesis, unconfirmed and publishable either way once tested: divergence concentrates on auto-fail and compliance criteria, where a human who knows the agent, the account, and that the call was otherwise fine exercises discretion, and the AI does not. If that holds, the AI is not harsher in general. It is unforgiving in the specific places where humans routinely forgive, and reasonable people can disagree about which of those is the error.

Practically: a raw AI average is not comparable to your historical manual average. A nine-point drop on switching is two different scorers, not a performance collapse. Re-baseline, and expect your first calibration sessions to be arguments about auto-fails.

The specific failure modes

Subjective criteria

The weakest area by a wide margin, and we have unpleasant evidence of it in our own data. On one customer's scorecard, a block of soft-skill criteria (empathy, gratitude, courtesy, active listening, confidence, addressing the customer by name) sits at a 0.0% fail rate across more than 30,000 evaluations each. A 100% pass rate on seven soft-skill criteria across thirty thousand human conversations is not a plausible description of human behavior. Either the criteria are loose enough to auto-pass or the model is systematically lenient on subjective judgments, and we suspect both. We are investigating it. If a criterion never fails, it is not a criterion, it is decoration.

Sarcasm, tone, and delivery

Transcripts lose prosody, and "that's great, thanks so much" scores the same whether it was warm or withering. Acoustic signals such as talk-over and dead air partially recover this, and sentiment analysis helps in aggregate, but a criterion resting on how something was said is on shaky ground.

Domain jargon, proper nouns, and multi-speaker calls

Product SKUs, drug names, legal terminology, and regional accents degrade transcription accuracy, and everything downstream inherits the error. Transfers, three-way calls, and supervisor takeovers break speaker attribution, at which point "did the agent say X" is unanswerable and the system will still return an answer. Test both on your own audio before you buy.

Ambiguous criteria

A rubric problem rather than an AI problem that lands the same way, and the most common cause of "the AI got it wrong" in practice.

Absence of a rubric entry

Auto-QA answers only what you asked, so a systemic problem no criterion covers is invisible at 100% coverage just as it was at 3%. This is why exploration over the same transcripts matters: exploration finds the problem, evaluation measures it.

Sentiment analysis recovers some of what transcripts lose in aggregate, and conversation intelligence over the same transcripts is what covers the problems no criterion anticipated.

Methodology note

Every figure above comes from our customers' own scorecards, applied to their own calls, in production. Criteria and strictness vary enormously between companies, so cross-company score comparison compares rubrics at least as much as performance: a team averaging 91 on a lenient scorecard is not outperforming a team averaging 72 on a strict one. Sample sizes and windows are stated with each figure. Each finding is from a single customer account, anonymized by descriptor, and presented as an illustration of a mechanism rather than an industry benchmark.

Evidence is the whole product

A score without evidence is an opinion with a decimal point.

If the system says the disclosure was missed, it should show you the 40 seconds where it should have been. If it says discovery was weak, it should quote the three questions that were asked. Without that, three things become impossible: coaching (the manager cannot point at anything), disputing (the agent cannot argue with a number), and auditing (nobody can reconstruct the score six months later).

Evidence also changes what a wrong score costs: with the quote attached, a mis-scored criterion is visible in five seconds and gets corrected, and without it the error is invisible and accumulates.

The demo test

Pick a criterion the system marked as failed and ask it to show you why, then pick one it passed. If either answer is a paraphrase of the criterion rather than a quote from the call, keep looking.

You can run that test on us in about five minutes without talking to anyone. Score one of your own calls free, open the criterion you care most about, and see whether the evidence is a quote from the call or a restatement of the question. That single check tells you more than a feature comparison will.

Score One of Your Own Calls Free

Agent trust under 100% coverage

Going from 8 reviewed calls a month to 500 is not a change in degree from the agent's chair. It reads as surveillance, and how you introduce it matters more than the accuracy of the model.

Full coverage does remove the most corrosive grievance in traditional QA, which is call selection: "you picked my worst calls" is sometimes true, and with a sample of four there is no principled way to refute it. What replaces it is a different objection, "the AI does not understand my calls," which is often legitimate and is why the dispute path has to exist and be used. What we have seen work:

Show agents their own scores, in full, before managers use them for anything

A silent scoring period of two to four weeks, where agents see every score and every piece of evidence and nothing is escalated, converts the rollout from something done to them into something they can inspect. It also finds your broken criteria faster than any internal review, because the people being scored are motivated to find them.

Publish the failure modes

Tell the floor, in writing, that subjective criteria are weakest, that sarcasm does not transcribe, and that transfers confuse speaker attribution. Agents already know the system is imperfect; naming the imperfections first separates a credible program from a defensive one.

Make disputes cheap and visible

One click from the score, a named responder, a stated turnaround, and an overturn rate you report to the floor. An overturn rate of zero is not a sign of accuracy, it is a sign nobody trusts the process enough to use it.

Change what the score is used for

The biggest trust lever is not accuracy. Full coverage makes it possible to coach a behavior instead of ranking people. "You asked for the appointment on 61% of your calls, the floor is at 78%, here are four where you did not" is a conversation about a behavior; a leaderboard is a conversation about a person. The first survives an imperfect model, the second does not.

Give agents the tool too

Coverage that only flows upward reads as monitoring; coverage that flows both ways reads as feedback, and agents who can review their own scored calls self-correct before anyone coaches them.

That last point is the practical difference between AI call monitoring used as measurement and used as surveillance.

The governance questions buyers actually ask

Five questions come up in nearly every serious evaluation. None of the following is legal advice, and employment and privacy obligations vary by jurisdiction, so run your program past your own counsel.

Can I use this in performance management?

The defensible pattern is that AI scores drive coaching, trends, and prioritization automatically, while any adverse action stays behind a human review of the specific calls. The AI finds the calls and cites the evidence; a person decides and owns it. That survives scrutiny in a way "the system said 62" does not.

How do agents dispute a score?

A documented path, a defined responder, a turnaround commitment, a record of outcomes. Disputes are also your best rubric-quality signal: a criterion that generates repeated disputes is usually badly written rather than badly scored.

What about bias?

Transcription accuracy varies with accent and audio quality, which can disadvantage some agents on criteria that turn on catching specific words, and criteria rewarding a particular communication style can encode a cultural preference as a quality standard. Segment score distributions by team, site, shift, and language on a schedule, and look for gaps you cannot explain operationally.

What is auditable?

For every score: the rubric version in force at the time, criterion results, evidence quotes, any human override with author and reason, and timestamps. Ask specifically about rubric versioning, the piece most often missing, without which a score from March cannot be interpreted after an April edit.

What happens to the data?

Retention, whether your recordings or transcripts train any shared model, where processing happens, how deletion works. In the contract, not the sales call.

What humans still do

The job changes, it does not disappear. Rubric design and maintenance becomes the highest-leverage activity in the program, because a criterion change now propagates to 100% of calls instead of 3%. Calibration gains a new participant: score a call as a group, compare with the AI, and write down the ruling when you disagree, because that ruling is either a rubric fix or a known limitation. Dispute adjudication is irreducibly human and is where the program's legitimacy is earned or lost. Edge cases (escalations, complaints, saves, litigation-adjacent calls) are where analysts should be spending time, which they cannot do while they are also the scoring mechanism.

And then the part no model does. The energy retailer's 28,467 unmentioned-product calls are a fact. Turning that into a script change, a coaching plan, and a weekly mention-rate target is management work.

What actually changes in your week

Vendor pages describe outcomes. Here is the operational texture, which is what a QA manager is actually trying to picture.

Moment in the QA week Before After
Monday Pull the sample, assign calls to reviewers Open the exception queue: auto-fails, threshold breaches, outlier agents
Tue to Thu Listen, score, write up. Most of the week Read flagged calls only. The rest goes to coaching and rubric fixes
Coaching prep Find a call that illustrates the issue Filter every call where the behavior was missed, pick three
Agent 1:1 "Here are your four calls this month" "You did this on 61% of your calls, the floor is 78%, here are four"
Month end An average with an unstated margin of error The same average with the whole population behind it, plus behavior rates
A compliance question "Let me pull a sample and check" Filter the criterion across every call in the period

Your unit of analysis moves from the call to the behavior

The most useful number stops being "Maria scores 84" and becomes "Maria asked for the appointment on 61% of eligible calls." Behavior rates are coachable in a way composite scores never were, and they are what the energy retailer's data is really about.

The first month is worse before it is better

Expect a score drop that is a scoring difference rather than a performance change, a batch of criteria you discover are unscorable, more flags than your team can action until you tune thresholds, and a round of agent pushback. Budget three to six weeks for that and it is a rollout. Do not, and it reads as a failed implementation.

How to evaluate an auto-QA vendor

Test on your own audio, not the demo corpus. If you are still assembling the shortlist to run these checks against, our rundown of the best call center quality assurance software is where to build it.

  1. 1

    Run a matched-pair test

    Take 50 to 100 calls you have already scored by hand, run them through the platform, and compare criterion by criterion. You are not looking for identical scores. You are looking at where they disagree, and whether that is the AI being wrong or your rubric being ambiguous. A vendor unwilling to support this test is telling you something.

  2. 2

    Ask for evidence on a failed criterion, live

    The fastest disqualifier available.

  3. 3

    Ask how criteria are evaluated

    Independently, or as one pass over the call. Ask whether the system can return "not applicable" and "insufficient evidence" rather than being forced into pass or fail.

  4. 4

    Test your worst audio

    Hardest accents, noisiest lines, transfer-heavy queues. Clean-audio performance tells you nothing you need to know.

  5. 5

    Ask where they are weak

    A vendor who cannot name a criterion type their system scores badly has either not measured it or will not say.

  6. 6

    Check rubric versioning and flag routing

    Ask to see a score from before a scorecard edit, and ask how flags are prioritized and routed, because unroutable volume is the most common way these deployments die.

  7. 7

    Ask whether the same transcripts support open questions

    Scoring answers what you asked; you will also need to ask things the rubric never anticipated, which is what AI call overviews across a batch and asking questions of your calls directly are for.

Where we stand

Scoring and exploration, over the same transcripts

Voxjar reads the recordings your phone system already produces, transcribes them, scores every call against your scorecard with the reasoning and the transcript moment attached, and runs conversation intelligence over the same transcripts so the questions your rubric did not anticipate still have somewhere to go. Scoring without exploration measures the wrong things precisely; exploration without scoring produces insight nobody is accountable for.

We publish the human-versus-AI gap because the alternative is that you find it in month two. Auto-QA is a better instrument than sampling for most of what QA programs are asked to do, and it is wrong often enough that a human has to stay in the loop for anything consequential. A vendor who tells you only the first half is not describing the product you will operate.

Run a free AI evaluation on one of your own calls. Upload a recording, apply a scorecard, and look at the score, the reasoning, and the evidence. Then take the criterion you are least sure about and see whether the problem is the AI or the way the criterion is written.

Frequently asked questions

What is auto QA?

Auto QA is the automatic evaluation of recorded conversations against a defined scorecard, producing a per-call score with the reasoning and the supporting quote attached to each criterion. It replaces the manual review of a small sample with evaluation of every call, and keeps humans in the loop for calibration, disputes and anything consequential.

How do I automate call center agent scoring?

Connect the recordings your phone system already produces, encode your existing scorecard as criteria the AI can judge, then run every call against it. The work is in the criteria, not the automation: a question like "was the agent professional" cannot be scored consistently, while "did the agent state the required disclosure before discussing the account" can. Start by rewriting ambiguous criteria into observable ones.

Is AI call scoring accurate?

Accurate enough to be useful, and wrong often enough that a human has to stay in the loop. Across 2,424 calls scored by both a human reviewer and AI on the same scorecard, human averages ran about ten points higher, but humans scored higher on only 36 percent of calls. Agreement is far better than the averages suggest; divergence concentrates in a minority of calls, and disproportionately in subjective criteria.

Can I use AI call scores in performance management?

Only with evidence, a dispute path, and human review before any adverse action. A score with a quote from the transcript can be examined, argued with and overturned. A score without one cannot, which makes it unusable for anything that affects someone's employment. Calibrate the AI against your reviewers first, and treat unexplained scores as a red flag in any vendor.

What is the difference between auto QA and conversation intelligence?

Auto QA judges a conversation against a standard you defined, producing a score. Conversation intelligence explains what is happening across conversations, producing summaries, topics and searchable transcripts. Scoring answers the questions you knew to ask; conversation intelligence answers the ones you did not. Most serious platforms now do both over the same transcripts.

See what your scorecard looks like applied to every call.

Run a Free AI Evaluation

Get 120 AI Credits and Full Access

AI First QA Platform
No Credit Card Required
Start Scoring in Minutes