QA for call centers:
the complete guide to quality assurance
Of the calls your team took last week, how many met your standards? That question is what QA for call centers exists to answer. This guide covers the whole discipline: what call center quality assurance is, how to build a program step by step, the failure modes that quietly kill most programs, how AI rewrote the economics of coverage, and the metrics that tell you whether any of it is working.
Score one of your own calls while you read. No card, no demo.
Every call center says quality matters
Far fewer can answer that opening question. Done well, quality assurance turns thousands of unheard conversations into a picture of how your operation really performs: which agents are losing winnable deals, where compliance risk is hiding, which processes create repeat calls, and what your best people do that the rest do not. Done badly, it becomes a monthly ritual where a supervisor scores three random calls per agent, files the results in a spreadsheet, and nothing changes.
Worth saying plainly up front, because most QA writing assumes otherwise: quality assurance is not just a customer support function. The teams that get the most out of it are usually the ones where calls carry money or risk directly. A personal injury firm's intake desk, where one mishandled call is a lost case. A collections floor, where a missed disclosure is a regulatory problem. A sales team, where the gap between a good and bad call is an appointment that did or did not get set. Support teams benefit too, but they are not the whole story, and a QA program designed only around satisfaction scores will miss most of what matters to those operations.
Call center quality assurance is the practice of systematically evaluating customer conversations against a defined standard, then using those evaluations to improve agent performance, fix broken processes, and manage compliance risk. Three parts of that definition do the heavy lifting.
Systematically
QA is a repeatable process, not a manager occasionally listening in when something feels off. Calls are selected, scored, and reviewed on a schedule, the same way every time.
Against a defined standard
Every call is evaluated with the same criteria, usually captured in a scorecard. Without a shared standard, quality is just each reviewer's opinion.
To improve
The score is not the product. The coaching conversations, process fixes, and compliance interventions that come out of the score are the product. A QA program that produces numbers but no changes is theater.
In practice, a QA program has five moving parts. The rest of this guide is those five, in order.
| Quality standards | What good looks like |
| Scorecards | How you measure it |
| Review process | Who evaluates which calls, and how many |
| Calibration | Keeping reviewers consistent |
| Feedback loop | Coaching and process change driven by the results |
Where the payoff actually lands
The business case is not abstract. Your phone lines are where money is won or lost and where risk is created, hundreds of times a day, in conversations almost nobody reviews. Quality assurance is how you manage that at scale, and it pays off in four places.
Revenue on the calls you already paid for
If your calls are sales, intake, or booking, QA is the only way to see why one agent converts and another does not. Most teams discover the gap is not effort or volume: it is that the offer never got made, the appointment never got asked for, or the objection never got answered. You already paid to generate that call. QA tells you what happened to it.
Compliance and liability
For collections, healthcare, insurance, and legal intake, a missed disclosure or an improper statement is not a quality issue, it is a liability. Reviewing 2% of calls means you find out about the other 98% from a regulator or a plaintiff.
Agent performance and turnover
Agents improve fastest with specific, evidence-based feedback tied to their own calls. Vague feedback, or none at all, is a top driver of frustration in a role that already has notoriously high attrition.
Customer experience and retention
QA surfaces the behaviors that resolve issues on the first call and the ones that send customers to a competitor. You cannot coach what you cannot see.
Which of those four leads depends entirely on what your calls are for, and that shapes the whole program.
What to measure changes by operation type
The mechanics in this guide are the same everywhere: standards, scorecards, coverage, calibration, coaching. What changes is what you put on the scorecard and which number you are trying to move. Copying a customer support scorecard onto a sales floor is the single most common way a QA program ends up measuring the wrong thing accurately. Four broad patterns cover most operations.
Revenue calls
Where the outcome is a sale, an appointment, or a signed client. The scorecard should be weighted toward the moments that decide the outcome: whether the agent actually asked for the business, handled the objection, and confirmed next steps. Scoring tone and courtesy here feels productive and changes nothing.
The pattern for personal injury intake, real estate ISAs, insurance agencies, mortgage, solar and home improvement, and lead generation.
Compliance-critical calls
Where a single sentence creates liability. These scorecards need auto-fail gates rather than point deductions, because a call that omits a required disclosure is not an 85, it is a failure regardless of how well the rest went. Sampling is the weak point: reviewing 2% means most violations are never seen.
The pattern for collections, addiction treatment, and dental and medical front desks.
Dispatch and service coordination
Where the call sets off physical work. Errors are expensive downstream: a wrong address, a missed detail, an over-promised window. Scorecards should weight information capture and expectation setting.
The pattern for transportation dispatch, home services, moving companies, and waste management.
Outsourced and multi-client operations
Where QA is a deliverable. A BPO is not scoring calls only to improve; it is scoring them to prove performance to a client who is deciding whether to renew. That demands per-client scorecards, defensible evidence attached to every score, and reporting the client can audit.
The pattern for BPOs and outsourced contact centers.
If your operation spans two of these, and many do, build separate scorecards rather than one compromise scorecard. A single form that tries to serve a sales team and a compliance team serves neither.
Step 1
Define your quality standards
Whether you are starting from zero or rebuilding a program that has drifted, the sequence is the same, and it starts here. Before you can score anything, you need to decide what a good call actually is for your operation. Do not build this in a silo. Pull in supervisors, your best agents, and whoever owns compliance, and answer three questions concretely.
| What must happen on every call? | Identity verification, required disclosures, accurate information. |
| What should happen on most calls? | Discovery questions, empathy, clear next steps, first-contact resolution. |
| What must never happen? | Misleading claims, disclosing information to the wrong party, rudeness, hanging up on a customer. |
Ground the standards in reality. Listen to ten of your best calls and ten of your worst before writing anything. Standards written from imagination produce scorecards nobody believes in; standards written from real calls produce scorecards agents recognize as fair.
Step 2
Decide how many calls you will review
Now the uncomfortable math. A thorough manual evaluation takes roughly 15 to 20 minutes per five-minute call once you account for listening, scoring, and writing notes. A full-time reviewer can realistically get through 20 to 30 calls a day. If your center handles 10,000 calls a month, one dedicated reviewer covers about 2 to 4 percent of them.
That is why nearly every manual QA program is a sampling program. If sampling is your reality, do it well.
Randomize within strata
Pull calls across agents, times of day, call types, and call lengths. If reviewers pick calls by hand, they drift toward short calls and familiar agents.
Set a floor per agent
Three to five evaluations per agent per month is a common minimum. Below that, a single unlucky call swings an agent's monthly score by double digits.
Oversample where risk lives
New hires, agents on improvement plans, and regulated call types deserve more than their proportional share.
Be honest about what sampling can and cannot do. It can track team-level trends and support coaching. It cannot catch the one compliance violation in the 97 percent of calls nobody hears, and any per-agent score built on four calls has a wide error bar. Run your own numbers below, then compare the answer to what your team reviews today.
Calculate Your Sample Size
Per Agent Basis
calls each month
Full Team Basis
calls each month
Tired of manual call evaluations?
Instead of manually grading 660 calls, let Voxjar score 100% of your calls automatically.
Step 3
Build your scorecard
The scorecard translates your standards into something a reviewer, human or AI, can apply consistently. A good one has weighted categories (compliance, communication, process adherence, resolution), specific observable criteria under each, and auto-fail logic for the small set of violations that should zero out an otherwise decent score.
Two rules of thumb: keep it to 25 criteria or fewer, and write every criterion as a question about observable behavior ("Did the agent confirm the callback number?") rather than a judgment ("Was the agent professional?"). Observable criteria are what let two reviewers apply the same scorecard the same way.
Scorecard design is a deep enough topic that we wrote a separate guide with three complete, copy-and-paste call center scorecard templates for inbound support, outbound sales, and compliance-heavy operations. Start there rather than from a blank page.
Step 4
Calibrate your reviewers
Hand the same call and the same scorecard to three reviewers and you will often get three different scores. That variance quietly destroys trust: when an agent's score depends on who reviewed the call, agents stop treating scores as feedback and start treating them as luck.
Calibration is the fix. On a regular cadence, usually monthly, reviewers independently score the same call, compare results criterion by criterion, and argue out every disagreement until they converge on a shared interpretation. Document the rulings, because they become your scoring precedents. Track the spread between reviewers over time; if scores on a calibration call differ by more than about 5 percent, your criteria are too vague or your definitions have drifted.
Calibration sessions are also where your scorecard improves. A criterion that generates disagreement every single month is not a reviewer problem, it is a wording problem. Fix the criterion.
Step 5
Close the loop with coaching
An evaluation that never turns into a conversation is wasted work. The score exists to make coaching specific, and specificity is what separates coaching that changes behavior from coaching that fills a calendar slot. For session templates, talk tracks, and scenario walkthroughs, see our full call center coaching guide. Four things work consistently.
Coach from evidence, quickly
Feedback tied to a specific moment in a specific call, delivered within days, beats a monthly summary of averages every time.
Keep it short and frequent
Ten focused minutes on one behavior weekly outperforms an hour-long quarterly review of everything.
One behavior at a time
An agent handed six improvement areas will improve at none of them. Pick the highest-impact gap, fix it, then move to the next.
Study your top performers, not just your strugglers
Reviewing your best agents' calls tells you what excellence actually sounds like in your operation, gives you real examples for training, and keeps QA from feeling purely punitive. Praise high performers openly; a program that only ever surfaces problems will be resented.
And treat quality metrics as a guide, not a cage. An agent with outstanding customer satisfaction and slightly long handle times is usually an asset, not a problem to correct.
The six ways QA programs quietly fail
Most QA programs do not fail loudly. They fail in one of these predictable ways.
The checkbox program
Evaluations happen because a manager requires them, scores get filed, and nothing downstream changes. If you cannot name a process or behavior that changed last quarter because of QA, you have this problem.
Sample sizes too small to mean anything
Ranking agents, paying bonuses, or writing improvement plans based on three or four calls a month is statistics abuse. Small samples are fine for finding coaching examples and terrible for judging people.
The unfair-score spiral
No calibration leads to inconsistent scores, which leads to agents disputing evaluations, which leads to supervisors softening scores to avoid conflict, which leads to a scorecard where everyone gets a 94 and the numbers mean nothing.
Scoring what is easy instead of what matters
Greeting used, hold procedure followed, branded sign-off delivered: easy to verify, weakly correlated with whether the customer got help. If a call can score 95 while the customer leaves furious and unresolved, the scorecard is measuring the wrong things.
QA as gotcha
When evaluations only ever appear as ammunition in disciplinary conversations, agents learn to fear the program and game the scorecard. The programs that work frame QA as the engine of coaching and celebrate great calls as often as they flag bad ones.
Insights that never leave the QA team
Reviewers notice the same broken process causing calls week after week, but there is no channel to route that to operations or product. Half the value of QA is organizational learning; build the channel.
Step 6
Measure whether the program is working
Make the program accountable to results. Are coached behaviors actually improving on subsequent evaluations? Are compliance pass rates trending up? Is QA insight reaching operations, or dying in a spreadsheet? Review the program quarterly, and be willing to rewrite standards and scorecards as your products, scripts, and channels change. A QA framework is a living document; a scorecard nobody has revised in two years is measuring the operation you used to have.
Track a small set of numbers, and keep call-level quality metrics separate from program-level health metrics.
| Quality and outcome metrics |
|---|
| Overall quality score Trended by agent, team, and call type. The trend matters far more than the absolute number. |
| Compliance pass rate Tracked separately from the blended score, because a 90 percent average can hide a 6 percent violation rate. |
| First-contact resolution The quality metric customers actually feel. |
| Customer satisfaction (CSAT) Cross-referenced against QA scores. If they diverge for long, your scorecard is measuring something customers do not care about. |
| Average handle time Watched as a guardrail rather than a target, so quality gains are not coming from rushed calls, and coaching is not inflating talk time. |
| Program health metrics |
|---|
| Coverage rate The percentage of calls evaluated. This is the number AI changes most. |
| Calibration variance The score spread between reviewers on a shared call. Under 5 percent is the goal. |
| Coaching follow-through The share of flagged issues that got a coaching conversation, and whether the coached behavior improved on the agent's next evaluations. This is the single best indicator that QA is driving change rather than producing paperwork. |
| Score trend post-coaching If scores never move after coaching, either the coaching or the scorecard is broken. |
What to automate, and what changed
Everything above describes the discipline. What changed in the last few years is the cost structure underneath it. The traditional constraint on QA has always been reviewer time. At 15 to 20 minutes of review time per call, coverage beyond a few percent of volume requires a headcount investment most operations cannot justify. Every design decision in classic QA, sampling strategy, per-agent minimums, monthly cadences, is downstream of that single constraint.
AI evaluation removes it. Modern call center quality assurance software ingests the recordings your dialer or phone system already produces, transcribes them, and applies your scorecard to every single call, with each score backed by the transcript evidence behind it. The practical differences are stark.
| What changes | Why it matters |
|---|---|
| Coverage goes from 2 to 4 percent to 100 percent | Every call is evaluated, so compliance issues surface wherever they occur instead of only when a sampled call happens to contain one. |
| Per-agent scores become statistically real | An agent's monthly score built on 200 evaluated calls is a measurement; one built on 4 calls is an anecdote. |
| Feedback gets faster | Calls are scored within hours of the recording arriving, so coaching can reference yesterday's calls instead of three weeks ago. |
| Consistency is structural | The same model applies the same rubric to every call. You still calibrate, but now you calibrate the AI's rubric against your judgment, refining criteria until its scores match how your best reviewer would score, and then that consistency applies everywhere at once. |
| Human reviewers move up the stack | Instead of spending their hours listening to randomly selected, mostly fine calls, your QA people spend them on the flagged calls, the disputed scores, and the coaching conversations. |
In that last row, AI handles the call scoring and humans handle the judgment and the people. The same coverage also feeds conversation intelligence, so the themes running through every call are visible alongside the scores rather than in a separate system.
Two honest caveats. First, AI evaluation is only as good as the scorecard you give it; vague criteria produce vague scores no matter who or what applies them, which is why steps 1 and 3 above still come first. Second, this does not eliminate the QA role. It eliminates the listening-lottery part of the job and leaves the parts that actually needed a human.
If you want the mechanics rather than the economics, our guide to auto QA walks through how an AI actually applies a scorecard to a call step by step, what makes a criterion machine-scorable, and where the scoring gets things wrong. And if you are evaluating tools, we maintain an honest, verdict-first roundup of the best call center quality assurance software, including where each option fits and where it does not.
Where to start
If you are building or rebuilding a QA program, the path is short: define standards from your real calls, adapt a scorecard template rather than starting blank, be deliberate about coverage, calibrate monthly, and make coaching the output that everything else serves. When manual sampling stops being enough, AI QA software takes the same scorecard to 100% of your calls.
And if you want to see what 100% coverage looks like on your own calls before committing to anything, you can try Voxjar's AI evaluation free. Upload a few real recordings, apply a scorecard, and compare the AI's evaluations to your own. No sales call required; the calls you already have are the only demo that matters.
Frequently asked questions
What is call center quality assurance?
Call center quality assurance is the practice of systematically evaluating customer conversations against a defined standard, then using those evaluations to improve agent performance, fix broken processes and manage compliance risk. A program has five moving parts: quality standards, scorecards, a review process, calibration to keep reviewers consistent, and a feedback loop that turns scores into coaching and process change.
What are call center quality assurance best practices?
Define observable standards before building a scorecard, weight criteria by business impact rather than convenience, calibrate reviewers until score spread is under about five points, close the loop with coaching rather than filing scores in a spreadsheet, and measure the program itself. The most common failure is a program that produces accurate numbers nobody acts on.
How many calls should you review per agent?
More than most programs manage. At four calls per agent per month the confidence interval on that agent's score is roughly plus or minus 19 points, so a 75 and an 85 are statistically indistinguishable. Sampling works reasonably for team-level trends and poorly for individual judgments or rare events like compliance breaches, which is why full-coverage scoring changed the economics of QA.
What is a call center QA scorecard?
A structured evaluation form: a fixed set of criteria, each with a point value, grouped into weighted categories and applied identically to every reviewed call. Good scorecards separate binary compliance criteria, which either happened or did not and often carry auto-fail gates, from scaled behavioural criteria that admit degrees.
How often should you calibrate QA reviewers?
Monthly is a reasonable default, more often when a scorecard changes or a new reviewer joins. Run the session by having every reviewer score the same call independently, then discuss only the criteria where scores diverged. Track the spread over time; the goal is convergence, not agreement enforced by the loudest voice in the room.
Stop reviewing 2% of your calls.
Get 120 AI Credits and Full Access