How to evaluate customer service calls
A customer service evaluation only works if it happens. Human reviewers typically reach 2 to 3 percent of call volume, and the method below works exactly the same whether you apply it to that sample or to every call: pick criteria that match what the call is for, weight them, decide what is pass/fail and what is graded, and calibrate your scorers. Then we evaluate one complete call, criterion by criterion, so you can see the method on a real-shaped conversation.
Evaluation is a method, coverage is a choice
Most advice on evaluating customer service skips the part that decides whether any of it matters: how many conversations actually get evaluated. A careful manual review takes 15 to 20 minutes per five-minute call, so a dedicated reviewer covers a low single-digit percentage of a busy operation's volume. That listening is valuable work, it is where coaching evidence and process insight come from. There is simply far more of it worth doing than any team has hours for.
So treat the two decisions separately. This page is about the method: what to score and how to score it consistently. Coverage, whether the method runs on a 3 percent sample or on 100 percent of calls, is a second decision the method survives either way. Build the evaluation right and it works at any scale; build it wrong and no amount of coverage fixes it.
Step 1
Pick criteria that match what the call is for
The most common evaluation mistake happens before a single call is scored: using one generic form for every kind of conversation. A sales call, an intake call, a collections call, and a support call succeed at different things, and criteria borrowed from the wrong operation type measure the wrong thing accurately. Start by naming what the call exists to accomplish, then write criteria as questions about observable behavior.
Sales calls
The call exists to move a buyer forward. Criteria should sit on the moments that decide the outcome: discovery questions asked, the offer actually made, the objection answered rather than dodged, and a committed next step. A sales call can be warm, polite, and a complete failure.
- Asked discovery questions before pitching
- Made the offer or the ask explicitly
- Addressed the stated objection
- Secured a concrete next step with a time
Intake calls
The call exists to capture a new customer or case correctly. Criteria should weight information capture and qualification: the required questions asked, details confirmed back, urgency assessed, and the handoff or booking completed. A missed question here is a downstream failure someone else inherits.
- Collected every required qualifying field
- Confirmed contact details back to the caller
- Assessed urgency and routed correctly
- Booked or scheduled before ending the call
Collections calls
The call carries compliance weight on every sentence. Criteria split into two groups: required conduct, which is pass/fail with auto-fail gates (identity verification, required disclosures, prohibited statements avoided), and effectiveness, which is graded (payment discussion, arrangement offered, commitment secured).
- Verified identity before discussing the account
- Delivered required disclosures
- Avoided prohibited statements entirely
- Proposed an arrangement and asked for commitment
Support calls
The call exists to resolve an issue and keep the customer. Criteria should weight diagnosis and resolution over etiquette: the issue restated accurately, the real cause found, a fix delivered or a concrete committed next step, and the customer clear on what happens next. This is the operation type the worked example below uses.
- Acknowledged and restated the issue
- Diagnosed the actual cause, not the symptom
- Resolved it or committed to a dated next step
- Summarized the outcome before closing
Keep the full list to 25 criteria or fewer. Past that, reviewers rush, and every criterion gets a shallower look than it deserves.
Step 2
Weight criteria by business impact
Unweighted scorecards drift toward measuring whatever is easiest to verify. A branded greeting and a proper sign-off are simple to check, weakly connected to whether the customer got what they needed. Weighting is how you encode judgment: the criteria closest to the call's purpose carry the most points. Here is a reasonable weighting for a support operation, as category shares of a 100-point evaluation.
| Category | Weight | Why |
|---|---|---|
| Process and verification | 20% | Greeting, identity or account verification, correct system notes. Necessary, but not what the call is for. |
| Diagnosis and resolution | 40% | The heaviest weight goes to the reason the call exists: was the real problem found and fixed, or moved concretely toward fixed. |
| Communication | 25% | Acknowledgment, clear explanation in plain language, expectations set honestly. |
| Wrap-up | 15% | Outcome summarized, next steps confirmed, nothing left ambiguous. |
A sales operation would move that heaviest weight onto the offer and the committed next step; a collections operation onto required conduct. The principle is the same everywhere: if a call can score 90 while failing at the thing it exists to do, the weights are wrong. For complete, copy-and-paste starting points, our call center scorecard templates include weighted builds for inbound support, outbound sales, and compliance-heavy operations.
Step 3
Decide what is pass/fail and what is graded
Not every criterion should be scored the same way, and forcing one scoring mode onto everything is a quiet source of inconsistent evaluations. The split is simple: binary events get pass/fail, behaviors with degrees get a defined scale.
Pass/fail (binary)
Events that either happened or did not: identity verified, required disclosure delivered, callback number confirmed.
Reserve auto-fail gates for genuine violations, the small set of misses that should zero out an otherwise decent call. If a third of your criteria are auto-fails, ordinary calls fail for ordinary reasons and agents stop trusting the score.
Graded (scaled)
Behaviors that admit degrees: quality of diagnosis, clarity of explanation, how completely an objection was handled.
Define every point on the scale in writing. A 1-to-5 scale with no definitions is five different opinions wearing the same number. Three well-defined levels beat five vague ones.
Step 4
Calibrate your scorers
An evaluation form is only half of consistency. The other half is the people applying it. On a regular cadence, monthly is a sensible default, have every scorer evaluate the same call independently, then compare results criterion by criterion and discuss only the places where scores diverged. Write the rulings down; they become your scoring precedents, and they are the training material for every new reviewer.
Track the spread between scorers over time. If scores on a shared calibration call differ by more than about 5 percent, the problem is usually the criterion, not the reviewer: a question that generates disagreement every session is worded ambiguously, and the fix is rewriting it. Calibration is also how AI evaluation stays trustworthy; you calibrate the AI's rubric against your best reviewer's judgment the same way you calibrate reviewers against each other.
Worked example
One support call, evaluated criterion by criterion
Here is the method applied end to end. The call is invented for illustration: a customer calls about being billed twice for the same subscription plan. The agent verifies the account, refunds the duplicate charge, and closes the call in about six minutes. Sounds like a good call, and it mostly is. The evaluation shows precisely where it was not, which is the entire point: a score without the "what the reviewer heard" column is a number nobody can coach from.
| Criterion | Type | Score | What the reviewer heard |
|---|---|---|---|
| Verified the account before discussing details | Pass/fail | 8/8 | "Can I get the email on the account and the last four digits of the billing zip?" Verified both before opening the account record. Pass. |
| Acknowledged and restated the customer's issue | Graded | 12/12 | "So you were charged on the 3rd and again on the 9th for the same plan, and the second one should not be there." Restated in the customer's own terms and confirmed before moving on. Full marks. |
| Diagnosed the cause, not just the symptom | Graded | 14/20 | Found the duplicate charge and refunded it, but never checked why it occurred. The customer had two overlapping subscriptions from an earlier plan change, which means the same charge will recur next cycle. Partial credit: the symptom was handled, the cause was not. |
| Resolved the issue or committed to a dated next step | Graded | 20/20 | "The refund is processed on my end now, and you will see it in 3 to 5 business days." A concrete resolution with a specific timeline. Full marks. |
| Explained clearly, without internal jargon | Graded | 10/13 | Mostly plain language, but "I've escalated the proration discrepancy to tier two" left the customer asking what that meant. Minor deduction. |
| Set honest expectations | Graded | 12/12 | Gave the 3 to 5 day refund window and did not promise the follow-up team would call same-day. Honest and specific. Full marks. |
| Summarized the outcome before closing | Graded | 5/10 | Closed with "anything else I can help with?" but never recapped what was done or what the customer should watch for. The customer will not remember the refund window. Half credit. |
| Professional close, offer of further help | Pass/fail | 5/5 | Offered further help, thanked the customer by name. Pass. |
| Total | 86/100 | A solid call with two coachable gaps. | |
Notice what the evaluation produced beyond the 86. Two specific coaching points: diagnose the cause behind the symptom (the duplicate charge will recur, and so will the call), and recap the outcome before closing. Both are behaviors the agent can change this week, tied to moments in a call they can listen to. That is what a good customer service evaluation looks like: evidence first, score second.
The three ways evaluations go wrong
Evaluation programs rarely fail loudly. They fail in one of three predictable ways, and all three are fixable.
Vague criteria
"Was the agent professional?" cannot be scored consistently because professional means something different to every reviewer. Every criterion should be a question about observable behavior: did the agent confirm the callback number, did the agent restate the issue. If two reviewers can honestly disagree about whether it happened, the criterion is not ready.
Uncalibrated reviewers
Hand the same call and the same form to three reviewers and you will often get three different scores. Fix it with calibration on a cadence: everyone scores the same call independently, then discusses only the criteria where scores diverged, and the rulings get written down as precedent. When an agent's score depends on who reviewed the call, agents treat evaluations as luck, not feedback.
Sampling bias
When reviewers choose which calls to evaluate, they drift toward short calls, familiar agents, and calls that were flagged for a reason. The resulting scores describe the sample, not the operation. Randomize selection across agents, times of day, and call lengths, or better, remove the sampling decision entirely by evaluating every call.
From evaluating one call to evaluating every call
Everything above scales without modification. The same criteria, weights, scoring modes, and calibration discipline that produced the worked example can be encoded into call center quality assurance software and applied to 100 percent of your conversations instead of the 2 to 3 percent a human team can reach, with each score backed by the transcript evidence behind it, the same "what the reviewer heard" column you saw above, on every call. Your reviewers then spend their expertise where it counts most: the flagged calls, the disputed scores, and the coaching conversations.
If you are starting from a blank page, do not build the form from scratch. Our scorecard templates are complete weighted builds you can adapt in an afternoon, and this page's method is exactly how to adapt them. This spoke is part of our full call center quality assurance guide, which covers the surrounding program: standards, coverage math, coaching, and measuring whether the program itself works.
Frequently asked questions
What is a customer service evaluation?
A customer service evaluation is a structured review of a real customer conversation against a fixed set of criteria, producing a score and specific feedback. Each criterion describes an observable behavior, such as verifying the account or confirming next steps, so two reviewers applying the same form to the same call reach the same conclusion. The evaluation exists to drive coaching and process fixes, not to file a number.
How do you evaluate customer service calls?
Four steps: pick criteria that match what the call is for (a sales call, an intake call, a collections call, and a support call each succeed differently), weight the criteria by business impact, decide which criteria are pass/fail and which are graded on a scale, and calibrate your scorers so the same call gets the same score regardless of who reviews it. Then apply the form to real calls and turn the results into coaching.
What criteria should a customer service evaluation include?
Criteria that describe observable behavior on the call: greeting and identity verification, acknowledging the customer's issue in their own words, accurate diagnosis, a resolution or a concrete committed next step, and a clear close. Weight them by what the call exists to accomplish. Avoid judgment words like professional or helpful as criteria; they cannot be scored consistently.
Should customer service evaluations use pass/fail or graded scoring?
Both, applied to different criteria. Binary events, like a required disclosure or identity verification, either happened or did not, so score them pass/fail, with auto-fail gates for the small set of violations that should zero out an otherwise good call. Behaviors that admit degrees, like diagnosis quality or clarity of explanation, belong on a graded scale with each level defined in writing.
How many customer service calls should you evaluate?
As many as you can score consistently. A human reviewer typically reaches 2 to 3 percent of call volume, which works for team-level trends but leaves most individual calls, and most rare events, unseen. The evaluation method is the same at any coverage level; AI evaluation applies the same scorecard to 100 percent of calls, which is how the coverage constraint gets removed rather than managed.
Run this evaluation automatically on one of your own calls.
Get 120 AI Credits and Full Access