Guide

Call monitoring forms: designing a QA checklist
that actually measures quality

The anatomy of a call quality monitoring form, how to write criteria two reviewers answer the same way, weighting, auto-fails, calibration, and the design failures that quietly ruin call quality audits.

The form is the instrument your whole QA program reads from. This page is the design manual; if you want finished forms instead, the templates are free.

What a call monitoring form is

A call monitoring form is the instrument your QA program uses to measure a conversation: a structured set of criteria, grouped into sections and weighted by importance, that a reviewer answers for a specific call to produce a score.

Some teams call it a QA checklist, a call quality monitoring form, an evaluation form, or a scorecard. The names vary; the job does not. It converts "was this a good call" into a set of questions that can be answered consistently, whether the reviewer is a person or an AI, compared across agents, and tracked over time.

That makes it the highest-leverage document in your quality program, and the least examined. Everything downstream inherits its quality: your call quality audits, your coaching conversations, your compliance evidence, your agent rankings. A badly designed form does not produce slightly worse data. It produces confident, precise-looking numbers that measure nothing, and an operation that steers by them anyway.

This guide covers the design work. If you want finished forms rather than a design manual, our call center scorecard templates are free, complete, and written to the standard this page describes.

The anatomy of a call monitoring form

Every workable form has the same skeleton, whether it lives on paper, in a spreadsheet, or in QA software. Six parts.

Sections

Group criteria by phase or theme: greeting, discovery, compliance, resolution, closing. Sections exist for the reader, not the math. They let a manager see at a glance that an agent is strong through discovery and weak at close, which a single composite score hides.

Criteria

The questions themselves. Each one asks about a single observable behavior in the call: "Did the agent confirm the callback number before ending the call?" Criteria are where forms are won and lost, and most of this page is about writing them.

Weights

Weights encode business impact. Not every behavior matters equally, and a form that pretends otherwise tells agents that a warm greeting and a required legal disclosure are interchangeable. Weights can sit on individual criteria, on sections, or both.

Auto-fail gates

Some failures invalidate the call no matter how well everything else went: disclosing account details to an unverified caller, promising something the company cannot deliver, profanity, dishonesty. An auto-fail sits outside the weighted math and zeroes the score, because a call that breached compliance was not 82% good.

A not-applicable state

Plenty of criteria only apply to some calls. "Did the agent respond to the objection" has no answer on a call with no objection, and a form that forces pass or fail on it is manufacturing noise. Every criterion that can be irrelevant needs an N/A that removes it from the denominator.

An evidence field

Space for the reviewer to note the moment that justified the answer: a timestamp, a quote, a short note. Scores without evidence cannot be coached from, disputed, or audited, and the discipline of citing evidence keeps reviewers honest about answers they are guessing on.

On the scale for each criterion: default to binary. Pass, fail, or N/A. A 1-to-5 rating on "professionalism" feels more precise and is less: no two reviewers share a definition of a 3, so the added resolution is added noise. Use graded scales only where the gradations are themselves defined in behavioral terms, and reach for them rarely.

Writing criteria that are observable and falsifiable

This is the discipline that separates a measurement instrument from an opinion survey, and it comes down to one test.

The stranger test

A competent stranger, given only the recording or transcript, should answer the criterion the same way you would.

A criterion passes that test when it has three properties.

1

It is observable

It refers to something that happened in the conversation, not to an internal state, an impression, or a fact outside the call. "Did the agent show empathy" names a feeling. "Did the agent verbally acknowledge the customer's stated problem before moving to a solution" names a behavior that either appears in the transcript or does not. You lose some nuance in the trade. You gain a number that means the same thing every time it is produced, and that trade is the entire point of having a form.

2

It is falsifiable

There must be a realistic call that fails it. "Did the agent attempt to help the customer" fails no one; every reviewer can construe any call as an attempt. A criterion that cannot fail is not a criterion, it is decoration, and decoration is not harmless: it inflates every score it touches and buries the criteria that carry information.

3

It has one decision point

"Did the agent verify identity and explain the recording notice" is unanswerable when one happened and the other did not. Split compound questions until each criterion asks exactly one thing, even when that lengthens the form. Ten answerable questions beat six unanswerable ones.

Decoration is not a victimless flaw. Our auto-QA guide documents a real scorecard where a block of soft-skill criteria sat at a 0.0% fail rate across more than 30,000 evaluations each. Whatever those criteria were doing, it was not measurement. If a criterion never fails, it inflates every score it touches and buries the criteria that carry information.

State the pass condition, not the topic

"Hold procedure" is a topic; "Did the agent ask permission before placing the customer on hold" is a pass condition. A reviewer should never have to infer what counts.

Define N/A explicitly

Write into the criterion when it does not apply: "If no objection was raised, mark N/A." Reviewers forced to improvise applicability rules will improvise them differently.

The same discipline determines whether a form can be scored by AI. A criterion a stranger can answer from the transcript is a criterion a model can answer from the transcript; a criterion that relies on reviewer intuition fails both scorers, and manual review just hides the failure behind reviewer confidence. The auto-QA guide covers the AI-scorability side in depth, including a table of unscorable criteria and their rewrites and the failure modes that survive even good design. Write the form to the stranger test and you get a form that works on paper today and scores automatically tomorrow, without a rewrite in between.

Weighting criteria by business impact

Once the criteria are written, decide what each one is worth. The principle is simple: weight by the cost of the failure, not the frequency of the behavior.

A missed compliance disclosure can cost a lawsuit. A missed appointment ask costs the revenue of that call. A missed "thank you for calling" costs almost nothing. Yet forms routinely weight these within a few points of each other, because each section got a tidy share of 100 and the criteria split it evenly. The result is a form where an agent can fail everything that matters commercially and still score in the 80s on courtesy.

  1. 1

    Sort criteria into three tiers

    Ask of each one: if this fails, what does it cost? Tier one is legal exposure, lost revenue, or lost customers. Tier two is degraded experience or rework. Tier three is polish.

  2. 2

    Promote the worst of tier one to auto-fail

    If a failure invalidates the call, no weight is large enough; take it out of the arithmetic entirely and let it gate the score. Keep the gate list short and non-negotiable, because an auto-fail that gets waived when the call was otherwise good is not a gate, it is a suggestion.

  3. 3

    Distribute weight across what remains, unevenly

    Tier one criteria should individually outweigh tier three criteria by multiples, not percentage points. If deleting a criterion would change no decision anyone makes, its weight is telling you it does not belong on the form.

Then check the weights against reality. After a few weeks of scoring, look at your score distribution next to your outcomes. If agents who score 90 and agents who score 70 close at the same rate, resolve at the same rate, and generate the same complaints, your weights (or your criteria) are measuring something other than performance, and the form needs revision, not defense.

How many criteria before a form fails

Fewer than you have. Almost every form we see errs long, because forms accrete: every incident adds a criterion and none ever leaves. Length fails the form in three distinct ways.

Reviewers degrade

A human scoring 30 criteria per call does not evaluate 30 behaviors; they form an overall impression in the first minute and back-fill the checkboxes to match it, which means your 30-criterion form is measuring one criterion called "vibe" with 30 decimal places.

Agents disengage

Nobody can hold 30 targets in mind on a live call. Past a certain length the form stops shaping behavior and becomes something that happens to agents after the fact.

Coaching dissolves

A review that surfaces 11 improvement areas surfaces none; priority is the thing a long form cannot express.

Practical bounds

Hold a manually scored form to 15 criteria or fewer, and treat anything past 20 as a redesign trigger. Get there by deleting the never-fail criteria (they contribute score but no information), merging near-duplicates, and moving anything that exists only for reporting, rather than coaching, out of the form and into your analytics.

Automated scoring relaxes the reviewer-fatigue bound, since an AI scores criterion 27 with the same attention as criterion 1, and it evaluates each criterion independently rather than back-filling from an impression. It does not relax the other two bounds. Agents still cannot act on 30 priorities, and coaching still needs focus. Full coverage means a leaner instrument applied to every call, not a longer form; if you want the watch-everything list anyway, that is what conversation intelligence over the same transcripts is for, not the scorecard.

Forms by operation type

One form cannot serve every call type. Sales is not support, and an outsourcer running twelve programs needs twelve forms, because a criterion that is essential on one queue is noise on another. The skeleton stays constant; the criteria and the weighting center of gravity move.

Sales and outbound

Weight concentrates on the behaviors that produce the outcome: was discovery done, was the product actually offered, were objections addressed, was the close attempted, were next steps set with a date. The single most valuable criterion on most sales forms is the simplest: did the agent ask for the sale. It is the behavior most often missing and the one with the most direct revenue consequence.

Intake and appointment setting

The call exists to capture information and convert the caller, so the form checks completeness and momentum: required fields collected and verified, qualifying questions asked, the appointment or next step offered on this call rather than promised in a follow-up, urgency handled without pressure.

Compliance-critical operations

Collections, healthcare, financial services, insurance. The form inverts: gates dominate and style recedes. Required disclosures stated, identity verified before any account detail, prohibited language absent, dispute and do-not-contact requests handled per procedure. Most of these belong in the auto-fail list rather than the weighted body, and the evidence field stops being a coaching aid and becomes your audit trail.

Support and service

Weight lands on diagnosis and resolution: the problem restated and confirmed, ownership taken, an accurate solution or a concrete next action with an owner and timeframe, expectations set honestly. Resist importing satisfaction into the form; the customer's mood is an outcome the agent influences but does not control, and a form should hold agents to their own behavior, not the customer's day.

Build the variants by subtraction from the standard, not accumulation: start from your core form, remove what does not apply to the queue, and add the few criteria the operation genuinely requires. And if you want ones ready to use, our scorecard templates include finished forms for sales, intake, compliance, and support that you can take as they are or edit down.

Calibration: keeping scorers consistent

A form is only an instrument if it produces the same reading regardless of who holds it. Calibration is how you test that: multiple reviewers score the same call independently, compare results criterion by criterion, and argue out the differences until the rulings are written down. We cover the session mechanics, cadence, and variance math in the call calibration guide; what matters here is what calibration does to the form.

Score variance is a property of the criterion, not the reviewers. When one rep gets 3/10, 6/10, and a delighted customer survey on the same call's "tone of voice," the finding is not that two reviewers need retraining. It is that "tone of voice" is not defined tightly enough to be scored, and every spread like that is the form telling you which criterion to rewrite next. Run the sessions on a schedule, log every disagreement against the criterion that produced it, and treat the log as the form's bug tracker.

Calibration also keeps its central role when AI does the scoring. The AI becomes one more scorer in the session: score the call as a group, compare against the machine, and write down the ruling when you disagree. Each disagreement resolves to either a rubric fix or a documented limitation, and both are worth having in writing.

Common form failures

The same handful of failures accounts for most broken forms we encounter. In roughly the order of damage done.

  1. 1

    One form for every call type

    Sales scored on a support form, or one bloated universal form where most criteria are N/A or noise on any given call. Different operations need different instruments.

  2. 2

    Criteria that never fail

    Universal-pass criteria inflate every score and dilute every real signal. Audit fail rates quarterly; a criterion at or near a 100% pass rate either gets rewritten with teeth or deleted.

  3. 3

    Ambiguity

    "Was the agent professional" scored three ways by three reviewers. Every judgment word on the form (professional, appropriate, effective, awesome) is a defect until it is defined behaviorally or replaced.

  4. 4

    Scoring what the agent does not control

    Holding reps accountable for the customer's happiness rather than their own behavior. An angry customer can stay angry through a flawless call; the form should be able to say so.

  5. 5

    Over-scripting

    Requiring exact phrases at exact moments strangles conversation and optimizes recitation over outcome. Gate the few legally mandated wordings, and everywhere else score whether the job got done, not whether the incantation was performed.

  6. 6

    Length

    Covered above, and it compounds every other failure: a 30-criterion form is where never-fail criteria and ambiguity hide.

  7. 7

    Weights that ignore consequence

    Courtesy weighted level with compliance, so the score cannot distinguish a rude-but-safe call from a friendly liability.

  8. 8

    A frozen form

    The business changes scripts, products, and regulations; the form from 2023 keeps measuring 2023. A form is a living document with an owner, a version history, and a revision cadence. If nothing has changed in a year, that is not stability, that is neglect.

Where this goes next

From a designed form to a running program

A finished form is the entry ticket, not the program. It needs reviewers, a sampling or coverage strategy, calibration on a calendar, a dispute path, and reporting that turns scores into coaching. All of that is what call center quality assurance software exists to run, from the form itself through evaluations, disputes, and trend reporting.

And a form built to the standard on this page has one more property worth naming: it is ready for scale. Because every criterion is observable, falsifiable, and single-pointed, the form scores identically whether a person fills it in on eight calls a month or AI scores it on every call. Teams that skip the design work and automate anyway just discover their ambiguous criteria at 100% coverage instead of 3%. Teams that do the work get the same instrument at any volume.

Start here

Start wherever removes the most friction: take a free scorecard template built to this standard, or take the form you have, rewrite it against the stranger test, and run one call through it to see what the instrument can tell you.

Frequently asked questions

What is a call monitoring form?

A call monitoring form is a structured evaluation instrument used to assess the quality of a phone conversation. It contains criteria, usually grouped into sections such as greeting, compliance, and resolution, that a reviewer answers for a specific call, with weights that roll the answers up into a score. It is also called a QA checklist, call quality monitoring form, evaluation form, or call scorecard.

What should a call center QA checklist include?

Sections that follow the call's phases; criteria that each name one observable behavior with a stated pass condition; weights that reflect business impact; auto-fail gates for failures that invalidate the call, such as compliance breaches; a not-applicable option for criteria that do not apply to every call; and an evidence field tying each answer to a moment in the conversation.

How many criteria should a call monitoring form have?

Fifteen or fewer for manually scored forms, and treat more than 20 as a redesign trigger. Long forms push reviewers into impression-based scoring, overwhelm agents with more targets than anyone can act on, and dilute coaching. Automated scoring removes the reviewer-fatigue limit but not the others, so lean stays right even at 100% coverage.

What is an auto-fail on a QA form?

An auto-fail is a criterion that sits outside the weighted scoring and zeroes the call's score when breached, regardless of how well everything else went. Typical auto-fails are missed required disclosures, sharing account details without identity verification, dishonesty, and profanity. They exist because a call that breached compliance was not partially good.

What is the difference between a call monitoring form and a scorecard?

In practice they are the same artifact, and most teams use the words interchangeably. When a distinction is drawn, the form is the instrument (the questions, weights, and gates) and the scorecard is the scored result for a call or an agent. If you want ready-made ones, our scorecard templates are free and editable.

How often should you update a call monitoring form?

Review it quarterly and after any change to scripts, products, process, or regulation. Use calibration disagreements and criterion fail rates as the revision queue: criteria that reviewers split on need rewriting, and criteria that never fail need teeth or deletion. A form that has not changed in a year is measuring a version of the operation that no longer exists.

See what a well-built form finds in your own calls.

Run a Free AI Evaluation

Get 120 AI Credits and Full Access

AI First QA Platform
No Credit Card Required
Start Scoring in Minutes