Call Center Quality Assurance: How to Build a QA Program That Actually Improves Performance
Every call center has some form of quality assurance. Most of them are not working.
The typical QA program looks like this: a supervisor listens to a random sample of calls, scores them on a form, emails the scores to agents once a week, and files the results somewhere. Scores hover around the same range month after month. Performance does not change much. Managers wonder why they bother.
The problem is not QA itself. It is how it is set up. QA that actually improves performance looks very different from QA as a compliance checkbox.
This guide covers how to build a call center QA program that changes what agents do on calls — and how to connect QA data to the revenue metrics that actually matter.
What Is Call Center QA?
Call center quality assurance is the process of monitoring, evaluating, and improving the quality of agent interactions with customers or prospects. It involves listening to call recordings (or live calls), scoring them against a defined standard, and using that data to coach agents and improve processes.
QA serves three purposes:
Compliance protection: Agents who say the wrong thing — making promises you cannot keep, violating TCPA disclosures, or using prohibited language — create legal and regulatory risk. QA catches this before it becomes a problem.
Performance improvement: Agents who learn why specific call behaviors drive conversions perform better over time. QA is the feedback mechanism that creates this learning.
Process diagnosis: When QA data is aggregated across agents, it reveals systemic problems — scripts that do not work, objections the team is not handling well, or call stages where performance consistently breaks down.
Building Your QA Scorecard
The scorecard is the foundation of your QA program. A good scorecard measures behaviors that actually correlate with outcomes — not just whether the agent followed a checklist.
The Sections Every Outbound Scorecard Needs
1. Opening (15–20 points)
- Did the agent identify themselves and the company correctly?
- Was the opener under 15 seconds before asking permission to continue?
- Did the agent avoid robotic openers ("How are you today?" etc.)?
- Did the agent confirm they have the right person before pitching?
2. Compliance (20–25 points — auto-fail items) Compliance items should be binary: either compliant or not. A single compliance failure should trigger an automatic fail for the call, regardless of other scores. Common compliance items:
- Required disclosures made (e.g., "this call may be recorded")
- No prohibited claims or guarantees
- Do-not-call requests honored immediately
- Correct legal language used if required by industry
3. Discovery and Needs Identification (15–20 points)
- Did the agent ask at least one discovery question before pitching?
- Did the agent listen to the answer and respond specifically to it?
- Was the pitch tailored to the customer's situation or generic?
4. Objection Handling (15–20 points)
- Did the agent acknowledge objections before responding?
- Did the agent ask a clarifying question rather than immediately counter-arguing?
- Was the response specific to the objection or a canned speech?
5. Close (15–20 points)
- Did the agent ask for a specific next step (not a vague "would you be interested?")?
- Was the close clear and confident without being pushy?
- Did the agent offer a specific time/date for follow-up rather than "sometime next week"?
6. Professionalism (5–10 points)
- Tone and pace appropriate for the conversation
- No background noise or distractions
- Agent remained professional if the customer was difficult
Total Score Structure
A 100-point scorecard with auto-fail categories works well. Define performance tiers:
- 90–100: Exceeds standard
- 75–89: Meets standard
- 60–74: Needs improvement
- Below 60 or any auto-fail: Immediate coaching required
How Many Calls to Monitor
Sample size is where most QA programs go wrong — either monitoring too few calls to be meaningful, or monitoring so many that the QA team is overwhelmed and the data is not actioned.
Practical guidelines by team size:
| Team Size | Calls per Agent per Week | Total QA Calls per Week |
|---|---|---|
| 5 agents | 5–10 | 25–50 |
| 20 agents | 3–5 | 60–100 |
| 50 agents | 2–3 | 100–150 |
| 100+ agents | 1–2 | 100–200 |
For new agents (first 90 days), double the sample rate. For agents on a performance improvement plan, monitor 10+ calls per week.
Which calls to monitor:
- Random selection: the baseline — catches typical performance
- Short calls: calls under 60 seconds are often early hang-ups and reveal opener problems
- Converted calls: understand what high-performing conversations look like
- Failed close calls: identify the specific objections costing you conversions
Calibration: Making QA Scores Mean the Same Thing
The most common reason QA programs fail: different evaluators score the same call differently. When scores are inconsistent, they stop meaning anything — and agents stop trusting the feedback.
Calibration sessions are the fix. Once a week or every two weeks, the QA team (evaluators + supervisors) listens to the same 3–5 calls independently, scores them, then compares and discusses discrepancies.
The goal is not for everyone to agree on the exact number — it is for the disagreements to get smaller over time, and for the team to articulate shared criteria for judgment calls.
Signs your calibration needs work:
- The same agent scores 85 from one evaluator and 68 from another on similar calls
- Evaluators cannot explain why they gave a specific score
- Agents dispute scores regularly and evaluators cannot defend them with specifics
Giving Feedback That Agents Actually Act On
QA feedback that gets ignored is waste. Here is what makes feedback land:
Be Specific About the Moment
Bad feedback: "Your objection handling needs work." Good feedback: "At the 2:15 mark, when the customer said 'I already have a solution,' you said 'but our product is better.' That triggers resistance. Try: 'That makes sense — can I ask what false positive rate you're seeing?' — it keeps them talking."
Use the Recording
Play the specific moment you are discussing. Agents who hear themselves in context understand the feedback far better than agents who receive a written summary.
Focus on One or Two Things Per Session
Agents who receive a list of ten things to improve fix none of them. Pick the one or two highest-impact items and focus coaching there. Once those are improved, move to the next items.
Connect Behavior to Outcome
Agents change behavior faster when they understand why it matters. "When you start with a long company introduction, answer rates drop by about 30% because it signals a sales call. A 10-second opener gets more people to stay on the line — which is why your conversion rate is lower than agents with similar scripts but faster openers."
Follow Up Within 48 Hours
Feedback that is not reinforced quickly fades. Listen to two calls from the same agent within 48 hours of a coaching session. Did the behavior change? If yes, acknowledge it. If no, address it again before the pattern solidifies.
QA Software: Tools for Monitoring at Scale
For teams monitoring 50+ calls per week, manual QA becomes a bottleneck. Software helps:
Scorecard platforms (Playvox, Evaluagent, Klaus): Centralize scoring, store call recordings alongside scores, generate performance trend reports. Starting around $20–50 per agent per month.
AI-assisted QA (Gong, Chorus, Observe.ai): Automatically transcribes and scores calls based on keyword detection and conversation patterns. Useful for large teams — covers more calls than manual review. Still needs human calibration to ensure accuracy.
VICIdial built-in recording: Call recordings are stored automatically. Pair with a Google Sheet-based scorecard for small teams — it is not elegant, but it works and costs nothing.
The Connection Between AMD Accuracy and QA Scores
Here is something most QA programs never measure: AMD misclassification inflates bad performance metrics in ways that look like agent problems.
When AMD has a high false positive rate — dropping live humans before agents speak — two things happen that distort QA data:
Short calls inflate negatively: Agents show a higher rate of very short calls (under 30 seconds) because AMD-dropped calls sometimes briefly connect to an agent before the system detects the error. These look like failed openings in QA reports — but they were not agent failures.
Agent utilization drops: When agents are not in conversations because AMD is dropping calls, they appear disengaged or slow in performance reports.
Voicemail calls reach agents: AMD false negatives (classifying machines as humans) send voicemail recordings to agents. Agents who hear voicemail and handle it incorrectly — or mute it and wait — show up in call duration and engagement metrics as underperformers.
Before concluding an agent has a performance problem, check whether AMD classification issues are affecting their specific call sample. A sudden drop in one agent's QA scores is sometimes AMD degradation affecting their campaign, not an agent performance change.
For VICIdial operations, replacing the built-in Asterisk AMD with amdify.io reduces these classification errors to 1–3% — cleaning up the noise in your QA data and ensuring agents are evaluated on actual performance, not system errors.
Connecting QA to Revenue
QA programs that exist in isolation from revenue metrics never get the organizational attention they deserve. The most effective QA programs tie scores directly to outcomes.
Track by agent, by week:
- QA score
- Conversion rate
- Average handle time
- Revenue generated
Then answer: does QA score predict conversion rate? It should — if your scorecard measures the right behaviors. If there is no correlation, your scorecard is measuring the wrong things.
Once the correlation is established, the conversation with leadership changes: "Raising average QA score from 74 to 82 across the team is associated with a 12% improvement in conversion rate." That is a business case, not just a compliance report.
Building the Program: First 30 Days
Week 1: Build your scorecard. Get input from top-performing agents on what behaviors drive conversions.
Week 2: Calibrate. Have three evaluators score the same 10 calls. Discuss and resolve scoring differences.
Week 3: Begin monitoring. Score 3–5 calls per agent. Conduct first coaching sessions.
Week 4: Review aggregate data. Which scorecard sections have the lowest scores across the team? That is your first process improvement target.
Month 2 and beyond: Establish weekly rhythm. Calibrate monthly. Update the scorecard quarterly as your script and campaigns evolve.