Customer support QA: what it catches, and the gap it can't fill
Quality assurance is one of the most useful things a support team can do — and one of the most misunderstood. Done well, QA catches real problems and turns them into coaching. But teams lean on it to answer a question it structurally cannot: "which of my agents can actually handle the hard stuff?" QA reviews what already happened. It can't tell you what an agent will do with a problem they haven't met yet. Here's how to run QA properly — and where its edge genuinely is.
What QA is genuinely good at
QA inspects handled tickets against a standard, so you can catch drift, coach specifics, and keep quality consistent as the team grows. To get value from it rather than theatre:
- Write a rubric that scores substance, not vibes. Score the things that matter: was the customer's actual problem solved, was the information accurate, was the response clear and actionable, was policy applied correctly, was the tone right for the situation. Vague criteria like "professionalism" produce inconsistent scores and useless feedback.
- Sample deliberately, not randomly. A flat random sample misses the interesting cases. Mix in escalations, reopened tickets, low-CSAT interactions, and a baseline of normal ones. You learn most from the edges.
- Calibrate your reviewers. Have every reviewer score the same five tickets independently, then compare. If they disagree wildly, your rubric is subjective and your scores are noise. Recalibrate until reviewers land within a point of each other. This single step separates real QA from opinion.
- Close the loop with coaching. A QA score with no conversation changes nothing. The score is the start of a coaching moment, not the end of an audit. Agents should see their reviews and what specifically to do differently.
Do those four and QA earns its keep. Skip calibration especially, and you get a number that feels objective and isn't.
The three things QA structurally can't tell you
Even excellent QA has hard limits — not because it's done badly, but because of what it is:
- It's backward-looking. Every QA review is of a ticket that already shipped to a real customer. It's a lagging indicator — the impact already happened. QA can tell you an agent handled something poorly; it can't warn you before they do.
- It's sample-based. You review a fraction of tickets. An agent's reviewed tickets can look fine while their handling of the situations you didn't sample — the rare, hard, unfamiliar ones — stays invisible.
- It scores execution on familiar work, not capability on unfamiliar work. QA mostly reviews routine tickets, because most tickets are routine. It tells you an agent executes the known playbook well. It says almost nothing about whether they can investigate something genuinely new — which is exactly the skill that separates your top agents from the rest.
QA vs capability testing — complementary, not competing
These answer different questions, and a mature team uses both:
- QA asks: "Was this handled well?" Inspection of real output, after the fact, on a sample. Keeps live quality honest.
- Capability testing asks: "Can this agent handle the next unfamiliar one?" Measurement of the underlying skill, before it reaches a customer, on a controlled problem everyone faces identically. Predicts future performance.
QA is your quality control on production. Capability testing is your early-warning system and your development compass — it's the same method you'd use to assess candidates before hiring and to know when a new starter is truly ramped. One inspects the past; the other predicts the future. You want both pointed at the same team.
Where PRISM comes in — the honest bit
Keep running QA — nothing here replaces it, and the four practices above will make it sharper with no new tool. What QA can't give you is the capability layer: a read on whether each agent can investigate an unfamiliar problem, before it shows up as a bad ticket. That's what PRISM measures. Agents work realistic tickets in live systems and the PRISM scoring engine scores their thinking across eight dimensions — the same problem for everyone, scored consistently, with an escalation dependency score per agent. It's the forward-looking complement to your backward-looking QA, and the engine behind ongoing training.
Play one real ticket in a live system and watch the PRISM scoring engine score your thinking across all 8 dimensions — the forward-looking read QA can't give you. 15 minutes, no signup.
Try the live scenario →