Dograh

Automated QA for Voice Agents: Grade the Agent, Not the Caller

Automated QA for Voice Agents: Grade the Agent, Not the Caller
TechnicalOctober 6, 2026·10 min read

Automated QA for Voice Agents: Grade the Agent, Not the Caller

Abhishek Kumar
Abhishek Kumar·Co-founder, Dograh AI

Co-founder of Dograh, building the future of open-source voice AI agents. Own your voice AI stack.

Automated quality assurance (QA) for voice agents is an AI review of each call's transcript after the call ends. It flags calls where the agent left its script or misunderstood the caller, and it gives each step of the call a score and a sentiment label. Once live, review a 5-10% sample and read the flagged calls yourself.

This post is part of our guide to testing, monitoring and scaling AI voice agents. The guide covers the whole loop, from testing before launch to scaling after it. This one is about the review that runs after each call: what it checks, and how many calls it needs to see once you are live.

Key Takeaways

  • QA grades the agent's conduct. Its sentiment label is never the customer's verdict.
  • A 5-10% sample finds common failures within days. Go to 100% after big changes.
  • Most QA findings end up as changes to the call flow.

Once a voice agent takes real calls, nobody on the team can listen to all of them. We build Dograh, an open-source voice agent platform, and this is how we use its QA node to decide which calls a person should read.

Video walkthrough: how to set up the QA Analysis node in Dograh

What automated QA checks on each call

Dograh's QA node reads the transcript of a finished call and returns a structured review of how the agent did.

It has no link to the live flow, so it measures the call without changing it. The default review scores each step (node) of the call separately. Each step gets its own tags, a 1-10 score, a sentiment (positive, neutral or negative) and a short summary, and all the tags are also listed together for the call. There is no single score for the whole call. The default tags, and where we look first when one shows up:

TagWhat it flagsCheck first
DEAD_AIRLong silencesThe reasoning delay on that turn
HEARING_ISSUESTrouble hearing each otherSpeech-to-text and the phone audio
UNCLEAR_CONVERSATIONA muddled callThe transcript, turn by turn
ASSISTANT_IN_LOOPThe agent repeating itselfA pathway condition that never fires
ASSISTANT_REPLY_IMPROPERA reply that doesn't fit the momentThe node prompt and any tool response
ASSISTANT_LACKS_EMPATHYA cold reply to an upset callerThe tone rules in the global prompt
USER_FRUSTRATEDThe caller getting annoyedThe turns just before it
USER_NOT_UNDERSTANDINGThe caller not following the agentThe wording of the agent's question
USER_REQUESTING_FEATUREThe caller asking for something the agent can't doA missing flow or tool, a backlog item
USER_DETECTS_AIThe caller realising it is an AILong pauses and the voice choice

There is no separate adherence checker or miscommunication checker. Both come from these tags or from your own QA prompt. Adherence asks whether the agent stayed on the flow and said what it had to. ASSISTANT_REPLY_IMPROPER and ASSISTANT_IN_LOOP catch the obvious slips, and a custom prompt can check rules only your business has, such as reading a booking back before confirming it.

Miscommunication is the agent and the caller talking past each other. USER_NOT_UNDERSTANDING and UNCLEAR_CONVERSATION cover it, and so does a caller who has to correct the same detail twice. QA reads only the transcript, so a bad phone line shows up only through what it causes, such as a caller asking the agent to repeat itself.

Some questions should never go to the QA review. Whether a booking landed is a fact your CRM or the tool's reply already holds, and a model reading the transcript can only guess at it. We treat a successful CRM response as proof the data was written.

Sentiment tells you about the call, not the customer

The sentiment label is the QA review's reading of how a step of the call went, and it should never be reported as how the customer felt.

Tone and satisfaction come apart more often than people expect. A June 2026 preprint on 70,450 support conversations found that tone and customers' own 1-to-5 ratings lined up only weakly (a 0.36 correlation). Read it with care. It studies text chats on a fundraising platform rather than calls, and its author's company sells AI conversation analysis. The paper goes on to propose an AI estimate of satisfaction. We use only its finding that the two line up weakly.

A common view in support analytics treats sentiment as what the customer felt and ties QA scores to CSAT (customer satisfaction). For a voice agent, that mixes up two jobs. The QA review exists to grade the agent. If you want to know how customers felt, ask them and record the answer, as we describe in our post on claims CSAT surveys.

Sentiment still earns its place as a sorting signal. A call with a negative step tagged USER_FRUSTRATED goes to the top of the reading pile. A neutral step with a 9 out of 10 score can wait.

Dograh

Open Source Alternative to Vapi / Retell

Self-hosted voice agent platform — no per-minute fees

dograh-hq/dograh

Star on GitHub

Why a 5-10% sample is enough once you are live

In production we set the QA node's Sample Rate to 5-10%, and at 1,000 calls a day the maths shows it finds common problems in a day and rarer ones within a week.

The case for scoring every call is easy to make, because any call you skip might hide a problem. Every QA review is a large language model (LLM) call you pay for, though, and the real limit is people. Someone has to read each flagged call and decide what to change, and a team that gets hundreds of flags a day reads few of them closely.

Here is the maths for an illustration: an agent taking 1,000 calls a day. These are probabilities, not Dograh figures.

SampleCalls reviewedChance of seeing a problem that hits 1% of callsIf you see none, the true rate is likely below
5%, one day5039%6%
10%, one day10063%3%
5%, one week35097%0.9%
10%, one week70099.9%0.4%

A problem on one call in a hundred has only a 39% chance of showing up in a single day at 5%. Over a week at the same 5%, it rises to 97%. A failure on one call in ten almost certainly shows up on day one at any of these rates. The last column uses the rule of three: if a problem never appears in n calls, its true rate is very likely below 3 divided by n.

These numbers also assume the QA review flags the problem whenever it sees one. It won't always, which is why you check its grades against your own.

That maths assumes the sampled calls are a fair mix. Dograh picks the sampled calls at random, which is what this maths assumes. Still check that the reviewed runs spread across the hours and call types you care about.

Human support teams review far less. In a survey of 500 support agents by Solidroad, a company that sells QA software, 81% said most of their conversations are never reviewed for quality. The report does not say when the survey ran. A steady 5-10% of your voice agent's calls, graded against the same rules each time, already beats that.

Some periods still call for 100%. Keep the Sample Rate at 100% for the first weeks after launch and for a week after any flow change, while you still expect new failures. Calls that must include a required disclosure stay at 100% for as long as the script runs.

Setting up the QA node in Dograh

The QA node sits in the workflow with no edges, and it runs for every eligible call, test calls from the editor included.

Add it to the workflow, leave it unconnected and publish. Its settings decide which calls get reviewed:

SettingDefaultWhat we set
EnabledOnOn
Use Workflow's LLMOnOn grades with your agent's own model; turn it off to use Dograh's default QA model.
Minimum Call Duration (seconds)1515, so hang-ups and wrong numbers are skipped
Include Voicemail CallsOffOff
Sample Rate (%)100100 for the first weeks, then 5-10

QA results show on each run's Run Detail page, under QA Analysis, and in the annotations field of the after-call webhook, so you can send them to any dashboard you already use.

QA also runs on test calls from the editor, so you can see its review before launch. For many scenarios at once, let a testing tool such as Tuner call your agent with generated scenarios. Our guide to testing voice agents before you go live covers that order.

The default prompt is a fair start, and the checks that matter most are the ones only you can write. The prompt below, written for a clinic booking agent, keeps the default tags and adds rules for adherence in read_back_before_booking and transfer_offered_when_asked, a miscommunication check in caller_corrected_detail, and a list of unplanned requests in unplanned_request:

You review transcripts of calls handled by a dental clinic's booking agent.
Grade the agent, never the caller. Return JSON with these fields:

tags: any that apply from DEAD_AIR, USER_FRUSTRATED, ASSISTANT_IN_LOOP,
  ASSISTANT_REPLY_IMPROPER, USER_NOT_UNDERSTANDING, HEARING_ISSUES,
  UNCLEAR_CONVERSATION, USER_REQUESTING_FEATURE, ASSISTANT_LACKS_EMPATHY,
  USER_DETECTS_AI
read_back_before_booking: true if the agent repeated the date and time
  and the caller agreed before the agent confirmed the booking
transfer_offered_when_asked: true, false or "not_asked"
caller_corrected_detail: true if the caller had to correct a name, date
  or number the agent got wrong
unplanned_request: what the caller asked for in a few words, if the flow
  had no answer for it, otherwise null
score: 1-10 for how the agent handled this step
sentiment: positive, neutral or negative, for this step
summary: one or two sentences

What this does: Each run now says whether the agent read the booking back and offered a transfer when asked, and caller_corrected_detail marks calls where agent and caller talked past each other. unplanned_request builds a running list of things callers wanted that the flow never planned for.

A sample booking call transcript with three highlighted moments, each matched to what the QA review returns: caller_corrected_detail, read_back_before_booking and USER_FRUSTRATED with negative sentiment
A missed read-back or a wrongly heard date never shows up as an error in your logs, so it is caught only if a tag or a QA prompt field names it. Use negative sentiment to decide which calls to read first, never as a satisfaction score.

Before you trust the scores, label a few dozen calls yourself and compare. LLM graders are uneven. In the Judge's Verdict study (2025), only half of 54 AI graders agreed with human graders as well as humans agree with each other, on a factual-answer task, not calls. A study of 242 retail and telecom voice-agent calls by Sprinklr, a customer experience (CX) software vendor, found an LLM judge's reliability depends on what it is judging. Safety and recovery checks were the least reliable, and the authors say human review still matters there. Rewrite the prompt until your labels and the grader mostly agree, and check again after every prompt change.

Join the Dograh Community

Dograh is an OSS alternative to Vapi. Join our Slack community for queries, releases, best practices & community interactions.

From flagged calls to flow changes

QA earns its keep after go-live, when real callers do things nobody planned for.

We see it on almost every launch. Callers ask questions the team never thought of, and they talk to an agent differently than they would to a person. A 2026 study from City University of Hong Kong in the International Journal of Human-Computer Interaction found people speak to conversational agents in a more task-focused way, with shorter sentences, than they do with other people. A flow written from experience with human calls misses a lot of that.

Read the flagged calls in the sample, starting with the lowest step scores. For any call where the reason isn't obvious from the transcript, open its trace to see which layer failed, as we walk through in reading call traces. Then change the flow, and put the Sample Rate back to 100% for a week to see whether the fix held.

The tags often point straight at the change. Repeated ASSISTANT_IN_LOOP flags often mean a pathway condition is too vague, so the agent can't tell when to move on. A growing unplanned_request list, or a run of USER_REQUESTING_FEATURE tags, tells you which new branch to build next. This is the same loop we describe in our post on restaurants and home services: QA flags the off-script calls, a person reads them and the agent gets fixed.

If you are setting this up now, start with the default prompt at 100% for your first weeks. Add the checks only your business needs, then drop the Sample Rate to 5-10% once the flags settle down.

Glossary

Sample Rate (%)
The share of eligible calls the QA node reviews. It defaults to 100, and we drop it to 5-10 once an agent is stable.
LLM judge
A language model asked to grade a transcript against written rules. Its grades need checking against human labels before anyone acts on them.
Rule of three
If a problem never appears in n reviewed calls, its true rate is very likely below 3 divided by n, so 350 clean calls put it under about 0.9%.
Default tags
Labels such as DEAD_AIR or USER_NOT_UNDERSTANDING that the QA review attaches to each run, so flagged calls can be found and read later.

Frequently Asked Questions

Get started with Dograh

Build, deploy, and scale AI agents with Dograh. Join the community of developers building the future.