How to Test Live Voice Agents with Google ADK
Google ADK now supports native live evaluation for voice agents. Learn how to build a test workflow, simulate spoken users, score conversations, inspect audio, and create regression tests.
On this page
A voice agent can sound perfect in a demo and still fail when a caller interrupts it, changes direction, or triggers a tool at the wrong moment. Googleβs Agent Development Kit (ADK) now has native live evaluation, so developers can test spoken, multi-turn conversations with simulated users instead of relying only on manual calls. This tutorial shows how to build a small live workflow, create evaluation cases, score its behavior, run the tests, and inspect the resulting audio and transcripts.
What Google ADKβs live evaluation actually changes
Traditional agent testing often compares a user prompt with an expected text response. That approach becomes weaker for voice systems because a spoken conversation has timing, interruptions, turn-taking, tool calls, and context that a single text answer cannot capture. ADKβs live evaluation adds an audio-based user simulator to the existing evaluation loop, allowing a live agent to receive spoken turns and then be graded across the conversation. Googleβs example uses three agents connected in a workflow: one greets the caller, another verifies a date of birth through a tool, and a third provides appointment information.
The important part is that the simulator does not simply read a fixed transcript aloud. A scenario can give the simulated user a goal and persona, allowing the model to decide how the conversation develops within the boundaries of that scenario. You can also use fixed conversations when you need to reproduce a known incident exactly. This gives developers two useful testing modes: controlled regression tests for known behavior and variable scenarios for discovering weaknesses. The result is closer to testing the behavior of the actual voice agent rather than testing whether it can produce one expected sentence.
Build a small live agent before testing it
For the tutorial, create a simple three-stage workflow rather than starting with a large production application. The first agent confirms the callerβs identity, the second asks for a date of birth and calls a validation tool, and the third provides appointment information only after verification succeeds. This arrangement is useful because it contains both conversation requirements and an action that must happen in the correct order. It also gives the evaluation system something meaningful to check beyond whether the agent sounds natural.
In your ADK project, define the live model and create the agents and workflow in your Python application. A simplified structure looks like this:
from google.adk.agents.llm_agent import Agent
from google.adk.tools.tool_context import ToolContext
from google.adk.workflow import START, Workflow
LIVE_MODEL = "gemini-live-2.5-flash-native-audio"
def validate_date_of_birth(dob: str, tool_context: ToolContext) -> dict:
match = dob == "1985-07-12"
tool_context.state["dob_verified"] = match
return {"match": match}
greeter_agent = Agent(
model=LIVE_MODEL,
name="greeter_agent",
mode="task",
instruction=(
"You are Sam, a friendly care-team assistant. "
"Confirm the caller's name before sharing anything else."
),
)
dob_verifier_agent = Agent(
model=LIVE_MODEL,
name="dob_verifier_agent",
mode="task",
tools=[validate_date_of_birth],
instruction=(
"Ask for the caller's date of birth, read it back to confirm, "
"then call validate_date_of_birth in YYYY-MM-DD format."
),
)
goals_agent = Agent(
model=LIVE_MODEL,
name="goals_agent",
mode="task",
instruction=(
"Identity is verified. Share the upcoming appointment and "
"answer the caller's questions."
),
)
root_agent = Workflow(
name="live_workflow",
edges=[
(START, greeter_agent),
(greeter_agent, dob_verifier_agent),
(dob_verifier_agent, goals_agent),
],
)
The example deliberately keeps the tool simple so the evaluation behavior is easy to understand. In a real application, the validation function could check an authenticated customer record or another protected data source. The key design principle is that the workflow makes the required sequence explicit: identity first, verification second, protected information afterward. That sequence becomes a testable contract rather than something you merely hope the model follows.
Create an evaluation case that describes the userβs goal
Next, create an evaluation set containing the conversations you want the agent to survive. ADK supports scenario-based cases in which you describe what the simulated user wants, as well as fixed conversations containing specific user turns. Scenario testing is particularly useful for voice agents because the simulator can vary how it reaches the same goal. For example, a novice persona can provide only broad information and wait for the agent to ask the right follow-up question.
A scenario can be structured around a simple objective such as confirming identity, providing a date of birth, asking a question about an appointment, and ending the call. The evaluation case should describe the desired behavior rather than dictating every sentence. That distinction matters because a good voice agent may legitimately phrase the same response in many different ways. What you really want to measure is whether it completed the task, respected the required order, and used its tools correctly.
{
"eval_id": "appointment_scenario",
"conversation_scenario": {
"starting_prompt": "Hello?",
"conversation_plan": "You are John Doe. Confirm your name when greeted. When asked for your date of birth, give July 12th, 1985. Listen to the appointment details, ask what you should bring, then end the call.",
"user_persona": "NOVICE"
},
"session_input": {
"app_name": "live_workflow",
"user_id": "test_user_id",
"state": {}
}
}
Keep the first test small. One strong scenario can reveal more than a large collection of predictable conversations because you can study exactly where the workflow breaks. Once that case works consistently, add scenarios for interruptions, incomplete answers, unexpected questions, tool failures, and users who change their minds. Treat every failure you discover as a candidate for a permanent regression test.
Configure the simulated user to speak instead of typing
The live configuration is where the normal text evaluation loop becomes an audio evaluation. ADK separates the model that controls the simulated userβs turn-taking from the model that converts those turns into speech. This means the simulator can decide what the user should say while a text-to-speech model produces the actual audio sent to the live agent. You can also select a voice and language configuration so the same scenario can be tested with different spoken characteristics.
{
"criteria": {
"rubric_based_multi_turn_trajectory_quality_v1": {
"threshold": 0.7,
"judge_model_options": {
"judge_model": "gemini-3.7-flash"
},
"rubrics": [
{
"rubric_id": "verifies_identity_first",
"rubric_content": {
"text_property": "The agent verifies identity before disclosing appointment details."
}
}
]
}
},
"live_model_config": {
"timeout_seconds": 300
},
"user_simulator_config": {
"type": "llm_audio",
"model": "gemini-3.7-flash",
"max_allowed_invocations": 10,
"audio_model": "gemini-3.1-flash-tts-preview",
"audio_model_configuration": {
"response_modalities": ["AUDIO"],
"speech_config": {
"voice_config": {
"prebuilt_voice_config": {
"voice_name": "Kore"
}
},
"language_code": "en-US"
}
}
}
}
The model and audio_model settings serve different purposes, so changing one does not automatically change the other. The invocation limit gives a dynamic scenario a predictable upper bound, preventing a poorly defined conversation from running indefinitely. The rubric threshold of 0.7 in Googleβs example is a demonstration value, not a universal quality requirement. For a high-risk workflow, you should choose thresholds based on the consequences of failure and combine model-based scoring with deterministic checks.
Write rubrics around behavior instead of vague quality
A rubric is most useful when it describes something that can actually be observed in the conversation. βThe agent is helpfulβ is too broad because different evaluators can interpret helpfulness differently. βThe agent confirms the callerβs identity and validates the date of birth before disclosing appointment detailsβ is much stronger because it defines an ordering requirement. You can create additional criteria for task completion, tool usage, disclosure rules, or the quality of the overall conversation.
For example, a voice appointment agent might need to satisfy four separate conditions: it must verify identity, call the validation tool, avoid revealing protected information before verification, and complete the appointment task. Separating those requirements makes failures easier to diagnose. A conversation could receive a good overall score while still making one dangerous tool call too early, so high-risk actions should have their own explicit checks. The goal is not to produce one impressive number; it is to learn exactly what failed and why.
Run the live evaluation from the command line
Once the agent, evaluation set, and configuration are ready, install the evaluation extras required by the project and configure the credentials needed for the Live API and Gemini text-to-speech service. Googleβs example runs the evaluation with the ADK command-line interface. The command points to the live workflow, the evaluation set, and the test configuration file.
uv pip install -e ".[eval]"
uv run adk eval \
contributing/samples/live/live_workflow \
contributing/samples/live/live_workflow/live_workflow.evalset.json \
--config_file_path contributing/samples/live/live_workflow/test_config.json
The important result is not simply whether the command finishes successfully. Look at the evaluation scores, the individual criteria, and the conversation trajectory. If the agent fails to verify identity, you have a workflow or instruction problem; if it verifies identity but calls the wrong tool, you have a tool-use problem; if it performs the correct actions but handles interruptions badly, the weakness is conversational. This separation turns debugging from guesswork into a repeatable process.
Inspect the audio and transcript in ADK Web
Text logs alone can hide problems that are obvious when you hear the interaction. ADK Web can reconstruct the live evaluation into a transcript with the corresponding audio available for individual turns. That lets you inspect not only what the agent said but also how the simulated conversation developed. A transcript might look correct while the audio reveals awkward turn-taking, unnecessary delays, or a response that starts after the simulated user has already moved on.
Use the transcript to locate the exact turn where the behavior changed, then listen to the corresponding audio before editing the agent. This is especially useful for multi-agent workflows because the conversation can cross several internal stages while remaining continuous to the caller. If a handoff loses context, the transcript can show where information disappeared. After fixing the workflow, rerun the same evaluation case rather than replacing it, so the original failure remains part of your regression suite.
Turn one test into a useful voice-agent test suite
Once the basic case passes, expand the suite around realistic failure modes instead of generating dozens of happy-path conversations. Test a caller who interrupts the agent while it is speaking, a caller who gives incomplete information, and a caller who changes an answer after confirmation. Add cases where a required tool fails or returns an unexpected result. For workflows involving sensitive actions, also test that the required steps happen in the correct order.
It is useful to separate conversational tests from operational tests. Rubrics can evaluate whether the agent understood the conversation and completed the intended goal, while explicit tool checks can verify that important actions occurred correctly. You should also vary voices and language settings where those variations are relevant to your users. The resulting suite becomes a release gate: changes to prompts, models, tools, or workflow logic can be evaluated against the same known scenarios before they reach users.
Know what live evaluation cannot prove
Native evaluation removes much of the manual testing work, but it does not make a voice agent automatically reliable. The simulated user and model-based evaluator are themselves model-driven, so their results are evidence rather than mathematical proof. A passing score does not guarantee that a real person with an unusual speaking style, noisy microphone, poor network connection, or unexpected request will behave the same way. High-risk applications should combine automated evaluation with deterministic assertions, human review, and production monitoring.
There is also a difference between synthesized speech and real-world audio. A generated voice can test the agentβs handling of spoken turns, but it cannot reproduce every microphone, background-noise, latency, packet-loss, or device condition. That means live evaluation should sit alongside infrastructure and device testing rather than replace it. The strongest workflow is layered: fast automated tests catch obvious regressions, live audio evaluations test conversational behavior, and human or production testing covers conditions that simulation cannot faithfully reproduce.
Use the evaluation loop before the next model change
The most useful way to adopt Google ADKβs live evaluation is to treat every important voice behavior as a testable contract. Start with one realistic scenario, define the actions and ordering that must be correct, run it with a simulated audio user, inspect both the score and the recording, then preserve every meaningful failure as a regression case. As the agent changes, the same suite tells you whether the change actually improved the system or merely made one demo sound better. That is the point where voice-agent development moves from repeated demonstrations to measurable engineering.
Written by


