Voice Evaluation & Simulation — Test and Score AI Voice Agents Automatically
What is Voice Evaluation?
Langoedge gives every Voice Graph a built-in QA suite with two halves that feed each other:
- Simulation generates the calls. A synthetic AI caller — with a persona and a goal — dials into a live session with your real voice agent, so you can exercise a flow before a single human ever hears it.
- Evaluation grades the calls. LLM-powered "judges" read the resulting transcripts against criteria you define, scoring empathy, accuracy, compliance, and task completion without anyone listening to a recording.
Definition — Voice Evaluation: An automated process where LLM-powered "judges" review voice call transcripts against your defined quality criteria — measuring empathy, accuracy, compliance, and task completion.
The Core Pillars
- **Synthetic Testing:** Launch simulated callers against your agent on demand, before you go live. - **Post-Call Analysis:** Every call — simulated or real — is transcribed and stored for later review. - **LLM-as-Judge:** Automated judges score transcripts against your custom criteria. - **Observability:** Track **Intelligence Score**, **Success Rate**, and **Audit Coverage** over time across agent versions.
Simulated and production calls land in the same session history and are scored by the same evaluators, so a pre-launch test run and a real customer call are directly comparable.
Where to Find It
Open a Voice Graph and switch from Designer to the Evals tab in the main header. That opens the Voice Intelligence Center, which has four sub-tabs:
| Sub-tab | What it's for |
|---|---|
| Dashboard | Aggregate quality metrics across every session. |
| Evals | The session-by-session transcript and audit view; also where you run the judges. |
| Evaluators | Create, auto-generate, and tune your judges. |
| Simulate | Generate personas and scenarios, then launch synthetic test calls. |
The Testing Loop
Both halves bootstrap themselves from the same source: the Instructions (system prompts) on your graph's nodes. You do not have to write personas or rubrics from a blank page.
Part 1 — Simulation: Generating Test Calls
Go to Evals → Simulate.
Step 1: Get your scenarios
The first time you open the tab, Langoedge reads the system prompts from every node in your Voice Graph and generates a set of test cases — typically 3 to 6, chosen based on how complex the agent is, covering main paths, edge cases, and failure scenarios.
Each generated case has:
| Field | Description |
|---|---|
| ID | A short slug, e.g. refund_delay, booking_inquiry. |
| Name | A descriptive title, e.g. "Booking a Demo — Tech Executive". |
| Persona | Who is calling: name, communication style, emotional state, background, technical level, and how they react when interrupted or when the agent makes a mistake. |
| Scenario Goal | The caller's objective, context, constraints, and specific mock data (order IDs, emails, dates) to use so the run is verifiable. |
Scenarios are saved on the Voice Graph, so they persist between sessions.
Step 2: Refine, or write your own
You have three ways to shape a test case:
Edit in place
Regenerate with your own prompt
Write a Custom Case
Step 3: Launch
Click Launch Simulation. Langoedge creates a live session for your Voice Graph and drops a synthetic caller into it, carrying the persona and goal as its instructions. Your real agent answers — same nodes, same tools, same transitions as production.
The simulated caller is built to sound like a person rather than a script: it speaks in short turns, uses fillers and mid-sentence self-corrections, varies its acknowledgements, and spells out numbers and emails the way a caller would over the phone.
While a run is active you can:
- Listen — join the room as a silent spectator and hear the conversation live. Toggle Mute to drop the audio without ending the call.
- Stop — cancel the run immediately.
- Launch another — simulations run in parallel, so you can fire several scenarios at once and let them play out together.
Step 4: Let it finish
A simulation runs for up to three minutes, and ends early when:
- your agent fails to join within the first 15 seconds (usually a deployment or configuration problem), or
- your agent drops out mid-call and does not reconnect within 5 seconds.
When the call ends, the transcript is persisted exactly like a production call — so it shows up in Evals ready to be scored.
Part 2 — Evaluators: Defining Your Judges
Go to Evals → Evaluators.
An evaluator is an LLM instruction set that reads a call transcript and returns a score plus its reasoning.
Auto-generate them from your graph
Click Auto Generate. Langoedge reads the system prompts from every node and writes a set of 3 to 6 rubrics covering core requirements, task success, policy adherence, and conversational flow — each with a name, description, scoring rubric, and a sensible passing threshold. This is the fastest way to get useful coverage on a new agent; treat the output as a starting point and tune from there.
Or configure one by hand
| Setting | Description | Example |
|---|---|---|
| Evaluator Name | The metric this judge owns. Shows up on the dashboard breakdown. | "Instruction Following", "Empathy", "Compliance" |
| Description | Optional note on what the judge measures. | "Checks the agent verified identity before disclosing details." |
| AI Model Vendor | Which model does the judging. | gpt-5-mini |
| Passing Threshold | Minimum score for a "Pass", set with a slider in 5% steps. | 70% |
| System Instructions | The rubric itself. | See the templates below. |
How scoring works
The judge is asked to return a score from 0 to 100 together with free-text reasoning. That score is compared against your threshold to decide Pass or Fail, and displayed as a percentage throughout the UI. Write your rubrics on the 0–100 scale.
Example Prompt:
"Score 90 if the agent successfully booked the appointment, 50 if they tried but failed, and 10 if they were rude or unhelpful."
Part 3 — Running the Judges
Go to Evals → Evals and click Run Intelligence Suite in the header.
This runs every evaluator on the graph against every stored session — simulated and real alike — rather than scoring one call at a time. It is a sweep, not a single-shot action, so expect it to take longer and cost more the deeper your session history goes. The button stays disabled until you have at least one evaluator.
When it finishes, select any session in the sidebar to see:
- The full timestamped conversation thread.
- The Intelligence Audit Report — one row per judge, with a Passed or Failed badge and the score as a percentage.
- Expand any row for that judge's reasoning: why it scored the call the way it did.
Part 4 — Reading the Dashboard
Go to Evals → Dashboard for the aggregate view across all sessions.
| Metric | What It Means |
|---|---|
| Intelligence Score | The average score across every evaluation on every session. |
| Success Rate | The percentage of evaluations that met their evaluator's threshold. |
| Audit Coverage | How many of your sessions have actually been scored — if this is low, your metrics are being drawn from a thin slice of calls. |
| Evaluator Performance | Per-judge average, ranked, so you can see which quality dimension is dragging. |
| Intelligence Health | A rollup verdict: Healthy at 80%+, Optimization Needed at 60–79%, and needs attention below that. |
Key Takeaway: The free-text reasoning is where the value is. It tells you not just that an agent failed, but why — enabling targeted prompt improvements rather than random trial and error. Use the Evaluator Performance ranking to decide which node's instructions to fix first.
Best Practices for Effective Testing
Simulate Before You Ship
Run the generated scenario set after every meaningful change to a node's instructions, and score the results before the change reaches a real caller. This is the cheapest place to catch a broken transition.
Test the Hard Calls
Write Custom Cases for the situations that actually hurt: the caller who interrupts, refuses to give details, changes their mind halfway, or is angry before the agent says a word.
Iterate on Judge Prompts
If a judge is too lenient, add specific "negative examples" to its instructions. If too strict, add positive examples.
Track Over Time
Re-run the suite regularly and watch Intelligence Score and Success Rate move. Keep an eye on Audit Coverage so you know the trend is built on the whole history, not a handful of calls.
Use Multiple Judges
Create separate judges for different concerns: one for empathy, one for task completion, one for compliance. This gives you multidimensional visibility.
Listen to a Few Yourself
Use the **Listen** control on a live simulation now and then. Judges score the transcript, so they cannot hear an agent talking over people, pausing awkwardly, or mangling a phone number.
Recommended Evaluator Templates
Here are proven evaluator prompts you can adapt for your use case (note that they should instruct the judge to score on a scale of 0 to 100):
Task Completion Judge
"Did the agent successfully complete the user's primary request? Score 100 if fully completed, 50 if partially completed, 10 if not attempted or failed."
Empathy & Tone Judge
"Was the agent empathetic, patient, and professional throughout? Score based on language used, acknowledgment of user emotions, and tone consistency. Score 100 for excellent empathy, 50 for neutral, 10 for rude or dismissive."
Compliance Judge
"Did the agent make any false promises, share unverified information, or violate compliance rules? Score 100 for full compliance, 0 for any violation detected."
Handoff Quality Judge
"When the agent transferred the call to a human or another department, was the handoff smooth? Did the agent provide context to the recipient? Score from 0 to 100 based on transition quality."
Instruction Adherence Judge
"Compare the agent's behaviour against its brief: it must verify the caller's date of birth before discussing any appointment. Score 100 if verification happened before any detail was disclosed, 40 if it happened late, 0 if it never happened."
Cost & Performance
The two halves of the suite cost very differently, and it is worth knowing which one you are spending on:
- Evaluations only consume LLM tokens for the judge assessment. Scoring a typical 5-minute transcript costs a few cents. Because Run Intelligence Suite sweeps every evaluator across every session, the cost scales with
evaluators × sessions— prune judges you no longer act on. - Simulations are real voice calls. They consume speech-to-text, LLM, and text-to-speech on both sides of the conversation for up to three minutes per run, plus your agent's own tool calls — which will hit real integrations unless you have pointed them at test credentials.
Neither affects a live agent's performance or its ability to handle production calls.