← Back to all guides

Voice Evaluation & Simulation — Test and Score AI Voice Agents Automatically

Langoedge Team13 min read

What is Voice Evaluation?

Langoedge gives every Voice Graph a built-in QA suite with two halves that feed each other:

  • Simulation generates the calls. A synthetic AI caller — with a persona and a goal — dials into a live session with your real voice agent, so you can exercise a flow before a single human ever hears it.
  • Evaluation grades the calls. LLM-powered "judges" read the resulting transcripts against criteria you define, scoring empathy, accuracy, compliance, and task completion without anyone listening to a recording.

Definition — Voice Evaluation: An automated process where LLM-powered "judges" review voice call transcripts against your defined quality criteria — measuring empathy, accuracy, compliance, and task completion.

The Core Pillars

- **Synthetic Testing:** Launch simulated callers against your agent on demand, before you go live. - **Post-Call Analysis:** Every call — simulated or real — is transcribed and stored for later review. - **LLM-as-Judge:** Automated judges score transcripts against your custom criteria. - **Observability:** Track **Intelligence Score**, **Success Rate**, and **Audit Coverage** over time across agent versions.

Simulated and production calls land in the same session history and are scored by the same evaluators, so a pre-launch test run and a real customer call are directly comparable.


Where to Find It

Open a Voice Graph and switch from Designer to the Evals tab in the main header. That opens the Voice Intelligence Center, which has four sub-tabs:

Sub-tab What it's for
Dashboard Aggregate quality metrics across every session.
Evals The session-by-session transcript and audit view; also where you run the judges.
Evaluators Create, auto-generate, and tune your judges.
Simulate Generate personas and scenarios, then launch synthetic test calls.

The Testing Loop

flowchart LR A[Node system prompts] --> B[Auto-generate scenarios] A --> C[Auto-generate evaluators] B --> D[Launch simulation] D --> E[Transcript saved to session history] F[Real production calls] --> E E --> G[Run Intelligence Suite] C --> G G --> H[Scores and judge reasoning] H --> I[Tune node prompts] I --> A

Both halves bootstrap themselves from the same source: the Instructions (system prompts) on your graph's nodes. You do not have to write personas or rubrics from a blank page.


Part 1 — Simulation: Generating Test Calls

Go to Evals → Simulate.

Step 1: Get your scenarios

The first time you open the tab, Langoedge reads the system prompts from every node in your Voice Graph and generates a set of test cases — typically 3 to 6, chosen based on how complex the agent is, covering main paths, edge cases, and failure scenarios.

Each generated case has:

Field Description
ID A short slug, e.g. refund_delay, booking_inquiry.
Name A descriptive title, e.g. "Booking a Demo — Tech Executive".
Persona Who is calling: name, communication style, emotional state, background, technical level, and how they react when interrupted or when the agent makes a mistake.
Scenario Goal The caller's objective, context, constraints, and specific mock data (order IDs, emails, dates) to use so the run is verifiable.

Scenarios are saved on the Voice Graph, so they persist between sessions.

Step 2: Refine, or write your own

You have three ways to shape a test case:

1

Edit in place

Select a scenario card and click **Edit** on either **Persona description** or **Scenario Goal**. Click **Save** to persist your changes back to the graph.
2

Regenerate with your own prompt

Open the scenario-prompt dialog to regenerate the whole set. You can start from the default generation prompt and modify it — for example, to change the mix of difficulty, to specify the demographics and speaking styles you want represented, or to focus entirely on compliance edge cases. Include the `{nodes_text}` placeholder if you want to control exactly where your agent's node instructions get injected; otherwise they are appended automatically.
3

Write a Custom Case

Select the **Custom Case** card and type a persona and a scenario goal by hand. Both fields are required before the run will start.
The stock generation prompt is deliberately opinionated, and it will not cover every population your agent has to serve. If accent, gender, age, or pacing diversity matters for your vertical, override it with your own prompt in the regenerate dialog rather than relying on the default set.

Step 3: Launch

Click Launch Simulation. Langoedge creates a live session for your Voice Graph and drops a synthetic caller into it, carrying the persona and goal as its instructions. Your real agent answers — same nodes, same tools, same transitions as production.

The simulated caller is built to sound like a person rather than a script: it speaks in short turns, uses fillers and mid-sentence self-corrections, varies its acknowledgements, and spells out numbers and emails the way a caller would over the phone.

While a run is active you can:

  • Listen — join the room as a silent spectator and hear the conversation live. Toggle Mute to drop the audio without ending the call.
  • Stop — cancel the run immediately.
  • Launch another — simulations run in parallel, so you can fire several scenarios at once and let them play out together.

Step 4: Let it finish

A simulation runs for up to three minutes, and ends early when:

  • your agent fails to join within the first 15 seconds (usually a deployment or configuration problem), or
  • your agent drops out mid-call and does not reconnect within 5 seconds.

When the call ends, the transcript is persisted exactly like a production call — so it shows up in Evals ready to be scored.


Part 2 — Evaluators: Defining Your Judges

Go to Evals → Evaluators.

An evaluator is an LLM instruction set that reads a call transcript and returns a score plus its reasoning.

Auto-generate them from your graph

Click Auto Generate. Langoedge reads the system prompts from every node and writes a set of 3 to 6 rubrics covering core requirements, task success, policy adherence, and conversational flow — each with a name, description, scoring rubric, and a sensible passing threshold. This is the fastest way to get useful coverage on a new agent; treat the output as a starting point and tune from there.

Or configure one by hand

Setting Description Example
Evaluator Name The metric this judge owns. Shows up on the dashboard breakdown. "Instruction Following", "Empathy", "Compliance"
Description Optional note on what the judge measures. "Checks the agent verified identity before disclosing details."
AI Model Vendor Which model does the judging. gpt-5-mini
Passing Threshold Minimum score for a "Pass", set with a slider in 5% steps. 70%
System Instructions The rubric itself. See the templates below.

How scoring works

The judge is asked to return a score from 0 to 100 together with free-text reasoning. That score is compared against your threshold to decide Pass or Fail, and displayed as a percentage throughout the UI. Write your rubrics on the 0–100 scale.

Example Prompt:

"Score 90 if the agent successfully booked the appointment, 50 if they tried but failed, and 10 if they were rude or unhelpful."

Do not include a `{transcript}` placeholder in your prompt. The engine appends the full conversation transcript to the end of your instructions automatically at run time.

Part 3 — Running the Judges

Go to Evals → Evals and click Run Intelligence Suite in the header.

This runs every evaluator on the graph against every stored session — simulated and real alike — rather than scoring one call at a time. It is a sweep, not a single-shot action, so expect it to take longer and cost more the deeper your session history goes. The button stays disabled until you have at least one evaluator.

When it finishes, select any session in the sidebar to see:

  1. The full timestamped conversation thread.
  2. The Intelligence Audit Report — one row per judge, with a Passed or Failed badge and the score as a percentage.
  3. Expand any row for that judge's reasoning: why it scored the call the way it did.

Part 4 — Reading the Dashboard

Go to Evals → Dashboard for the aggregate view across all sessions.

Metric What It Means
Intelligence Score The average score across every evaluation on every session.
Success Rate The percentage of evaluations that met their evaluator's threshold.
Audit Coverage How many of your sessions have actually been scored — if this is low, your metrics are being drawn from a thin slice of calls.
Evaluator Performance Per-judge average, ranked, so you can see which quality dimension is dragging.
Intelligence Health A rollup verdict: Healthy at 80%+, Optimization Needed at 60–79%, and needs attention below that.

Key Takeaway: The free-text reasoning is where the value is. It tells you not just that an agent failed, but why — enabling targeted prompt improvements rather than random trial and error. Use the Evaluator Performance ranking to decide which node's instructions to fix first.


Best Practices for Effective Testing

Simulate Before You Ship

Run the generated scenario set after every meaningful change to a node's instructions, and score the results before the change reaches a real caller. This is the cheapest place to catch a broken transition.

Test the Hard Calls

Write Custom Cases for the situations that actually hurt: the caller who interrupts, refuses to give details, changes their mind halfway, or is angry before the agent says a word.

Iterate on Judge Prompts

If a judge is too lenient, add specific "negative examples" to its instructions. If too strict, add positive examples.

Track Over Time

Re-run the suite regularly and watch Intelligence Score and Success Rate move. Keep an eye on Audit Coverage so you know the trend is built on the whole history, not a handful of calls.

Use Multiple Judges

Create separate judges for different concerns: one for empathy, one for task completion, one for compliance. This gives you multidimensional visibility.

Listen to a Few Yourself

Use the **Listen** control on a live simulation now and then. Judges score the transcript, so they cannot hear an agent talking over people, pausing awkwardly, or mangling a phone number.


Here are proven evaluator prompts you can adapt for your use case (note that they should instruct the judge to score on a scale of 0 to 100):

Task Completion Judge

"Did the agent successfully complete the user's primary request? Score 100 if fully completed, 50 if partially completed, 10 if not attempted or failed."

Empathy & Tone Judge

"Was the agent empathetic, patient, and professional throughout? Score based on language used, acknowledgment of user emotions, and tone consistency. Score 100 for excellent empathy, 50 for neutral, 10 for rude or dismissive."

Compliance Judge

"Did the agent make any false promises, share unverified information, or violate compliance rules? Score 100 for full compliance, 0 for any violation detected."

Handoff Quality Judge

"When the agent transferred the call to a human or another department, was the handoff smooth? Did the agent provide context to the recipient? Score from 0 to 100 based on transition quality."

Instruction Adherence Judge

"Compare the agent's behaviour against its brief: it must verify the caller's date of birth before discussing any appointment. Score 100 if verification happened before any detail was disclosed, 40 if it happened late, 0 if it never happened."


Cost & Performance

The two halves of the suite cost very differently, and it is worth knowing which one you are spending on:

  • Evaluations only consume LLM tokens for the judge assessment. Scoring a typical 5-minute transcript costs a few cents. Because Run Intelligence Suite sweeps every evaluator across every session, the cost scales with evaluators × sessions — prune judges you no longer act on.
  • Simulations are real voice calls. They consume speech-to-text, LLM, and text-to-speech on both sides of the conversation for up to three minutes per run, plus your agent's own tool calls — which will hit real integrations unless you have pointed them at test credentials.

Neither affects a live agent's performance or its ability to handle production calls.


Frequently Asked Questions

Are simulations still supported?
Yes. Synthetic simulation is a supported, first-class part of the Voice Intelligence Center, available under the **Simulate** tab. Earlier documentation described it as deprecated; that is no longer accurate.
Does a simulated call get scored like a real one?
Yes. A simulation runs against your live agent and its transcript is persisted to session history exactly like a production call, so **Run Intelligence Suite** picks it up alongside everything else. That also means simulated calls count toward your dashboard metrics — keep that in mind when reading Intelligence Score for a production agent.
Will a simulation trigger my real integrations?
Yes. The agent under test is the real one, with its real tools. If a node books into Cliniko, writes to a CRM, or sends an SMS, a simulated call will do that for real. Point your tools at sandbox or test credentials before running a scenario that writes data.
How long does a simulation run?
Up to three minutes. It ends earlier if the agent never joins (15-second grace period) or disconnects mid-call without reconnecting within 5 seconds. You can also stop a run manually at any time.
Can I run more than one simulation at once?
Yes. Each launch opens its own session, and they appear as separate active runs you can listen to or stop individually.
Can I hear a simulation while it happens?
Yes. Click **Listen** on an active run to join the room as a silent spectator. You will not be heard by either party.
Can I create custom evaluation criteria?
Yes. Each evaluator is defined by a prompt where you control the scoring criteria. You can create unlimited evaluators for any metric — empathy, accuracy, compliance, brand voice, etc. **Auto Generate** gives you a first set derived from your own node instructions.
Do I need to include a transcript placeholder in my prompt?
No. The evaluation engine automatically appends the conversation transcript to the end of your evaluator prompt at runtime. You do not need to include any `{transcript}` placeholders in your custom prompt.
What scale should my judge score on?
0 to 100. The judge is instructed to return a score in that range; your evaluator's passing threshold is set as a percentage on the slider, and the two are compared for you.
How often should I run evaluations?
Simulate and score after any significant change to your voice agent's prompts, model selection, or node configuration — and sweep production sessions weekly to catch drift.

LT

Langoedge Team

The Langoedge engineering team builds AI agent infrastructure that empowers businesses to deploy reliable, observable AI staff. Follow Langoedge Team on LinkedIn for product updates and architectural deep dives.