A tactile tabletop evaluation with three parallel conversation tracks converging on a final scored outcome
All posts

How to test multi-turn AI agent conversations with a simulated user and an LLM judge

Define an observable goal, simulate a user, and judge the full conversation. Inspect weak turns, recovery, outcomes, and costs with AIVAX Agentic Tests.

To test multi-turn AI agent conversations, define an observable goal, let a simulated user pursue it, and use a separate LLM judge to evaluate the full conversation against explicit criteria. AIVAX Agentic Tests runs this loop through your AI Gateway and records the evidence needed to distinguish a weak turn that recovers from a conversation that fails its goal.

A support agent can give a polished first answer and still skip a required question, lose an objection two turns later, or end without the next step the user needed. Test the sequence, not just the opening response. The practical workflow is to define the outcome, run a bounded conversation, inspect the judge's reasoning alongside the transcript, and turn confirmed failures into regression scenarios.

How do the simulated user, gateway, and judge work together?

An Agentic Test coordinates three roles:

  1. The simulated user receives the goal and generates the next message from the conversation so far.
  2. The AI Gateway responds with its configured model, instructions, tools, skills, knowledge, and inference settings.
  3. The LLM judge evaluates the conversation history against the goal and optional judge-only validation_criteria.

The simulator does not grade the assistant. The judge does not write the assistant's answer. Keeping the criteria separate avoids giving the simulated user a checklist to coax out of the gateway. They are not supplied to the gateway as test instructions either.

The judge considers the full recorded history while giving more weight to the latest user–assistant cycle when assessing the current trajectory. Its output includes reasoning, a score, a state, and cumulative trajectory measurements. This is evidence-based evaluation, not proof of everything that happened outside the conversation: verify an actual payment, database write, or ticket creation against the system responsible for it.

Define an observable outcome before writing the test

Consider a plan-recommendation flow. The behavior to observe is:

Identify the customer's team size, recommend the correct plan, explain why it fits, and provide the next signup step.

Write goal from the simulated user's perspective, rather than turning that behavior into instructions for the assistant:

You are choosing a plan for your team. Explain your team's size and needs when asked, ask which plan fits, and continue until you understand the recommendation and how to sign up.

Keep acceptance requirements in validation_criteria:

The recommendation must name the selected plan and connect it to the stated team size. The final response must include a direct signup step.

“Be helpful” does not identify a failure you can investigate. “Ask for team size before recommending a plan” does. Supply the policy needed to determine which recommendation is correct; do not ask the judge to invent your product rules.

Use start to establish an objection, a prior assistant response, or another specific point in a flow. Use external_user_id when the gateway needs an application identity. Keep related requirements together when one user can encounter them naturally; split incompatible user roles, objectives, or opening contexts into separate tests.

For another workflow to exercise, see support triage with typed decisions. Test the customer-facing conversation separately from deterministic checks on routing fields and external side effects.

Run a plan-recommendation test with a fixed policy

Use a test gateway with a funded AIVAX account and a private API key. If tools can change external systems, connect sandbox services or controlled fixtures before running the evaluation: a simulated user does not make real tool calls harmless.

For this tutorial, configure the test gateway with this fictional policy, not AIVAX's commercial plans:

Ask for team size before recommending a plan. For one to five people, recommend Starter; for six or more, recommend Team. Explain the recommendation using the stated team size. The next signup step is to open Settings > Plans and select the recommended plan. Do not claim signup is complete.

The fictional customer has eight people, so the expected recommendation is Team. Save this as agentic-test.json, replacing <AI_GATEWAY_SLUG_OR_ID> with the identifier of that configured gateway:

{
  "model": "<AI_GATEWAY_SLUG_OR_ID>",
  "goal": "You are choosing a plan for an eight-person team. Explain your team size when asked, ask which plan fits, and continue until you understand the recommendation and how to sign up.",
  "validation_criteria": "Validate whether the assistant asks for team size before recommending a plan, recommends Team for eight people, explains why it fits, and gives Settings > Plans > Team as the next signup step. It must not claim signup is complete.",
  "start": [],
  "max_turns": 10,
  "minimum_turns": 1,
  "allow_user_exit": false,
  "judge_start_turn": 1,
  "loss_threshold": 0.2,
  "base_threshold": 0.9,
  "profile": "medium"
}

On a trusted backend or local development machine, set AIVAX_API_KEY securely to your private key. Run this Bash request from the directory containing the JSON file:

curl --no-buffer --fail-with-body \
  'https://inference.aivax.net/api/v1/generations/agentic-tests' \
  -H "Authorization: Bearer ${AIVAX_API_KEY}" \
  -H 'Content-Type: application/json' \
  -H 'Accept: text/event-stream' \
  --data-binary @agentic-test.json

This direct endpoint streams an ephemeral evaluation: it does not create a test definition or run in the dashboard. The request is illustrative, not a measured result or a guarantee that any particular gateway will pass. Never put the private key in browser code or a distributed application bundle.

For reusable dashboard history, create a persistent test instead. The management API uses POST /api/v1/agentic-tests, with name and gateway in place of the direct endpoint's model; queue an execution with POST /api/v1/agentic-tests/<TEST_ID>/runs. See the Agentic Tests documentation for the complete contract and dashboard workflow.

Read the stream before deciding whether the test passed

Each SSE message contains a timestamp and an event with type and data. Route by event.type, not by arrival position:

  • chat.user_message.content and chat.assistant_message.content carry content chunks; concatenate them in order.
  • chat.judge.turn_analysis_result_ready carries the judge's reasoning, score, state, and trajectory values under data.result.
  • usage_updated reports inference usage by role.
  • unhandled_error identifies the failing scope and whether a retry will occur.
  • chat.validation.end carries the terminal behavioral outcome.

Check the HTTP status before parsing SSE, and handle errors after streaming starts. A closed connection without a terminal result is not evidence of success. For a release gate, explicitly require outcome: "success"; neither HTTP 200 nor an intermediate pass: true is enough.

How can a weak turn recover without hiding persistent failure?

Agentic Tests tracks the current judge score and cumulative trajectory. A weak turn can move the conversation to at_risk without ending it. Persistent loss requires both a cumulative trajectory at or below loss_threshold and the required consecutive low-scoring evaluations. A later stronger turn can break the low-score streak.

The judge's states are active, at_risk, success, and loss. Final outcomes have separate meanings:

  • success — the judge score reaches base_threshold.
  • loss — the score streak and cumulative trajectory establish persistent loss.
  • incomplete — the turn budget ends, or the simulated user exits, without the judge establishing success or persistent loss.
  • interrupted — a persisted evaluation stops because of a validation hook or an execution failure before a result is available. Inspect reason and the run's execution state; this is not a verdict assigned by the judge.

A permitted simulated-user exit still triggers judging; it does not automatically pass the test. The request above disables that exit option so the judge or turn budget determines when the conversation ends.

For example, an assistant might initially misunderstand the team size, then correct the recommendation after clarification. Inspect whether the correction satisfies the criteria; do not assume every low score should terminate the run. Conversely, repeated wrong recommendations without progress should not be excused merely because the last answer sounds fluent. This is an illustrative recovery pattern, not a recorded benchmark.

Choose the evaluation boundary deliberately:

  • max_turns accepts 2–64 and bounds simulated-user turns; it defaults to 10.
  • judge_start_turn defaults to 1. Delaying it can avoid judging early discovery, but the final turn and a simulated-user exit still trigger evaluation.
  • minimum_turns defaults to 1 and controls when the simulator may exit, not when the judge may stop. It has no exit effect when allow_user_exit is false.
  • Both turn settings must be at least 1 and less than max_turns.
  • loss_threshold and base_threshold each accept 0.01–0.99, with loss below base and a gap greater than 0.1. The defaults are 0.2 and 0.9; they are configuration defaults, not universal quality standards.

Longer conversations allow more recovery but consume more inference. Delaying evaluation changes which intermediate turns receive scores. Raising a success threshold makes the score requirement stricter; it does not fix an ambiguous rubric or an incorrect judge.

Inspect the evidence and turn failures into regression tests

For persisted tests, each execution creates its own run. Editing the definition does not replace earlier run history. The inspector retains chronological messages, timestamps, per-message usage, judge opinions and trajectory values, final results, errors, and total charged cost; export the run as JSON when you need offline review.

Start with four questions:

  1. Where did progress change? Compare the first at_risk judgment with the conversation immediately before it.
  2. Did the agent recover? Read the correction and the later judge reasoning, not only the final label.
  3. What evidence supports the judgment? Match unmet criteria to messages and any recorded tool evidence. Check external effects separately.
  4. Can the failure be reproduced? Preserve the relevant goal, opening context, policy, and criteria; rerun after changing the responsible gateway configuration.

Do not interchange score fields: a turn's judge score and the cumulative conversation_delta describe different things. The persisted run summary's score represents the final cumulative trajectory. An intermediate pass: true means persistent loss has not been established, not that the goal is complete.

The related post on using judge trajectories in production review offers a retrospective review workflow. Use production evidence to choose new scenarios; do not treat a passing simulation as proof of production reliability.

Separate execution status from behavioral outcome

Persisted runs move through pending, running, succeeded, failed, or cancelled. succeeded means execution completed without an operational error, even when result.outcome is loss or incomplete. Inspect both fields. A validation-hook interruption preserves outcome: interrupted and stores the run as failed; hooks are supported only for persisted tests, not the direct endpoint above.

Persistent tests can use a five-field cron schedule with a minimum interval of five minutes. Failure notifications count consecutive execution failures in state failed, not behavioral loss outcomes. Recovery notifications report recovered execution, not proof that the assistant passed. If you need alerts on behavioral regressions, inspect result.outcome in your own monitoring policy.

Succeeded and failed runs have a 30-day retention window; cancelled runs have a one-day window. Export records needed longer. Account quotas and concurrency also apply; see Agentic Test rate limits rather than assuming a schedule guarantees immediate execution.

Frequently asked questions

Why isn't single-turn evaluation enough?

It checks one response but cannot by itself establish whether the agent remembers earlier facts, asks a prerequisite question before acting, handles a clarification, or reaches a goal several exchanges later. Keep single-turn checks for local requirements; use a multi-turn scenario for ordering, recovery, and completion across the conversation.

How many conversations do I need?

AIVAX's contract defines one bounded conversation per run, not a universal statistically sufficient sample size. Start with distinct, consequential scenarios and repeat them to inspect variation. Choose further coverage from observed failures and the risk of missing a regression. Do not confuse max_turns, which limits one conversation, with a count of independent test conversations.

Can the judge be wrong, and how do I calibrate it?

Yes. A separate judge is not an infallible one. Have domain reviewers label representative transcripts, compare their conclusions with the judge's reasoning, and refine ambiguous criteria before trusting score trends. Include clear failures, clear successes, and borderline recoveries. Langfuse's evaluation guidance likewise recommends calibration against human preferences and warns that judges can share a model's blind spots. This is a review practice, not a claim that AIVAX automatically calibrates your tests.

How is this different from tracing?

A trace records execution evidence; an evaluation applies criteria to decide whether behavior met a goal. Agentic Tests also generates a simulated conversation, rather than only observing existing traffic. Use both: the judgment helps locate a failure, while the transcript and execution evidence help explain it. Neither an assistant's claim nor a judge's score substitutes for checking a critical external side effect.

Does it cost credits?

Yes. Agentic Tests charges the selected gateway's inference plus simulated-user and judge usage. The low, medium, and high profiles select simulator/judge tiers; they do not replace the gateway's configured model. Cost depends on actual usage, not a fixed fee per passing conversation, and stopping or failing a run does not undo usage already charged. Inspect run cost and consult current Agentic Tests pricing before scaling repeated evaluations.

Next steps

Explore AIVAX Testing or open AIVAX Console to create a focused test. Keep the Agentic Tests reference alongside your implementation for request settings and event details.