OpenAI Evals shutdown: where do multi-turn agent tests go before November 30?
Manually recreate OpenAI Evals datasets and graders in Promptfoo, then choose between its multi-turn tests and AIVAX gateway conversations.
OpenAI Evals becomes read-only on 31 October 2026, and its dashboard and API shut down on 30 November 2026. Move dataset-based evals into a versioned configuration with assertions, following OpenAI's recommended Promptfoo migration. For multi-turn agent tests, decide what must happen across the conversation: Promptfoo supports simulated users, while AIVAX Agentic Tests exercises an AIVAX AI Gateway with a simulated user and an independent judge that tracks the conversation's progress.
The distinction matters when a support assistant answers a policy question correctly but fails to resolve the customer's problem. Keep the policy-answer check as a dataset row. Add a conversation test for the clarification, recommendation, objection, and next step. AIVAX Agentic Tests does not import OpenAI evals or replace single-turn datasets and graders.
What shuts down, and what should you preserve?
OpenAI announced the Evals platform deprecation on 3 June 2026. Graders documented for eval workflows are part of the transition. The two remaining deadlines affect different work:
- Before 31 October: preserve and reconstruct. Inventory the evals you still depend on. Export or otherwise save the datasets, prompts, grader definitions, and results you need to retain, using the access currently available to you. Recreate active tests before existing evals become read-only.
- Before 30 November: remove the runtime dependency. Switch jobs that call the Evals API to the replacement workflow, validate their results, and retain the historical evidence you need before the dashboard and API close.
This checklist does not assume a bulk-export feature. OpenAI's migration guide to Promptfoo explicitly describes manual recreation and says it does not require an OpenAI Evals export feature. A fresh Promptfoo run is separate from your old OpenAI runs; it does not transfer their history.
Which evals belong in code, and which need a conversation?
Start with the failure the test must detect, rather than selecting one tool for every existing eval.
| What you test | Where it belongs | What to carry over or define |
|---|---|---|
| One input and a verifiable output | Dataset plus assertions in Promptfoo, run locally or in CI | Inputs, expected outputs, prompts, provider settings, and scoring rules |
| An answer judged by an LLM | A recreated assertion or metric in Promptfoo | The rubric and representative outputs; validate the new grader before using it to block releases |
| A conversation against an application or custom provider | Promptfoo's simulated-user provider with the target configured explicitly | Scenario instructions, message format, state handling, stopping conditions, and conversation assertions |
| A goal-oriented conversation through an AIVAX AI Gateway | AIVAX Agentic Tests | User goal, judge-only criteria, starting context, turn budget, and the gateway configuration under test |
For the first two rows, OpenAI maps dashboard-managed evaluations to a config-file and CLI/CI workflow, and graders to assertions and metrics. Its guide warns that similarity scores need not be numerically identical across systems. Recreated LLM graders need validation before you trust them for regression decisions. Tool, agent, and custom-provider workflows also need explicit setup; copying test inputs alone does not recreate the application being tested.
For a conversation test, record more than the final answer. A correct recommendation after repeated wrong advice may require a different decision from a correct recommendation after a necessary clarification. Review the transcript and the reason for the verdict, not just a pass label.
Promptfoo already supports multi-turn simulated users
The Promptfoo simulated-user provider is promptfoo:simulated-user. It supports initialMessages, a stateful option, and assertions such as llm-rubric. Its default maxTurns is 10. The conversation stops at that limit, on an error, or when the simulated user emits ###STOP###.
There are integration details to check before choosing it for an existing agent. The target must accept messages in OpenAI chat format. Function-calling agents need their function definitions in the provider configuration. Confirm how your application carries session state, invokes tools, and exposes results to the evaluation.
By default, simulated-user responses come from Promptfoo-hosted conversation models. The documentation describes target execution as local; that does not imply that a configured remote model API becomes offline. Remote generation can be disabled with PROMPTFOO_DISABLE_REMOTE_GENERATION=true. Review the provider configuration and data flow before using private conversations.
If you want your evaluation definitions beside application code and already have a target adapter, that workflow may cover both single-turn and conversation tests. Moving to AIVAX is a separate choice about testing an AIVAX-configured gateway, not a workaround for a missing Promptfoo multi-turn feature.
What changes when the target is an AIVAX AI Gateway?
AIVAX Agentic Tests runs the gateway's model, instructions, tools, skills, knowledge, and inference settings together. The simulated user pursues goal; a separate judge evaluates the conversation against that same goal and optional validation_criteria.
The distinction between those fields is operationally useful. Put the customer's situation and intent in goal. Put acceptance requirements in validation_criteria, which only the judge receives. Otherwise, a simulated customer can end up prompting the assistant to perform the very checks you hoped to test independently.
The judge tracks whether the conversation reaches the goal, remains recoverable, or persistently moves away from it. The default success threshold is base_threshold: 0.9, and the loss threshold is loss_threshold: 0.2; these are scoring settings, not measured accuracy claims. The default turn budget is 10, configurable from 2 to 64. The judge-trajectory explanation covers why a weak turn need not mean the conversation has failed.
Reproduce a support plan-selection scenario
Use the fictional policy from our multi-turn Agentic Tests walkthrough. Configure a test gateway with this policy first; these are not AIVAX's commercial plan rules:
Ask for team size before recommending a plan. For one to five people, recommend Starter; for six or more, recommend Team. Explain the recommendation using the stated team size. The next signup step is to open Settings > Plans and select the recommended plan. Do not claim signup is complete.
In the dashboard, create an Agentic Test, select that gateway, and copy the following goal and validation_criteria into their corresponding fields. Keep the default turn budget and thresholds. This is a settings excerpt, not a complete API request:
{
"goal": "You are choosing a plan for an eight-person team. Explain your team size when asked, ask which plan fits, and continue until you understand the recommendation and how to sign up.",
"validation_criteria": "Validate whether the assistant asks for team size before recommending a plan, recommends Team for eight people, explains why it fits, and gives Settings > Plans > Team as the next signup step. It must not claim signup is complete."
}
Run the test and inspect whether the assistant asks before recommending, connects Team to the stated team size, and gives the signup step without claiming signup has happened. The walkthrough includes a complete direct-API request if you prefer a streamed evaluation. This scenario is illustrative; no pass rate is claimed here.
Connect tools to sandbox services or controlled fixtures before running. Simulated users can still trigger real gateway tools. Verify external side effects, such as an actual signup or ticket creation, against the responsible system rather than treating the assistant's statement as proof.
Use starting messages for a specific point in an existing conversation. If the user and judge need shared reference material, resources accepts up to 16 Text or RemoteResource objects. Those resources do not add knowledge to the gateway under test. Supply its policy through its own instructions, knowledge, or tools, as in the example above.
Keep execution status separate from behavioral results
AIVAX run states are pending, running, succeeded, failed, and cancelled. A run marked succeeded completed execution; its behavioral outcome can still be loss or incomplete. Inspect the evaluation result before treating the test as passed. Execution-failure notifications do not flag those behavioral outcomes merely because they are undesirable.
For recurring checks, persistent tests support cron schedules with a minimum interval of five minutes. Manual, scheduled, and direct-API evaluations share the account's new-run quota. If a test also needs an external validator, consult the validation-hook contract.
Plan evidence retention along with the migration. Succeeded and failed AIVAX runs remain available for one month; cancelled runs for one day. Export JSON for records you need longer. Direct SSE evaluations are ephemeral and do not create dashboard tests or runs, so the calling application must retain any evidence it needs.
FAQ
Can I import my OpenAI evals into AIVAX Agentic Tests?
No. Recreate the relevant conversational scenarios with a gateway, goal, and validation criteria. Keep dataset-based single-turn checks in a dataset-and-assertion workflow. OpenAI's recommended Promptfoo migration is also a manual recreation, not a promise of transferred run history or equivalent grader scores.
Does Promptfoo support multi-turn agent tests?
Yes. Its simulated-user provider supports bounded conversations, starting messages, stateful targets, and assertions. Check the OpenAI-chat message-format requirement, tool configuration, and default hosted simulated-user generation described above before connecting your application.
How much do AIVAX Agentic Tests cost?
A run bills the selected gateway's inference plus simulated-user and judge usage at the selected profile's rates. Conversation length and configuration affect usage. Consult current Agentic Tests pricing rather than assuming a fixed price per test.
What should I migrate first?
Start with evals that currently gate a release or protect a known failure case. Save their inputs, rules, and historical evidence before the October read-only date. Validate the recreated assertions before switching release decisions to them, then add conversation scenarios for failures a single response cannot expose.