A row of weathered wooden stakes with pale bands stands in wet tidal sand at dusk, a lantern hanging from the nearest one, while driftwood and a small rowboat drift along curved currents
All posts

Should you pin an LLM model ID or use an alias? How to handle model deprecations

Pin an exact model ID where a parser, budget or customer promise depends on it; use aliases only with a check that catches retargeting and retirement dates.

Pin an exact, dated model ID when a change in the model could break a parser, exceed a budget or affect a customer commitment, and confirm that the provider treats that ID as a fixed snapshot. Use a moving alias only where you can tolerate a version change and will detect it. Pinning does not avoid migration: it replaces "whenever the alias is retargeted" with "the retirement date on the provider's schedule", which you can plan around. Either way you need an inventory of what you call, a replacement test and a recurring check.

This matters now because several retirements land within weeks of each other. Anthropic retires claude-sonnet-4-5-20250929 on 30 November 2026, and OpenAI removes the dated GPT-5 and o3 snapshots on 11 December 2026.

Which should you use: a pinned model ID or an alias?

The question is who decides when behavior changes. With an alias, the provider (or a router) decides, and nothing in your code changes when it does. With an exact ID, you decide until the retirement date, and after it requests fail. Anthropic's deprecation page puts it plainly: requests to models past the retirement date will fail.

What each way of naming a model asks of you
NamingWho decides when behavior changesWhat you must doFits
Exact or dated IDYou, until the provider's retirement dateTrack the retirement date and run a replacement test before itParsers, budgets, customer-facing promises, regression baselines
Provider aliasThe provider, on its own scheduleDetect retargeting yourself; record what the alias resolved toPrototypes, internal helpers where a different answer costs little
Router alias (AIVAX @model-router/...)The platform, when it changes the alias targetRe-run your checks when the target changesTier-based choices where you accept movement within a tier
Gateway nameYou, by editing the gatewayQualify the new model before changing the gatewayMany clients that should switch together without a redeploy

An alias with a nightly check is a defensible middle path, as long as you are honest about it: the provider still decides, you only hear about it sooner.

What do providers actually do with aliases and snapshots?

The rules differ by provider and change between model generations, so read the provider's page rather than inferring from the name.

  • Anthropic. For the 4.6 generation and later, the dateless ID is the canonical ID and maps to one fixed snapshot; Anthropic says it does not update the weights or configuration of an existing ID. For earlier models, a dateless alias such as claude-sonnet-4-5 points to the most recent dated snapshot of that minor version. Every ID, dated or not, has its own retirement schedule, and partner platforms such as Amazon Bedrock and Google Cloud set their own, so dates can differ there. See Model IDs and versioning and Model deprecations.
  • OpenAI. The December 2026 removals on the deprecations page are listed by dated snapshot, for example gpt-5-2025-08-07 and o3-2025-04-16. Check how the undated name behaves for your account before you rely on it.
  • Mistral. Its model lifecycle page warns that aliases switch to newer models automatically once they reach general availability, and recommends pinning a major.minor version for precise control.

A pinned ID is still subject to the provider's retirement date. Notice periods also differ: Anthropic promises at least 60 days for publicly released models, while OpenAI's page lists gpt-5.4-cyber as deprecated on 11 September 2026 and removed on 1 October, and gives six months for the April 2027 removals.

Which retirements are on the calendar for Q4 and early 2027?

These dates come from the official pages as read on 2026-10-05. Recheck them before you plan around them.

Announced retirements from the providers' own pages, read on 2026-10-05
DateModelListed replacement
2026-11-30Anthropic claude-sonnet-4-5-20250929claude-sonnet-5-5
2026-12-11OpenAI gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07, gpt-5-pro-2025-10-06, o3-2025-04-16, o3-pro-2025-06-10gpt-5.6-sol, gpt-5.6-terra or gpt-5.6-luna, depending on the model
2027-01-06OpenAI tts-1, tts-1-hd and the gpt-4o-mini-tts snapshotsgpt-realtime-2.1-mini
2027-01-20OpenAI gpt-realtime, gpt-audio and their -mini and gpt-4o variantsgpt-realtime-2.1, gpt-realtime-2.1-mini or gpt-audio-1.5
2027-02-26OpenAI whisper-1 and the gpt-4o transcription modelsgpt-transcribe or gpt-live-transcribe
2027-04-01OpenAI gpt-5.3-codex, gpt-5.4-nano, gpt-5.1gpt-6-sol or gpt-6-luna

Two notes. Anthropic lists claude-haiku-4-5-20251001 as active, "not sooner than October 15, 2026", which is a floor and not a retirement notice. And Google's deprecations page, last updated 2026-10-01, lists gemini-2.5-flash, gemini-2.5-flash-lite and gemini-2.5-pro with no shutdown date announced; third-party summaries that give a date for them do not match the official page.

How do AIVAX aliases and gateways change the trade-off?

AIVAX offers three ways to name what runs: an integrated model ID such as @openai/gpt-6-sol, a router alias such as @model-router/openai:mid, and a gateway name. They sit at different points in the table above.

Router aliases move. The alias resolves to a concrete model, and the public changelog records each retarget. @model-router/openai:mid was pointed at GPT-6 Sol on 22 September and at GPT-6.1 Sol on 2 October; the 2 October entry also moved @model-router/claude:mid to Claude Sonnet 5.5. Each entry warns that applications using the aliases may see changes in response quality, latency and cost, and that existing explicit model identifiers remain unchanged. Our routing guide covers when tier aliases are worth that movement.

Explicit IDs hold still, until upstream retires them. The catalog flags some models as deprecated, and in the current implementation AIVAX can email an account that used a flagged model in the previous 30 days. The alert is enabled by default and debounced to once every seven days, but it is not described in the public docs, so check your account's notification settings before relying on it. Treat it as a prompt to look, not as the calendar: keep your own retirement dates from the provider pages above.

A gateway is a place to make the change once. The AI Gateway docs describe a gateway as a persistent configuration that applications call by name, so its behavior can change without redeploying the caller. That makes it a natural switch: clients keep calling the gateway while you qualify and change the model behind it. Two limits apply. We found no documented automatic substitution when a model becomes unavailable, so treat an unavailable model as an error your application must handle. And a gateway only moves the decision to you; you still own the replacement test.

How do you monitor alias targets and deprecation flags?

Keep a lock file of every model name your code or gateways use, plus the concrete model each router alias should resolve to. A CI job compares it with the live catalog. The catalog is served at https://inference.aivax.net/api/v1/information/models.json, and we called it without an API key on 2026-10-05. We did not find this endpoint in the documentation, so treat its shape as unversioned. The script below reports an entry as failing when the fields it reads are missing.

{
  "@model-router/openai:mid": "@openai/gpt-6-sol",
  "@openai/gpt-6.1-sol": null,
  "@openai/o3": null,
  "@acme/retired-model": null
}
import { readFile } from "node:fs/promises";

const lock = JSON.parse(await readFile("models.lock.json", "utf8"));
const response = await fetch("https://inference.aivax.net/api/v1/information/models.json");
const { data } = await response.json();
const catalog = new Map(data.flatMap((group) => group.models).map((model) => [model.name, model]));

let failures = 0;

for (const [name, expectedTarget] of Object.entries(lock)) {
  const model = catalog.get(name);
  const problems = [];

  if (!model) {
    problems.push("not in catalog");
  } else {
    if (model.stability === undefined || model.flags === undefined) problems.push("unexpected catalog shape");
    if ([model.stability].flat().includes("Offline")) problems.push("offline");
    if (model.flags?.isDeprecating) problems.push("flagged as deprecating");
    if ((model.routingModel ?? null) !== expectedTarget) {
      problems.push(`resolves to ${model.routingModel ?? "nothing"}, expected ${expectedTarget ?? "nothing"}`);
    }
  }

  failures += problems.length > 0;
  console.log(`${problems.length ? "FAIL" : "ok  "} ${name}${problems.length ? ` (${problems.join("; ")})` : ""}`);
}

process.exit(failures ? 1 : 0);

Run with bun check-models.js on 2026-10-05, this deliberately stale lock produced:

FAIL @model-router/openai:mid (resolves to @openai/gpt-6.1-sol, expected @openai/gpt-6-sol)
ok   @openai/gpt-6.1-sol
FAIL @openai/o3 (flagged as deprecating)
FAIL @acme/retired-model (not in catalog)

The exit code was 1. The stability field arrives as an array, which is why the script flattens it. Run it nightly and before each deploy, and change the lock deliberately, in a pull request, when you accept a retarget. This lock format holds only model targets or null, so track retirement dates and owners separately, for example in a calendar or an inventory next to each model's owner. A nightly run can still miss a retarget that happens between two runs; it narrows the window and does not close it.

What should a replacement run check?

Freeze a set of real requests and their accepted outputs before you touch anything, and run the candidate model against it. Compare the things that break quietly:

  • Parse and schema failures. If a parser or schema sits behind the call, see where structured-output healing stops before assuming the new model fails the same way.
  • Request parameters. Models differ in what they accept. Anthropic's page says a non-default temperature, top_p or top_k returns a 400 from Claude Opus 4.7 onward, and the AIVAX gateway docs note that some integrated models reject assistant prefill, temperature, stop sequences or reasoning effort.
  • Tool calls and multi-turn behavior. A single-turn comparison can miss a model that stops calling a tool several turns in. Agentic Tests run a simulated user and a judge against an AI Gateway, so create a second gateway with the replacement model and run the same scenarios against both. They complement a single-turn regression set and do not replace one; the Agentic Tests docs list the quotas that apply.
  • Latency and cost per completed task, not per token. A model that needs fewer retries can cost less per task at a higher token price. Our post on judging agent trajectories covers production signals for whole runs.

When the candidate passes, change the gateway or the pinned ID in one place, keep the old value in the lock file's history, and run the check again.

Frequently asked questions

Is a -latest alias ever safe in production?

When a changed answer is cheap, a human reads the output, and you record what the alias resolved to at the time. For a parser, a budget or a customer commitment, pin.

Does pinning protect me from deprecation?

No. A fixed snapshot prevents alias retargeting, but it does not guarantee identical outputs or availability until retirement, and the retirement date becomes the date that matters. Anthropic and OpenAI both publish those dates, and notice periods differ.

What happens in AIVAX when a model I use is deprecated?

The catalog flag changes, and an account that used the model in the last 30 days can receive the email alert described above. We found no documented automatic substitution, so plan to change the model yourself.

Should I put the model name in application code or in a gateway?

Use a gateway when several clients must change together or when you want to switch without redeploying. Use an exact ID in code when one service owns its own behavior and tests. Either way, record the choice in the lock file.

Start with one production call: write down the exact model it uses, who owns it and the date that model retires, and add it to the lock file.