Search the live AIVAX catalog before a workload commits to a provider,
context window or capability set. Every row below comes from the
public catalog during this build.
Available models
153
Provider groups
28
Fetched
Jul 22, 2026
Two ways to run
Integrated models use the AIVAX balance.
BYOK keeps the provider contract and model bill with
you.
Live decision data.
Models are ordered by published release date, newest first. Prices
are USD per million tokens; From marks a context-dependent tier.
Available AIVAX models ordered by release date, with capabilities,
context, token pricing and lifecycle status
@google/gemini-3.5-flash-lite
Gemini 3.5 Flash-Lite is Google's most cost-efficient general availability model, optimized for high-volume agentic tasks, translation, and simple data processing.
Laguna S 2.1 is Poolside's coding agent model for software engineering and long-horizon agentic workflows, with 118B total parameters and 8B active parameters.
Released
New
Capabilities
ThinkingTool Calling
Context window
1M
131.1K max output
Price per 1M
Input $0.1Output $0.2
@meituan/longcat-2.0
LongCat 2.0 is Meituan's sparse mixture-of-experts language model for coding, repository-level changes, long-horizon problem solving, and agentic workflows.
Released
New
Capabilities
ThinkingTool Calling
Context window
1M
262.1K max output
Price per 1M
Input $0.3Output $1.2
@thinkingmachines/inkling
Inkling is Thinking Machines Lab's open-weight multimodal mixture-of-experts model for general-purpose reasoning, coding, agentic workflows, tool use, and retrieval-augmented generation.
Released
New
Capabilities
Audio InputImage InputThinkingTool Calling
Context window
524.3K
524.3K max output
Price per 1M
Input $1Output $4.05
@model-router/kimi:latest
Route between latest Kimi models.
Released
New
Capabilities
Image InputThinkingTool CallingModel Router
Context window
1M
1M max output
Price per 1M
Input $3Output $15
@moonshotai/kimi-k3
Kimi K3 is Moonshot AI's ultra-large-scale, open-weight multimodal reasoning model for complex coding, knowledge work, and long-horizon agentic workflows.
Released
New
Capabilities
Image InputThinkingTool Calling
Context window
1M
1M max output
Price per 1M
Input $3Output $15
@kwaipilot/kat-coder-air-v2.5
KAT-Coder-Air V2.5 is an efficient agentic coding model designed to autonomously locate, modify, and complete end-to-end software tasks.
Released
New
Capabilities
Tool Calling
Context window
256K
80K max output
Price per 1M
Input $0.15Output $0.6
@kwaipilot/kat-coder-pro-v2.5
KAT-Coder-Pro V2.5 is a flagship agentic coding model designed to autonomously locate, modify, and complete end-to-end software tasks.
Grok 4.5 is xAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Released
New
Capabilities
Image InputThinkingTool CallingFile Input
Context window
500K
500K max output
Price per 1M · from
Input $2Output $6
2 context tiers
@aion-labs/aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.
Released
Capabilities
ThinkingTool Calling
Context window
131.1K
32.8K max output
Price per 1M
Input $3Output $6
@aion-labs/aion-3.0-mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models.
Released
Capabilities
ThinkingTool Calling
Context window
131.1K
32.8K max output
Price per 1M
Input $0.7Output $1.4
@tencent/hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent, built for reasoning, agentic workflows, and real-world production use.
Released
Capabilities
ThinkingTool Calling
Context window
268.3K
131.1K max output
Price per 1M
Input $0.14Output $0.58
@poolside/laguna-xs-2.1
Laguna XS 2.1 is Poolside's compact coding agent model in the 33B-A3B category, combining tool calling and reasoning for agentic software engineering tasks.
Released
Capabilities
ThinkingTool Calling
Context window
268.3K
32.8K max output
Price per 1M
Input $0.1Output $0.2
@anthropic/claude-5-sonnet
Claude Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.
Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI for coding, tool use, deep research, and long-horizon agentic workflows.
Released
Capabilities
Image InputThinkingTool Calling
Context window
262.1K
262.1K max output
Price per 1M
Input $0.025Output $0.1
@model-router/glm:latest
Route between latest GLM models.
Released
Capabilities
ThinkingTool CallingModel Router
Context window
1M
131.1K max output
Price per 1M
Input $1.4Output $4.4
@z-ai/glm-5.2
GLM-5.2 is Z.ai's flagship model for long-horizon tasks, with a usable 1M-token context window for project-level engineering context and long-running agent workflows.
Released
Capabilities
ThinkingTool Calling
Context window
1M
131.1K max output
Price per 1M
Input $1.4Output $4.4
@moonshotai/kimi-k2.7-code
Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
262.1K
262.1K max output
Price per 1M
Input $0.95Output $4
@anthropic/claude-fable-5
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1M
131.1K max output
Price per 1M
Input $10Output $50
@nex-agi/nex-n2-pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total, built for coding, tool use, and long-horizon agentic workflows.
Released
Capabilities
Image InputThinkingTool Calling
Context window
262.1K
262.1K max output
Price per 1M
Input $0.25Output $1
@nvidia/nemotron-3-ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture.
Released
Capabilities
ThinkingTool Calling
Context window
1M
16.4K max output
Price per 1M
Input $0.5Output $2.5
@qwen/qwen3.7-plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities.
Released
Capabilities
Image InputTool CallingFile Input
Context window
1M
65.5K max output
Price per 1M · from
Input $0.4Output $1.6
2 context tiers
@minimax/m3
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use.
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window.
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks, and long-horizon autonomous agents.
Released
Capabilities
Tool Calling
Context window
1M
65.5K max output
Price per 1M
Input $2.5Output $7.5
@x-ai/grok-build-0.1
Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding agents, tool use, and multi-step development.
Released
Capabilities
Image InputTool Calling
Context window
256K
256K max output
Price per 1M · from
Input $1Output $2
2 context tiers
@google/gemini-3.5-flash
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.
Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual logic.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1M
131.1K max output
Price per 1M · from
Input $1.25Output $2.5
2 context tiers
@qwen/qwen3.6-27b
Qwen3.6 27B is a dense 27-billion-parameter model from Alibaba Qwen, designed for agentic coding, long-context reasoning, and multimodal workflows.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
268.3K
82.9K max output
Price per 1M
Input $0.32Output $3.2
@deepseek/deepseek-v4-flash
An efficiency-optimized Mixture-of-Experts model from DeepSeek designed for fast inference and high-throughput workloads.
Released
Capabilities
Tool Calling
Context window
1M
384K max output
Price per 1M
Input $0.14Output $0.28
@deepseek/deepseek-v4-pro
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model designed for advanced reasoning, coding, and long-horizon agent workflows with 1.6T total parameters.
Released
Capabilities
Tool Calling
Context window
1M
131.1K max output
Price per 1M
Input $4.45Output $5.5
@model-router/deepseek:latest
Route between latest DeepSeek models.
Released
Capabilities
Tool CallingModel Router
Context window
1M
131.1K max output
Price per 1M
Input $4.45Output $5.5
@openai/gpt-5.5
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning and higher reliability.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1.1M
131.1K max output
Price per 1M · from
Input $5Output $30
2 context tiers
@openai/gpt-5.5-pro
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window.
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks.
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro.
Released
Capabilities
ThinkingTool Calling
Context window
1.1M
131.1K max output
Price per 1M · from
Input $1Output $3
2 context tiers
@moonshotai/kimi-k2.6
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration.
Released
Capabilities
Image InputTool Calling
Context window
134.1K
32.8K max output
Price per 1M
Input $1Output $4
@anthropic/claude-4.7-opus
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1M
131.1K max output
Price per 1M
Input $5Output $25
@z-ai/glm-5.1
GLM-5.1 represents a major advance in coding ability, with especially notable improvements in tackling long-horizon tasks.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $1.4Output $4.4
@moonshotai/kimi-k2.5:fast
Kimi K2.5 turbo is a high-performance variant of the Kimi K2.5 model, optimized for faster processing and enhanced capabilities.
Released
Capabilities
Image InputTool Calling
Context window
134.1K
32.8K max output
Price per 1M
Input $2.5Output $5
@google/gemma-4-26b-a4b-it
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
262.1K
262.1K max output
Price per 1M
Input $0.07Output $0.35
@google/gemma-4-31b-it
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input. Strong on coding, reasoning, and document understanding tasks with multilingual support across 140+ languages.
Released
Capabilities
Image InputThinkingTool Calling
Context window
262.1K
262.1K max output
Price per 1M
Input $0.13Output $0.38
@qwen/qwen3.6-plus
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
1M
65.5K max output
Price per 1M · from
Input $0.5Output $3
2 context tiers
@arcee-ai/trinity-large-thinking
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks.
Released
Capabilities
ThinkingTool Calling
Context window
2M
2M max output
Price per 1M
Input $0.25Output $0.9
@x-ai/grok-4.20-multi-agent
Grok 4.20 Multi-Agent Beta is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows.
Released
Capabilities
Image InputThinking
Context window
2M
2M max output
Price per 1M · from
Input $2Output $6
2 context tiers
@x-ai/grok-4.20-reasoning
Grok 4.20 Beta is xAI's newest flagship model with industry-leading speed and agentic tool calling capabilities.
Released
Capabilities
Image InputThinkingTool Calling
Context window
2M
2M max output
Price per 1M · from
Input $2Output $6
2 context tiers
@reka/reka-edge
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs.
Released
Capabilities
Image InputVideo Input
Context window
16.4K
16.4K max output
Price per 1M
Input $0.1Output $0.1
@openai/gpt-5.4-mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
400K
128K max output
Price per 1M
Input $0.75Output $4.5
@openai/gpt-5.4-nano
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
400K
128K max output
Price per 1M
Input $0.2Output $1.25
@z-ai/glm-5-turbo
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $1.2Output $4
@x-ai/grok-4.20
Grok 4.20 Beta is xAI's newest flagship model with industry-leading speed and agentic tool calling capabilities.
Released
Capabilities
Image InputTool Calling
Context window
2M
2M max output
Price per 1M · from
Input $2Output $6
2 context tiers
@qwen/qwen3.5-9b
Qwen3.5 9B is an efficient multimodal model from the Qwen3.5 family for low-cost reasoning, coding, and visual understanding workloads.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
268.3K
82.9K max output
Price per 1M
Input $0.04Output $0.15
@openai/gpt-5.4
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
Released
Capabilities
Image InputThinkingTool Calling
Context window
1.1M
128K max output
Price per 1M · from
Input $2.5Output $15
2 context tiers
@openai/gpt-5.4-pro
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1M
128K max output
Price per 1M
Input $30Output $180
@inception/mercury-2
The fastest reasoning LLM and Inception most powerful model.
Released
Capabilities
ThinkingTool CallingDiffusion
Context window
131.1K
65.5K max output
Price per 1M
Input $0.25Output $0.75
@openai/gpt-5.3-chat
GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful.
Released
Deprecating
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.75Output $14
@bytedance/seed-2.0-mini
Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
262.1K
134.1K max output
Price per 1M
Input $0.1Output $0.4
@qwen/qwen3.5-27b
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $0.3Output $2.4
@qwen/qwen3.5-flash
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
1M
65.5K max output
Price per 1M
Input $0.1Output $0.4
@google/gemini-3.1-pro
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows.
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1M
131.1K max output
Price per 1M · from
Input $3Output $15
2 context tiers
@qwen/qwen3.5-397b-a17b
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $0.6Output $3.6
@qwen/qwen3.5-plus
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
1M
65.5K max output
Price per 1M
Input $0.5Output $3
@minimax/m2.5
MiniMax-M2.5 is a state-of-the-art large language model built for real-world productivity.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $0.3Output $1.2
@minimax/m2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $0.3Output $1.2
@z-ai/glm-5
GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $1Output $3.2
@anthropic/claude-4.6-opus
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
1M
131.1K max output
Price per 1M · from
Input $5Output $25
2 context tiers
@moonshotai/kimi-k2.5
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm.
Released
Capabilities
Image InputTool Calling
Context window
262.1K
16.4K max output
Price per 1M
Input $0.55Output $3
@z-ai/glm-4.7-flash
As a SOTA 30B-class model, GLM-4.7-Flash provides a new option that balances efficiency and performance.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $0.07Output $0.4
@openai/gpt-5.2-codex
GPT-5.2-Codex is an enhanced version of GPT-5.1-Codex, optimized for software engineering and coding tasks.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.75Output $14
@openai/gpt-5.3-codex
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model. It pairs the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.75Output $14
@minimax/m2.1
MiniMax-M2.1 is a cutting-edge, lightweight large language model designed for coding, agentic workflows, and modern application development.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $0.3Output $1.2
@z-ai/glm-4.7
GLM-4.7 is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.
Released
Capabilities
ThinkingTool Calling
Context window
204.8K
131.1K max output
Price per 1M
Input $0.6Output $2.2
@google/gemini-3-flash
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.
GPT-5.2 Pro is OpenAI's most advanced model, featuring significant upgrades in agentic coding and long-context capabilities compared to GPT-5 Pro.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $21Output $168
@openai/gpt-5.2
GPT-5.2 is the newest frontier-level model in the GPT-5 line, providing enhanced agentic abilities and better long-context performance than GPT-5.1.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.75Output $14
@openai/gpt-5.2-chat
GPT-5.2 Chat (also known as Instant) is the fast and lightweight version of the 5.2 family, built for low-latency chatting while maintaining strong general intelligence.
Released
Deprecating
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.75Output $14
@openai/gpt-5.1-codex-max
GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.25Output $10
@deepseek/v3.2
DeepSeek-V3.2 is a large language model optimized for high computational efficiency and strong tool-use reasoning.
Released
Capabilities
ThinkingTool Calling
Context window
166.9K
16.4K max output
Price per 1M
Input $0.6Output $1.7
@mistral/large-2512
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total).
Released
Capabilities
Tool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $0.5Output $1.5
@anthropic/claude-4.5-opus
Claude Opus 4.5 is Anthropic’s latest reasoning model, developed for advanced software engineering, complex agent workflows, and extended computer tasks.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
200K
65.5K max output
Price per 1M
Input $5Output $25
@openai/gpt-5.1
GPT-5.1 is the newest top-tier model in the GPT-5 series, featuring enhanced general reasoning, better instruction following, and a more natural conversational tone compared to GPT-5.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.25Output $10
@openai/gpt-5.1-chat
GPT-5.1 Chat (also known as Instant) is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.25Output $10
@openai/gpt-5.1-codex
GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.25Output $10
@moonshotai/kimi-k2-thinking
Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.
Released
Capabilities
ThinkingTool Calling
Context window
262.1K
16.4K max output
Price per 1M
Input $0.6Output $2.5
@anthropic/claude-4.5-haiku
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, offering near-frontier intelligence with much lower cost and latency than larger Claude models.
Released
Capabilities
Image InputTool Calling
Context window
200K
65.5K max output
Price per 1M
Input $1Output $5
@model-router/claude:budget
Route between latest budget-tier Claude models.
Released
Capabilities
Image InputTool CallingModel Router
Context window
200K
65.5K max output
Price per 1M
Input $1Output $5
@z-ai/glm-4.6
GLM‑4.6 is a high‑capacity LLM with a 200K‑token context window, strong coding and reasoning abilities, and enhanced tool‑use capabilities.
Released
Deprecating
Capabilities
ThinkingTool Calling
Context window
204.8K
16.4K max output
Price per 1M
Input $0.6Output $2
@anthropic/claude-4.5-sonnet
Claude Sonnet 4.5 is the newest model in the Sonnet series, offering improvements and updates over Sonnet 4.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
200K
65.5K max output
Price per 1M
Input $3Output $15
@openai/gpt-5-codex
GPT-5-Codex is a specialized version of GPT-5 tailored for software engineering and coding tasks.
DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report.
Released
Deprecating
Capabilities
ThinkingTool Calling
Context window
166.9K
16.4K max output
Price per 1M
Input $0.27Output $1
@moonshotai/kimi-k2
Model with 1tri total parameters, 32bi activated parameters, optimized for agentic intelligence.
Released
Capabilities
Tool Calling
Context window
262.1K
16.4K max output
Price per 1M
Input $1Output $3
@openai/gpt-5
OpenAI's newest flagship model for coding, reasoning, and agentic tasks across domains.
Released
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.25Output $10
@openai/gpt-5-chat
GPT-5 snapshot currently used by OpenAI's ChatGPT.
Released
Capabilities
Image InputTool Calling
Context window
400K
128K max output
Price per 1M
Input $1.25Output $10
@openai/gpt-5-mini
GPT-5 mini is a faster, more cost-efficient version of GPT-5.
Released
Capabilities
Image InputTool Calling
Context window
400K
128K max output
Price per 1M
Input $0.25Output $2
@openai/gpt-5-nano
OpenAI's fastest, cheapest version of GPT-5.
Released
Capabilities
Image InputTool Calling
Context window
400K
128K max output
Price per 1M
Input $0.05Output $0.4
@anthropic/claude-4.1-opus
Claude Opus 4.1 is Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
200K
65.5K max output
Price per 1M
Input $15Output $75
@openai/gpt-oss-120b
OpenAI's flagship open source model, built on a Mixture-of-Experts (MoE) architecture with 120 billion parameters and 128 experts.
Released
Capabilities
ThinkingTool Calling
Context window
131.1K
32.8K max output
Price per 1M
Input $0.15Output $0.75
@openai/gpt-oss-20b
OpenAI's flagship open source model, built on a Mixture-of-Experts (MoE) architecture with 20 billion parameters and 128 experts.
Released
Capabilities
ThinkingTool Calling
Context window
131.1K
32.8K max output
Price per 1M
Input $0.1Output $0.5
@google/gemini-2.5-flash-lite
A Gemini 2.5 Flash model optimized for cost efficiency and low latency.
Google's best model in terms of price-performance, offering well-rounded capabilities. 2.5 Flash is best for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases.
The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528.
Released
Capabilities
ThinkingTool Calling
Context window
163.8K
16.4K max output
Price per 1M
Input $0.5Output $2.15
@anthropic/claude-4-sonnet
Anthropic's mid-size model with superior intelligence for high-volume uses in coding, in-depth research, agents, & more.
Released
Capabilities
Image InputThinkingTool CallingFile Input
Context window
200K
65.5K max output
Price per 1M
Input $3Output $15
@openai/o3
A well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks.
Released
Deprecating
Capabilities
Image InputThinkingTool Calling
Context window
200K
100K max output
Price per 1M
Input $2Output $8
@openai/o4-mini
Optimized for fast, effective reasoning with exceptionally efficient performance in coding and visual tasks.
Released
Deprecating
Capabilities
Image InputThinkingTool Calling
Context window
200K
100K max output
Price per 1M
Input $1.1Output $4.4
@openai/gpt-4.1
Versatile, highly intelligent, and top-of-the-line. One of the most capable models currently available.
Released
Deprecating
Capabilities
Image InputTool Calling
Context window
1M
32.8K max output
Price per 1M
Input $2Output $8
@openai/gpt-4.1-mini
Fast and cheap for focused tasks.
Released
Deprecating
Capabilities
Image InputTool Calling
Context window
1M
32.8K max output
Price per 1M
Input $0.4Output $1.6
@openai/gpt-4.1-nano
The fastest and cheapest GPT 4.1 model.
Released
Deprecating
Capabilities
Image InputTool Calling
Context window
1M
32.8K max output
Price per 1M
Input $0.1Output $0.4
@cohere/command-a
Command A is Cohere's most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.
Released
Capabilities
Image InputTool Calling
Context window
262.1K
65.5K max output
Price per 1M
Input $2.5Output $10
@perplexity/sonar-pro
Sonar Pro is an enterprise-grade API from Perplexity, built for advanced, multi-step queries with added extensibility.
Released
Capabilities
Thinking
Context window
131.1K
131.1K max output
Price per 1M
Input $3Output $15
@openai/o3-mini
o3-mini provides high intelligence at the same cost and latency targets of previous versions of o-mini series.
Released
Deprecating
Capabilities
ThinkingTool Calling
Context window
200K
100K max output
Price per 1M
Input $1.1Output $4.4
@perplexity/sonar
Sonar is Perplexity’s lightweight, affordable, and fast question-answering model, now featuring citations and customizable sources.
Released
Capabilities
Text
Context window
131.1K
131.1K max output
Price per 1M
Input $1Output $1
@metaai/llama-3.3-70b
Previous generation model with many parameters and surprisingly fast speed.
Released
Capabilities
Tool Calling
Context window
131.1K
32.8K max output
Price per 1M
Input $0.59Output $0.79
@amazon/nova-lite
A very low cost multimodal model that is lightning fast for processing image, video, and text inputs.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
307.2K
65.5K max output
Price per 1M
Input $0.06Output $0.24
@amazon/nova-pro
A highly capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks.
Released
Capabilities
Image InputVideo InputThinkingTool Calling
Context window
307.2K
65.5K max output
Price per 1M
Input $0.8Output $3.2
@metaai/llama-3.1-8b
Cheap and fast model for less demanding tasks.
Released
Capabilities
Tool Calling
Context window
131.1K
131.1K max output
Price per 1M
Input $0.05Output $0.08
@openai/gpt-4o-mini
Smaller version of 4o, optimized for everyday tasks.
Released
Deprecating
Capabilities
Image InputTool Calling
Context window
128K
16.4K max output
Price per 1M
Input $0.15Output $0.6
@openai/gpt-4o
Dedicated to tasks requiring reasoning for mathematical and logical problem solving.
Released
Deprecating
Capabilities
Image InputTool Calling
Context window
128K
16.4K max output
Price per 1M
Input $2.5Output $10
@amazon/nova-micro
A text-only model that delivers the lowest latency responses at very low cost.
ReleasedNot published
Capabilities
Tool Calling
Context window
131.1K
65.5K max output
Price per 1M
Input $0.04Output $0.14
@minimax/m2
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows.
ReleasedNot published
Capabilities
ThinkingTool Calling
Context window
131.1K
8.2K max output
Price per 1M
Input $0.3Output $1.2
@mistral/nemo-12b-it-2407
12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.
ReleasedNot published
Capabilities
Tool Calling
Context window
131.1K
16.4K max output
Price per 1M
Input $0.02Output $0.04
@model-router/complexity
Model Router: chooses the best models according to the complexity of the conversation.
ReleasedNot published
Capabilities
Model Router
Context window
Not published
Price per 1M
Input $0.5Output $1
@openai/gpt-5.1-codex-mini
GPT-5.1-Codex-Mini is a more compact and faster variant of GPT-5.1-Codex.
ReleasedNot published
Capabilities
Image InputThinkingTool Calling
Context window
400K
128K max output
Price per 1M
Input $0.25Output $2
@qwen/qwen3-32b
32B-parameter LLM with a 131K-token context window, offering advanced chain-of-thought reasoning, seamless tool calling, native JSON outputs, and robust multilingual fluency.
ReleasedNot published
Capabilities
ThinkingTool Calling
Context window
131.1K
41K max output
Price per 1M
Input $0.29Output $0.59
@qwen/qwen3-coder-480b-a35b-it
Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Browser-Use and other foundational coding tasks, achieving results comparable to Claude Sonnet.
ReleasedNot published
Capabilities
Tool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $0.29Output $1.2
@qwen/qwen3-coder-plus
Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming.
ReleasedNot published
Capabilities
Tool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $1Output $5
@qwen/qwen3-next-80b-a3b-it
An 80 B-parameter instruction model with hybrid attention and Mixture‑of‑Experts, optimized for ultra‑long contexts up to 262 k tokens.
ReleasedNot published
Capabilities
Tool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $0.14Output $1.4
@qwen/qwen3-next-80b-a3b-think
A 80 B‑parameter “thinking‑only” model with hybrid attention and high‑sparsity MoE, designed for deep reasoning over ultra‑long contexts.
ReleasedNot published
Capabilities
ThinkingTool Calling
Context window
268.3K
16.4K max output
Price per 1M
Input $0.14Output $1.4
@anthropic/claude-3-haiku
Claude 3 Haiku is Anthropic's fastest model yet, designed for enterprise workloads which often involve longer prompts.
ReleasedNot published
Deprecating
Capabilities
Image InputTool Calling
Context window
200K
65.5K max output
Price per 1M
Input $0.25Output $1.25
No models match this combination of search and filters.
Cached-input and audio-input rates are published only when they apply.
Open the source record before final cost modelling, especially for models
with multiple token thresholds.