A dense cluster of tiny pink glass seedpods narrows toward three larger coral pods on a rippled sage-green surface.
All posts

Too many MCP tools: how to reduce tool-definition tokens in your agent's context

Measure what MCP tool definitions cost, then cut it with allowlists, deferred tool search, or a shell. What each option needs and what AIVAX supports.

When an MCP client includes every tool definition in the model's request, unused names, descriptions, and JSON Schemas consume input tokens on each call. Reduce that overhead by exposing only the tools a task needs, shortening redundant definitions, deferring discovery through tool search, or moving selected tools behind a shell. Measure task completion, total input tokens, and latency before choosing.

Which option fits depends on where the agent runs. OpenAI and Anthropic offer deferred loading in their own APIs. An AIVAX AI Gateway does not have tool search today; it has a shell option that turns tools into commands, and a hide flag whose behavior is narrower than its name suggests. Anthropic's documentation puts a five-server, 58-tool setup at roughly 55k tokens before the agent does any work; its figures for tool search are vendor measurements, not independent ones.

How many tokens do your MCP tools use?

Start by listing what the server publishes. The script below connects to a Streamable HTTP MCP server, pages through tools/list, and ranks tools by the size of their serialized definition. It targets @modelcontextprotocol/sdk 1.x: install it with npm install @modelcontextprotocol/sdk@1, save the script as tool-audit.mjs, and run node tool-audit.mjs https://your-server.example/mcp. If the server needs authentication, set MCP_AUTHORIZATION to the full Authorization header value, including its scheme.

import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js";

const headers = process.env.MCP_AUTHORIZATION
  ? { Authorization: process.env.MCP_AUTHORIZATION }
  : {};

const client = new Client({ name: "tool-audit", version: "1.0.0" });
await client.connect(
  new StreamableHTTPClientTransport(new URL(process.argv[2]), {
    requestInit: { headers },
  }),
);

const tools = [];
let cursor;
do {
  const page = await client.listTools({ cursor });
  tools.push(...page.tools);
  cursor = page.nextCursor;
} while (cursor);

const rows = tools
  .map((t) => ({
    name: t.name,
    characters: JSON.stringify({
      name: t.name,
      description: t.description,
      inputSchema: t.inputSchema,
    }).length,
  }))
  .sort((a, b) => b.characters - a.characters);

console.table(rows);
console.log(`${rows.length} tools, ${rows.reduce((n, r) => n + r.characters, 0)} characters`);
await client.close();

The script ranks tools by serialized length, which is not a token count. To estimate the incremental input-token overhead, compare the first model request with and without the tool definitions, keeping the model, messages, system instructions, and other settings identical, and compare the input tokens the provider reports (including cached tokens when they are reported separately). Billed cost and multi-turn latency need their own comparison. When conversation logging is enabled, AIVAX conversation records include the tools, their input schemas, and usage, so you can compare before and after without extra instrumentation.

Long descriptions and deeply nested schemas usually dominate the ranking. Trimming them is the cheapest cut, and it needs no platform feature. A proposal to standardize this in the protocol, SEP-1576 on token bloat (schema deduplication, adjustable response detail, embedding-based tool selection), was still labeled SEP-requested and dormant when I checked on 8 October 2026, so do not wait for the spec to solve it.

Which approach fits your agent?

Options for reducing tool-definition tokens, what the model sees, and where each is documented
OptionWhat the model seesTrade-offWhere it is documented
Allowlist tools at the sourceOnly the listed tools of each serverStatic: you decide in advance which tools a task can useOpenAI Responses API allowed_tools. No per-source filter in an AIVAX gateway
Defer loading with tool searchThe search tool plus server labels; full definitions load when the model asksOne extra search step; success depends on finding the right toolAnthropic defer_loading; OpenAI tool_search. Not in an AIVAX gateway
Move tools behind a shellOne shell function; tools are commands with generated --helpThe model must discover commands; output is capped at 4,096 charactersAIVAX shell. Command-line wrappers exist elsewhere but are vendor-specific
Trim descriptions and schemasThe same tools, with smaller definitionsShorter text can lower selection accuracy if it removes the "when not to call" ruleAny server you control
Split by taskA different small tool set per agent or gatewayMore configurations to maintainAny platform

How do OpenAI and Anthropic defer MCP tools?

Both providers keep every definition available to the API while omitting the deferred ones from the model's initial context. The search tool, any non-deferred tools, and the server label or description the model uses to decide what to search remain visible.

  • Anthropic: mark tools, or a whole MCP server through its default configuration, with defer_loading: true and include a tool search tool. The documentation recommends it from about 10 tools or 10k tokens of definitions, and suggests keeping the 3–5 most used tools non-deferred. Anthropic states that tool search typically reduces definition tokens by over 85 percent and, in internal MCP evaluations, raised accuracy from 49% to 74% for Opus 4 and from 79.5% to 88.1% for Opus 4.5 (Advanced tool use). Treat those as the vendor's own results.
  • OpenAI: add {"type": "tool_search"} and set defer_loading: true on the MCP server tool. In the Responses API this requires gpt-5.4 or later (tool search guide). The MCP guide separately documents allowed_tools to import only a subset of a server's tools.

OpenAI's own caveat applies to every option here: compare task completion, input tokens, and latency on representative requests before choosing a default.

What can you do in an AIVAX gateway today?

An AIVAX gateway acts as an MCP client: it lists each configured server's tools and exposes every one to the model as a function named mcp_<source>_<tool> with the server's description and schema (MCP docs). As of 8 October 2026, I found no tool search or deferred-loading option in the gateway code or docs, and an MCP source takes no tool filter: its settings are the name, URL, headers, cache duration, and two switches for server instructions and remote skills. Three things are available.

Move the tools into the shell. With enableBash, the shell can take selected tools out of the model's function list and expose them as commands. The selected tools are replaced by one shell function, while unselected tools stay direct functions. The model runs help to list commands and gets a --help page for each command generated from the tool's JSON Schema.

{
  "mcpSources": [
    {
      "name": "crm",
      "url": "https://mcp.example.com/mcp",
      "headers": { "Authorization": "<TOKEN>" }
    }
  ],
  "enableBash": true,
  "bashOptions": {
    "toolList": ["mcp_crm_*"],
    "toolExclusionMode": "WhiteList"
  }
}

toolList accepts exact names and * or ? wildcards, and in WhiteList mode the listed tools become commands while everything else stays a direct function. Design for the limits: each command is limited to 60 seconds and 4,096 characters of output, and the shell description names only the last ten commands, so the model relies on help and on predictable tool names.

Hide tools with a static allowlist. hideToolsWithoutSkill removes every tool except read_skill and the names in alwaysVisibleTools. The option is designed to reveal tools through skills, but the documentation says the always-visible list is the reliable path, and in my reading of the gateway code a skill's allowedToolsNames does not bring a hidden tool back. Treat the flag as a fixed allowlist, not as progressive disclosure. If you combine it with the shell, put shell in alwaysVisibleTools or the filter will hide that too.

Split the catalog across gateways. A support gateway that connects only the CRM server carries only the CRM tools. This keeps unrelated tools out of the request entirely, at the cost of maintaining several gateway configurations.

None of these options removes the need to assess what a server publishes: descriptions are untrusted input whether they are loaded up front, searched, or printed by --help. MCP is a trust boundary, not just a tool catalog covers that side, and MCP skills explains what read_skill loads and when.

How do you confirm the cut did not break the agent?

Fewer definitions can remove the wrong tool. Anthropic names wrong tool selection and incorrect parameters as the most common failures with large catalogs, and a cut that hides a needed tool produces the same symptoms. Before and after each change, replay the same set of tasks and record three numbers: whether the task finished, input tokens, and latency. For multi-step flows, an AIVAX Agentic Test can run the same simulated conversation against the gateway before and after the change and score the outcome.

FAQ

Does reducing tools lower the cost of every request? It lowers the input tokens that definitions add to each request. If the model needs an extra search or help turn, that turn costs tokens and latency too, so measure the end-to-end task.

Will changing the tool list break prompt caching? It can. OpenAI notes that changing a definition can invalidate a cached prefix, and describes tool search as designed to preserve the model's cache. Keep the tool list stable within a session.

Does AIVAX support tool search? Not as of 8 October 2026. The options are the shell, the fixed allowlist, and splitting by gateway.

How many tools is too many? Anthropic recommends tool search from about 10 tools or 10k tokens of definitions, and says selection accuracy degrades beyond roughly 30–50 tools. Those are vendor guidelines; test your own model and tasks.

What this post did not verify

I read the gateway paths described above in the source and the public docs, but I did not run a gateway to measure tokens saved, so no AIVAX token figures are given. The provider figures are quoted from their documentation and were not reproduced. The measurement script was not executed against a live MCP server; run it against a server you control and treat its output as a ranking by length, not as token counts.