How to generate images and speech from Claude Code, Cursor or VS Code with MCP
Configure AIVAX's hosted Media Generation MCP, discover models and prices, and generate images or MP3 speech with explicit tool and privacy limits.
To generate images and speech from an MCP-capable coding client, configure AIVAX's hosted Media Generation MCP with an AIVAX API key, call list_models to choose a model and inspect its price, then call generate_image or generate_speech. The tools return public media URLs as MCP text. The connection uses Streamable HTTP, with no local MCP server process to run.
The same endpoint covers both generation tasks under one AIVAX key. It suits a workflow where an assistant needs an illustration or a spoken version of prepared text. A hosted dependency and public output URLs are the trade-offs. The Media Generation MCP reference defines the complete tool contract; our guide to AIVAX MCPs explains how these service connections fit into an agent's tools.
How do you connect your coding client?
Use this configuration from the AIVAX documentation as the connection specification:
{
"servers": {
"aivax-media": {
"type": "http",
"url": "https://inference.aivax.net/v1/mcp/media-generation",
"headers": {
"Authorization": "Bearer <AIVAX_API_KEY>",
"X-Mcp-Enabled-Tools": "list_models, generate_image, generate_speech"
}
}
}
}
Follow the remote HTTP MCP setup for Claude Code, Cursor, or VS Code. Reuse the URL and headers above; the surrounding JSON depends on the client.
Keep the real key in the client's secure secret storage. The placeholder above documents the required header; it is not an invitation to commit a credential. Never put the key in a repository, shared prompt, browser-side code or log. A positive AIVAX account balance is required, including for access to the model-listing tool.
Once connected, the client discovers enabled tools through tools/list and invokes them through tools/call. Client setup and secret handling should be checked before any paid generation is approved.
How do you enable only images or only speech?
Set X-Mcp-Enabled-Tools to the tools this client needs. For images, use list_models, generate_image. For speech, use list_models, generate_speech. Keep discovery enabled so the assistant can select an actual model rather than guess a name.
Names are case-insensitive, and surrounding spaces are ignored. Omitting the header exposes all three tools. An empty value, or a value containing no recognized names, exposes none. Refresh or reconnect the server after changing the header because the client may cache tool discovery.
This allowlist controls the tools exposed on that connection. Treat the API key as a credential even when the list is narrow: a client holding it can send a different header. For the broader review of permissions and tool exposure, see MCP as a trust boundary.
How do you discover a model and generate an image?
First call list_models with these arguments:
{
"type": "image"
}
The response lists model names, descriptions and prices, plus whether each image model accepts references. Deprecated image models are omitted. Listing models has no charge. Use a returned name in place of <IMAGE_MODEL_NAME> below, then call generate_image:
{
"prompt": "An editorial photograph of red fabric folded into a sail, charcoal background, no text",
"model": "<IMAGE_MODEL_NAME>",
"count": 1
}
prompt must be non-empty. count is an integer from 1 to 4 and defaults to 1. These blocks are tool arguments, not a raw HTTP request body or a transcript of a completed generation.
To guide the result with an existing image, choose a model whose listing explicitly supports references. Add reference_images containing up to four public HTTP(S) image URLs; each must pass the server's URL safety checks. For example, replace the illustrative address before submitting:
{
"prompt": "Use the reference object's silhouette in a restrained editorial composition, no text",
"model": "<IMAGE_MODEL_NAME>",
"count": 1,
"reference_images": [
"https://example.com/approved-reference.jpg"
]
}
References guide the result without guaranteeing identity preservation. Review the returned images before using them. See Image Generation for reference behavior and prompt guidance.
How do you turn prepared text into speech?
Call list_models again, this time with audio rather than image:
{
"type": "audio"
}
Choose a returned speech model and pass its name to generate_speech:
{
"input": "The exhibition opens on Thursday. Please collect your ticket at the entrance.",
"model": "<SPEECH_MODEL_NAME>"
}
input must be non-empty. Omitting voice uses the model's default voice. You can add voice with a value supported by that model; do not assume voices are interchangeable across models.
The optional instructions field accepts guidance such as "Friendly and calm, moderate pace." Some model paths rewrite the input with auditory hints through an additional billed inference step before synthesis. Omit instructions when you want to avoid that text-transformation step, as in the minimal example. Confirm the selected model's behavior and review usage when adding instructions.
The result is MCP text containing a public MP3 URL. This tool has no output-format argument and does not return inline audio. For WAV, OGG or binary delivery, use the Speech Generation API, or convert the downloaded MP3 if that fits your workflow.
What will generation cost, and which limits apply?
Generation is charged to the authenticated AIVAX account at the corresponding service tariffs. Discover current model prices with list_models and consult Pricing for billing rules rather than copying a rate or model name from an old example.
Image billing uses the model's fixed price per delivered output, plus its reference-image price for each reference sent with each output. Prompt processing is included. Image generation adds no separate AIVAX markup or account/plan multiplier. Reference charges therefore matter when requesting variations; check both parts of the tariff.
Speech synthesis uses the same shared service and records the returned speech usage with the API key attached, as the Speech Generation API does. Optional instruction processing can add the inference charge described above.
Image generation shares its service quota with the Image Generation API. Speech shares the account's text-to-speech rate-limit counter with the Speech Generation API; the MCP also checks the general service rate limit. Connecting a second client does not create a separate quota. See Plans and Limits for the applicable account limits.
How should you handle the public URLs?
Anyone with a generated URL can open the image or play the audio. Authenticating the generation request does not make the resulting asset private.
Download the outputs and store your working copies where your application controls access and retention. Doing so does not make the original public URL private. Avoid generating personal, confidential or otherwise sensitive content through this workflow, and do not distribute sensitive reference images at public URLs.
Review assets for accuracy and suitability before publishing. Keep returned URLs out of places where you would not want the assets shared. Plan your own storage rather than treating a generated URL as your archival copy.
When should you choose another approach?
A local model with a local MCP server may be a better fit when you need offline operation or want to avoid metered generation charges. That still requires suitable hardware and maintenance. Running a local wrapper around a paid provider does not remove that provider's bill.
A single-provider server may also fit better when your workflow already depends on that provider and needs options beyond this MCP's tool schema. Compare the required parameters before switching.
What else should you check? FAQ
Why are no tools visible after connecting?
Inspect X-Mcp-Enabled-Tools: an empty header or a list containing no recognized tool names exposes no tools. Correct the allowlist and refresh discovery or reconnect. Also check authentication and positive balance; leaving the header out enables all tools.
What should you do if model discovery fails?
Restore access to list_models and use a name from its response. An unavailable name is rejected by generation validation; substituting a familiar provider name from another integration does not establish availability here.
How should I prepare abbreviations and dates for speech?
Write them as they should be spoken and include punctuation for pauses. Review pronunciation with representative content before adopting a voice, particularly for names and specialized terms. The Speech Generation guide explains text preparation.