Build a retrieval pipeline on AIVAX Free
The expanded Free plan includes separate daily allowances for RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. Here is what each covers—and what remains metered.
A useful retrieval prototype has more than a vector search box. It needs to read source material, index it, find relevant passages, and sometimes rerank them before a model can answer. The expanded AIVAX Free plan now includes daily allowances for four of those building blocks: RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. These are separate allowances for specific metered work, not a promise that every step in an AI application is free.
Start with RAG: index, then search
The RAG embedding allowance covers eligible document insertion and search embeddings on Free. Both operations draw from the same RAG allowance. An embedding served from the query cache does not consume it. That lets you build a small collection and test whether the passages returned for real questions are useful before adding answer generation.
RAG is the emphasis here because retrieval quality needs its own iteration loop: prepare sensible document chunks, index them, issue representative queries, and inspect the matches. Adding an LLM too early can hide poor retrieval behind fluent prose. The allowance covers embeddings, not the model inference used to write a RAG answer. Collection count, storage, insertion volume, and search rate also have separate Free-plan limits; check Plans and limits before sizing a workload.
Rerank with Reflex, decide with Julia-1
Reflex can reorder candidate passages after retrieval; it can also rerank without a collection through its own endpoint. Free includes a separate daily Reflex allowance shared by those uses. Cached and uncached input both count toward it. Reflex does not generate an answer, and its separate processing-time cap and request limits still apply.
For routing or classification steps, the Decisions API evaluates named questions against a shared state and returns typed answers rather than chat prose. The @supersonic-labs/julia-1 model is served through AIVAX infrastructure and is eligible for the Free semantic-decision allowance. Other decision models are not covered by that allowance. Julia-1 has request and context limits, and account balance and rate-limit checks still apply even when a decision is included.
Turn source material into text with Fetch and OCR
Fetch and OCR extracts readable text from supported web pages, documents, and images. Its Free daily extraction allowance is separate from the RAG, Reflex, and decision allowances, and extraction usage is measured in Processing Units. It can help prepare source material before indexing; it does not itself create a RAG collection or answer a question. Optional schema-guided JSON conversion is charged separately and is not part of the extraction allowance.
Understand the boundary before you build
Each service has its own daily allowance. Coverage is checked for each metered item: a request can contain included work and separately billed work. When an item is not covered, normal rates apply. Allowances do not transfer between services or bypass positive-balance requirements, rate limits, storage limits, or the Reflex processing-time cap. LLM inference and RAG answer generation are currently outside subscription coverage. Reseller accounts do not receive these subscription allowances.
Use the account's subscription usage indicators to track consumption and reset status, and consult Pricing for uncovered rates and Plans and limits for the current technical boundaries. That distinction makes it practical to test retrieval, reranking, decisions, and extraction independently without mistaking a covered building block for an entirely free production pipeline.
Create an AIVAX account to try the Free plan with your own documents and queries.