Five cut bamboo sections stacked beside curled paper in a misty bamboo grove
All posts

How to convert a PDF to Markdown for RAG and split it into chunks

Convert a PDF to Markdown with OCR, split it into chunks for retrieval, and export JSONL for a RAG collection. Code and a free demo included.

To prepare a PDF for RAG, do three things in order: extract its text as Markdown (with OCR for scanned pages), split that Markdown at topic boundaries into chunks, and write the chunks as JSONL with a stable ID per line so you can import and re-import them. With AIVAX that is two API calls, POST /api/v1/web/fetch and POST /api/v1/generations/segment, plus a short script that writes the JSONL. If you would rather see it work first, AIVAX Toys runs both calls in a browser, free and without sign-up. The toy enables sanitization by default; the script below leaves it off.

A PDF becomes Markdown, then chunks, then JSONL lines in a collection Four stages from left to right. A PDF, image or link goes to the Fetch endpoint, which returns Markdown. The segment endpoint splits the Markdown into chunks. Your code writes one JSONL line per chunk with a stable docid, and the JSONL import loads them into a collection. PDF, image or link up to 10 MB per item scans are read with OCR 1. Fetch POST /web/fetch returns Markdown 2. Segment POST /generations/segment returns chunks 3. JSONL import one line per chunk stable docid, __ref, __tags, __meta
Fetch extracts the Markdown and Segment splits it. Your application calls both endpoints, writes the JSONL and imports the documents.

What does the pipeline look like end to end?

The three steps, the call behind each, and what to check after it
StepCallOutputWhat to check
ExtractPOST /api/v1/web/fetchextractedText as Markdown, plus processingUnitsThe per-item error field, then headings and table headers against the source page
SplitPOST /api/v1/generations/segmentsegments, an array of strings in source orderChunk sizes and whether any chunk starts or ends mid-thought
ImportPOST /api/v1/collections/{id}/documentsDocuments queued for indexingThat re-running the import updates documents instead of duplicating them

Each step can be swapped. If you already parse PDFs locally, start at step two. If your text is already one idea per unit, such as FAQ answers, skip step two. The Text Segmentation docs say so explicitly: each unnecessary split adds indexing work and can separate statements that only make sense together.

Can you try it without writing code?

Yes. Media → Markdown is the first toy in AIVAX Toys, a set of small, open-source demos of AIVAX features. You paste a web page link or drop a PDF or image (PNG, JPEG, WebP, TIFF or BMP), read the Markdown, then split it and download the JSONL.

What is worth knowing before you use it:

  • The visitor is not billed. The Worker uses AIVAX's own key. Each visitor gets 5 documents per minute and 20 per hour, enforced per IP address by a Durable Object, and a Cloudflare Turnstile check runs before each call. Files are capped at 10 MB, the Fetch API's per-item limit.
  • The page shows what the same work would cost you. A "For developers" drawer lists every API call with its latency and processing units, and an estimate based on public list prices. You can switch the plan used for OCR pricing.
  • The Worker code is open. The repository, aivaxlabs/toys, is Apache-2.0. The two upstream calls live in src/worker.js, and the drawer on the page shows the same calls as cURL, JavaScript and Worker code. The Worker does not write the upload anywhere; it forwards it to the Fetch API and returns the result.

The toy is a demonstration, not a production ingestion service. The rest of this post is about what you would build yourself.

How do you convert a PDF to Markdown?

Send the file to the Fetch API as a base64 data URI with its real MIME type. Web pages go in as plain URLs, and one request can mix several items.

curl https://inference.aivax.net/api/v1/web/fetch \
  -H "Authorization: Bearer $AIVAX_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"contents\": [\"data:application/pdf;base64,$(base64 -w0 manual.pdf)\"],
    \"returnErrors\": true
  }"

The response has one entry in results per item, matched by index:

{ "data": { "results": [ {
  "index": 0,
  "extractedText": "# Title\n\nParagraph...",
  "processingUnits": 7,
  "jsonProcessingUnits": 0,
  "error": null
} ] } }

The values above are placeholders that show the shape. Three behaviors from the Fetch and OCR docs shape how you handle the result:

  • Digital and scanned pages are both handled. The docs describe text from digital PDFs and OCR text from scanned pages, and mixed PDFs can combine both.
  • Failures arrive per item. With returnErrors true, a failed item has error set and extractedText null. Check it before the next step: a failed extraction says nothing about whether the PDF contains the answer.
  • Layout is not preserved exactly. The docs state that layout, table structure, charts and embedded images are not guaranteed to be reproduced, and that exact numbers and names may need review after OCR. Read a sample of your own pages before you trust the output; the ingestion quality checklist covers how.

Extraction is metered in processing units (PUs), not tokens. The docs advise against estimating cost from text length; read processingUnits from the response. For the broader picture of what Fetch reads, see Fetch turns the messy web into text your code can use.

How do you split the Markdown into RAG chunks?

Pass the Markdown to the segmentation endpoint as an element of documents:

curl https://inference.aivax.net/api/v1/generations/segment \
  -H "Authorization: Bearer $AIVAX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "documents": ["# Title\n\nParagraph..."], "sanitize": false }'
{ "data": {
  "result": [ { "index": 0, "count": 8, "segments": ["...", "..."] } ],
  "usage": { "processing_units": 2516, "cost": 0 }
} }

Again, the numbers are placeholders. What the docs and the implementation commit to:

  • Without sanitize, boundaries follow the document. Segments target about 300 tokens, cover the whole document in source order, and do not overlap. A document of about that size or smaller comes back as one segment.
  • sanitize trades completeness for cleanliness. It is processed by a language model, takes longer, and omits content it judges useless for retrieval, such as front matter and link references. The API default is false; the toy's checkbox is on, because a demo benefits from dropping clutter such as page headers and link lists. Use it only when what is omitted is genuinely irrelevant. When exact wording or full traceability matters, leave it off and review the source.
  • Clean the Markdown before segmenting. The docs list the failure modes: tables lose their headers, OCR drops line structure, and exported documents repeat running headers on every page. Restore table headers and remove repeated page headers first, rather than expecting the segmenter to work around them.
  • Segmentation does not fix OCR mistakes, and it does not embed anything. It returns strings to your application and stores nothing.

If you need overlap between chunks, add it in your own code; the endpoint does not produce it. Segmentation has its own account quota that counts every entry in documents, so a batch of documents consumes it faster than a single request suggests. See Plans and limits.

What should the JSONL look like for a RAG collection?

A collection import takes one JSON object per line. These are the documented fields, with the values the toy writes:

JSONL fields accepted by the collection import, and what the toy puts in them
FieldMeaningValue in the toy
docidStable document name; existing documents are matched by itmanual.pdf#3 (source name plus chunk number)
textThe text that is embeddedThe chunk
__refGroups chunks of one source; stored up to 64 charactersThe source name, cut to 64 characters
__tagsTags for filtering and maintenanceThe source kind (pdf, image or link) and aivax-toys
__metaMetadata returned with results; not embedded{ "source": ..., "index": ... }

Put anything people will search for in text. Tags and metadata are for filtering, audits and source links, and metadata is not part of the semantic text.

Upload the file as the multipart field documents:

curl https://inference.aivax.net/api/v1/collections/$COLLECTION_ID/documents \
  -H "Authorization: Bearer $AIVAX_API_KEY" \
  -F "[email protected]"

What happens when you import the same PDF again?

Because docid is the match key, the import is an upsert. Per the collection docs, a line with an existing docid and changed text updates that document and queues it for reindexing; a line with unchanged text is skipped, and a line that changes only __ref or __tags is skipped too. Indexing is billed on document text tokens when documents are created or their text changes, so re-importing an unchanged file should not trigger reindexing.

The source#number scheme has two consequences you should plan for. First, the example uses the file name as the source, so two different files that share a name collide; use a unique source identifier within the collection. Second, if a new version of the PDF splits into different boundaries, the numbering shifts and many chunks change text at once. If the new version has fewer chunks, the old extra ones stay in the collection. The import accepts an insert-mode form field set to sync, which deletes documents in the collection whose names are not in the uploaded file. Use it only when the collection holds nothing but that one source, because it compares against every document in the collection, not just those with the same __ref.

Per-request line limits and daily insertion limits depend on your plan; split large files into several imports.

A script for the whole pipeline

This script reads a PDF or image, calls both endpoints, and prints JSONL to stdout, with the chunk and PU counts on stderr. It runs on Bun or Node 18+ and uses no dependencies.

import { readFile } from "node:fs/promises";
import { basename, extname } from "node:path";

const API = "https://inference.aivax.net/api/v1";
const MAX_BYTES = 10 * 1024 * 1024;
const MEDIA_TYPES = {
  ".pdf": "application/pdf",
  ".png": "image/png",
  ".jpg": "image/jpeg",
  ".jpeg": "image/jpeg",
  ".webp": "image/webp",
  ".tif": "image/tiff",
  ".tiff": "image/tiff",
  ".bmp": "image/bmp",
};

const args = process.argv.slice(2);
const sanitize = args.includes("--sanitize");
const [path] = args.filter((arg) => !arg.startsWith("--"));
const key = process.env.AIVAX_API_KEY;

if (!path || !key) {
  console.error("Usage: AIVAX_API_KEY=... bun pdf-to-rag.mjs <file.pdf|image> [--sanitize] > output.jsonl");
  process.exit(1);
}

const mediaType = MEDIA_TYPES[extname(path).toLowerCase()];
if (!mediaType) throw new Error("Use a PDF or a PNG, JPEG, WebP, TIFF or BMP image.");

const bytes = await readFile(path);
if (bytes.length > MAX_BYTES) throw new Error("Fetch accepts at most 10 MB per item. Split the file first.");

async function post(route, body) {
  const response = await fetch(`${API}${route}`, {
    method: "POST",
    headers: { authorization: `Bearer ${key}`, "content-type": "application/json" },
    body: JSON.stringify(body),
  });
  const payload = await response.json().catch(() => ({}));

  if (!response.ok) throw new Error(`${route} failed with HTTP ${response.status}: ${payload.message ?? "no message"}`);
  return payload.data;
}

const fetched = await post("/web/fetch", {
  contents: [`data:${mediaType};base64,${bytes.toString("base64")}`],
  returnErrors: true,
});
const [item] = fetched.results;
if (item.error) throw new Error(`Fetch could not read the file: ${item.error}`);

const segmented = await post("/generations/segment", { documents: [item.extractedText], sanitize });
const [{ segments }] = segmented.result;

const source = basename(path);
segments.forEach((text, index) => {
  console.log(
    JSON.stringify({
      docid: `${source}#${index + 1}`,
      text,
      __ref: source.slice(0, 64),
      __tags: ["pdf-to-rag"],
      __meta: { source, index: index + 1 },
    }),
  );
});

console.error(`${segments.length} chunks | extraction ${item.processingUnits} PU | segmentation ${segmented.usage.processing_units} PU`);

Run it with AIVAX_API_KEY=... bun pdf-to-rag.mjs manual.pdf > manual.pdf.jsonl, then upload the file as shown above. The script is deliberately short: it handles one file, stops on the first error, and has no retries.

We ran it against simulated API responses to confirm the request bodies, the sanitize flag and the shape of the JSONL lines. We did not run it against live OCR for this post, so the chunk counts and PU values you see on your own files are the ones to trust. Keep your API key on a server or in your shell; never ship it in browser code.

What should you check before you index the chunks?

Open the Markdown next to the PDF for a handful of pages, and open the chunks next to the Markdown:

  • Headings in the PDF are headings in the Markdown, so the chunk boundaries have structure to follow.
  • Table headers still sit with their rows, and numbers match the page.
  • Running headers and footers are not repeated inside the text.
  • Chunk sizes are in the same range, with no chunk that is a single orphaned line or a whole chapter.
  • Pick three questions whose answers you know, search the collection, and confirm the right chunk comes back before you attach it to anything in production.

The PDF ingestion checklist turns this into an acceptance matrix, and A vector database is not a RAG system places the preparation step among the other jobs a RAG system has.

When is a different approach a better fit?

The Fetch and segment route is hosted and metered; it is not the right answer for every corpus.

Three ways to prepare a PDF for a RAG collection
ApproachYou controlChoose it when
Fetch, segment, JSONL importChunk text is the extracted Markdown (with sanitize off); IDs, tags and metadata are yoursYou want deterministic IDs, application-controlled metadata and chunks that stay close to the extracted text, without running OCR yourself. Verify OCR against the PDF when exact wording matters
Media InjectorProcessing context; output is generated documentsThe source has several independent topics and you prefer AIVAX to write focused documents; review them afterward, since wording is not preserved
A local parser such as Docling, Marker or PyMuPDFThe whole parsing stack and where documents are processedPDFs must stay inside your infrastructure, you need layout-level tuning, or you would rather run the compute yourself

These can be combined. A local parser can produce the Markdown and the AIVAX segmentation endpoint can split it, or you can chunk with your own splitter and use only the JSONL import. We have not benchmarked the Markdown quality of these tools against each other, so this table is about control and operations, not accuracy.

FAQ

How do I convert a scanned PDF to Markdown for RAG?

Send it to the Fetch API like any other PDF. The docs say scanned pages are read with OCR and mixed PDFs can combine direct text extraction with OCR. Check the result on a few pages, since scan quality and complex layouts affect it.

What chunk size should I use for RAG?

Without sanitize, the segmenter targets about 300 tokens per segment and follows the document structure rather than a fixed length. It is a boundary-finding step, not a size parameter; if you need a specific size or overlap, add a pass of your own after it.

Does the toy store my PDF?

The Worker code does not persist uploads: it reads the file, forwards it to the Fetch API and returns the Markdown. You can verify that in src/worker.js. Do not upload documents you are not allowed to send to a third-party service, in the toy or in your own pipeline.

How much does it cost to run my own pipeline?

Extraction is billed in processing units at a rate that depends on your plan, after daily allowances. At the time of writing the list rate for uncovered extraction is $0.15 per 1,000 PUs on Free, $0.05 on Pro and $0.02 on Max, and text segmentation is $0.30 per million tokens; check Pricing for current values. Indexing the resulting documents is billed separately. The toy's session panel applies these list prices to the PUs of each call, which is a quick way to estimate your own corpus from a few samples.

Can the same pipeline read web pages and images?

Yes. The Fetch API accepts links and PNG, JPEG, WebP, TIFF and BMP images as well as PDFs, and the toy accepts all three. The script above handles local PDFs and images; for a web page, put its URL directly in contents instead of reading and encoding a file.

Why not import the whole PDF as one document?

One embedding per document blurs unrelated topics together, and a long manual then matches many queries weakly. Segmenting gives each topic its own vector. For text that is already one idea per unit, the docs recommend importing it directly.