Sunlit open-air concrete room with blank paper sheets clipped to slatted walls, diagonal bands of yellow light across them, and a blue ring on a stand in the foreground
All posts

How to check PDF and scan ingestion quality before you embed for RAG

Before embedding, sample real pages, write the facts you expect, and grade each generated document for wrong numbers, lost table headers, and missing facts.

To check ingestion quality before embedding, take a small sample of real pages, write down by hand the facts a user will ask about, run the import into a scratch collection, and grade every generated document against those facts. Treat one wrong number as a blocking defect. A missing fact is a coverage gap you can fix by changing the source or the processing context. Search tuning comes after this check, because no retriever or reranker can recover a value that was misread on the way in.

The rest of this guide turns that check into a procedure for Media Injector output and for text segmented with Text Segmentation. It describes a method, not a measurement: it contains no accuracy figures for AIVAX, and the acceptance thresholds are yours to set.

Inspect the evidence before you embed the document

Once a misread total sits in a collection, it is retrieved with the same confidence as a correct one, and the answer built on it cites a real document. A vector database is only one part of a RAG system; the text that goes into it decides what the system can honestly say.

Developer reports point to the same failures. One r/LangChain write-up describes a table split across two pages whose second chunk has no header information, and a later thread on PDF parsing for RAG shows people still comparing setups. These are user reports, not measurements. They do show why the check has to run on your pages.

What does a benchmark tell you about your own PDFs?

OmniDocBench is a public reference for what to measure. Its README describes 1,651 PDF pages across 10 document types, with annotations for text, tables, formulas, and reading order, and it scores parsers on four things: text (normalized edit distance), tables (TEDS), formulas (CDM), and reading order. The README dates its current version, v1.6, to April 2026.

It has two limits. It ranks parsers that turn a page into Markdown; it does not measure retrieval, and it says nothing about Media Injector or any AIVAX component. And its scores rely on a page-level ground truth that your invoices and policies do not have. Use the four axes as a vocabulary, then build a ground truth for a few of your own pages by hand.

How does Media Injector turn a file into documents?

Media Injector accepts PDFs, images, audio, and video from the dashboard and produces RAG documents. It is not a page-by-page transcription. The current implementation runs an interview: a model reads the file and your optional processing context, proposes factual questions, and each question is answered from the source into one self-contained document. It stops when it judges that no materially new knowledge remains, and it is instructed to ignore pagination, decorative text, boilerplate, and repeated summaries.

Three consequences shape the check:

  • The output is a rewrite, not an excerpt. Each document is a concise factual answer in the source's language, so the wording is the model's even when the values come from the page. Verify facts against the source; do not diff text.
  • Missing facts are easy to overlook. The interview may decide a table row is not worth a question. A wrong number is visible; an absent one is not, unless you wrote down what you expected.
  • Processing context steers priority but is not evidence. The docs state that it cannot add facts absent from the file. If a document contains something only the context said, treat it as a defect.

Generated documents are tagged autogenerated and interview, plus a tag derived from the source file name. Their metadata records the source file name and the question that produced them, so each document can be traced to what it was meant to answer.

Build the sample and write the expected facts first

Pick pages that break things, not the cleanest ones: a scanned page, a table that crosses a page break, a two-column page, a chart or photo with values, and a page dense with repeated headers. Fifteen to twenty pages spread across those kinds is a workable start; the number is a judgment call, not a statistic.

For each page, write the expected facts before you import anything. A fact is one sentence a user could plausibly ask about, with its exact values:

invoice-0412.pdf, p.2: the invoice total is 4,180.00 BRL, due 2026-10-15.
invoice-0412.pdf, p.2: line 3 (freight) is 320.00 BRL and is not taxed.

Write them before importing: once you have read the generated documents, it is easy to accept them as the answer key.

Grade each document against four axes

Import the sample into a scratch collection, wait for indexing to finish, then open the generated documents. The first three rows are the OmniDocBench axes, restated for retrieval; the last is specific to indexing.

Acceptance matrix for generated or segmented documents
AxisWhat to checkFailure that blocks indexing
Fact fidelityEvery name, number, date, unit, and currency matches the pageAny wrong or transposed value, or a value the page does not contain
Table integrityEach value is attached to its row and column header, including on the continuation pageA figure without its label, or a header applied to the wrong column
Reading orderFacts from one column or sidebar are not merged into another's subjectA statement that combines text from two unrelated blocks
Retrievable unitThe document names its subject, answers one question, and stands aloneA pronoun with no referent, a fragment that needs the previous page, or several policies in one block

For the retrievable-unit axis, the RAG best-practices guide gives concrete signals. Keep most documents between 20 and 700 words, which is a target and not an API rule. Documents under about 10 tokens may be too small to retrieve reliably. Above about 1,562 tokens, the indexer truncates the text used for embedding to about 5,000 characters and records a warning. The Get Document endpoint returns each document's character and word counts, so you can read the smallest and largest first.

Then score coverage. For every expected fact, mark it as found and correct, found and wrong, or missing, and record any generated fact that is neither on your list nor on the page. To list the generated documents, browse the collection with a filter; -t filters by tag and -c by content:

curl -G "https://inference.aivax.net/api/v1/collections/COLLECTION_ID/documents" \
  -H "Authorization: Bearer YOUR_AIVAX_API_KEY" \
  --data-urlencode "filter=-t autogenerated -c 4,180.00"

If a fact is found but wrong, decide whether the fault is in the source or in the extraction. A faint scan, a mislabeled extension, or a stamp over a total is a source problem: fix or replace the file. A correct source with a wrong result is a signal to retry with a sharper processing context, or to use the text route in the next section.

When should you segment text yourself instead?

Use Media Injector when you want a managed pipeline that turns a file into question-shaped knowledge units. Choose a text route when exact wording matters: the docs advise retaining and reviewing the source text when legal fidelity or full traceability is required. A common shape is Fetch and OCR to extract text, your own review or cleanup, Text Segmentation to propose boundaries, and then an import of the segments you accept.

The segmenter changes what you must check. In the implementation, the model does not rewrite anything. It returns line ranges, and the endpoint joins the original lines for each range. Segments are therefore excerpts of your text, with empty lines dropped and each line trimmed. But the endpoint silently skips any range it cannot parse or that falls outside the text, and it does not verify that the ranges cover every line or avoid overlapping. With sanitize left at false, the instruction is to segment the whole document, so a line-count shortfall means lines were lost. Test for it:

const source = ocrText;

const response = await fetch(
  "https://inference.aivax.net/api/v1/generations/segment",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.AIVAX_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ documents: [source], sanitize: false }),
  },
);

const { segments } = (await response.json()).data.result[0];

const countLines = (text) =>
  text.split(/\r?\n/).map((line) => line.trim()).filter(Boolean).length;

const expected = countLines(source);
const actual = countLines(segments.join("\n"));

if (actual !== expected) {
  console.warn(`Line count differs: source ${expected}, segments ${actual}`);
}

This check catches dropped and duplicated lines, not a segment boundary that cuts a table between its header and its rows. That needs a manual read; the docs list lost table headers first among the failure modes of messy sources. Segmentation returns text to your application; it neither embeds nor stores it, so nothing is indexed until you send the accepted segments to the collection.

What should you do with a failed sample?

Delete the affected documents with the Delete Document endpoint, or the whole scratch collection, fix the cause, and rerun the same pages so results stay comparable. Failed or cancelled Media Injector jobs can be retried while their uploaded data is still available. Keep the expected-facts file with your ingestion setup: when you change the processing context, the file type mix, or the segmenter settings, rerun it as a regression check before the production import.

Once the sample passes, semantic search and reranking can be evaluated on documents you have already verified.

FAQ

Can I use a public benchmark score to pick an ingestion method?

Use it to choose what to measure, not to rank a method for your files. OmniDocBench scores parsers on its own annotated pages and documents its known limits, such as inconsistent symbol recognition and text evaluation limited to Chinese and English. Your documents, languages, and scan quality differ.

Does Text Segmentation fix OCR errors?

No. It chooses boundaries and returns your lines grouped into segments. A misread character in the input stays in the output, so review the OCR text before segmenting.

Why did Media Injector leave out a fact I expected?

It generates knowledge units from questions it considers materially useful and stops when it finds no new ones. Add the fact to your processing context as something to prioritize, rerun the sample, and check whether it appears. If the source does not state it, context cannot supply it.

How large should each generated document be?

The docs give a practical target of 20 to 700 words for most documents, which is not a hard rule. Prefer one answerable topic per document, with its subject named in the text.

Should tags carry the important meaning?

No. Only document text is embedded for semantic matching, while tags and metadata are for filtering, audits, and source links. If users search for "refund policy", those words belong in the text.