When a RAG system gives a strange answer, most teams suspect the embedding model, the vector database, or the prompt first.
But if the PDF text was misread at the start, no search or generation downstream can produce the right answer.
Mistral's OCR 4 and 4.1 line, released in 2026, reads more than plain text. It handles structure block by block — titles, lists, tables, images, equations, code — and returns confidence information with it.
Use this model as a reason to revisit the front of your RAG pipeline.
Search quality starts with "did we read the document correctly?" — before chunking.
A PDF is not a text file
It looks like one page to you, but a PDF mixes many elements inside.
- Body text
- Two-column layouts
- Headers and footers
- Tables
- Captions
- Equations
- Code blocks
- Scanned images
If the parser gets the reading order wrong, sentences get scrambled.
The screen can show this:
Left column: experiment conditions
Right column: results
But line-by-line extraction can interleave the two:
experiment conditions results first condition accuracy 92% second condition ...
And the meaning breaks.
CodeBridge mini experiment: inspect one PDF before any search
Pick one PDF you plan to put into RAG. Before searching, compare just five parts by eye.
- Title and section order
- One table
- One number-heavy paragraph
- A footnote or header
- One figure caption
Then attach this checklist to the extraction result.
[ ] Reading order is correct
[ ] Table row/column relationships are preserved
[ ] Headers and page numbers stay out of the body
[ ] Numbers and units are not lost
[ ] Image captions are not glued to the wrong paragraph
If problems already show here, fix parsing before tuning embedding parameters.
Why structure information matters
Cutting all text into fixed character counts is easy to implement. But it can split against the document's logic.
For example, a document shaped like this:
## Refund policy
### Within 30 days
...
### Digital goods exception
...
can separate the Digital goods exception heading from its body when cut by length. Then search results lose meaning.
If the OCR stage gives you headings, lists, and tables as structure, your chunking step can use that information.
Treat confidence scores as recheck points, not auto-reject
OCR 4.1 supports confidence granularity at page, block, and word levels. Instead of deleting low-confidence content outright, use it like this.
- Reprocess low-confidence pages with an image-based pass
- Send important low-confidence numbers to human review
- Cross-check table regions with a second parser
- Flag answers that cite low-confidence source text
In short, confidence is not the answer itself. It is a signal telling you where to look again.
When RAG fails, trace in this order
When a question goes unanswered, tracing in this order makes the cause easier to find.
1. Is the answer actually in the source document?
2. Did the OCR or parser extract that part correctly?
3. Does the chunk keep the needed context?
4. Did the retriever fetch that chunk?
5. Did the LLM use the retrieved result correctly?
Many teams start at steps 4 and 5, but the real problem can sit at step 2.
This flow also helps when you read what RAG is and the internal document chatbot post.
Mistakes teams repeat in practice
Processing every PDF with the same parser
Text PDFs, scanned PDFs, slide-style PDFs, and table-heavy reports behave differently. Evaluate samples per document type instead.
Never looking at OCR output with your own eyes
Loading vectors straight into the database hides parsing errors. Compare the original and the extraction side by side for at least a few pages.
Evaluating only answer quality
When the final answer is wrong, you cannot tell which stage caused it. Check parsing, retrieval, and generation separately.
Conclusion: good RAG starts one step before good search
What document models like Mistral OCR 4.1 show is not a matter of a few percent of OCR accuracy.
It is that documents are starting to be treated as structured data, not plain strings.
When RAG behaves oddly, do not swap the model first. Put the original and the parse result side by side.
No LLM can revive information that broke into an unsearchable shape.
Further reading
References
Go deeper with a course
If you want to build the full pipeline from document intake to retrieval to cited answers, a guided course takes you through each stage hands-on.