
Mistral releases OCR 4 with bounding boxes and block classification
Mistral AI unveiled OCR 4, a document-parsing model that returns bounding boxes, typed block classification and confidence scores alongside extracted text, supports 170 languages and can run self-hosted in a single container.
Mistral AI has released OCR 4, a document-parsing model that returns a structured representation of a page — bounding boxes, typed block classification and confidence scores — alongside the extracted text. The company announced the model on June 23, 2026, and positions it as an ingestion component for enterprise search, retrieval-augmented generation and domain-specific retrieval pipelines.
What is new
Previous OCR generations mainly converted a page into clean text and tables. OCR 4 returns structure: each block is localized with a bounding box, classified by type — titles, tables, equations, signatures and more — and accompanied by confidence scores generated per page and per word. According to Mistral, bounding boxes were the most-requested capability, enabling in-context highlighting and more reliable data pipelines, while block types and confidence scores support source-grounded citations, redactions and human-in-the-loop verification.
The model accepts common enterprise formats, including PDF, DOC, PPT and OpenDocument, and supports 170 languages across 10 language groups. Mistral says the gains are measurable in specialized and low-resource languages where several competing systems degrade. Because it is compact enough to run in a single container, OCR 4 can be deployed fully self-hosted, keeping document data inside an organization's own infrastructure.
Benchmarks and pricing
In a head-to-head human evaluation covering more than 600 documents in over 12 languages, independent annotators preferred OCR 4 over every leading OCR and document-AI system tested, with win rates averaging 72%, according to the company. The model also reports the top overall score among the systems tested on the public OlmOCRBench (85.20) and a score of 93.07 on OmniDocBench. Mistral notes that both benchmarks have known scoring limitations — from incorrect ground-truth annotations to equivalent LaTeX notation counted as a mismatch — and treats aggregate scores as directional.
OCR 4 through the API is priced at $4 per 1,000 pages, with a 50% Batch API discount that lowers the price to $2 per 1,000 pages. Document AI, which adds schema-based structured output on the same endpoint, costs $5 per 1,000 pages. The model is available via API through Mistral Studio, Amazon SageMaker and Microsoft Foundry, with Snowflake Parse Document coming soon.
Where it fits
Mistral describes OCR 4 as an ingestion component of its Search Toolkit, an open-source, composable search framework announced at the AI Now Summit. Early users apply it to turn invoices into structured fields, digitize company archives, extract text from technical and scientific reports, and power enterprise search. Mistral also lists what the model is not intended for: medical diagnosis, legal advice or judgment, high-stakes financial decisions, safety-critical systems, real-time processing and non-document inputs such as raw audio or video.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.