France's Mistral AI has released OCR 4, a new version of its document-parsing model. The headline change is that OCR 4 does more than read characters: it returns the location of every element on the page as coordinates, classifies whether each block is a heading, table, or equation, and structures the entire document. It supports 170 languages and offers a self-hosted deployment that runs entirely on a company's own servers, aimed at organizations that cannot let confidential documents leave their premises. Pricing starts at 4 USD (about 640 yen) per 1,000 pages.

From "Turning Pages into Text" to "Extracting Structure"

Until now, OCR (optical character recognition) has mainly served to convert paper or PDFs into clean text and tables. OCR 4 goes a step further and treats a document as a single map of meaning. For each block, such as a paragraph or a figure, it adds a bounding box (a coordinate frame that encloses the target), classifies the block by type such as title, table, equation, signature, or header, and returns a confidence score at both the page and word level.

According to Mistral, bounding boxes were the most-requested capability. Without location data, you cannot trace an extracted number or phrase back to where it came from in the original document. For retrieval-augmented generation (RAG, a setup that references external material to assemble answers) and for audit-related work, being able to later trace "where did this figure come from?" is essential, and OCR 4 is designed to close that gap. Block classification ties directly into real workflows as well: a region judged to be a table can be routed to numerical processing, while a region judged to be a signature can be routed to redaction. Confidence scores support workflows where only low-confidence areas are sent to human reviewers while high-confidence areas pass automatically.

These capabilities are not new in isolation, but obtaining them directly as OCR output, without building a separate layout-analysis stage, lightens the development burden for enterprises. OCR 4 is the fourth generation in roughly 15 months, a sign of how much Mistral has invested in the document space.

Languages, Availability, and Pricing

OCR 4 supports 170 languages across 10 language groups and accepts PDF, DOC, PPT, and OpenDocument formats. It is offered through the Mistral API, Document AI within Mistral Studio, Amazon SageMaker, and Microsoft Foundry, with Snowflake support arriving soon.

Pricing is 4 USD (about 640 yen) per 1,000 pages, dropping to 2 USD (about 320 yen) with a batch-processing discount. At that level, digitizing an internal archive of 100,000 pages costs roughly 200 USD (about 32,000 yen), making large-scale digitization feasible on a realistic budget. Mistral also offers a configuration that bundles the whole model into a single container so it can run entirely inside a company's own infrastructure. This is pitched primarily at heavily regulated industries that do not want to entrust confidential documents to clouds operated by overseas providers.

※ Exchange rates (as of June 25, 2026): 1 USD = 161 JPY, 1 EUR = 185 JPY

Benchmark Results, and How to Read Them

Mistral says that in a human evaluation using more than 600 real-world documents across over 12 languages, OCR 4 achieved an average win rate of 72 percent against competitors. On automated metrics, it reports 85.20 on the public OlmOCRBench and 93.07 on OmniDocBench.

What stands out is that Mistral itself takes a cautious stance on these numbers. The company disclosed the scoring issues it found, such as errors in the ground-truth data and cases where equations that are equivalent but written differently were judged as mismatches, and it said the aggregate score should be read as a rough trend rather than a strict ranking. It is unusual for a vendor to reveal information unfavorable to itself at a product launch. In fact, on the public benchmark leaderboard, OCR 4 sits around third, with open models such as Chandra OCR 2 ranking higher in places. For companies weighing adoption, what matters more than leaderboard figures is which model produces the fewest errors under their own conditions of documents, languages, budget, and processing speed.

Early adopters have offered favorable feedback. One AI company in finance reported that, on chart-heavy data, it kept comparable accuracy while cutting costs to about one-eighth and response latency to about one-seventeenth. An intellectual-property management firm said processing was about four times faster per page than its previous service. All of these are accounts from the providers themselves, and the true performance needs to be verified by each company on its own documents.

Competition Centered on "Sovereign AI"

The ability to run entirely on a company's own servers is a product-level expression of the "sovereign AI" idea Mistral champions, namely keeping data and infrastructure under one's own national or corporate control. The backdrop includes reports of major AI models suddenly becoming unavailable to overseas users on the basis of U.S. export controls, which has heightened interest among enterprises in not being shut off at a provider's discretion. In Europe, penalty provisions of AI regulation are set to take effect in August, and that trend may provide a tailwind.

The competition is intensifying. The day before this announcement, Baidu released a lightweight open model, free of charge, that can read long documents in one pass without splitting them. Beyond that, document-parsing services from Google, Amazon, and Microsoft, specialist vendors such as ABBYY, and a host of open models all crowd the field. Market research puts the document-processing market at about 4.4 billion USD (about 710 billion yen), with annual growth projected above 33 percent through 2030. Mistral is a company of roughly 1,000 people, and going head-to-head with U.S. players that raise enormous sums is no easy task. Even so, it is pursuing a strategy of capturing European demand with an enterprise suite built around sovereignty, structuring, and workflow automation, and reports suggest a new funding round at a valuation of about 20 billion EUR (about 3.7 trillion yen). OCR 4 can be seen as a move positioned at the entrance to that strategy.

Summary

Mistral's OCR 4 is a document-parsing model that does not merely read characters but structures documents by adding location, type, and confidence. With support for 170 languages, a self-hosted deployment option, and pricing of 4 USD per 1,000 pages (2 USD in batch mode), it aims to advance how enterprises put their documents to use. At the same time, Mistral itself urges caution in reading the benchmark numbers, and with many competitors in the field, the final verdict will only settle once each company verifies the model on its own data. How far Mistral can raise its profile in document AI, using sovereignty and structuring as its weapons, is the point to watch.