AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Baidu has open-sourced Unlimited-OCR, a 3-billion-parameter AI model capable of parsing entire multi-page PDFs in a single pass. Its innovative architecture offers significant performance advantages over existing models, especially for long documents.

Baidu has open-sourced Unlimited-OCR, a 3-billion-parameter AI model designed to parse entire multi-page documents in a single forward pass. This development, announced on June 22, 2026, introduces a novel constant-memory architecture that enables faster and more accurate reading of lengthy PDFs on local hardware, potentially transforming document processing workflows.

The model, released under an MIT license and available on Hugging Face, builds upon Baidu’s existing DeepSeek-OCR lineage. It employs a new Reference Sliding Window Attention (R-SWA) mechanism, replacing traditional attention layers to maintain a fixed memory footprint regardless of document length. This allows the model to process dozens of pages simultaneously without the typical slowdowns caused by growing memory requirements.

According to the technical report published alongside the release, Unlimited-OCR achieves a throughput of approximately 5,580 tokens per second on OmniDocBench, outperforming Baidu’s previous DeepSeek-OCR model by about 12.7%. Its accuracy metrics on various benchmarks show improvements in text, formula, and table recognition, with overall scores reaching over 93 points on OmniDocBench v1.5 and v1.6. Notably, it can parse long documents—up to 40 pages—with an error rate below 0.11, a significant advance for long-form OCR tasks.

Despite viral claims of high download numbers, Baidu’s model has approximately 8,400 downloads in July 2026, far below the 1.9 million figure circulating online, indicating a more modest adoption rate so far.

At a glance
reportWhen: announced June 2026, released June 22,…
The developmentBaidu released Unlimited-OCR, a new AI OCR model that can process multi-page PDFs in a single forward pass using a constant-memory mechanism, marking a technical breakthrough in document reading.

Impact of Constant-Memory Architecture on Document OCR

This development matters because it addresses the longstanding challenge of processing lengthy documents efficiently. Traditional OCR models slow down or require splitting documents into pages, which can cause errors in cross-page references and reading order. Baidu’s approach maintains constant memory usage, enabling faster, more accurate processing of entire documents in a single pass. This could streamline workflows in legal, academic, and enterprise settings, reducing manual effort and improving accuracy in long-form document digitization.

Amazon

AI OCR software for PDF processing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Advances in OCR and Baidu’s Model Lineage

Prior to this release, most OCR systems relied on page-by-page processing, which limits accuracy for multi-page documents and complicates tasks like table recognition and cross-referencing. Baidu’s previous models, such as PaddleOCR-VL and DeepSeek-OCR, achieved high accuracy but still depended on splitting documents. The new Unlimited-OCR builds on Baidu’s DeepSeek-OCR architecture, introducing a specialized attention mechanism that enables a true single-pass, multi-page parsing capability. This approach aligns with broader trends toward more efficient, scalable AI models capable of handling complex documents on local hardware.

“Unlimited-OCR’s constant-memory architecture allows parsing of entire multi-page documents in a single forward pass, significantly improving speed and accuracy for long documents.”

— Baidu Research Team

NetumScan 13MP Book Document Camera for Teachers with Windows,Mac OS,Linux

NetumScan 13MP Book Document Camera for Teachers with Windows,Mac OS,Linux

  • Automatic Image Correction: One-key skew correction and auto scanning
  • High-Quality Image Capture: 1300W CMOS sensor for clear images
  • Universal Compatibility: Supports Windows, Mac OS, Linux

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects and Performance Limits

While the technical report presents promising results, it remains unclear how Unlimited-OCR performs across diverse, real-world datasets outside Baidu’s internal tests. Its accuracy relative to top models like PaddleOCR-VL and Zhipu’s GLM-OCR is slightly lower in some benchmarks, and broader adoption and integration into workflows are still developing. Additionally, the long-term robustness and handling of complex layouts or poor-quality scans need further validation.

Amazon

local hardware OCR scanner

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Deployment and Industry Adoption of Unlimited-OCR

Next steps include broader testing by third parties, integration into commercial OCR pipelines, and potential updates to improve accuracy on diverse document types. Baidu may also expand its open-source ecosystem, encouraging community contributions. Monitoring real-world performance and adoption rates over the coming months will be key to understanding its impact on the OCR landscape.

Amazon

long document OCR software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does Unlimited-OCR differ from existing OCR models?

It employs a constant-memory attention mechanism that enables processing entire multi-page documents in a single pass, unlike traditional models that process pages separately, improving speed and accuracy for long documents.

Can I run Unlimited-OCR on my own hardware?

Yes, the model is open-sourced under an MIT license and supports deployment via Docker, Transformers, and community quantizations, making it accessible for local use.

What are the main limitations of Unlimited-OCR?

Its performance on very diverse or poor-quality documents remains to be fully tested, and it currently offers slightly lower peak accuracy than some page-by-page models, trading off some accuracy for better long-document handling.

Will this model replace existing OCR solutions?

It is likely to complement rather than replace current models, especially in scenarios requiring processing of lengthy documents where speed and coherence are critical.

Source: ThorstenMeyerAI.com

You May Also Like

DARPA, U.S. Air Force Fly AI-controlled F-16

DARPA and the U.S. Air Force successfully flew an F-16 fighter jet operated entirely by artificial intelligence in a recent demonstration.

RHEO On The Web: Find Your Flow

Discover RHEO’s web version, offering instant, private, browser-based fluid simulations for relaxation, creativity, and ambient displays, with no downloads or sign-up.

Why Investors See Anthropic’s Series H as a Compute Power Play

Discover why Anthropic’s $965B valuation is less about valuation and more about massive compute capacity — reshaping AI infrastructure and hardware supply chains.

ALIA. The Spanish answer.

Spain launches ALIA-40B, a €240M public-funded multilingual LLM, demonstrating strategic positioning and operational capabilities, but below Llama 2 benchmarks.