Unlimited OCR Works
- Published
- Source
- arXiv
- Paper number
- 468
- Field
- Computer Vision
- arXiv ID
- 2606.23050
Key points
- Reference Sliding Window Attention (R-SWA) keeps the KV cache constant by giving reference tokens full attention and output tokens a sliding window of 128.
- On OmniDocBench v1.5, it improves by 6% over DeepSeek OCR, reaching 93%, with an MoE structure that uses only 0.5B active parameters.
- It can process more than 40 pages at once and is 35% faster when generating 6,000 tokens.
- Replacing full standard attention with R-SWA does not reduce single-page OCR accuracy.
- It was developed by Baidu, and the code and weights are released on GitHub.
Paper links
External research summaries. These are not HDATF publications or measured product results.