MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
- Published
- Source
- arXiv
- Paper number
- 610
- Field
- Computer Vision
- arXiv ID
- 2607.11562
Key points
- It builds MonkeyDoc v2, the largest document image pretraining corpus, with 113 million images across 17 languages.
- It jointly learns image-to-text generation and pixel-level document reconstruction to achieve character-level visual recognition.
- It improves CRNN recognition accuracy from 58.7 percent to 67.3 percent, and the 110M UniMERNet-T outperforms the 325M UniMERNet-B.
- A 0.7B document parsing model achieves open-source SOTA on MDPBench, beating the previous best, the 3B dots.mocr, by 2.8 percent.
- With the vision encoder frozen, it achieves stronger document understanding than large VLMs such as Qwen3-VL-235B and GPT-5.2.
- It shows consistent gains over CLIP, DINO, and SAM across 8 benchmarks.
Paper links
External research summaries. These are not HDATF publications or measured product results.