MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Published
Source
arXiv
Paper number
610
Field
Computer Vision
arXiv ID
2607.11562

Key points

  • It builds MonkeyDoc v2, the largest document image pretraining corpus, with 113 million images across 17 languages.
  • It jointly learns image-to-text generation and pixel-level document reconstruction to achieve character-level visual recognition.
  • It improves CRNN recognition accuracy from 58.7 percent to 67.3 percent, and the 110M UniMERNet-T outperforms the 325M UniMERNet-B.
  • A 0.7B document parsing model achieves open-source SOTA on MDPBench, beating the previous best, the 3B dots.mocr, by 2.8 percent.
  • With the vision encoder frozen, it achieves stronger document understanding than large VLMs such as Qwen3-VL-235B and GPT-5.2.
  • It shows consistent gains over CLIP, DINO, and SAM across 8 benchmarks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)