OvisOCR2 Technical Report

Published
Source
arXiv
Paper number
638
Field
Computer Vision
arXiv ID
2607.13639

Key points

  • Given a document-page image, it generates text, equations, tables, and visual regions as Markdown in a natural reading order in one pass.
  • It combines real document annotations with synthetic pages whose images and ground truth are generated from the same HTML, applying supervised learning, reinforcement learning, on-policy distillation, and model merging.
  • The 0.8B model achieved the highest scores on both benchmarks: 96.58 on OmniDocBench v1.6 and an Avg3 of 75.06 on PureDocBench.
  • By outperforming pipeline approaches with a single small model, it suggests the possibility of simplifying deployment configurations for document-retrieval and knowledge-building systems.
  • However, validation is limited to two public benchmarks and an in-house benchmark for converting page images to Markdown, so it does not guarantee performance in other document-processing workflows.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)