OvisOCR2 Technical Report
- Published
- Source
- arXiv
- Paper number
- 638
- Field
- Computer Vision
- arXiv ID
- 2607.13639
Key points
- Given a document-page image, it generates text, equations, tables, and visual regions as Markdown in a natural reading order in one pass.
- It combines real document annotations with synthetic pages whose images and ground truth are generated from the same HTML, applying supervised learning, reinforcement learning, on-policy distillation, and model merging.
- The 0.8B model achieved the highest scores on both benchmarks: 96.58 on OmniDocBench v1.6 and an Avg3 of 75.06 on PureDocBench.
- By outperforming pipeline approaches with a single small model, it suggests the possibility of simplifying deployment configurations for document-retrieval and knowledge-building systems.
- However, validation is limited to two public benchmarks and an in-house benchmark for converting page images to Markdown, so it does not guarantee performance in other document-processing workflows.
Paper links
External research summaries. These are not HDATF publications or measured product results.