baidu/Unlimited-OCR
- Source
- GitHub
- First trending
- Category
- Documents
- GitHub stars
- 25,740
- Main language
- Python
This page introduces an external open-source repository. It is not an HDATF product.

What it does
An OCR model aimed at one-shot long-horizon parsing, intended to push DeepSeek-OCR one step further. It runs inference with Hugging Face Transformers on NVIDIA GPUs or with vLLM, and supports training with ms-swift.
How it helps ATF
Worth comparing as an OCR step for turning scanned company documents and papers into text that Company Brain or LabChin could search.
License
MIT Permissive. Commercial use and changes are allowed if the copyright notice is kept.
More in this category
- langchain-ai/openwiki
OpenWiki is a CLI in which an agent reads a codebase or personal sources, writes a linked Markdown wiki and keeps it current on each change. It also tracks key facts back to versioned source evidence. - opendatalab/MinerU
Converts PDFs, Office files, scanned document images and other formats into Markdown or JSON for LLM and agent workflows. Version 4.0 combines parsing, a local document library and service tools, with parsing tiers from fast previews to demanding layouts. - iOfficeAI/OfficeCLI
A single-binary command-line tool that lets AI agents read, edit and automate Word, Excel and PowerPoint files without installing Office. It renders .docx, .xlsx and .pptx to HTML or PNG so an agent can check and fix its output. - run-llama/liteparse
A PDF parser that runs locally and extracts text with spatial positions and bounding boxes, without cloud services or LLM features. It bundles Tesseract OCR, can call external OCR servers, and checks whether a document needs heavier parsing. - chuspeeism/dashi-ppt-skill
An agent skill that turns documents into web presentations in 12 visual themes. Each page has in-browser controls for editing text, layout and media, and the result can be exported to HTML, PDF or editable PPTX.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.