run-llama/liteparse
- Source
- GitHub
- First trending
- Category
- Documents
- GitHub stars
- 12,317
- Main language
- Rust
This page introduces an external open-source repository. It is not an HDATF product.

What it does
A PDF parser that runs locally and extracts text with spatial positions and bounding boxes, without cloud services or LLM features. It bundles Tesseract OCR, can call external OCR servers, and checks whether a document needs heavier parsing.
How it helps ATF
Could be a reference for how Company Brain reads PDF material, and for LabChin when pulling text from searched sources. Its check for which documents need OCR is worth comparing for mixed scanned and digital files.
License
Apache-2.0 Permissive, with a patent grant. Commercial use is allowed; keep the notices and state your changes.
More in this category
- opendatalab/MinerU
Converts PDFs, Office files, scanned document images and other formats into Markdown or JSON for LLM and agent workflows. Version 4.0 combines parsing, a local document library and service tools, with parsing tiers from fast previews to demanding layouts. - docusealco/docuseal
An open-source, self-hosted tool for creating PDF forms that people fill and sign online. It offers a WYSIWYG field builder, multiple submitters per document, PDF signature verification, and an API with webhooks. - baidu/Unlimited-OCR
An OCR model aimed at one-shot long-horizon parsing, intended to push DeepSeek-OCR one step further. It runs inference with Hugging Face Transformers on NVIDIA GPUs or with vLLM, and supports training with ms-swift. - iamgio/quarkdown
Quarkdown extends Markdown with functions, variables and a standard library for layout, math, conditions and loops. One project can compile into a book, academic paper, knowledge base, presentation or website. - langchain-ai/openwiki
OpenWiki is a CLI in which an agent reads a codebase or personal sources, writes a linked Markdown wiki and keeps it current on each change. It also tracks key facts back to versioned source evidence.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.