opendataloader-project/opendataloader-pdf
- Source
- GitHub
- First trending
- Category
- Documents
- GitHub stars
- 29,300
- Main language
- Java
This page introduces an external open-source repository. It is not an HDATF product.

What it does
Extracts Markdown, JSON with bounding boxes, and HTML from PDFs, using a deterministic local mode or an AI hybrid mode with OCR for complex pages. It can also tag untagged PDFs so screen readers can read them.
How it helps ATF
For Company Brain, the bounding boxes in its JSON output could be a reference for pointing an answer back to the exact spot in a source PDF. Worth comparing with MinerU for parsing company documents.
License
Apache-2.0 Permissive, with a patent grant. Commercial use is allowed; keep the notices and state your changes.
More in this category
- microsoft/markitdown
MarkItDown is a Python utility that converts files such as PDF, Word, PowerPoint, Excel, images, audio, HTML and ZIP into Markdown for LLMs and text analysis, keeping structure like headings, lists, tables and links. - google/langextract
LangExtract is a Python library that uses LLMs to pull structured information from unstructured text according to user instructions and examples. Each extracted item is mapped to its exact location in the source, and results can be viewed interactively. - Cocoon-AI/architecture-diagram-generator
A Claude skill that turns a text description of a system into a dark-themed architecture diagram, saved as a standalone HTML file with SVG and buttons to copy or export PNG and PDF. - usememos/memos
Memos is an open-source, self-hosted note-taking app for quick capture. Daily notes, links, work logs and snippets go into a chronological Markdown timeline with tags, search and pins, and each memo can stay private or be shared. - refactoringhq/tolaria
A desktop app for macOS, Windows and Linux for managing markdown knowledge bases. Notes stay plain markdown files, every vault is a git repository, and it works offline with no accounts.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.