yifanfeng97/Hyper-Extract
- Source
- GitHub
- First trending
- Category
- RAG and search
- GitHub stars
- 4,068
- Main language
- Python
This page introduces an external open-source repository. It is not an HDATF product.
What it does
Hyper-Extract is a Python library that extracts facts and relations from documents into structured knowledge such as knowledge graphs, and it ships the he CLI and the he-mcp server. It is meant for developers and analysts who want to turn scattered documents into knowledge that supports search and question answering. It reads txt and md files directly, and an optional install extra based on markitdown adds PDF, DOCX, PPTX and XLSX. Results are stored in structures such as lists, sets, graphs, hypergraphs, temporal graphs, spatial graphs, spatio-temporal graphs and document collections, and they can be exported to Obsidian, GraphML, CSV, JSON-LD and Cypher. Users choose among domain YAML templates such as finance, legal, medicine and industry and extraction methods such as GraphRAG, LightRAG, Hyper-RAG and KG-Gen. Extraction needs an LLM that supports function calling, a way of receiving answers in a fixed format, together with an OpenAI-compatible embedding model. According to the provider documentation, Anthropic, Gemini and DeepSeek need a separately paired embedding model, and Bailian qwen-max and deepseek-v3 support only the json_object format and cannot be used. Each document's source and content hash are recorded so only changed documents are processed again, and the content that came from one document can be rolled back on its own. The he-mcp server only reads and exports, offering tools for listing templates, viewing information, search, question answering and export. Knowledge templates accept only Chinese or English as the language option, directory processing reads only the files directly inside the given folder, and PDFs without a text layer need OCR first.
License
Custom license Custom or non-standard license. Read the license file before any reuse.
More in this category
- mongodb-developer/mongodb-jvm-showcase
Independent MongoDB example projects for Java and Kotlin. Examples cover drivers, frameworks, RAG and vector search. - hydra-db/hydradb
A Rust distributed graph database with durable S3-compatible object storage, OpenCypher queries, Bolt connectivity, and an HTTPS API. Query-serving nodes and indexers operate as separate compute roles. - realZachi/pg-jev
A PostgreSQL extension that filters, ranks and classifies rows using natural-language conditions evaluated by TypeSafe's Jev. It uses returned probabilities rather than generated text. - asciimoo/hister
A private search engine that indexes the full contents of visited pages and local files. Search is available through a web interface, terminal or an AI assistant connected via MCP. - Tencent/WeKnora
WeKnora turns raw documents into RAG-based Q&A and a ReAct agent that uses retrieval, MCP tools and sandboxes. Its Wiki Mode has agents build an interlinked, editable markdown knowledge base with revision history.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.