deepseek-ai/DeepGEMM
- Source
- GitHub
- First trending
- Category
- Model inference
- GitHub stars
- 8,159
- Main language
- Cuda
This page introduces an external open-source repository. It is not an HDATF product.

What it does
A CUDA tensor core kernel library for language-model computations, including matrix multiplication and fused MoE. Compiles kernels at runtime through DeepJIT.
How it helps ATF
LabChin could reference its kernel implementations when planning experiments on GPU computation for language models. It is a low-level library, not a research-note or workflow tool.
License
MIT Permissive. Commercial use and changes are allowed if the copyright notice is kept.
More in this category
- Ebony-Vinyl/dsh-our-free-model
This dsh plugin connects an agent to externally hosted models rather than running model weights locally. It sends conversation messages, tool results and optional images to upstream gateways and returns model answers or tool calls. Its adapter refreshes model catalogs, probes availability and translates several response protocols into the host's streaming format. It also offers a local OpenAI-compatible forwarding service and local usage records. Starting requires a compatible dsh host, a supported Node.js runtime and network access. Kilo officially allows anonymous requests to free models, but applies IP-based rate limits. The OpenCode route additionally imitates client fingerprints, so its authorization and service terms need separate review rather than treating every route as an officially supported public API. Other account channels and the desktop EAC route have separate login or authorization requirements, and the free gateways must not be treated as unlimited, private or suitable for sensitive production data. - Edge0-AI/Edge0
Edge0 is a streaming MoE inference framework using SSD expert offloading, Recover-LoRA and routing prediction. Its README distinguishes released platform engines from a unified framework planned for Q4 2026. - Vibra-Ingenn/Janus
A single Go binary that runs gguf models on a local machine, using llama.cpp through Vulkan on AMD, Intel or NVIDIA graphics with a CPU fallback, and exposes an OpenAI compatible API for chat completions and models. It swaps models without a restart, splits reasoning output into its own field, and reads the chat template from gguf metadata, with no Python and no Docker required. - magnitudedev/magnitude
An inference engine for agents that compiles and tunes kernels on the user's hardware. It runs open models on Apple Silicon, NVIDIA, AMD or CPU and connects to existing agents through a desktop app and CLI. - ProjectDMX/DMI
A decoupled, asynchronous observability backend for LLM inference and training that exposes internal model states. It is a research preview with HuggingFace, vLLM and Megatron-LM integrations, and its APIs may change.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.