lyogavin/airllm

Source
GitHub
First trending
Category
Model inference
GitHub stars
34,422
Main language
Jupyter Notebook

This page introduces an external open-source repository. It is not an HDATF product.

lyogavin/airllm

What it does

AirLLM lowers the GPU memory needed for large language model inference by loading model layers one at a time, so a 70B model can run on a single 4GB GPU without quantization, distillation or pruning. Training support was added too.

How it helps ATF

Worth comparing for AX consulting when a customer site wants to run open models on its own small GPUs. Speed and fit for real workloads would need to be checked separately.

License

Apache-2.0 Permissive, with a patent grant. Commercial use is allowed; keep the notices and state your changes.

More in this category

  • FareedKhan-dev/kimi-k3-in-c
    A portable C99 engine that runs Kimi K3 inference on one CPU with no GPU, BLAS or framework, streaming the model from disk. The readme reports an 8.24 GB peak memory use, with more RAM only adding speed.
  • MoonshotAI/Kimi-K3
    Kimi K3 is an open-weight multimodal Mixture-of-Experts model with 2.8T total and 104B activated parameters and a 1-million-token context window. It understands text, images and video and targets long-horizon coding, knowledge work and reasoning.
  • unslothai/unsloth
    Unsloth is a desktop app for running and training LLM, diffusion, embedding and audio models, including GGUF and MLX formats. Local models can also be used with Claude Code, Codex and MCP, together with web search and RAG.
  • drumih/turbo-fieldfare
    A Swift and Metal runtime that runs Gemma 4 26B-A4B on Apple Silicon Macs in about 2 GB of RAM. It keeps the shared core and KV cache in memory and streams only the needed experts from SSD.
  • cactus-compute/needle
    Needle 2 is an open 45M-parameter model for tool calling, device use and structured extraction, packed into a single 14MB engine. The Python package covers inference, LoRA fine-tuning and export, returning JSON tool calls with a confidence score.

Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.

View on GitHub (opens in a new tab)