FareedKhan-dev/kimi-k3-in-c

Source
GitHub
First trending
Category
Model inference
GitHub stars
8,000
Main language
C
Website
medium.com/@fareedkhandev/building-kimi-k3-in-c-to-run-a-2-8t-model-on-consumer-hardware-a5792cbf3b59 (opens in a new tab)

This page introduces an external open-source repository. It is not an HDATF product.

FareedKhan-dev/kimi-k3-in-c

What it does

A portable C99 engine that runs Kimi K3 inference on one CPU with no GPU, BLAS or framework, streaming the model from disk. The readme reports an 8.24 GB peak memory use, with more RAM only adding speed.

How it helps ATF

An example of running a large model on a single CPU, which could be a loose reference for AX consulting on sites without GPUs. The readme itself says generation is slow.

License

Apache-2.0 Permissive, with a patent grant. Commercial use is allowed; keep the notices and state your changes.

More in this category

  • unslothai/unsloth
    Unsloth is a desktop app for running and training LLM, diffusion, embedding and audio models, including GGUF and MLX formats. Local models can also be used with Claude Code, Codex and MCP, together with web search and RAG.
  • lyogavin/airllm
    AirLLM lowers the GPU memory needed for large language model inference by loading model layers one at a time, so a 70B model can run on a single 4GB GPU without quantization, distillation or pruning. Training support was added too.
  • cactus-compute/needle
    Needle 2 is an open 45M-parameter model for tool calling, device use and structured extraction, packed into a single 14MB engine. The Python package covers inference, LoRA fine-tuning and export, returning JSON tool calls with a confidence score.
  • MoonshotAI/Kimi-K3
    Kimi K3 is an open-weight multimodal Mixture-of-Experts model with 2.8T total and 104B activated parameters and a 1-million-token context window. It understands text, images and video and targets long-horizon coding, knowledge work and reasoning.
  • FlashML-org/FreeToken
    A serving engine for running large open-weight Mixture-of-Experts models on personal hardware by splitting work across GPUs, CPUs and memory. It exposes Anthropic- and OpenAI-compatible APIs so coding agents can connect.

Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.

View on GitHub (opens in a new tab)