NVIDIA/Model-Optimizer
- Source
- GitHub
- First trending
- Category
- Model inference
- GitHub stars
- 4,079
- Main language
- Python
This page introduces an external open-source repository. It is not an HDATF product.

What it does
A model optimization library with quantization, pruning, distillation and speculative decoding. It exports optimized checkpoints for downstream inference frameworks.
How it helps ATF
For Harness, it can inform comparisons of model optimization options for the inference infrastructure behind AI execution.
License
Apache-2.0 Permissive, with a patent grant. Commercial use is allowed; keep the notices and state your changes.
More in this category
- Low-Zi-Hong/ESP32s3-LLM-Cluster
A distributed language-model inference project using seven ESP32-S3 nodes and 1.58-bit quantization. A master handles tokenization and embeddings while SPI-connected nodes process model layers. - Taichu-AI/ZDTaichu5.0-9B
A multimodal model supporting text, images, and video for visual understanding, spatial reasoning, agent tool use, and embodied-AI research. It combines a Qwen3.5-9B language backbone with a C-RADIOv4-H vision encoder. - Niko1221/Strata
A local inference engine for Qwen3.8-Flash-Next on Windows and Linux PCs with NVIDIA GPUs. The repository describes OpenAI- and Anthropic-compatible localhost APIs and optional image input. - mizorewww/laya-coreml
A port of Laya typed decision models to Apple Core ML and the Neural Engine, running on device without generating tokens, published with speed and energy benchmarks and a Snake demo. - firelex/jeff
Fine-tuned Qwen3.5 and Gemma 4 models classify situations against options described in plain language and return probabilities without generating text.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.