huggingface/transformers
- Source
- GitHub
- First trending
- Category
- Model inference
- GitHub stars
- 167,264
- Main language
- Python
This page introduces an external open-source repository. It is not an HDATF product.

What it does
Transformers is Hugging Face's Python library that gathers model definitions for text, image, audio, video and multimodal models for use in inference and training. Its main users are researchers, ML engineers and developers who want to download and run or fine-tune open models, and more than 1 million checkpoints in this format are published on the Hub. It serves as the shared reference whose model definitions are also used by training tools such as Axolotl, Unsloth and DeepSpeed, inference engines such as vLLM, SGLang and TGI, and llama.cpp and mlx. Each model consists of three classes, a configuration, a model and a preprocessor, and Pipeline runs text generation, speech recognition, image classification and visual question answering in a few lines of code. Trainer supports mixed precision, torch.compile, FlashAttention and distributed training, and generate offers streaming and several decoding strategies. Getting started requires Python 3.10 or later and PyTorch 2.6 or later, downloaded models are stored in the HF_HUB_CACHE folder, and offline use means downloading in advance with snapshot_download and then setting HF_HUB_OFFLINE=1. transformers serve, available through the serving extra, opens an OpenAI SDK compatible server at localhost:8000 with endpoints such as chat completions, responses, audio transcription and model listing. Tool calling works with models whose tokenizer declares tool call tokens, and the official documentation presents this server for evaluation, experiments and moderate-load deployment while recommending vLLM or SGLang for large-scale production. The official README presents the library as a model-definition-centered tool, and because its training API is tuned to the PyTorch models Transformers provides, it recommends Accelerate for generic training loops and adapting the example scripts to each situation. The README still carries wording about moving models between PyTorch, JAX and TF2.0, but the current setup.py (5.19.0.dev0) and installation docs cover PyTorch only, and while the library is Apache-2.0, each checkpoint can carry its own license and access conditions.
License
Apache-2.0 Permissive, with a patent grant. Commercial use is allowed; keep the notices and state your changes.
More in this category
- Tencent-Hunyuan/Hy-MT2
Hy-MT2 is a family of translation language models released by the Tencent Hunyuan team in three sizes, 1.8B, 7B and 30B-A3B, the last being a mixture-of-experts (MoE) model. The models translate among 33 languages, and Korean and Japanese are in the supported language table. Users give the source text and translation instructions as a prompt and receive the translation, and the README provides prompt templates for terminology, style, personalization, delimiters and structured data. The maximum context length is 8192 tokens and there is no default system prompt, so long documents need to be split before translation. FP8 and GGUF quantized versions, which shrink the model files, go down to 2-bit and 1.25-bit formats, and the README says the 1.25-bit version of the 1.8B model needs about 440 MB of storage. Running it requires transformers 5.6.0 or later with trust_remote_code, a setting that executes code from the model repository, or vLLM and SGLang built from source, and the README says low-bit formats in llama.cpp need the STQ kernel from pull request 22836. LoRA or full fine-tuning runs through LLaMA-Factory with DeepSpeed, and the training guide asks for one GPU with 24 GB or more for 1.8B, one 80 GB GPU for 7B LoRA and eight 80 GB GPUs for 30B. The IFMTBench benchmark for instruction-following translation is published under CC-BY-4.0 and scores six constraint types with rules and an LLM judge. The model license is Apache 2.0 according to LICENSE.txt, and the README benchmark comparisons were not independently checked here. - Ebony-Vinyl/dsh-our-free-model
This dsh plugin connects an agent to externally hosted models rather than running model weights locally. It sends conversation messages, tool results and optional images to upstream gateways and returns model answers or tool calls. Its adapter refreshes model catalogs, probes availability and translates several response protocols into the host's streaming format. It also offers a local OpenAI-compatible forwarding service and local usage records. Starting requires a compatible dsh host, a supported Node.js runtime and network access. Kilo officially allows anonymous requests to free models, but applies IP-based rate limits. The OpenCode route additionally imitates client fingerprints, so its authorization and service terms need separate review rather than treating every route as an officially supported public API. Other account channels and the desktop EAC route have separate login or authorization requirements, and the free gateways must not be treated as unlimited, private or suitable for sensitive production data. - deepseek-ai/DeepGEMM
A CUDA tensor core kernel library for language-model computations, including matrix multiplication and fused MoE. Compiles kernels at runtime through DeepJIT. - Edge0-AI/Edge0
Edge0 is a streaming MoE inference framework using SSD expert offloading, Recover-LoRA and routing prediction. Its README distinguishes released platform engines from a unified framework planned for Q4 2026. - Vibra-Ingenn/Janus
A single Go binary that runs gguf models on a local machine, using llama.cpp through Vulkan on AMD, Intel or NVIDIA graphics with a CPU fallback, and exposes an OpenAI compatible API for chat completions and models. It swaps models without a restart, splits reasoning output into its own field, and reads the chat template from gguf metadata, with no Python and no Docker required.
Only repositories in the ranked Trendshift lists are included, and the lists are used only to find candidates. We do not copy their ranks. Descriptions, licenses and star counts come from each GitHub repository. The notes are our own reading. We have not tested these projects, and a place on a trending list does not prove quality.