BitNet b1.58 2B4T Technical Report
- Published
- Source
- arXiv
- Paper number
- 061
- Field
- LLMs / Efficiency
- arXiv ID
- 2504.12285
Key points
- Large language models are compute-intensive, requiring substantial memory and energy and causing high latency.
- These resource demands constrain LLM deployment on edge devices and in real-time applications.
- Previous 1-bit quantization methods either hurt performance, as in post-training quantization, or failed to scale to competitive model sizes.
- It implements a modified Transformer architecture with custom BitLinear layers for 1.58-bit weights and 8-bit activations.
- It uses a multi-stage training approach that begins with an unprecedented 4 trillion-token pretraining run and continues with supervised fine-tuning and Direct Preference Optimization.
- It develops and open-sources a specialized inference library with custom kernels for efficient W1.58A8 operations on both GPUs and CPUs.
Paper links
External research summaries. These are not HDATF publications or measured product results.