BitNet b1.58 2B4T Technical Report

Published
Source
arXiv
Paper number
061
Field
LLMs / Efficiency
arXiv ID
2504.12285

Key points

  • Large language models are compute-intensive, requiring substantial memory and energy and causing high latency.
  • These resource demands constrain LLM deployment on edge devices and in real-time applications.
  • Previous 1-bit quantization methods either hurt performance, as in post-training quantization, or failed to scale to competitive model sizes.
  • It implements a modified Transformer architecture with custom BitLinear layers for 1.58-bit weights and 8-bit activations.
  • It uses a multi-stage training approach that begins with an unprecedented 4 trillion-token pretraining run and continues with supervised fine-tuning and Direct Preference Optimization.
  • It develops and open-sources a specialized inference library with custom kernels for efficient W1.58A8 operations on both GPUs and CPUs.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)