TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Paper number
1159
Field
Machine Learning

Key points

  • Nearly eliminated optimizer state by using a sparse geometry that updates only the sign of the largest-magnitude coordinate in each column.
  • Reduced optimizer state memory by 174x (27.7GB to 0.16GB) and peak training memory by 2.9x (80.6GB to 27.5GB) versus AdamW8bit (OPT-13B).
  • Achieved 94.22% accuracy on OPT-13B SST-2, close to Adam's 95.3%, while cutting memory from 254.8GB to 27.5GB.
  • Enabled full-parameter fine-tuning of 30-32B models on a single 80GB H100 GPU across multiple model families and tasks.
  • Unlike Muon, it shows little performance degradation when applied to AdamW-pretrained models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)