TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
- Paper number
- 1159
- Field
- Machine Learning
Key points
- Nearly eliminated optimizer state by using a sparse geometry that updates only the sign of the largest-magnitude coordinate in each column.
- Reduced optimizer state memory by 174x (27.7GB to 0.16GB) and peak training memory by 2.9x (80.6GB to 27.5GB) versus AdamW8bit (OPT-13B).
- Achieved 94.22% accuracy on OPT-13B SST-2, close to Adam's 95.3%, while cutting memory from 254.8GB to 27.5GB.
- Enabled full-parameter fine-tuning of 30-32B models on a single 80GB H100 GPU across multiple model families and tasks.
- Unlike Muon, it shows little performance degradation when applied to AdamW-pretrained models.
Paper links
External research summaries. These are not HDATF publications or measured product results.