Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
- Published
- Source
- arXiv
- Paper number
- 1072
- Field
- AI / General
- arXiv ID
- 2609.05275
Key points
- Across more than 2,400 training runs, it finds the optimum is stronger dropout for deeper layers, gradually annealed as training progresses.
- At the same training budget, validation loss matches or beats the no-dropout dense baseline while saving up to 25% of compute.
- An 8.2-billion-parameter model with an aggressive schedule peaking at 99% dropout reached loss 1.663, better than the dense baseline's 1.732.
- Dropout-trained models stay robust when layers are skipped, enabling self-speculative decoding—where the model drafts and verifies quickly—for up to 1.55x speedups.
- In contrast, dense models trained without dropout blow up to loss 6.446 when layers are skipped, making them effectively unusable.
Paper links
External research summaries. These are not HDATF publications or measured product results.