Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding
- Published
- Source
- arXiv
- Paper number
- 476
- Field
- LLMs / NLP
- arXiv ID
- 2606.21906
Key points
- The Guess-Refine-Perturb pattern shows that early layers guess, middle layers refine, and final layers are disturbed by alignment bias.
- It uses entropy-based conservative backward search to dynamically select the best nearby final layer on a token-by-token basis.
- It improves reasoning benchmarks such as GPQA-Diamond, Omni-MATH, and HLE consistently across both dense and MoE architectures.
- It can be dropped into existing inference pipelines with no extra memory and less than 2 percent latency overhead.
- The stronger the alignment, the larger the disturbance in the final layer, which means the effect grows with model size.
Paper links
External research summaries. These are not HDATF publications or measured product results.