Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

Published
Source
arXiv
Paper number
476
Field
LLMs / NLP
arXiv ID
2606.21906

Key points

  • The Guess-Refine-Perturb pattern shows that early layers guess, middle layers refine, and final layers are disturbed by alignment bias.
  • It uses entropy-based conservative backward search to dynamically select the best nearby final layer on a token-by-token basis.
  • It improves reasoning benchmarks such as GPQA-Diamond, Omni-MATH, and HLE consistently across both dense and MoE architectures.
  • It can be dropped into existing inference pipelines with no extra memory and less than 2 percent latency overhead.
  • The stronger the alignment, the larger the disturbance in the final layer, which means the effect grows with model size.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)