Full-bandwidth transformer

Published
Source
arXiv
Paper number
878
Field
AI / General
arXiv ID
2608.08888

Key points

  • We introduce latent feedback decoding, which feeds the full top hidden state from the previous token directly into the next input.
  • It adds less than 1% extra inference cost while preserving the standard Transformer structure and KV cache.
  • At 1B parameters and 400B training tokens, it consistently improves validation loss, math and code generation, and instruction following.
  • It saves about 1.5x training tokens for the same performance and in some tasks approaches models trained with 5x more tokens.
  • When latent feedback is enabled, the reasoning trace becomes shorter while accuracy is maintained or improved.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)