Full-bandwidth transformer
- Published
- Source
- arXiv
- Paper number
- 878
- Field
- AI / General
- arXiv ID
- 2608.08888
Key points
- We introduce latent feedback decoding, which feeds the full top hidden state from the previous token directly into the next input.
- It adds less than 1% extra inference cost while preserving the standard Transformer structure and KV cache.
- At 1B parameters and 400B training tokens, it consistently improves validation loss, math and code generation, and instruction following.
- It saves about 1.5x training tokens for the same performance and in some tasks approaches models trained with 5x more tokens.
- When latent feedback is enabled, the reasoning trace becomes shorter while accuracy is maintained or improved.
Paper links
External research summaries. These are not HDATF publications or measured product results.