Latent Thought Flow: Efficient Latent Reasoning in Large Language Models

Published
Source
arXiv
Paper number
444
Field
AI / General
arXiv ID
2606.16222

Key points

  • This is a new approach that solves the language-space bottleneck of Chain-of-Thought, meaning the token-decoding overhead, by reasoning in latent space.
  • A continuous GFlowNet learns a distribution over latent reasoning trajectories proportional to reward, preserving diverse accurate and efficient paths.
  • Entropy-Weighted Subtrajectory Balance propagates terminal rewards to intermediate steps, which addresses the sparse-supervision problem.
  • Reference-prior regularization provides stable early exploration, and annealing then shifts the model toward reward-based trajectories.
  • Compared with CoLaR and ReGuLaR, it improves fine-tuning accuracy by 12.9 percent and shortens reasoning length by 34.5 percent.
  • In transfer-learning settings, it also improves accuracy by 6.0 percent and shortens reasoning length by 19.9 percent, which demonstrates generality.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)