CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts

Published
Source
arXiv
Paper number
540
Field
Computer Vision
arXiv ID
2606.31986

Key points

  • The problem is that text CoT must generate thousands of tokens autoregressively, which causes large inference latency, and the natural-language form can erase alternative reasoning paths.
  • It uses three forms of step-level supervision: forward decoding from latent to text, backward decoding from text to latent, and internal supervision for consistency between latent steps.
  • At inference time it removes the decoder and internal supervision, so reasoning is completed with only 3 latent vectors and no extra parameter overhead.
  • Across 8 benchmarks on Qwen3-VL-8B, including MMStar, MathVista, and SeedBench, it consistently outperforms both text CoT and previous latent reasoning methods.
  • For speed, text CoT takes 7.24 seconds and 142 tokens, whereas CoLT takes 0.32 seconds with 3 latent vectors, which is a 22.6x decoding speedup and a 10.1x end-to-end speedup.
  • It is also robust to noise: under visual and textual noise, Direct Answer drops by 7.2 to 19.0 percent, Text CoT drops by 5.0 to 13.8 percent, but CoLT drops only 2.8 to 9.0 percent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)