CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
- Published
- Source
- arXiv
- Paper number
- 540
- Field
- Computer Vision
- arXiv ID
- 2606.31986
Key points
- The problem is that text CoT must generate thousands of tokens autoregressively, which causes large inference latency, and the natural-language form can erase alternative reasoning paths.
- It uses three forms of step-level supervision: forward decoding from latent to text, backward decoding from text to latent, and internal supervision for consistency between latent steps.
- At inference time it removes the decoder and internal supervision, so reasoning is completed with only 3 latent vectors and no extra parameter overhead.
- Across 8 benchmarks on Qwen3-VL-8B, including MMStar, MathVista, and SeedBench, it consistently outperforms both text CoT and previous latent reasoning methods.
- For speed, text CoT takes 7.24 seconds and 142 tokens, whereas CoLT takes 0.32 seconds with 3 latent vectors, which is a 22.6x decoding speedup and a 10.1x end-to-end speedup.
- It is also robust to noise: under visual and textual noise, Direct Answer drops by 7.2 to 19.0 percent, Text CoT drops by 5.0 to 13.8 percent, but CoLT drops only 2.8 to 9.0 percent.
Paper links
External research summaries. These are not HDATF publications or measured product results.