Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
- Published
- Source
- arXiv
- Paper number
- 014
- Field
- Reasoning
- arXiv ID
- 2501.04682
Key points
- Meta-CoT is defined as p(y|x)=∫p(y|x,z)p(z|x)dz, where z denotes a latent reasoning and search trajectory rather than a directly supervised final explanation.
- It targets the limits of standard CoT: one-shot generation, weak self-verification, poor backtracking, and weak exploratory reasoning on hard math, logic, and planning tasks.
- The paper proposes a training pipeline that internalizes search-like behavior using process supervision, process reward models, synthetic search traces from algorithms such as MCTS or A*, instruction tuning, and reinforcement-learning post-training.
- It reports qualitative and empirical signs that larger frontier models can approach more complex solutions with more test-time reasoning compute, while highlighting unresolved scaling, verifier, and consistency issues.
Paper links
External research summaries. These are not HDATF publications or measured product results.