VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
- Published
- Source
- arXiv
- Paper number
- 429
- Field
- AI / General
- arXiv ID
- 2606.16140
Key points
- With only 3 billion parameters, it reaches 94.3 on AIME26 and 97.1 with CLR, at the level of DeepSeek V3.2 671B and GLM-5 744B.
- On the latest LeetCode contests from 2026-04 to 2026-05, it achieves a 96.1% acceptance rate, outperforming GPT-5.2 at 95.3% and Kimi K2.5 at 90.6%.
- We propose a staged post-training pipeline consisting of curriculum SFT, MGPO RL, offline self-distillation, and Instruct RL.
- We introduce the Parametric Compression-Coverage Hypothesis, which says that verifiable reasoning can be compressed, while open-ended knowledge requires coverage.
Paper links
External research summaries. These are not HDATF publications or measured product results.