VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

Published
Source
arXiv
Paper number
429
Field
AI / General
arXiv ID
2606.16140

Key points

  • With only 3 billion parameters, it reaches 94.3 on AIME26 and 97.1 with CLR, at the level of DeepSeek V3.2 671B and GLM-5 744B.
  • On the latest LeetCode contests from 2026-04 to 2026-05, it achieves a 96.1% acceptance rate, outperforming GPT-5.2 at 95.3% and Kimi K2.5 at 90.6%.
  • We propose a staged post-training pipeline consisting of curriculum SFT, MGPO RL, offline self-distillation, and Instruct RL.
  • We introduce the Parametric Compression-Coverage Hypothesis, which says that verifiable reasoning can be compressed, while open-ended knowledge requires coverage.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)