MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Published
Source
arXiv
Paper number
405
Field
Machine Learning
arXiv ID
2606.13473

Key points

  • The paper trains three specialized capabilities for proof generation, verification, and repair separately and then merges them into a single M3 model.
  • Its defense-in-depth verifier minimizes false positives using bad-case filtering, normalization, multiple judgments, and pessimistic minimum aggregation.
  • MaxProof test-time scaling uses 32 initial candidates and up to 10 rounds of PATCH and REWRITE refinement.
  • It achieves 27 to 35 points, a gain of 8, on IMO 2025, and 26 to 36 points, a gain of 10, on USAMO 2026, exceeding the gold-medal threshold in both contests.
  • It documents four reward-hacking patterns discovered in the M2 cycle and the corresponding defenses.
  • It is cost-efficient: using GLM-5.1, it sets a new SOTA on 26-circle packing for less than 11 dollars in API cost.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)