Towards Robust Mathematical Reasoning
- Published
- Source
- arXiv
- Paper number
- 092
- Field
- Math Reasoning / Benchmarks
- arXiv ID
- 2511.01846
Key points
- Existing mathematical reasoning benchmarks, such as GSM8K, MATH, and AIME, are approaching saturation, which limits their usefulness for distinguishing advanced model capability.
- Many current benchmarks rely mainly on final-answer matching, which can let models guess or memorize answers without demonstrating robust multi-step reasoning.
- There is a lack of a comprehensive evaluation framework for assessing deeper mathematical understanding, such as the ability to generate and rigorously evaluate proofs.
- The paper introduces IMO-Bench, a comprehensive collection of three benchmarks: IMO-AnswerBench for robust problem solving, IMO-Proof Bench for rigorous proof writing, and IMO-GradingBench for proof evaluation.
- IMO-AnswerBench uses 400 olympiad problems across four categories and applies explicit hardening techniques, such as paraphrasing and numerical changes, to prevent data memorization.
- IMO-Proof Bench provides 60 IMO-level proof problems split into foundational and advanced sets, and it is evaluated mainly by human experts on a 0 to 7 scale that mimics traditional IMO grading.
Paper links
External research summaries. These are not HDATF publications or measured product results.