Towards Robust Mathematical Reasoning

Published
Source
arXiv
Paper number
092
Field
Math Reasoning / Benchmarks
arXiv ID
2511.01846

Key points

  • Existing mathematical reasoning benchmarks, such as GSM8K, MATH, and AIME, are approaching saturation, which limits their usefulness for distinguishing advanced model capability.
  • Many current benchmarks rely mainly on final-answer matching, which can let models guess or memorize answers without demonstrating robust multi-step reasoning.
  • There is a lack of a comprehensive evaluation framework for assessing deeper mathematical understanding, such as the ability to generate and rigorously evaluate proofs.
  • The paper introduces IMO-Bench, a comprehensive collection of three benchmarks: IMO-AnswerBench for robust problem solving, IMO-Proof Bench for rigorous proof writing, and IMO-GradingBench for proof evaluation.
  • IMO-AnswerBench uses 400 olympiad problems across four categories and applies explicit hardening techniques, such as paraphrasing and numerical changes, to prevent data memorization.
  • IMO-Proof Bench provides 60 IMO-level proof problems split into foundational and advanced sets, and it is evaluated mainly by human experts on a 0 to 7 scale that mimics traditional IMO grading.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)