Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Published
Source
arXiv
Paper number
768
Field
LLMs / NLP
arXiv ID
2607.28568

Key points

  • It builds the full-stack OpenMLE system for AI-for-AI improvement and creates 5,758 executable task environments.
  • It trains the meta-evolution agent Frontis-MA1-35B by using reinforcement learning to learn four evolutionary operators: drafting, improving, debugging, and crossover.
  • On MLE-Bench Lite, the medal average rises from 39.4 percent to 71.2 percent, which beats GPT-5.5 plus Codex at 68.2 percent.
  • A separate benchmark, NatureBench Lite, also confirms transfer learning, which shows that the result is not simple benchmark memorization.
  • It releases the model weights, task environments, and training code, providing a reproducible basis for recursive self-improvement research.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)