The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

Published
Source
arXiv
Paper number
512
Field
Machine Learning
arXiv ID
2606.26294

Key points

  • Controlled utility evolution, with a fixed criterion within each epoch and replacement only at epoch boundaries, preserves self-improvement guarantees even under nonstationary utility.
  • On coding tasks such as Polyglot, adding agent-as-a-judge code review signals improves pass rate by 1.8 percentage points over the previous SOTA while using 1.35x to 1.72x fewer tokens.
  • In the paper-writing setting, the co-evolved writer's acceptance rate rises from 21.8 percent to 40.5 percent, which is a 1.86x improvement.
  • In IMO-level proof grading, the co-evolved grader reaches higher accuracy with three times less search cost than the fixed baseline.
  • By injecting adversarial goals at epoch boundaries, it corrects reviewer self-preference bias toward AI-generated papers, where AI papers are accepted 1.91 times too often.
  • As the evaluator gets stronger, it also creates a curriculum effect for the task agent, which leads to archive reranking and gradual strengthening.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)