Can AI agents conduct open-ended AI research? Early evidence from two case studies

Published
Source
arXiv
Paper number
769
Field
AI / General
arXiv ID
2607.27191

Key points

  • It proposes a new method called shadow evaluation, in which AI is given the research questions from unpublished papers and the original authors do the grading.
  • The same failure patterns were reproduced for both Opus 4.8 and GPT-5.6 Sol.
  • The agents were good at engineering work, such as coding and running experiments, but they lacked strong judgment and creativity in research design.
  • The five core failures were weak research-level standards, failure to back away from dead ends, poor resource awareness, non-creative responses to feedback, and drift away from the instructions.
  • They used less than half of the API budget, still had time left over, and never received a pass in AI self-review.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)