Can AI agents conduct open-ended AI research? Early evidence from two case studies
- Published
- Source
- arXiv
- Paper number
- 769
- Field
- AI / General
- arXiv ID
- 2607.27191
Key points
- It proposes a new method called shadow evaluation, in which AI is given the research questions from unpublished papers and the original authors do the grading.
- The same failure patterns were reproduced for both Opus 4.8 and GPT-5.6 Sol.
- The agents were good at engineering work, such as coding and running experiments, but they lacked strong judgment and creativity in research design.
- The five core failures were weak research-level standards, failure to back away from dead ends, poor resource awareness, non-creative responses to feedback, and drift away from the instructions.
- They used less than half of the API budget, still had time left over, and never received a pass in AI self-review.
Paper links
External research summaries. These are not HDATF publications or measured product results.