Agent Laboratory: Using LLM Agents as Research Assistants
- Published
- Source
- arXiv
- Paper number
- 011
- Field
- Research Agents
- arXiv ID
- 2501.04227
Key points
- In pipeline terms, human research ideas are transformed into a curated arXiv literature review, an experimental plan, executable ML code, and a final research report and code repository.
- A method-specialized LLM agent team collaborates across PhD student, postdoc, ML engineer, and professor roles. mle-solver uses replacement and edit operations, execution feedback, LLM grading, self-reflection, and a pool of high-scoring programs.
- As a result, o1-preview scored highest overall, co-pilot mode improved human-evaluated quality from 3.8 to 4.38 out of 10, and gpt-4o execution was reported at about $2.33 per paper, showing an 84% cost reduction versus prior autonomous research methods.
- Notably, automatic NeurIPS-style reviews averaged 6.1 out of 10, while human reviewers averaged 3.8, so the system is better viewed as a research assistant than a reliable autonomous scientist.
Paper links
External research summaries. These are not HDATF publications or measured product results.