Agent Laboratory: Using LLM Agents as Research Assistants

Published
Source
arXiv
Paper number
011
Field
Research Agents
arXiv ID
2501.04227

Key points

  • In pipeline terms, human research ideas are transformed into a curated arXiv literature review, an experimental plan, executable ML code, and a final research report and code repository.
  • A method-specialized LLM agent team collaborates across PhD student, postdoc, ML engineer, and professor roles. mle-solver uses replacement and edit operations, execution feedback, LLM grading, self-reflection, and a pool of high-scoring programs.
  • As a result, o1-preview scored highest overall, co-pilot mode improved human-evaluated quality from 3.8 to 4.38 out of 10, and gpt-4o execution was reported at about $2.33 per paper, showing an 84% cost reduction versus prior autonomous research methods.
  • Notably, automatic NeurIPS-style reviews averaged 6.1 out of 10, while human reviewers averaged 3.8, so the system is better viewed as a research assistant than a reliable autonomous scientist.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)