CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation

Published
Source
arXiv
Paper number
048
Field
Research Agents
arXiv ID
2503.22708

Key points

  • Existing autonomous scientific discovery, or ASD, systems often focus on only part of the research pipeline or rely on domain-specific languages, which limits the range of possible discoveries.
  • There is a need for an end-to-end semi-automated system that uses large language models for scientific discovery while emphasizing code-based experimentation.
  • Evaluating the reliability and reproducibility of discoveries made by LLM-based ASD systems is difficult because of the inherent variability of language model outputs and the potential unreliability of generated code.
  • CODESCIENTIST implements a semi-automated end-to-end scientific discovery workflow consisting of ideation, planning, experiment design and execution, reporting, and meta-analysis.
  • It uses a mutator-style LLM genetic search approach for ideation, combining fragments of research papers and code blocks to generate diverse research ideas.
  • The system generates and runs code in an instrumented sandbox, uses iterative generate run reflect debug loops, and performs multiple independent executions for each experiment to ensure reproducibility.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)