CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
- Published
- Source
- arXiv
- Paper number
- 048
- Field
- Research Agents
- arXiv ID
- 2503.22708
Key points
- Existing autonomous scientific discovery, or ASD, systems often focus on only part of the research pipeline or rely on domain-specific languages, which limits the range of possible discoveries.
- There is a need for an end-to-end semi-automated system that uses large language models for scientific discovery while emphasizing code-based experimentation.
- Evaluating the reliability and reproducibility of discoveries made by LLM-based ASD systems is difficult because of the inherent variability of language model outputs and the potential unreliability of generated code.
- CODESCIENTIST implements a semi-automated end-to-end scientific discovery workflow consisting of ideation, planning, experiment design and execution, reporting, and meta-analysis.
- It uses a mutator-style LLM genetic search approach for ideation, combining fragments of research papers and code blocks to generate diverse research ideas.
- The system generates and runs code in an instrumented sandbox, uses iterative generate run reflect debug loops, and performs multiple independent executions for each experiment to ensure reproducibility.
Paper links
External research summaries. These are not HDATF publications or measured product results.