Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

Published
Source
arXiv
Paper number
392
Field
LLMs / NLP
arXiv ID
2606.11926

Key points

  • Autonomous Optimization, or AO, is defined as a task in which an initial output is improved through repeated experiments without step-by-step human supervision.
  • The hypothesis tree implements cumulative research state by binding each node to a hypothesis, a version of the output, experimental evidence, and refined insight.
  • It separates the coordinator, which handles the global strategy, from the executor, which runs individual experiments in isolated worktrees.
  • Across six real research tasks, including model training, harness work, and data synthesis, it reports more than 2.5x average relative gains over Codex and Claude Code.
  • On MLE-Bench Lite, it achieves 86.36 percent Any Medal with GPT-5.5, and ablations confirm that HTR and insight propagation are key.
  • It demonstrates generalization by transferring the BrowseComp optimization harness to a previously unseen search-agent task.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)