FARS: A Fully Automated Research System Deployed at Scale

Published
Source
arXiv
Paper number
541
Field
AI / General
arXiv ID
2606.31651

Key points

  • The system uses a fully automated four-stage pipeline: Ideation, Planning, Experiment, and Writing, with a specialized agent at each stage and a shared workspace.
  • It generated 166 papers across 67 fine-grained AI and ML topics and preserved the full output, including successful, failed, and negative results, rather than only selected samples.
  • It collected 282 structured volunteer reviews covering 140 papers, including ratings, soundness, presentation, contribution, integrity checks, and disclosure of LLM use.
  • In a cross-comparison with the Stanford Agentic Reviewer, FARS scores 5.00 out of 10 across 165 papers, which surpasses existing systems such as DeepScientist at 4.13 and ScientistOne at 4.00.
  • Human reviews average 3.23 out of 10, so they are stricter than automated reviewers but still aligned in direction, and 11.4 percent of papers reach the accept threshold under automated review criteria.
  • Repeated failure modes include narrow experimental scope, methodological limitations, and integrity issues such as claim-artifact mismatch, which makes contribution quality, experimental sufficiency, and faithful reporting the true bottlenecks rather than polished presentation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)