FARS: A Fully Automated Research System Deployed at Scale
- Published
- Source
- arXiv
- Paper number
- 541
- Field
- AI / General
- arXiv ID
- 2606.31651
Key points
- The system uses a fully automated four-stage pipeline: Ideation, Planning, Experiment, and Writing, with a specialized agent at each stage and a shared workspace.
- It generated 166 papers across 67 fine-grained AI and ML topics and preserved the full output, including successful, failed, and negative results, rather than only selected samples.
- It collected 282 structured volunteer reviews covering 140 papers, including ratings, soundness, presentation, contribution, integrity checks, and disclosure of LLM use.
- In a cross-comparison with the Stanford Agentic Reviewer, FARS scores 5.00 out of 10 across 165 papers, which surpasses existing systems such as DeepScientist at 4.13 and ScientistOne at 4.00.
- Human reviews average 3.23 out of 10, so they are stricter than automated reviewers but still aligned in direction, and 11.4 percent of papers reach the accept threshold under automated review criteria.
- Repeated failure modes include narrow experimental scope, methodological limitations, and integrity issues such as claim-artifact mismatch, which makes contribution quality, experimental sufficiency, and faithful reporting the true bottlenecks rather than polished presentation.
Paper links
External research summaries. These are not HDATF publications or measured product results.