ASI-Bench: At the Dawn of Artificial Superintelligence
- Published
- Source
- arXiv
- Paper number
- 943
- Field
- AI / General
- arXiv ID
- 2608.17271
Key points
- The benchmark evaluates the same task under four levels of guidance: providing the full procedure, naming only the method, allowing autonomous method selection, and adding distracting information. This separates instruction following from autonomous research ability.
- Across 18 agent and model combinations, the average score fell from 50.91 with the full procedure to 29.10 when only the method was specified and 26.62 when method selection was autonomous.
- Scores varied substantially by harness, even for the same model. MiMo V2.5 Pro rose from 16.17 with its own harness to 23.25 with the Claude Code harness.
- Providing only the method name was the most expensive condition, consuming an average of 6.91 million tokens per task. Incomplete guidance increased the cost of exploration.
- The best-performing combination scored 51.60%, while a lower-cost combination delivered similar performance at one quarter of the cost, revealing a wide spread in cost efficiency.
Paper links
External research summaries. These are not HDATF publications or measured product results.