EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Published
Source
arXiv
Paper number
567
Field
LLMs / NLP
arXiv ID
2607.05155

Key points

  • Environment learning average performance follows a log-sigmoid scaling law, with R² = 0.998, and the pattern is stable across all 6 task families.
  • It studies 134 real-world tasks, each requiring 12 or more continuous hours of execution, with large tasks that take human experts an average of 57.2 hours.
  • Frontier agent learning speed roughly doubles every three months, based on models released after September 2025.
  • Continuous experience is stronger than independent restarts, and accumulated experience strongly determines long-horizon performance.
  • The theory derives the log-sigmoid shape as a frontier-expansion process on a latent task graph.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)