SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Published
Source
arXiv
Paper number
035
Field
Benchmarks / Code Agents
arXiv ID
2502.12115

Key points

  • Existing LLM coding benchmarks rely on unit tests, which makes them vulnerable to grader hacking and insufficient for validating complex full-stack software engineering tasks.
  • Earlier evaluations lack direct economic metrics, so they cannot quantify the real monetary value of LLM outputs in realistic software development scenarios.
  • Current benchmarks are often too narrow, focusing on isolated programming problems or specific repositories rather than the diverse, ambiguous, and end-to-end nature of commercial freelance software engineering work.
  • We developed SWE-Lancer, a benchmark built from more than 1,400 real freelance software engineering tasks at Expensify with up to $1 million in actual monetary rewards.
  • It includes two task types: individual-contributor (IC) SWE tasks, which require code patch generation, and SWE manager tasks, where the model must choose the best solution among multiple proposals.
  • We adopted a rigorous five-stage pipeline that uses end-to-end (E2E) tests written in Playwright and triple-validated by expert engineers to ensure robust, hack-resistant evaluation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)