SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
- Published
- Source
- arXiv
- Paper number
- 035
- Field
- Benchmarks / Code Agents
- arXiv ID
- 2502.12115
Key points
- Existing LLM coding benchmarks rely on unit tests, which makes them vulnerable to grader hacking and insufficient for validating complex full-stack software engineering tasks.
- Earlier evaluations lack direct economic metrics, so they cannot quantify the real monetary value of LLM outputs in realistic software development scenarios.
- Current benchmarks are often too narrow, focusing on isolated programming problems or specific repositories rather than the diverse, ambiguous, and end-to-end nature of commercial freelance software engineering work.
- We developed SWE-Lancer, a benchmark built from more than 1,400 real freelance software engineering tasks at Expensify with up to $1 million in actual monetary rewards.
- It includes two task types: individual-contributor (IC) SWE tasks, which require code patch generation, and SWE manager tasks, where the model must choose the best solution among multiple proposals.
- We adopted a rigorous five-stage pipeline that uses end-to-end (E2E) tests written in Playwright and triple-validated by expert engineers to ensure robust, hack-resistant evaluation.
Paper links
External research summaries. These are not HDATF publications or measured product results.