AI & agents

GitHub open-sources ReviewBench, a benchmark for AI code review

Published
Source
GitHub blog

Summary

GitHub has launched ReviewBench, an open benchmark for comparing AI code review agents. Its corpus of 219 public pull requests across 19 languages mirrors the distribution of 103.9 million real GitHub pull requests, and its golden set combines human reviewers, frontier LLMs, and static analysis. According to GitHub, senior engineers independently agreed with the golden set in 96.6% of cases, and the benchmark helped offline evaluations of Copilot code review better anticipate production experiments.

Why it matters for our work

Teams choosing or building code review agents can now compare what each system catches and misses, with evidence. Open evaluation standards also tend to speed up agent improvement across the field.

Translated from the Korean original. Summaries may be translated and edited. Commentary reflects our perspective; forecasts remain the source’s views.

Read original (opens in a new tab)