Delivery and mobility / Ride-hailing and delivery

Uber: Reviewing software code

Company
Uber
Country
United States
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

As AI began helping to write code, the volume of code to review grew. Reviewers did not have enough time to find subtle bugs and security problems and to keep to internal rules. Uber saw this limit leading to missed errors, production incidents and late releases.

Technology and data

At the heart of uReview is Commenter, a generative AI review system split into several stages. It builds a prompt that puts the changed code together with surrounding functions, class definitions and import statements. Review runs in three branches. The Standard Assistant looks for bugs, wrong exception handling and logic flaws. The Best Practices Assistant refers to a shared style rule repository and checks internal coding rules. The AppSec Assistant looks at application-level security vulnerabilities. A separate prompt evaluates the quality of the comments produced, assigns a confidence score and merges overlapping comments. Among the models, Anthropic Claude-4-Sonnet was best at generating comments and OpenAI o4-mini-high at scoring them. Developers mark each comment Useful or Not Useful and leave a note.

Results

Uber said uReview analyses more than 90 percent of roughly 65,000 diffs a week. It processes more than 10,000 commits each week. Engineers who used the tool marked 75 percent of the comments as useful, and more than 65 percent of the comments posted were acted on. Uber calculated this as about 1,500 hours saved per week and explained it comes close to 39 developer years annually. In an internal audit, only 51 percent of comments written by people were accepted by the author as bugs and fixed in the same change.

Limits and open questions

uReview can only see the code. It has no access to past PRs, feature flag settings, database schemas or technical documents. It cannot judge whether the overall design is right.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)