Delivery and mobility / Ride-hailing and delivery

Uber: Running AI agents across the software development lifecycle and managing their cost

Company
Uber
Country
United States
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

Uber said that between February and August 2026, weekly active users grew sevenfold and weekly agentic requests grew 9.4 times, so it had to control rising cost while keeping quality. With over 100 tools installed, standard MCP preloading added roughly 50,000 to 70,000 tokens of tool schemas to the initial prompt, and that overhead was sent again on every context turn. An agent without internal context could not see the table it needed and spent 20 minutes inspecting service code and spawning two subagents.

Technology and data

Uber said managed agents that are not started by humans handle code review, automatic fixing of failed CI runs, completing end to end pull requests with visual validation, triaging on call alerts, debugging and maintenance, while humans review the results and take escalations. Models are chosen with benchmarks built from real work, and subagents with well specified inputs default to a cheaper model. The prompt cache TTL for interactive sessions was raised from 5 minutes to 1 hour, and automatic compaction starts at 400,000 tokens. A gateway in front of more than 1,000 MCP servers handles authentication and policy centrally. An AI context graph holds 24 million nodes and 80 million edges drawn from more than 30 internal systems. Cost is made visible through a live cost counter in the status line, Slack nudges at 50%, 80% and 100% of expected spend, manager sign off for tier upgrades and a dashboard that flags 16 wasteful usage patterns.

Results

Uber said more than 70% of pull requests are attributed to local or cloud agents. Engineers have built more than 3,600 skills, which run more than 30,000 times a day. Total spend has stayed relatively stable since April. Cost per 1,000 model requests fell almost 34% from its peak between February and July, and cost per session fell 52% from its June peak. The company said usage grew sevenfold while unit costs fell and quality held. With internal context, an agent queried historical usage, found the table used by over 50 analysts and answered in 38 seconds.

Limits and open questions

The 70% figure counts pull requests attributed to agents, not pull requests merged without human review, and no bug or quality metrics are given. Uber notes that its cost figures are unique to its environment, and most cover February to July. No adoption month is given, so the record is ordered by the publication date.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)