Information technology / Cloud infrastructure

Cloudflare: Reviewing code and design documents against internal standards with AI

Company
Cloudflare
Country
United States
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

Cloudflare manages its engineering standards as a set of RFCs called the Codex. Requirements are written with the SHOULD and MUST keywords defined by RFC 2119. The source said that with more than 60 RFCs already and counting, feeding the entire Codex to a large language model as is would put a lot of stress on the context window and hurt the results. Having people find every standards violation across all merge requests and design documents was also hard because of the volume.

Technology and data

Rather than feeding in the whole Codex, the company built a structure that selects only the RFCs relevant to what is being reviewed. The AI code reviewer reads merge requests and flags standards violations. Findings from approved RFCs are reported without blocking, while MUST violations under enforced RFCs cause approval to be withheld. TypeScript was the first language to receive Codex linter support and standardised on oxlint for execution speed. Rust is under development and Go follows. The spec reviewer runs on the Developer Platform. It runs as a Cloudflare Worker, stores results and state in D1, routes model requests through AI Gateway, and starts scanning for new specs through a Cron Trigger. The incident report reviewer looks for gaps such as missing follow up action items, incomplete timelines and omitted detection signals. Developers can also run the same review in their own environment through a CLI on OpenCode based agents.

Results

The company said that since the Codex began earlier this year the AI code reviewer has flagged close to 230,000 standards violations, of which almost 16,000 led to approval being withheld. It said the spec reviewer has reviewed almost 600 unique open specs since the beginning of May 2026, and that counting reruns triggered by spec changes or on demand, review invocations passed 3,200. It said the severity of findings was major for 65% and minor for 29%, with critical the smallest group at 6%. It said the incident report reviewer has assessed more than 200 reports since May 2026, and that 93% of them covered incidents that were low impact, internal only or declared preemptively.

Limits and open questions

The number of flagged violations and withheld approvals is given, but there is no figure for the false flag rate or for how often developers accepted the findings. The source also gives no figures for how incident counts changed after the review tools were introduced.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)