LLM-Based Automated Diagnosis Of Integration Test Failures At Google
- Published
- Source
- arXiv
- Paper number
- 148
- Field
- Agents / Coding / Production
- arXiv ID
- 2604.12108
Key points
- Diagnosing integration test failures is highly time-consuming, often taking hours or days, and imposes a substantial cognitive burden on developers.
- Modern distributed systems produce huge volumes of logs, averaging 11,058 lines across 26 files per failure from many distributed, dynamically named sources, which makes manual inspection difficult.
- Logs are often semi-structured and heterogeneous, with a low signal-to-noise ratio, so many warnings and errors unrelated to the real root cause obscure the key information.
- The authors implemented Auto-Diagnose, an LLM-centered automation pipeline that collects and aggregates logs from the test driver and system under test (SUT) components when an integration test fails.
- It uses Gemini 2.5 Flash together with carefully designed prompt templates that include step-by-step guided reasoning, strict negative constraints to prevent guessing, and a precise output format for diagnosis.
- The LLM output is post-processed into readable Markdown and the identified relevant log lines are made clickable, integrating diagnosis directly into Google's internal code review system, Critique, for a smooth developer workflow.
Paper links
External research summaries. These are not HDATF publications or measured product results.