LLM-Based Automated Diagnosis Of Integration Test Failures At Google

Published
Source
arXiv
Paper number
148
Field
Agents / Coding / Production
arXiv ID
2604.12108

Key points

  • Diagnosing integration test failures is highly time-consuming, often taking hours or days, and imposes a substantial cognitive burden on developers.
  • Modern distributed systems produce huge volumes of logs, averaging 11,058 lines across 26 files per failure from many distributed, dynamically named sources, which makes manual inspection difficult.
  • Logs are often semi-structured and heterogeneous, with a low signal-to-noise ratio, so many warnings and errors unrelated to the real root cause obscure the key information.
  • The authors implemented Auto-Diagnose, an LLM-centered automation pipeline that collects and aggregates logs from the test driver and system under test (SUT) components when an integration test fails.
  • It uses Gemini 2.5 Flash together with carefully designed prompt templates that include step-by-step guided reasoning, strict negative constraints to prevent guessing, and a precise output format for diagnosis.
  • The LLM output is post-processed into readable Markdown and the identified relevant log lines are made clickable, integrating diagnosis directly into Google's internal code review system, Critique, for a smooth developer workflow.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)