Retail / Fashion e-commerce

Zalando: Analysing incident reports and setting investment priorities

Company
Zalando
Country
Germany
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

Zalando has kept a postmortem document after each incident and relied on people reading them to learn. Reading one properly took 15 to 20 minutes. Even with full concentration, about four a hour was the limit for one person. Meanwhile the archive had grown to thousands of documents, so finding repeated causes across teams by human effort alone was difficult.

Technology and data

The analysis procedure was split into four stages: summarising, technology classification, per-incident analysis and extraction of overall patterns. The input is the postmortem archive accumulated internally. At the classification stage a technology list is supplied alongside, and the model is restricted to returning only technology names with a confirmed direct link in the document. At first an open source model on LM Studio was used; now it uses Claude Sonnet 4 on AWS Bedrock. While the procedure was being built, 100 percent of output batches were reviewed by people; once the system settled, 10 to 20 percent of each batch is sampled at random for review.

Results

Even in the period when a map-fold structure was used, a full year of analysis could already be finished within 24 hours. The latest version processes one document in about 30 seconds with Claude Sonnet 4. Insights from the analysis led to automated change validation for infrastructure code, and Zalando said this measure prevents 25 percent of subsequent datastore incidents. The repeated failure types it identified were the absence of automated change validation, inconsistent change management, the absence of gradual rollout, underestimating traffic volume, and scaling up later than demand.

Limits and open questions

Zalando said that even with recent models such as Claude Sonnet 4, about 10 percent of cause attributions remain wrong because they lean on surface clues. It wrote that it has not achieved the accuracy needed to extract figures such as GMV or EBIT loss from postmortems.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)