Incident Reviews Are an Investment, Not a Cost

Last year, the team responsible for the reliability of Spotify for Artists (S4A) took a close look at every incident that affected the service in 2021. We formed hypotheses, then analyzed each incident to build a set of quantitative and qualitative metrics. The findings confirmed some instincts and challenged others, but the overall takeaway is that studying failure—while expensive in the short term—pays dividends down the road.

Every unplanned event carries a cost, but the most frequently overlooked one is productivity. In our analysis, 55% of incidents required at least one responder to spend the better part of a day on the problem. In 23% of cases, the productivity blast radius was even larger. That sounds like wasted effort on the surface, but the return comes when you invest in a proper incident review. Reading code and docs can tell you how a system should work; analyzing a real failure corrects your mental model and brings it closer to how the system actually behaves. That improved understanding is how you buy future productivity with today's downtime.

Most Failures Were Technically Avoidable

We scored every incident on a preventability scale from 1 ("almost impossible to prevent") to 5 ("we saw this coming and let it happen"). No incidents scored a 5, but most landed at 3 or above: localized failures that known, well-understood actions could have prevented or mitigated.

That doesn't mean prevention is free. But defining a service level objective (SLO) gives you an error budget. If you can stay within that budget without adding a preventative measure, you don't need one—you can spend that effort improving the user experience instead of defending against scenarios that aren't costing you reliability.

The Paperwork Problem

Two of the DORA "golden metrics" are Mean Time To Recover and Change Fail Rate. These are useful as a map, not a destination, so we reconstructed a clear timeline for every S4A incident. The data exposed a gap between intent and practice: around 50% of incidents had no recorded start and end times at all, and when responders did attempt to record them, we had to adjust the times by more than five minutes in 81% of cases.

This wasn't operator negligence—it was a system failure. We asked engineers to fill out paperwork without pairing it with any incentive to do so. Unsurprisingly, they spent their time on tasks with clearer rewards.

Synthetic Tests Slash Recovery Time

Recovery time varied widely across incidents, and we struggled to correlate it with any single characteristic of our systems. The one exception was synthetic testing. We evaluated whether a synthetic test could have plausibly detected each outage, then compared recovery times for incidents that were detected by tests versus those that weren't covered.

Incidents involving coverable features that had a synthetic test recovered 10 times faster than those without. That result was more striking than we expected, and it's reshaping our priorities. Synthetic testing is now a heavier focus for us because getting services back up quickly matters that much.

Treat Incidents as Data

The most valuable practice we've adopted is treating incidents as a source of real-world data—whether analyzed as a collection or examined individually in your own context. It's tempting to skip the investigation and get back to building, but there's no substitute for learning from what actually broke. Investing time in failure analysis is essential if you want to grow and sustain your service's uptime.