Outage Report: Global Service Disruption on April 16, 2025

On April 16, 2025, between 12:20 and 15:45 UTC, Spotify experienced a worldwide outage. The incident disrupted traffic for the majority of users, with the exception of the Asia Pacific region, which remained unaffected due to lower traffic volumes at the time. Successful requests on the perimeter dropped sharply across most regions, while the Asia Pacific regional line remained steady.

Root Cause Analysis

Spotify's networking perimeter is built on Envoy Proxy, extended with custom filters developed in-house, including one for rate limiting. On the day of the incident, the team changed the order of Envoy filters. The change was classified as low risk and was therefore applied simultaneously across all regions.

That reordering triggered a bug in one of the custom filters, causing Envoy to crash. Unlike typical isolated failures, this crash occurred on every Envoy instance at once. The immediate consequence was a mass restart of all perimeter instances, which, combined with client-side retry logic, produced an unprecedented load spike.

The sudden traffic surge then highlighted a misconfiguration. Envoy's max heap size was set higher than the Kubernetes memory limit. When a new Envoy instance started, it immediately received a large volume of traffic, exceeded the allowed Kubernetes memory limit, and was shut down by Kubernetes. The cycle repeated continuously, preventing the perimeter from stabilizing.

The Asia Pacific region escaped the incident because its traffic at that time of day was significantly lower. Regional Envoy memory usage never reached the Kubernetes limit, so the crash-restart loop did not take hold there.

Timeline of Events

  • 12:18 UTC — Envoy filter order changed; all Envoy instances crash.
  • 12:20 UTC — Alarms triggered indicating a significant drop in incoming traffic.
  • 12:28 UTC — Situation escalated; no traffic worldwide except Asia Pacific.
  • 14:20 UTC — Traffic fully recovered in European regions.
  • 15:10 UTC — Traffic fully recovered in US regions.
  • 15:40 UTC — All traffic patterns back to normal.

Mitigation and Remediation

Mitigation was achieved by increasing total perimeter server capacity. This allowed Envoy servers to operate below the Kubernetes memory limits, which stopped the continuous server cycling and restored service.

The engineering team has since taken corrective action on several fronts:

  • Fixed the bug that caused Envoy to crash.
  • Fixed the configuration mismatch between Envoy heap size and Kubernetes memory limits.
  • Improving the rollout procedure for perimeter configuration changes.
  • Enhancing monitoring capabilities to detect these issues earlier.

Spotify states it will continue to publish incident reports on similar occasions to maintain transparency and support ongoing improvements to its services.