DNS Failure Takes Down Spotify for Over Three Hours

Spotify experienced a significant service outage on January 14, lasting from approximately 00:15 UTC to 03:45 UTC. Impact escalated gradually over the first hour, eventually degrading most functionality, including music playback, before full restoration. The incident was triggered by routine maintenance on Spotify's internal Git repository (GitHub Enterprise, or GHE), which cascaded into a failure of the company's internal DNS resolvers.

Root Cause: A New Failure Mode in a Known Dependency

Spotify's internal DNS resolvers pull configuration from several sources, including the GHE repository. The team was aware of this dependency and had safeguards in place to handle outages in the Git service. However, a previously introduced change in the DNS component created a new, unforeseen failure mode. While GHE's maintenance downtime did not itself cause the outage, the process of bringing GHE back online triggered the application of invalid configurations across the resolver fleet over the course of roughly an hour. Each resolver that received the bad configuration entered a crash loop and became unable to answer internal DNS queries.

As DNS resolution failures proliferated, production traffic and internal services began to falter. The operational impact extended to employee VPN access and internal tooling, forcing the incident response team to improvise troubleshooting methods and prolonging triage time. The team was first alerted at 00:40 UTC, identified the root cause by 02:00 UTC, and had mitigation measures in place for both internal tooling and production traffic shortly thereafter.

Timeline of Events

Traffic to a core service during the incident (green) and one week earlier (blue).

All times are UTC.

  • 00:15 – DNS resolvers begin to fail.
  • 00:36 – Roughly 30% of the DNS resolver fleet is down; early signs of user impact appear.
  • 00:40 – Spotify platform engineers are alerted by automatic monitoring.
  • 01:15 – The entire DNS resolver fleet is down, causing severe service impact.
  • 02:00 – Root cause is identified.
  • 02:15 – Mitigation deployed for internal tooling.
  • 02:33 – Mitigation deployed for production traffic.
  • 03:45 – All core services are fully recovered.

Post-Incident Actions

Spotify has taken several steps to address the failure and reduce the likelihood of recurrence:

  • The underlying bug in the DNS component that enabled the invalid configuration has been fixed.
  • Changes have been made across multiple services and infrastructure layers to improve resilience against future DNS outages.
  • The failed DNS component is slated for decommissioning as part of an ongoing project to simplify Spotify's DNS infrastructure, a project that has now been reprioritized.