When Slack Goes Down: Anatomy of a Major Outage
On May 12th, 2020, Slack experienced a total service disruption that began at 4:45 PM Pacific. At peak impact, the company found itself unable to use its own communication platform to coordinate the response. This is the story of how the incident was handled—and the incident response machinery, both human and automated, that Slack has built to manage such crises.
Two Incidents in One Day
Earlier that morning, during Slack's normal deployment cycle, a set of pull requests pushed to production triggered a surge in requests against the Vitess database tier that stores user data. Customers began experiencing sporadic errors. The API service automatically scaled up by 80% to absorb the unexpected demand. Within 13 minutes, the problematic code was rolled back and host counts were returned to normal.
The day appeared to resume as usual—until 4:45 PM, when a second, more serious problem emerged.
A Process Built for Scale
Slack's current incident response model stands in stark contrast to its earlier approach. Historically, a single operations rotation called AppOps was responsible for every page and every fix. As Slack grew, that model became unsustainable: a limited group of engineers could not maintain operational awareness of an ever-expanding service portfolio or withstand the toil of constant high-stress paging.
In 2018, Reliability Engineering replaced that structure with a process modeled on FEMA's Incident Command System. Development teams now own the on-call rotations for their own services. A separate Major Incident Command, staffed by trained facilitators, coordinates response efforts. Incident Commanders are deliberately not hands-on-keyboard; their role is facilitation, information gathering, resource coordination, and decision-making under pressure.
An incident moves through a defined lifecycle, from detection through assembly, verification, dispatch, mitigation, repair, and finally review. The Incident Commander works from Conditions, Actions, and Needs (CAN) reports to dispatch the right subject matter experts and keep the response focused.
4:45 PM: The Outage Begins
The trouble actually started 19 minutes before the full outage. During routine deploy operations, engineers noticed an uptick of errors in deployment dashboards. Alerts fired, and someone invoked the custom Incident Bot with an /assemble slash command, paging critical responders in Engineering and Customer Experience via PagerDuty.
At the time, responders from the edge and load balancing team and the web infrastructure team were still gathering signals, unsure they were heading toward a full outage. Customer Experience reported incoming tickets describing real-time customer problems. All signs converged on the main web service tier serving Slack's API. Load balancers were timing out on health checks and removing API instances from service. Instrumentation showed web workers heavily loaded and response times climbing.
Chaos and Coordination
At 4:51 PM, a responder declared they were about to go offline: all main_www instances in the webapp pool were lost. The moment marked the transition to full crisis mode.
Because Slack itself was unreliable, the response followed prescribed alternative channels. The team moved to an incident Zoom meeting. The severity was elevated to Sev-1, and executives were paged. A company-wide email with links to department-specific runbooks for full site outages was sent. Real-time client instrumentation, which sends telemetry to an isolated endpoint, confirmed that success rates had dropped to abysmal levels.
The Zoom call filled with engineers from multiple teams, scribes recording the conversation for later dissemination. The Incident Commander established multiple parallel workstreams, assigning each to investigate a different hypothesis. The goal was to balance speed against risk in restoring service. The flurry of activity continued until a resourceful engineer found the right answer.
Root Cause and Repair
By 5:17 PM, investigation pointed to the configuration for Slack's HAProxy load balancing tiers. Slack uses a combination of consul-template and a bespoke HAProxy management service to maintain the manifest of active backend servers. The manifest contained servers that shouldn't have been there. It became clear that linting errors had prevented HAProxy configuration files from being re-rendered, leaving them stale.
An engineer proposed stopping HAProxy, deleting its state file, restarting consul-template to refresh the manifest, and then restarting HAProxy. The hypothesis was tested on a control portion of servers, which quickly returned to healthy status and began receiving traffic. With the fix confirmed, automation safely executed the command sequence across availability zones.
Under Control, Then All Clear
At 5:33 PM, 48 minutes after the outage began, remediation succeeded and service was restored. Customer Experience reported a wave of happy customer reports. The Incident Commander declared "Under Control," signaling that the incident was contained and the team could begin working toward the final state.
Declaring All Clear required more than observing green metrics. The Incident Commander reviewed current signals and posed pointed questions: What are the risks of recurrence? Is appropriate alerting in place? The all-clear work also uncovered an unexpected side effect: the outage had caused a regression in Slack's desktop client that required a restart to restore service. Frontend Foundations engineers consulted with Customer Experience on the workaround, and the team notified affected customers.
Once most customer tickets were resolved, metrics showed stability, and engineers agreed recurrence was unlikely, the Incident Commander called All Clear.
Learning After the All Clear
The All Clear marks the end of the response, but not the end of the process. The cornerstone of Slack's incident response is the feedback loop: failure, remediation, learning, and repair. An incident review provides a psychologically safe space for engineers to discuss what happened. Action items are then created, assigned to teams, prioritized, and completed on schedule.
A successful incident response process converts a single painful event into durable institutional knowledge. It deepens understanding of how complex systems fail, drives down technical debt and change risk, and ultimately reduces the blast radius of the next incident. For an organization that is itself a critical tool for millions of users, that continuous improvement is not optional—it is the point.



