Two Job Queuing Incidents Degrade GitHub Services in February
GitHub experienced two service disruptions in February, both stemming from its background job queuing infrastructure. The first incident occurred on February 26 and lasted 53 minutes, starting at 18:34 UTC. The second, on February 29, began at 09:32 UTC and extended for 142 minutes, with the bulk of user-facing delays concentrated in a 22-minute window between 11:05 and 11:27 UTC.
February 26: Capacity Constraints and Failover Failure
The February 26 incident was triggered by capacity limits in the job queuing service combined with a malfunctioning automated failover system. This led to noticeable delays across Webhooks, GitHub Actions, and UI updates, such as those on pull request pages. GitHub engineers manually redirected traffic to the secondary cluster to resolve the issue, with no data lost in the process.
February 29: Restoration Misstep Compounds Delays
The second incident also impaired Webhooks, GitHub Actions, and GitHub Issues. While automated failover initially directed traffic correctly at 09:32 UTC, an improper switch back to the primary cluster at 10:32 UTC caused queued jobs to pile up sharply. The error was corrected at 11:21 UTC, allowing healthy services to process the backlog until full restoration was achieved by 11:27 UTC.
Response and Remediation
In response to both events, GitHub has already deployed three short-term fixes aimed at improving automation, strengthening the reliability of the fallback process, and expanding capacity for background job queuing services. A longer-term initiative focused on enhancing the overall scalability and resilience of the job processing platform is also underway.



