September Incidents: Actions, Pages, and Codespaces See Intermittent Disruptions

GitHub reported three separate incidents during September 2024 that caused service degradation across several core products. Each event was isolated, fully resolved, and followed by internal fixes to prevent recurrence.

Runner Misconfiguration Delays Actions and Pages Deployments

On September 16, degraded performance struck GitHub Actions and GitHub Pages, lasting 57 minutes from 21:11 UTC to 22:08 UTC. Customers deploying Pages from a source branch encountered delayed runs. The underlying cause was a misconfiguration in the service that manages runner connections, which led to CPU throttling and poor performance in that service.

The impact was significant: Actions jobs experienced average delays of 23 minutes, with some pushing past 45 minutes. At the peak, 80% of runs had delays exceeding five minutes, though over the whole incident that figure was 17%.

A quick mitigation began just five minutes after detection at 21:16 UTC by diverting runner connections from the faulty nodes. Beyond the immediate configuration fix, GitHub says it improved general monitoring to make automated detection and reaction faster in future.

Codespaces Hit by Network Connectivity and Capacity Issues

Two separate Codespaces incidents occurred later in the month. The first ran from 08:20 UTC to 09:04 UTC on September 24 (44 minutes total) and saw roughly a 25% error rate caused by interrupted network connectivity. The interruption traced back to Source Network Address Translation (SNAT) port exhaustion that surfaced after a deployment, breaking individual codespace connections to the service.

GitHub mitigated by increasing port allocations to handle the spike in outbound connections that follow deployments. The team intends to scale outbound connectivity and add stronger network-capacity monitoring as longer-term safeguards.

The second incident was region-specific. On September 30, from 10:43 UTC to 11:26 UTC (43 minutes), customers in Central India could not create new codespaces, though resumes were unaffected and other regions saw no impact. Storage capacity constraints were the culprit, and GitHub temporarily redirected create requests to other regions while adding storage to Central India.

An additional bug also came to light during the investigation: some available capacity was not being utilized, which artificially limited capacity and triggered premature creation halts. That bug has been fixed so capacity scales in line with projections.