January Double Header: Three Outages Degrade GitHub Services
GitHub reported three separate incidents in January that led to degraded performance across its services. Two of the outages were brief and contained, while one stretched on for several hours and impacted a significant portion of Codespaces users.
Git Backend Latency Spikes on January 9
The first incident began at 12:20 UTC on January 9 and lasted 140 minutes. During this window, services in one of GitHub's three sites experienced elevated connection latency, resulting in a sustained period of timed-out requests across multiple services, including the Git backend. At its peak, 10% of requests failed with a 5xx response or timed out, with an average failure rate of 5%.
The root cause was traced to a host upgrade that temporarily reduced capacity as it rolled through the fleet. While the hosts themselves had ample capacity to handle the load, a configuration issue with the connection limit—set lower than it should have been—caused the failures. GitHub has since increased that limit. The company also identified broader improvements to monitoring connection limits and plans to adjust procedures to reduce the risk that future host upgrades cause similar capacity shortfalls.
Codespaces Creation and Resume Failures on January 21
A more prolonged incident struck on January 21 at 02:01 UTC, affecting customers using GitHub Codespaces. Operational problems with compute and storage resources led to errors creating and resuming Codespaces across multiple regions. Roughly 25% of customers were impacted, primarily in East US and West Europe.
GitHub re-routed traffic for new Codespace creations to less affected regions, but existing Codespaces in the impacted areas may have been unable to resume during the outage. Connectivity was restored to all regions except West Europe by 7:30 UTC; that region's recovery was extended due to increased load. Full service—normal creation and resuming of Codespaces across all regions—was not achieved until 9:34 UTC, meaning the incident lasted 7 hours and 3 minutes total.
Going forward, GitHub says it is working to improve alerting and build more resiliency into the system to reduce the duration and blast radius of region-specific failures.
Load Balancer Change Causes IP Allow List Failures on January 31
The final incident of the month began at 12:30 UTC on January 31 and ran for 147 minutes. It stemmed from an infrastructure update to load balancers, part of a longer-term effort to enable IPv6 at GitHub.com. The change was deployed to a subset of the global edge sites and had an unintended side effect: IPv4 addresses began arriving as IPv4-mapped IPv6-compatible addresses (e.g., 10.1.2.3 became ::ffff:10.1.2.3). While the IP Allow List feature was built to handle IPv6, it was not designed to process these mapped addresses, so it began blocklisting requests it incorrectly determined were outside the allowed set. At the peak, 0.23% of all requests were failing.
GitHub deployed fixes to remediate the immediate issue and said it has since taken steps to improve testing and monitoring in order to catch similar problems earlier in the future.



