GitHub has published its explanation of the large-scale outage that hit the platform on August 17. The trigger was not a code or configuration change but infrastructure that failed to keep up with a record traffic peak. What dragged the recovery out was a chain of retries, as failed requests were mechanically thrown back at the platform.
Seven hours and 47 minutes of disruption
The incident ran from 13:28 to 21:15 UTC on August 17, lasting 7 hours and 47 minutes. Issues, pull requests, APIs, Actions, and Copilot were all affected. At peak, error rates reached roughly 20 percent for web experiences and API traffic, and roughly 50 percent for archive downloads and raw repository content.
Authentication took a deep hit as well. SAML and OIDC authentication, SCIM, and Team Sync were all disrupted. In GitHub Enterprise Cloud with data residency, Actions workflows that depend on public workflow step definitions hosted on GitHub.com stopped working.
Recovery came in stages. Most services returned around 16:36 UTC as the Central US data center recovered, Actions remained degraded until about 18:03 UTC, and the Copilot Token Service was fully restored at 21:02 UTC.
It started with an autoscaling policy for sidecars
The immediate cause was network saturation on load balancers in Central US. The path that led there is the most technically interesting part of the explanation.
It began with the sidecar pods that Istio, the service mesh, runs alongside each service. Those pods hit their concurrency limits but did not scale out, because the scaling policy watched metrics on the host service and not the sidecar's own limits. If the bottleneck forms somewhere the policy is not watching, autoscaling concludes that nothing is happening.
One choke point led to the next. Eventually four HAProxy nodes exhausted their flow limits, degrading the gateway authentication path and producing widespread authentication latency and failures. Optimistic retry logic, which simply resends a request when it fails, then piled additional load onto internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.
A retry storm that multiplied traffic tenfold
Some of the failing traffic was shifted from Central US to Northern Virginia. That move surfaced a different problem. Delayed replies from one internal endpoint triggered a latent retry bug in VS Code, amplifying traffic by roughly 10 times.
The numbers show the scale of it. Copilot Token Service traffic, normally 7,000 to 9,000 requests per second, spiked to between 70,000 and 100,000 requests per second. A single failed token operation could generate a large volume of extra requests and settle into a retry loop.
Containing it took two moves. GitHub first shipped a pull request that temporarily reduced gateway retry logic, then blocked inbound token requests at the load balancers with a 403 response before gradually ramping traffic back up site by site. Scraping attacks against codeload endpoints during the recovery window made the work harder still.
Monthly commits have doubled since April
GitHub says neither this outage nor the Actions failure on August 6 was caused by a code or configuration change, and that both were capacity failures at their core. In its own words, the company failed to scale critical components before demand exceeded their capacity.
The load figures explain the pressure. Monthly commits have grown from 1.4 billion to 2.9 billion since April. Against that, GitHub has added more than 3 million CPU cores, 120 petabytes of high-speed storage, and significant network capacity. It installed as much hardware as available power allowed in existing data centers while accelerating its migration to Azure.
Azure now serves roughly 58 percent of GitHub's platform load and half of all Git operations, up from 12 percent of platform load in May. That is a fast pace of migration. The next milestone is an architecture that scales read capacity linearly with the number of readers, rolled out gradually beginning with the largest monorepos.
Remediation steps, and the work still outstanding
The published follow-up actions map closely onto each point of failure: correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity, auditing Istio request, concurrency, and scaling limits across affected services, reviewing retry limits and backoff behavior across gateways and clients, addressing the VS Code behavior that amplified Copilot token traffic, and improving load-balancer capacity monitoring and regional failover safeguards.
On top of that, GitHub is applying consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions. Think of it as a shared brake against retry storms and cascading load. The company is also reviewing lower-priority CPU and memory alerts to identify components that could fail during sudden traffic spikes.
GitHub CTO Vlad Fedorov writes that scale is not the only challenge. As the pace and complexity of change increased, existing operational practices did not keep up, and while teams and resources have been redirected toward stronger testing, safer rollouts, better observability, and more effective alerting, the work is not complete. GitHub is also isolating critical systems and removing shared dependencies between them.
Summary
The August 17 GitHub outage lasted 7 hours and 47 minutes and affected Issues, pull requests, APIs, Actions, and Copilot. The cause was not a code or configuration change but a capacity shortfall that began when Istio sidecar autoscaling failed to engage, cascading into HAProxy saturation and degraded authentication. The delayed recovery was driven largely by a VS Code retry bug that amplified traffic roughly tenfold. Behind it all is a load curve that took monthly commits from 1.4 billion in April to 2.9 billion, and GitHub is responding with added CPU cores and storage, an accelerated Azure migration, and unified retry controls.
