
GitHub's August 17 outage: 7 hours 47 minutes and a capacity failure, not a code change
GitHub has published an update on the August 17 outage that disrupted github.com, authentication, Actions, APIs and Copilot for 7 hours 47 minutes, calling it a capacity failure and promising faster scaling work.
On August 17 GitHub suffered an outage that lasted 7 hours and 47 minutes and disrupted github.com, authentication, GitHub Actions, APIs, pull requests, issues and Copilot for developers and organizations around the world.
It was the company's second significant incident in August, after an Actions failure on August 6. GitHub had already published reliability updates in March and April; the new incidents, it says, make clear that this work has to accelerate.
What happened
The investigation found that the outage began when traffic reached a new peak and a critical infrastructure component in the Central US data center failed to scale with it. The resulting capacity pressure spread through GitHub's systems, causing authentication failures and disrupting several services at once.
Recovery took several coordinated actions: teams rerouted traffic, isolated affected infrastructure and restored services in stages. Most services came back earlier that day, but some Copilot services took longer — errors there triggered a client-side retry loop that added traffic during recovery, which had to be mitigated before traffic could be safely restored. A detailed technical timeline is published in the root cause analysis.
Capacity, not code
GitHub says neither outage was caused by a code or configuration change: both were capacity failures, in which critical components were not scaled before demand exceeded their capacity. Since April, monthly commits on the platform have grown from 1.4 billion to 2.9 billion. The company has added more than 3 million CPU cores, 120 petabytes of high-speed storage and significant network capacity, installing as much hardware as available power allowed while accelerating its migration to Azure. Azure now serves roughly 58% of GitHub's platform load and half of all Git operations, up from 12% of platform load in May.
The work ahead
Its priorities are adding capacity, improving efficiency and removing architectural bottlenecks. The next milestone is an architecture that scales read capacity linearly with the number of readers, enabling unlimited read operations, to be rolled out gradually starting with the largest monorepos. GitHub is also isolating critical systems and removing shared dependencies between them, and investing in stronger testing, safer rollouts, better observability and more effective alerting.
Two immediate changes followed the August incidents: consistent retry limits, retry budgets and variable timeouts across service-to-service interactions to prevent retry storms and cascading load, and a review of lower-priority CPU and memory alerts to identify components that could fail during sudden traffic spikes.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.