Back
GitHub Actions and Pages Hit by Hours of Degraded Availability
SiTech Team2 წთ. საკითხავი

GitHub Actions and Pages Hit by Hours of Degraded Availability

A routine internal deployment cascaded into failed and delayed workflow runs on August 6; at peak, 71 percent of Actions runs hit infrastructure failures, GitHub's postmortem says.

GitHub Actions was degraded for most of August 6, 2026, with workflow runs failing or waiting in queues for extended periods, and the impact spilling over to GitHub Pages, Copilot features and enterprise migration tooling. GitHub's status page logged the incident from 15:05 UTC on August 6 until 00:14 UTC the following day.

What went wrong

According to GitHub's postmortem, the trigger was a routine deployment to an internal Actions service that processes events and generates jobs. The deployment exposed a pre-existing capacity and concurrency weakness: as pods were replaced, the remaining capacity became saturated, services crashed, and the failure cascaded across multiple clusters and downstream services. At peak, 71 percent of workflow runs experienced infrastructure failures, and 75 percent of the remaining runs were delayed by more than five minutes. Both GitHub-hosted and self-hosted runners were affected.

A second bug and hours of throttling

Engineers expanded capacity, throttled webhook-triggered work so the system could recover, and added processing power for the backlog; the core services came back around 17:00 UTC. By then a second problem had surfaced: a latent bug in a service that assigns jobs to runners meant runners were handed jobs that were no longer valid and got stuck retrying them. At the low point, roughly 30 to 40 percent of queued jobs were succeeding; after fixes were deployed, success rates for starting workflows climbed to 97 percent and then 99 percent, and the accumulated queues drained overnight.

The aftermath

Recovery left edges to clean up. Some Actions Runner Controller pods stayed stuck in an idle state and needed manual recovery; a mitigation deployed mid-incident had to be rolled back, and automatic recovery is promised in upcoming Runner and ARC releases. GitHub also warned that push and pull request events dropped during the outage cannot be replayed automatically, so some workflows must be re-triggered by hand. To prevent a repeat, the company says it is improving deployment and capacity safeguards, monitoring for the conditions that preceded the incident, and the resilience of queued work and runner assignment.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.