ChangeIntelIT change radar
Public mode

No tenant, Graph, or device access. Every item links to its source. What this means

Some sources or documents need attention: 215/216 feeds/APIs · 387/387 docs · synced 01:36 UTC Customize Public modeDiscuss ChangeIntel on Discord

Service status

What Microsoft, GitHub, Cloudflare, Ubiquiti, Cisco Meraki, Palo Alto Networks, GitLab, OpenAI, Anthropic, and Google post on their public status pages, read every 5 minutes while something is open. Each incident shows what else was happening and what changed in the same products just before.

Make this page yours

Choose the providers you depend on and the regions you run in, and what affects you comes first
Providers you depend onall providers
Providers you depend on

None chosen means all.

Regions you run inevery region
Regions you run in

Incidents elsewhere fold away; incidents that name no region always show.

Azure · Americas16
Azure · Europe20
Azure · Asia Pacific18
Azure · Middle East and Africa6
Azure Government6
Azure China6
Azure · Jio2
Azure DevOps8
Cloudflare datacenters6
Google Cloud2
On the overviewincidents that touch my providers and regions
On the overview
Changes apply as you make them.

Ended in the last 7 days

53 incidents · newest first

Wed 26 Aug1

  • GitHub Incident with Actions last seen 18:01 UTC unknown

    Posted 26 Aug 15:11 UTCUpdated 43d ago

    Actions

    On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. 3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs. Customers using concurrency groups saw longer impact due to a separate issue w…

    10 earlier updates
    1. 26 Aug 18:00 UTCUpdateAll inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.
    2. 26 Aug 17:54 UTCMonitoringThe degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
    3. 26 Aug 17:32 UTCUpdateWe are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.
    4. 26 Aug 16:50 UTCUpdateWe are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.
    5. 26 Aug 16:49 UTCUpdatePages is operating normally.
    6. 26 Aug 16:14 UTCUpdateWe believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.
    7. 26 Aug 15:48 UTCUpdateprimary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues
    8. 26 Aug 15:23 UTCUpdateWe've identified an issue with a database primary and are failing over to a replica immediately
    9. 26 Aug 15:12 UTCUpdatePages is experiencing degraded performance. We are continuing to investigate.
    10. 26 Aug 15:11 UTCInvestigatingWe are investigating reports of degraded availability for Actions
    Around it1 other incident at the same time · 5 earlier incidents

    Elsewhere at the same time

    Earlier incidents in the 14 days before

Yesterday9

Wed 7 Oct10

Tue 6 Oct17

Mon 5 Oct11

Fri 2 Oct6

Bars compare how long each incident lasted, up to three days; a lighter bar is a lower bound and a hatched one is unknown. Older incidents are kept for 90 days on the timeline.

Azure post-incident reviews

What went wrong in major incidents and what Microsoft is changing
ChangeIntel

An IT change radar: releases, security, known issues, retirements, documentation changes, and service status from public sources. Every item links to supporting evidence; dates and statuses can change after they are read.

Sources read 9 Oct 01:36 UTC · 215 of 216 readable · documentation 387/387 current