ChangeIntelIT change radar
Public mode

No tenant, Graph, or device access. Every item links to its source. What this means

Some sources or documents need attention: 215/216 feeds/APIs · 387/387 docs · synced 00:35 UTC Customize Public modeDiscuss ChangeIntel on Discord

Service status

What Microsoft, GitHub, Cloudflare, Ubiquiti, Cisco Meraki, Palo Alto Networks, GitLab, OpenAI, Anthropic, and Google post on their public status pages, read every 5 minutes while something is open. Each incident shows what else was happening and what changed in the same products just before.

Make this page yours

Choose the providers you depend on and the regions you run in, and what affects you comes first
Providers you depend onall providers
Providers you depend on

None chosen means all.

Regions you run inevery region
Regions you run in

Incidents elsewhere fold away; incidents that name no region always show.

Azure · Americas16
Azure · Europe20
Azure · Asia Pacific18
Azure · Middle East and Africa6
Azure Government6
Azure China6
Azure · Jio2
Azure DevOps8
Cloudflare datacenters6
Google Cloud2
On the overviewincidents that touch my providers and regions
On the overview
Changes apply as you make them.

Ended in the last 7 days

53 incidents · newest first

Thu 24 Sep1

  • GitHub Incident across several services last seen 04:55 UTC unknown

    Posted 23 Sep 10:11 UTCUpdated 14d ago

    GitHub

    Starting at 07:57 UTC on September 23, GitHub experienced elevated 500 and 404 responses across several application pages. This caused failures when installing GitHub Apps, creating organizations, and making some organization membership changes. Customers also experienced delayed label updates and stale search results in Projects. The infrastructure failure was isolated to our Azure Central US region. The API errors were mitigated by 10:58 UTC on September 23. Projects' processing continued to recover while an accumulated backlog was drained, and full service was restored at 04:55 UTC on September 24. The incident was caused by a failed planned maintenance operation on a primary database. Automated recovery initiated an emergency database failover, after which several replicas in the affected region were unable to resume replication correctly. This reduced available database capacity and caused the API errors and downstream Projects processing delays. We have mitigated the immediate failure mode. We are also improving maintenance safety checks, database failover handling, post-failover replica validation, and downstream processing resilience to reduce the likelihood and impact of similar incidents.

    15 earlier updates
    1. 24 Sep 04:55 UTCMonitoringThe degradation has been mitigated. We are monitoring to ensure stability.
    2. 24 Sep 03:14 UTCUpdateWe are continuing to process the backlog of issue label updates for Projects. Label changes may still be delayed. All other services are operating normally.
    3. 24 Sep 02:00 UTCUpdateWe are continuing to process the backlog of issue label updates for Projects. Users may still see delays before label changes are reflected in Projects. All other services are operating normally.
    4. 24 Sep 00:16 UTCUpdateWe've deployed a change intended to accelerate processing of the backlog of issue label updates in Projects. A sizable backlog still remains and we continue working through it. All other services are operating normally. We will provide another update within the next hour.
    5. 23 Sep 21:39 UTCUpdateUpdates to issue labels may be delayed in being reflected in Projects. We are continuing to deploy a change that will accelerate processing of the backlog of label updates. All other services are available. We will provide another update within the next hour.
    6. 23 Sep 20:26 UTCUpdateWe are preparing to deploy a change that will mitigate the impact.
    7. 23 Sep 18:42 UTCUpdateContinuing to investigate the lag that may be experienced in issue labels being accurately reflected in Projects. We are working on alternate solutions to process the backlog of label updates.
    8. 23 Sep 17:30 UTCUpdateWe will post another update in approximately one hour to share our progress.
    9. 23 Sep 17:01 UTCUpdateUpdates to issue labels may be delayed in being reflected in Projects by about ~10 minutes. We have added some capacity to work through the backlog more quickly, but it'll likely be a few hours to complete processing the full backlog of messages. All other services are available.
    10. 23 Sep 13:35 UTCUpdateUsers may experience stale Project search results. We are working to increase indexing speed. All other services are available.
    11. 23 Sep 11:38 UTCUpdateWe are seeing recovery for Projects. Users may experience stale search results for Projects while indexing catches up.
    12. 23 Sep 10:58 UTCUpdateThe degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.
    13. 23 Sep 10:57 UTCUpdateDatabase replicas have been restored. Org creation and the GitHub API are no longer degraded.
    14. 23 Sep 10:20 UTCUpdateDatabase replicas have detached. We're working to restore the database replicas. Users may experience issues beyond creating organizations and a degraded experience with the GitHub API and Projects.
    15. 23 Sep 10:11 UTCInvestigatingWe are investigating reports of degraded performance for API Requests
    Around it6 other incidents at the same time · 5 earlier incidents

    Elsewhere at the same time · within 12 hours of its start

    Earlier incidents in the 14 days before

Yesterday9

Wed 7 Oct10

Tue 6 Oct17

Mon 5 Oct11

Fri 2 Oct6

Bars compare how long each incident lasted, up to three days; a lighter bar is a lower bound and a hatched one is unknown. Older incidents are kept for 90 days on the timeline.

Azure post-incident reviews

What went wrong in major incidents and what Microsoft is changing
ChangeIntel

An IT change radar: releases, security, known issues, retirements, documentation changes, and service status from public sources. Every item links to supporting evidence; dates and statuses can change after they are read.

Sources read 9 Oct 00:35 UTC · 215 of 216 readable · documentation 387/387 current