Provider-report summary
Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally
Google: some customers' monitoring on Google Cloud also failed, leaving them without a signal
Google's fixes: fail open, audit globally replicated data, feature-flag critical changes, enforce backoff, improve external communications
Three clocks, not one
Google's mini report states the duration as 3 hours
All Google Cloud products recovered by 01:27 UTC on 13 June (Vertex AI last)
Customer impact ran 17:51–21:10 UTC (3 h 19 min): errors from 17:51, alert 17:52, baseline 21:10
How the story changed
Then: Root cause
Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally
Then: A BGP routing failure is behind the outages
Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally
Then: Google has identified the root cause and applied mitigations, without disclosing the cause
Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally
Then: The outage is a result of the ConnectWise compromise
Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally
Now: Google's mini report states the duration as 3 hours
Now: All Google Cloud products recovered by 01:27 UTC on 13 June (Vertex AI last)
Now: Google: some customers' monitoring on Google Cloud also failed, leaving them without a signal
Now: Cloudflare: its outage was not the result of an attack
Now: Google's fixes: fail open, audit globally replicated data, feature-flag critical changes, enforce backoff, improve external communications
Statement checks
7/7 still hold. Nothing in this retrospective needs correcting.
Customer impact lasted 3 hours 19 minutes, from 17:51 to 21:10 UTC.
We detected it at 17:52, 54 minutes before Google's first public update.
Google's report gives a 3-hour duration, but some Google products took until 01:27 UTC.
A holding statement went out at 17:56.
We run in a single Google Cloud region, so we couldn't fail over.
Staff login runs on Cloudflare Access, which failed at the same time.
Approve budget for a second region for login and API.
Our actions
Budget for a second region for login and API
Incident lead · 13 Jun 09:00 UTC · Open
Move our external monitoring off Google Cloud
Engineering lead · 13 Jun 09:00 UTC
Add randomised backoff to client retries
Engineering lead · 13 Jun 09:00 UTC