Provider-report summary

CONFIRMED

Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally

CONFIRMED

Google: some customers' monitoring on Google Cloud also failed, leaving them without a signal

CONFIRMED

Google's fixes: fail open, audit globally replicated data, feature-flag critical changes, enforce backoff, improve external communications

Three clocks, not one

CONFIRMED

Google's mini report states the duration as 3 hours

CONFIRMED

All Google Cloud products recovered by 01:27 UTC on 13 June (Vertex AI last)

CONFIRMED

Customer impact ran 17:51–21:10 UTC (3 h 19 min): errors from 17:51, alert 17:52, baseline 21:10

How the story changed

CONFIRMEDUnknown → Confirmed

Then: Root cause

Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally

UNCONFIRMEDMerged into C-003

Then: A BGP routing failure is behind the outages

Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally

CONFIRMED
CONFIRMEDMerged into C-003

Then: Google has identified the root cause and applied mitigations, without disclosing the cause

Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally

CONFIRMED
UNCONFIRMEDMerged into C-003

Then: The outage is a result of the ConnectWise compromise

Now: Root cause: a new Service Control quota-policy feature had no error handling and no feature flag; a policy change with blank fields crashed it globally

CONFIRMED
CONFIRMEDNew · Confirmed

Now: Google's mini report states the duration as 3 hours

CONFIRMEDNew · Confirmed

Now: All Google Cloud products recovered by 01:27 UTC on 13 June (Vertex AI last)

CONFIRMEDNew · Confirmed

Now: Google: some customers' monitoring on Google Cloud also failed, leaving them without a signal

CONFIRMEDNew · Confirmed

Now: Cloudflare: its outage was not the result of an attack

CONFIRMEDNew · Confirmed

Now: Google's fixes: fail open, audit globally replicated data, feature-flag critical changes, enforce backoff, improve external communications

Statement checks

7/7 still hold. Nothing in this retrospective needs correcting.

✓ Still holds

Customer impact lasted 3 hours 19 minutes, from 17:51 to 21:10 UTC.

CONFIRMED
✓ Still holds

We detected it at 17:52, 54 minutes before Google's first public update.

CONFIRMEDCONFIRMED
✓ Still holds

Google's report gives a 3-hour duration, but some Google products took until 01:27 UTC.

CONFIRMEDCONFIRMED
✓ Still holds

A holding statement went out at 17:56.

✓ Still holds

We run in a single Google Cloud region, so we couldn't fail over.

CONFIRMED
✓ Still holds

Staff login runs on Cloudflare Access, which failed at the same time.

CONFIRMEDCONFIRMED
✓ Still holds

Approve budget for a second region for login and API.

Our actions

Budget for a second region for login and API

Incident lead · 13 Jun 09:00 UTC · Open

Move our external monitoring off Google Cloud

Engineering lead · 13 Jun 09:00 UTC

Add randomised backoff to client retries

Engineering lead · 13 Jun 09:00 UTC