1. Causes
The main cause was a failure in Google Cloud’s infrastructure that affected cold data storage. Google Cloud uses a system called “Service Control” to manage APIs. It handles functions like request authorization, policy and quota checks, and global metadata replication. According to Google’s official report, on May 29, 2025, a new feature was added to Service Control to perform additional quota checks. However, this feature had a critical vulnerability: it lacked proper error handling for cases where policy data was empty. Moreover, the code was not protected by “feature flags,” which would allow for gradual rollout during testing. During deployment, the bug did not appear because the code path leading to failure wasn’t activated. It was only discovered on June 12 at 10:45 AM PDT when a policy change containing unexpected empty fields was written into the global Spanner database. These data instantly replicated across all regions. When Service Control tried to process the empty fields, a null pointer exception occurred, crashing binaries in every region. Restarting Service Control in large regions (e.g., us-central1) triggered a surge of simultaneous requests to Spanner, which overloaded the infrastructure.
Unexpected Dependency: Cloudflare’s Reliance on Google Cloud
Cloudflare, known for its self-owned infrastructure, was surprisingly dependent on Google Cloud. Its Worker KV service — critical for many other products — used Google Cloud for cold storage. This meant that rarely-used data was stored on Google servers, and when their infrastructure failed, Worker KV lost access to that data. Until this incident, many believed Cloudflare was entirely autonomous. Even AI systems analyzing the situation claimed Cloudflare did not use Google Cloud.
2. Domino Effect
Direct Effect: The Google Cloud failure led to the shutdown of services that directly relied on its APIs and resources. While the risk of a single service failing is expected, the scale of impact here was critical due to the centralized nature of the Service Control failure. Since all Google cloud services depend on this system to process requests, the outage quickly triggered cascading consequences. Within minutes, platforms began returning HTTP 503 errors or became completely unavailable, halting authentication, API calls, and dependent services worldwide.
The impact extended far beyond Google’s infrastructure. Services built on Google Cloud faced timeouts, complete failures, or performance degradation, affecting entire ecosystems — from development environments and IoT dashboards to consumer apps like Twitch, Shopify, Discord, and many other smaller projects and organizations partially or fully reliant on GCP.
Indirect Effect: Cloudflare, positioned as an autonomous infrastructure, turned out to be vulnerable due to its reliance on Google Cloud in a critical component — Worker KV. This caused a cascading failure. The inability to access cold storage led to KV errors.
Since Worker KV was used for configuration, authentication, and caching in other Cloudflare services, its failure paralyzed the entire ecosystem. The Google Cloud outage stopped services that knowingly relied on its APIs and resources. Though the risk of individual service failure is expected, the scale was critical due to Service Control’s centralization. As all Google cloud services depend on this system, the failure quickly spread. Within minutes, platforms returned HTTP 503 errors or became inaccessible, halting authentication, API calls, and dependent services globally.
Technical Consequences
Authorization Issues: Due to a failure in the Cloudflare Access service, users couldn’t log into websites. Both regular login and login via Google or SSO didn’t work.
File Upload Failure: Outages in the key data store (Worker KV) caused 100% failure rates for image and file uploads.
Real-Time Problems: Services related to streaming and chats (especially Stream) were partially or completely down.
Cloudflare Dashboard Inaccessibility: Even Cloudflare’s internal teams couldn’t work properly with their dashboard due to its reliance on the same failed storage.
This incident clearly demonstrated how deeply interconnected modern internet services are. The Google Cloud failure was the first domino that caused the second — a widespread Cloudflare outage. Many users didn’t even realize their Cloudflare-based services indirectly depended on Google infrastructure. It highlights the importance of distributed architectures, critical dependency redundancy, and transparency in communications between service providers and clients.
3. Response and Changes
Google began notifying users about the problem an hour after the outage started. Since monitoring infrastructure was also affected, the first updates were posted to the Cloud Service Health Dashboard — the only accessible source of client information. It was updated continuously with recovery status.
Google took the following actions to resolve the issue:
Root Cause Identification: Engineers discovered the crash was caused by invalid quota policy data globally replicated.
Disabling the Problematic Path: A “red button” was activated to disable the policy handling path that caused the failure, allowing recovery in most regions within 40 minutes.
Regional Recovery: Additional issues in the us-central1 region arose due to Spanner overload. Google mitigated it by limiting task creation and redirecting traffic to multi-region databases.
Full Recovery: Most services were restored by 6:18 PM Pacific Time, though some products continued experiencing residual effects until 10:00 PM.
The official report also proposed several preventative measures:
- Service Control Architecture Modularization: Isolating functionalities to avoid cascading failures.
- Auditing Globally Replicated Systems: Preventing untested changes from spreading instantly.
- Use of Feature Flags: Allowing gradual activation of new features.
- Enhanced Testing and Error Handling: Additional code checks for null pointers and invalid data.
- Randomized Exponential Backoff: Avoiding “herd” effects during service restarts.
- Improved Communication: Ensuring rapid, accurate customer updates during incidents.
Google’s response was swift and structured. They resolved the issue and offered concrete improvements to prevent recurrence. Their communication was transparent and clear.
CLOUDFLARE
Cloudflare demonstrated high accountability by publishing a detailed incident report just four hours after resolving the issue. The report stated:
- Duration: 2 hours and 28 minutes of global downtime.
- Affected Services: Worker KV, Access, Gateway, Stream, Workers AI, Turnstile, etc., with direct error rate data.
- Cause: Outage in the Worker KV cold data store, which was found to depend on Google Cloud.
They openly accepted fault, even though the outage stemmed from another provider’s failure. Cloudflare emphasized that they are responsible for choosing dependencies and designing system architecture — a rare level of openness among major tech companies.
Planned Changes:
Migration to R2 Storage: Worker KV’s cold storage will be moved to Cloudflare’s own R2 service (similar to Amazon S3), eliminating third-party dependency.
Reducing Single-Provider Reliance: The company aims to make its infrastructure more distributed so no external service becomes a single point of failure.
Improved Recovery Tools: Cloudflare is developing tools that will allow partial service restoration during outages — for example, gradually reactivating data segments (namespaces) instead of waiting for a full system recovery.
Cloudflare became a model of transparency and responsibility. They not only acknowledged the issue promptly but also detailed its cause — Worker KV’s reliance on Google Cloud. Their report included technical details, a timeline, error charts, and, most importantly, a clear plan to fix systemic weaknesses. This openness allowed clients and the community to understand the scale of the problem and trust their solutions.
Conclusion: Cloudflare responded more quickly and transparently, publishing a detailed report within 4 hours and openly admitting the problem. Google also acted professionally but more technically. Both companies proposed solutions, but Cloudflare’s approach was more client-oriented.
4. Conclusions
This incident revealed how closely interconnected modern cloud services are, where failure in one component can paralyze many others. It emphasizes a core issue: the internet, which we perceive as decentralized, actually relies on a few “pillars.” If one collapses, it affects millions of users and businesses. Even large companies like Cloudflare sometimes rely on competitors’ infrastructure, making systems less resilient.
How to fix this?
Multicloud Strategy: Don’t depend on a single provider. For example, if using Cloudflare KV, add a backup option (like your own database or another service).
Failure Testing: Check how your product behaves when a critical service (like KV storage) is temporarily unavailable. Can it operate in limited mode?
Disaster Plan: Use automated monitoring (e.g., status pages) and be ready to switch to backups quickly.



Comments (0)
Add a comment