On September 1, 2026, portions of Google Cloud’s us-central1-b and us-central1-f zones experienced network degradation for four hours and 11 minutes. Customers saw elevated packet loss and errors across a wide range of products, from BigQuery and Cloud SQL through to GKE, Cloud Run and App Engine. Google traced the incident to a procedural error during routine router maintenance that resulted in the accidental disconnection of redundant fiber-optic connections. Service was restored, and Google outlined additional maintenance safeguards to prevent a recurrence.

Incidents scoped to part of a zone are a normal property of operating infrastructure at this scale, and every major provider has them. That is precisely why providers give us zones and regions in the first place — so a fault in one location does not have to become a fault in our application. Whether it does is a design question on our side of the shared responsibility line. September 1 is a good prompt to revisit that design, because the affected-product list was long enough to catch a lot of teams off guard.

Why a Partial-Zone Fault Touches So Many Services

Only portions of the affected zones were impacted, yet nearly twenty products appeared on the status page. That happens because network fabric sits beneath everything else.

A zone is not a single flat unit; it is several clusters, each with its own internal network fabric. Managed services are not exempt — Cloud SQL, Spanner and Bigtable run on machines attached to that fabric like anything else. Systems relying on affected components can experience latency spikes that turn into timeouts. Network dependencies can affect both control-plane operations and serving traffic, so provisioning requests may fail at the same moment as application requests. And layered services recover last: Cloud Run and App Engine were still stabilizing after the infrastructure beneath them had already recovered.

The practical takeaway is that a localized network fault produces a wide symptom surface. Planning for “one zone degrades” is more useful than planning service by service.

Start by Checking Where the Nodes Actually Run

The recommended workarounds during the incident were service-layer retries, or failover to alternate zones. Both assume we have somewhere else to go.

On GKE there is a specific gap worth verifying. A regional cluster replicates the control plane across zones. Although GKE distributes nodes across multiple zones by default, custom node pool configurations can leave workloads concentrated in a single zone. A regional cluster with every node pool pinned to a single zone gives us a highly available API server sitting on top of a single point of failure for the workloads — and the console still reports that cluster as Regional.

Check directly rather than trusting the label:

kubectl get nodes -L topology.kubernetes.io/zone

Expanding node locations on an existing regional cluster is a CLI operation; the console’s default node zones field applies to zonal clusters only. Plan the cost before running it — when a fixed node count is configured per zone, adding two zones triples the node count unless the pool is resized. Setting the autoscaler location policy to BALANCED helps keep nodes distributed evenly across zones rather than prioritizing utilization.

Four Things to Get Right After That

Spread the pods, not just the nodes. The scheduler can place every replica of a service in one zone. Topology spread constraints on topology.kubernetes.io/zone help distribute replicas across available zones, while pod disruption budgets protect availability during voluntary disruptions. Together, they help turn multi-zone capacity into multi-zone availability.

Audit stateful workloads for zonal disks. A zonal persistent disk only attaches to nodes in its own zone, so those pods stay pinned however many zones we add. Moving to regional replication may require creating replacement volumes and transferring data, including through snapshot-and-restore procedures — worth scheduling as real work rather than assuming it is a configuration change.

Make retries safe. Retry guidance assumes exponential backoff with jitter and a healthy target. Without backoff, a partial failure becomes a thundering herd and the recovery adds load instead of relieving it.

Get incident visibility scoped to our own projects. A zone-level status notice tells us a zone had trouble. It does not tell us whether our resources sat in the affected cluster. Provider tooling that maps incidents to specific projects, with alerting and API access, is worth enabling before we need it.

Then Test It

Zone redundancy that has never been exercised is a hypothesis. Draining a node pool in one zone during a planned window, then watching how traffic, storage and downstream dependencies behave, turns it into a fact. This is the step most teams skip, and it is the one that surfaces the zonal disk nobody remembered.

And Make Single-Zone a Deliberate Choice

Multi-zone is not free. It means more nodes, inter-zone egress charges and more moving parts to reason about. For a development environment or a batch job that can wait, single-zone may well be the right answer — made knowingly, with the trade-off understood.

The trouble comes when it is not deliberate: when a cluster labelled Regional is assumed to be spread, and nobody checks until a four-hour window makes it obvious. Localized faults come with the territory of large distributed infrastructure. Which of them reach our users is something we get to decide in advance.