Region us-central1 partial degradation
Aug 19, 10:43 UTCAug 20, 08:16 UTC
Duration
21h 33m
Impact
Major
Root cause
Power
Nebius, 90 days
22 incidents
Affected
Compute CloudObject StorageVirtual Private Cloud (Networking)Managed Service for Kubernetes®Token FactoryManaged Service for MLflowManaged Service for PostgreSQL®Managed Service for SkyPilotServerless AIUS-CENTRAL1 - Compute CloudUS-CENTRAL1 - Object StorageUS-CENTRAL1 - Virtual Private Cloud (Networking)US-CENTRAL1 - Managed Service for Kubernetes®US-CENTRAL1 - Token Factory
Lesson: Data center cooling failures can lead to rapid thermal shutdowns, requiring robust disaster recovery plans for whole-region power loss and manual recovery processes.
What happened
[https://nebius.com/blog/posts/incident-post-mortem-analysis-us-central1-service-disruption-on-august-19](https://nebius.com/blog/posts/incident-post-mortem-analysis-us-central1-service-disruption-on-august-19)
Timeline
- Postmortem · Sep 1, 09:15 UTC
[https://nebius.com/blog/posts/incident-post-mortem-analysis-us-central1-service-disruption-on-august-19](https://nebius.com/blog/posts/incident-post-mortem-analysis-us-central1-service-disruption-on-august-19)
- Resolved · Aug 20, 08:16 UTC
This incident has been resolved.
- Identified · Aug 19, 21:59 UTC
Most critical issues have been resolved and the region is operational, though a subset of nodes remains unavailable.
- Identified · Aug 19, 20:50 UTC
The Managed Service for Kubernetes is nearly fully restored, though a subset of nodes remains unavailable.
- Identified · Aug 19, 19:13 UTC
Most nodes are operating normally. The Managed Service for Kubernetes® is still degraded, but the situation is improving.
- Identified · Aug 19, 17:26 UTC
Most nodes are operating normally. The Managed Service for Kubernetes® is experiencing a partial outage, and the team is actively working to restore it.
- Identified · Aug 19, 16:00 UTC
Services are still gradually continuing to recover. Many of the virtual machines are already running, but not all of them.
- Identified · Aug 19, 15:06 UTC
Services are still gradually continuing to recover. Some virtual machines are already running, but not all of them.
- Identified · Aug 19, 13:49 UTC
Services are gradually continuing to recover.
- Identified · Aug 19, 12:52 UTC
The root cause was fixed, and the situation is stabilizing. The services are recovering.
- Identified · Aug 19, 12:20 UTC
We are still experiencing temperature issues in the us-central1. Team is working on it
- Identified · Aug 19, 11:18 UTC
The root cause has been identified as overheating hardware. GPU performance is impacted, and some nodes may be unavailable.
- Investigating · Aug 19, 10:43 UTC
We are currently investigating this issue.
More from Nebius
Full historyMetrics unavailable in Management Console for Monitoring in eu-west1lasted 11mProblems with power for several dozens of nodeslasted 14h 45mIssues with preempted Managed Kubernetes nodes in all regionslasted 24hInfiniBand connectivity issues for GPU clusters in eu-west2lasted 6h 28mCompute VM creation and Object storage are partial unavailablelasted 9mNetwork issues in eu-west2lasted 3h 12m