Power issues during planned maintenance
Mar 10, 15:47 UTCMar 11, 00:49 UTC
Duration
9h 2m
Impact
Major
Root cause
Not disclosed
Nebius, 90 days
22 incidents
Affected
Compute CloudObject StorageVirtual Private Cloud (Networking)Managed Service for Kubernetes®MonitoringToken FactoryManaged Service for MLflowManaged Service for PostgreSQL®Standalone ApplicationsUS-CENTRAL1 - Compute CloudUS-CENTRAL1 - Object StorageUS-CENTRAL1 - Virtual Private Cloud (Networking)US-CENTRAL1 - Managed Service for Kubernetes®US-CENTRAL1 - Monitoring
Final update
This incident has been resolved and all services and clusters should be operational now.
Timeline
- Resolved · Mar 11, 00:49 UTC
This incident has been resolved and all services and clusters should be operational now.
- Monitoring · Mar 10, 23:37 UTC
Maintenance has been completed, and we do not expect any further power-related issues. We are now recovering services and closely monitoring system status.
- Investigating · Mar 10, 22:00 UTC
We are continuing to investigate this issue.
- Investigating · Mar 10, 21:40 UTC
We are continuing to investigate this issue.
- Investigating · Mar 10, 20:39 UTC
We are currently experiencing another service disruption due to ongoing on-site maintenance. Our engineers are actively working to stabilize the affected clusters and maintain service throughout the maintenance period.
- Monitoring · Mar 10, 18:14 UTC
We have identified failed switches and nodes. All major services were recovered, we are now collecting faulty instances and recovering them.
- Identified · Mar 10, 18:12 UTC
We are continuing to work on a fix for this issue.
- Identified · Mar 10, 15:59 UTC
We are continuing to work on a fix for this issue.
- Identified · Mar 10, 15:54 UTC
We are continuing to work on a fix for this issue.
- Identified · Mar 10, 15:47 UTC
During planned maintenance we experienced unexpected issues with power supply. We have identified the faulty circuits and already bringing nodes that are not part of planned maintenance back to work.
More from Nebius
Full historyMetrics unavailable in Management Console for Monitoring in eu-west1lasted 11mProblems with power for several dozens of nodeslasted 14h 45mIssues with preempted Managed Kubernetes nodes in all regionslasted 24hInfiniBand connectivity issues for GPU clusters in eu-west2lasted 6h 28mCompute VM creation and Object storage are partial unavailablelasted 9mNetwork issues in eu-west2lasted 3h 12m