Skip to content
GitLab ยท Developer toolsJan 31, 2017, 23:00 UTC

Primary database data accidentally deleted, 18-hour restore

CriticalOperatorUpdated 38h ago
Jan 31, 23:00 UTCFeb 1, 18:00 UTC
Duration
19h
Impact
Critical
Root cause
Operator
GitLab, 90 days
0 incidents
Affected
GitLab.comGlobal

Lesson: A backup is only real once a restore has been tested; destructive commands on production hosts need an unmistakable prompt.

What happened

While fighting replication lag under spam load, an engineer accidentally removed data from the primary database server. Backups had not been working, so GitLab.com was restored from a six-hour-old staging snapshot on slower disks, losing projects, issues and comments made in that window.

Also caused by operator action

All
StartedIncidentDuration
Feb 608:14 UTCR2 object storage disabled during a phishing report remediationCloudflare1h 22m
Apr 507:38 UTCMaintenance script deletes 883 customer sitesAtlassian12d 16h
Feb 2817:37 UTCMistyped command removes S3 index servers in US-EAST-1AWS4h 17m

Sources: vendors' own status pages, published postmortems and SEC 8-K Item 1.05 filings, read daily. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Outages by email

Saturday mornings: the week's major outages, new postmortems and disclosed breaches, only in weeks that had some.

Double opt-in. Unsubscribe any time.