qualitylab

the station

Postmortem of database outage of January 31

tier III/2017/GitLab engineering blog

https://about.gitlab.com/blog/postmortem-of-database-outage-of-january-31/

Method

none stated

Population

none stated

What it does not show

A single incident with no base rate for how common the pattern is, and no counterfactual showing a prior restore drill would have caught the failing dump job. Does not report full financial or trust impact, or whether later drills prevented a repeat.

GitLab Inc.

Roughly 300 GB destroyed by a command run against the wrong host. Five redundancy mechanisms existed — periodic dumps to object storage, volume snapshots copied to staging, cloud disk snapshots, streaming replication, and a staging copy — and the write-up states the process of both finding and using backups failed completely. The dumps had been silently failing for months on a version mismatch, and their failure notifications were being rejected by mail policy. Recovery used a staging snapshot from six hours earlier, copied over a throttled link, taking about eighteen hours, with final loss estimated at 5,000 projects, 5,000 comments and about 700 users.

Tier III: Single-organisation incident report with concrete dated numbers for data destroyed, records lost and recovery time. One incident, no comparison group, no claimed rate.

Cited by