qualitylab

Reversibility · Observability

Restore from backup on a normal day, before you have to

tier III/cost to adopt: medium/active

A recovery path that has never been executed is a hypothesis, and the mechanisms most likely to fail are the ones that report success without ever having been asked to produce a usable restore.

Do this firstRolling back is one step, and it is practised

Take a real backup and restore it, on a Tuesday, with the clock running and somebody timing it.

The incident that makes the case is not a story about carelessness. Every sensible mechanism was in place; what was missing was any occasion on which one of them had been asked to do its job. Recovery came down to a single volume snapshot taken six hours earlier and an eighteen-hour copy over a throttled link.

The contrasting case is instructive for what it credits. Two recoveries of very large user-data losses at another organisation are attributed by their own write-up to the fact that similar situations had been simulated many times beforehand — one of them tooled and rehearsed weeks earlier during a scheduled disaster exercise.

Be clear about the evidence: both are case studies, and no large-N measurement of drilled against undrilled recovery appears to exist. The mechanism is plausible and the counting has not been done.

The decoy

Backups that report success. One organisation held five independent mechanisms — periodic dumps, volume snapshots, cloud disk snapshots, streaming replication, and a staging copy — and when it mattered, four had failed silently or were never enabled. The dumps had been broken for months and the failure emails were being rejected by the mail configuration.

Evidence

Last reviewed 2026-08-19.