qualitylab

Observability

Alert on what users feel, not on what machines feel

tier II/cost to adopt: medium/active

Paging on user-visible symptoms rather than internal causes reduces alert volume and raises the fraction of pages that correspond to a real problem, which is the number that determines whether anyone still reads them.

Do this firstEvery deploy leaves a mark in the telemetry

Page when requests are failing, when they are slow, when the thing a person came to do cannot be done. Graph the causes; do not wake anyone for them.

The mechanism is attention. An alert set that fires for conditions the system handles trains the on-call to close pages without reading them, and that habit does not distinguish between the noisy alerts and the one that mattered. Alert fatigue is not a morale problem, it is a detection failure with a human in it.

The decoy

A comprehensive set of resource alerts. CPU, memory, disk, queue depth — each one a cause that sometimes produces a symptom, firing at 3am for a condition the system absorbed on its own.

Evidence

Seen in the wild

Anonymised field observation. Illustration, not evidence — these carry no tier and cannot raise one.

Last reviewed 2026-08-19.