qualitylab

the station

Meaningful Availability

tier III/2020/USENIX NSDI '20

https://www.usenix.org/system/files/nsdi20-paper-hauer.pdf

Method

"We present availability data from production Google servers; in doing so we ignore unavailability due to issues with user devices or mobile networks. Since the actual availability data is Google proprietary, we linearly scale and shift all curves in our graphs."

Population

"Since last year, we have used windowed user-uptime for all G Suite applications (Calendar, Docs, Drive, Gmail, etc.)."

What it does not show

Does not show the per-user metric improves any downstream outcome — no faster detection, lower recovery time or fewer repeat incidents, only that the two metrics diverge and why. Single vendor, proprietary data with scaled values, not independently reproducible.

Tamás Hauer, Philipp Hoffmann, John Lunney, Dan Ardelean, Amer Diwan

A success-ratio availability metric is measurably biased toward the most active clients — 1% of users generate 62% of events — and that bias caused it to over- and under-state real outage severity compared with a per-user metric. What you count decides whether your alerting reflects what users experienced.

Tier III: Single-organisation evaluation of a metric against production traffic, with real curves but scaled values; no comparison group and no causal claim tested.

Cited by