qualitylab

Rights and compliance · Observability

Report accuracy per subgroup, never as one number

tier I/cost to adopt: medium/active

A single accuracy figure conceals disparities large enough to make a system unusable for some people — audited against a balanced benchmark, three commercial classifiers had error rates up to 34.7% for darker-skinned women against 0.0% to 0.8% for lighter-skinned men.

Do this firstNothing. This is a place to start.

Build or obtain a benchmark balanced across the groups your system acts on, then publish the error rate for each cell rather than the mean.

The reason this is a control and not a research finding is what the audit revealed about the benchmarks themselves. The disparity was not hidden by subtlety; it was hidden by construction, because the datasets everyone evaluated against were overwhelmingly light-skinned. Any team using the standard benchmark would have seen a good number and had no mechanism to discover otherwise.

Note the scope honestly. The audit covered three vendors and one task at one moment. What generalises is the method — balanced benchmark, intersectional cells, per-cell reporting — rather than the specific figures.

The decoy

Aggregate accuracy on a standard benchmark. The two benchmarks in common use when this was measured were 79.6% and 86.2% lighter-skinned subjects, so a model could post an excellent headline score precisely because the population it failed on was barely in the test set.

Evidence

Last reviewed 2026-08-19.