Rights and compliance · Observability
Report accuracy per subgroup, never as one number
A single accuracy figure conceals disparities large enough to make a system unusable for some people — audited against a balanced benchmark, three commercial classifiers had error rates up to 34.7% for darker-skinned women against 0.0% to 0.8% for lighter-skinned men.
Do this firstNothing. This is a place to start.
Build or obtain a benchmark balanced across the groups your system acts on, then publish the error rate for each cell rather than the mean.
The reason this is a control and not a research finding is what the audit revealed about the benchmarks themselves. The disparity was not hidden by subtlety; it was hidden by construction, because the datasets everyone evaluated against were overwhelmingly light-skinned. Any team using the standard benchmark would have seen a good number and had no mechanism to discover otherwise.
Note the scope honestly. The audit covered three vendors and one task at one moment. What generalises is the method — balanced benchmark, intersectional cells, per-cell reporting — rather than the specific figures.
The decoy
Aggregate accuracy on a standard benchmark. The two benchmarks in common use when this was measured were 79.6% and 86.2% lighter-skinned subjects, so a model could post an excellent headline score precisely because the population it failed on was barely in the test set.
Evidence
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification — IError rates for darker-skinned women were 20.8%, 34.5% and 34.7% across the three systems, against a maximum lighter-skinned-male error of 0.8% and two systems at 0.0% and 0.3%. All classifiers performed better on males (8.1-20.6% error difference) and on lighter faces (11.8-19.2%). The two standard benchmarks these systems were likely evaluated against were 79.6% and 86.2% lighter-skinned, so the disparity was invisible by construction.
Last reviewed 2026-08-19.