Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification
http://proceedings.mlr.press/v81/buolamwini18a/buolamwini18a.pdf
Method
"We evaluate 3 commercial gender classification systems using our dataset and show that darker-skinned females are the most misclassified group (with error rates of up to 34.7%)."
Population
"PPB consists of 1270 individuals"
What it does not show
Three vendors, one task — binary gender classification — at one moment in 2017-18. Does not test other attributes or vendors, does not establish the causal mechanism, and does not evaluate face verification. Whether the audited models changed afterwards is documented in a later follow-up, not here.
Joy Buolamwini, Timnit Gebru
Error rates for darker-skinned women were 20.8%, 34.5% and 34.7% across the three systems, against a maximum lighter-skinned-male error of 0.8% and two systems at 0.0% and 0.3%. All classifiers performed better on males (8.1-20.6% error difference) and on lighter faces (11.8-19.2%). The two standard benchmarks these systems were likely evaluated against were 79.6% and 86.2% lighter-skinned, so the disparity was invisible by construction.
Tier I: An audit of three real commercial production systems against a purpose-built, balanced benchmark with four pre-specified intersectional comparison groups. Meets the large-N-with-comparison-group clause rather than the controlled-experiment one.