qualitylab

the station

Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification

tier I/2018/PMLR 81 / FAT* 2018

http://proceedings.mlr.press/v81/buolamwini18a/buolamwini18a.pdf

Method

"We evaluate 3 commercial gender classification systems using our dataset and show that darker-skinned females are the most misclassified group (with error rates of up to 34.7%)."

Population

"PPB consists of 1270 individuals"

What it does not show

Three vendors, one task — binary gender classification — at one moment in 2017-18. Does not test other attributes or vendors, does not establish the causal mechanism, and does not evaluate face verification. Whether the audited models changed afterwards is documented in a later follow-up, not here.

Joy Buolamwini, Timnit Gebru

Error rates for darker-skinned women were 20.8%, 34.5% and 34.7% across the three systems, against a maximum lighter-skinned-male error of 0.8% and two systems at 0.0% and 0.3%. All classifiers performed better on males (8.1-20.6% error difference) and on lighter faces (11.8-19.2%). The two standard benchmarks these systems were likely evaluated against were 79.6% and 86.2% lighter-skinned, so the disparity was invisible by construction.

Tier I: An audit of three real commercial production systems against a purpose-built, balanced benchmark with four pre-specified intersectional comparison groups. Meets the large-N-with-comparison-group clause rather than the controlled-experiment one.

Cited by