Fairness Criteria Cannot All Hold
Adjust a classifier and watch demographic parity, equal opportunity, and calibration compete.
Controls
Group A
Fraction of group A that is truly qualified/positive.
Distance between the positive and negative score means.
Group B
Overridden while single threshold is on.
Uses tₐ for both groups instead of a group-specific threshold.
Group A score distribution
base rate 60%
Group B score distribution
base rate 25%
Group A confusion counts
n = 1000
Group B confusion counts
n = 1000
Fairness metrics by group
Compare each criterion across A and B; the gap is what a fairness audit reports.
Forcing demographic parity breaks the others
Holding tₐ fixed, group B's threshold is moved until its selection rate matches group A's exactly.
Base rate gap between groups is 35.0%. When base rates differ and thresholds are forced to equalize selection rates, equal opportunity and calibration cannot also hold — this is Kleinberg–Mullainathan–Raghavan / Chouldechova's impossibility result.
Demographic parity (equal selection rates), equal opportunity (equal TPR), and calibration (equal PPV / predictive parity) are three different, individually reasonable requirements. Chouldechova (2017) and Kleinberg, Mullainathan & Raghavan (2016) proved that whenever base rates differ across groups and the classifier is not perfect, no single threshold rule — and no single score — can satisfy all three at once except in degenerate cases. The gaps above are the quantitative proof: shrinking one to zero grows another.
