All labs

Fairness Criteria Cannot All Hold

Adjust a classifier and watch demographic parity, equal opportunity, and calibration compete.

Ch 1 1.8
Explorer
Foundations & Data Quality
20 mindifficulty 4/5

Controls

Group A

60%

Fraction of group A that is truly qualified/positive.

2

Distance between the positive and negative score means.

1000
0

Group B

25%
2
1000
0

Overridden while single threshold is on.

Single threshold for both groups

Uses tₐ for both groups instead of a group-specific threshold.

Group A score distribution

base rate 60%

00.050.100.150.200.250.30-6-4-20246tₐscoredensity

Group B score distribution

base rate 25%

00.050.100.150.200.250.30-6-4-20246t_bscoredensity
truly positivetruly negative

Group A confusion counts

n = 1000

TP
505
FP
63
FN
95
TN
337

Group B confusion counts

n = 1000

TP
210
FP
119
FN
40
TN
631

Fairness metrics by group

Compare each criterion across A and B; the gap is what a fairness audit reports.

A: selection rate
56.8%
A: TPR
84.1%
A: FPR
15.9%
A: PPV
88.8%
B: selection rate
32.9%
B: TPR
84.1%
B: FPR
15.9%
B: PPV
63.9%
Demographic parity gap
23.9%
|selection rate A − selection rate B|
Equal opportunity gap
0.0%
|TPR A − TPR B|
Equalized odds gap
0.0%
max(|TPR gap|, |FPR gap|)
Calibration gap
25.0%
|PPV A − PPV B| (calibration-in-the-large)

Forcing demographic parity breaks the others

Holding tₐ fixed, group B's threshold is moved until its selection rate matches group A's exactly.

Base rate gap between groups is 35.0%. When base rates differ and thresholds are forced to equalize selection rates, equal opportunity and calibration cannot also hold — this is Kleinberg–Mullainathan–Raghavan / Chouldechova's impossibility result.

Forced t_b
-0.84
was 0
Demographic parity gap
0.00%
driven to ≈0 by construction
Equal opportunity gap now
12.6%
was 0.0% before forcing
Calibration gap now
46.3%
was 25.0% before forcing
A fair model does not satisfy every fairness definition at once

Demographic parity (equal selection rates), equal opportunity (equal TPR), and calibration (equal PPV / predictive parity) are three different, individually reasonable requirements. Chouldechova (2017) and Kleinberg, Mullainathan & Raghavan (2016) proved that whenever base rates differ across groups and the classifier is not perfect, no single threshold rule — and no single score — can satisfy all three at once except in degenerate cases. The gaps above are the quantitative proof: shrinking one to zero grows another.