Why you cannot have all three
A result, not an engineering shortfall
In 2016 and 2017 two independent pieces of work established something that should be taught in every course on this subject: when two groups have different base rates for the outcome being predicted, a score cannot simultaneously be calibrated for both groups and have equal false positive and false negative rates. Not "is difficult to". Cannot. The only exceptions are a perfect predictor, which does not exist, or equal base rates, which you do not get to choose.
This is why the fairness debate never resolves, and why every attempt to build a system that satisfies everyone fails in the same place. It is worth working through the arithmetic once, because after that no amount of arguing will confuse you.
The arithmetic, in small numbers
Take a risk score with a single threshold: above it, "high risk"; below it, "low risk". Two groups of 1,000 people. In Group A, 300 of the 1,000 will actually reoffend — a base rate of 30%. In Group B, 600 will — a base rate of 60%. Set aside for a moment why the base rates differ; that question is treated below and it matters enormously.
Suppose the tool is calibrated: among everyone labelled high risk, 70% actually reoffend, in both groups. Now count.
Group A. To get 70% precision among the high-risk labels, suppose 300 people are labelled high risk. Then 210 of them reoffend and 90 do not. The 90 are false positives, drawn from the 700 people who do not reoffend, giving a false positive rate of 90 ÷ 700, about 13%.
Group B. Suppose 600 are labelled high risk, with the same 70% precision: 420 reoffend, 180 do not. Those 180 false positives come from the 400 people who do not reoffend — a false positive rate of 180 ÷ 400, which is 45%.
Same calibration. Same threshold rule. False positive rates of 13% and 45%. A person in Group B who will not reoffend is more than three times as likely to be branded high risk.
Try to fix it. Raise the threshold for Group B until its false positive rate falls to 13%: now the score no longer means the same thing in each group, and a judge reading "high risk" is reading two different quantities. You have traded one definition for the other. There is no third setting where both hold, because the base rates differ and the arithmetic will not permit it.
COMPAS: both sides were right
This is exactly what happened in the most-cited controversy in the field. In 2016 ProPublica analysed COMPAS, a commercial recidivism risk tool used in US courts, and reported that Black defendants who did not go on to reoffend were labelled high risk at roughly twice the rate of white defendants who did not reoffend. Northpointe, the vendor, replied that the score was equally predictive for both groups: a given score corresponded to a similar reoffending rate regardless of race.
Both claims were true of the same data. ProPublica measured error rates; Northpointe measured calibration. Because the observed rearrest base rates differed between the groups, the two could not both be equal, and each side reported the one its framing cared about. Academic work by Chouldechova and by Kleinberg, Mullainathan and Raghavan then proved that the conflict was structural, not a defect in this particular tool.
Notice what the proof does and does not settle. It settles that no vendor can deliver both properties. It settles nothing at all about whether the tool should be used.
The question underneath
The impossibility result assumes the base rates as given. The far more important question is what those base rates are measuring.
COMPAS predicts rearrest. Rearrest is not offending. It is offending filtered through where officers patrol, who gets stopped, who gets charged rather than warned, and who can afford a lawyer. If policing is unequal, the measured base rates differ partly because the measurement differs — and every fairness definition computed against that target inherits the distortion, however elegantly you balance them.
So the honest position has two layers. Given a target, you must choose which fairness property to enforce and accept losing the others; this is settled mathematics. Whether that target should be predicted at all, by a machine, in a court, is a prior question that no metric answers, and it is usually the one that matters more.
The practical rule that follows: when someone shows you a system that is "fair", ask which definition, and then ask what the label was actually measuring. If they have an answer to both, you are talking to someone serious.
The one thing to keep
When base rates differ between groups, calibration and equal error rates are mathematically incompatible, so every system silently chooses one — and the deeper question is whether the outcome being predicted is itself measured equally.
Before you move on
A vendor claims its new risk score is calibrated by group and also has equal false positive rates by group, on a population where the two groups have measurably different base rates. What should you conclude?
Pick the one you would defend. Nobody sees your answer.