Every subgroup can move one way while the total moves the other. Before you trust a single number, find out which question it is answering.
Since 2000, the median American wage went up. Adjusted for inflation, about one percent. Modest, but up.
Over the same years the median wage fell for high school dropouts. And for high school graduates. And for people with some college. And for people with a degree. Every education group got poorer. The average got richer.
| Group | Real median wage, change since 2000 |
|---|---|
| High school dropouts | −7.9% |
| High school graduates | −4.7% |
| Some college | −7.6% |
| College degree or more | −1.2% |
| All workers together | +0.9% |
Both columns are correct. The trick is in the last row. Over those years the share of workers with a degree grew, and degree holders earn more than the rest. So the workforce shifted toward its best-paid group, which dragged the overall median up, even as every group’s own median slipped. The headline is true. It is also useless to almost everyone it describes.
This is Simpson’s paradox. A trend that holds in every group can vanish or flip when you pool the groups together. It is not a rare freak. It is what happens whenever the groups differ in size or in some hidden way, and you average over the difference.
A cleaner case, with a courtroom attached. In 1973 the University of California, Berkeley looked at admissions to its six largest graduate departments. Across all six, men were admitted at 44% and women at 30%. A fourteen-point gap. The university worried it was about to be sued for discrimination.
| Dept | Men applied | Men admitted | Women applied | Women admitted |
|---|---|---|---|---|
| A | 825 | 62% | 108 | 82% |
| B | 560 | 63% | 25 | 68% |
| C | 325 | 37% | 593 | 34% |
| D | 417 | 33% | 375 | 35% |
| E | 191 | 28% | 393 | 24% |
| F | 373 | 6% | 341 | 7% |
| All six | 2691 | 44% | 1835 | 30% |
Read the columns. In four of the six departments women were admitted at the same rate or higher. The aggregate says bias against women. Every department says nothing of the kind. Both are computed from the same applicants.
The resolution is in the applicant columns. Women applied in large numbers to departments C, E, and F, which admit few people of either sex. Men crowded into A and B, which admit most applicants. The fourteen-point gap was not the admissions committees treating women worse. It was women applying to the harder departments. Bickel and his colleagues found no pattern of discrimination by the committees.
Here is the shape of it, in a setting you will meet at work. You plot the discount each customer was offered against their lifetime value. The line through all the points slopes down. Deeper discounts, lower value. The obvious read: discounts destroy value, so stop discounting.
The slider sets how strong the hidden grouping is. Drag it, then press the button to colour the two kinds of customer and draw a line through each. Watch what the overall line was hiding.
Loyal customers get small discounts and spend a lot. Bargain hunters get deep discounts and spend little. Within each kind, a little more discount nudges value up. Pool them, and discounting looks like poison.
The two groups both slope up. The combined line slopes down. Nothing in the data changed when you revealed the colours. The only thing that changed was whether you could see the lurking variable, which kind of customer you were looking at. Aggregate over it and the trend inverts.
A model reports one number and you are tempted to trust it. A churn model posts 92% accuracy. Strong. Ship it.
Slice it by plan tier. On the basic plan, where most of your customers sit, it scores 96%. On the enterprise plan, your highest-value accounts, it scores 54%, barely better than a coin flip. The 92% is a weighted average, and the basic plan outvotes everyone. The single number is high precisely because it is dominated by the segment you care least about losing.
A strong aggregate score can sit on top of a model that fails every group that matters. The fix is the same move that resolved Berkeley and the wage headline. Disaggregate. Report the metric per segment before you report it overall. The course returns to this when it splits accuracy into precision and recall, and again when it builds proper validation.
Each group’s overall rate is a weighted average of the per-department rates, weighted by where that group applied:
The second factor, the department’s admit rate, can be equal for men and women. The paradox lives entirely in the first factor. Women put more of their weight on departments with a low \(r_d\), so their weighted average comes out lower, even with equal or better treatment inside every department. Change the weights and you change the aggregate, without changing a single per-group rate. That is the whole mechanism, and it is why a single pooled number can contradict every part it is built from.
You run an A/B test on the checkout page. Variant B converts better overall, so you are about to ship it. A colleague slices the result by device. On mobile, B converts worse than A. On desktop, B also converts worse than A. Both.
Explain how B can win overall while losing on every device. Then decide which variant you ship, and say what you would need to check before you do.
You have now built models on aggregated data. This is the moment to distrust the aggregate. The same reflex, slice before you trust, comes back when the course separates accuracy into precision and recall, and when it builds validation that checks a model on the groups it will actually serve.