The discrimination case that reversed when you looked closer
Berkeley's 1973 graduate admissions show men admitted at 44% and women at 35%. Broken out department by department, most departments admitted women at a slightly higher rate than men. Both statements are arithmetically true of the same data.
Women applied disproportionately to departments that admitted few of anyone. Aggregate the departments and you see a gap favouring men; disaggregate and the gap mostly inverts. This is Simpson's paradox, and the Berkeley study is its canonical real-world instance — published in Science in 1975 by the university's own statisticians.
There is no raw version of a dataset. The level at which you aggregate is already a claim about where the cause lives, made before any analysis begins, usually without anyone noticing they made it.
The instinct is to say the disaggregated table is the correct one. It is not a general rule. Whether to condition on department depends entirely on whether department choice is a mediator of the discrimination or a confounder independent of it — and that is a question about the world, not about the numbers. Two identical tables can require opposite treatments.
The paper does not conclude that Berkeley was innocent. It observes that the departments women applied to were systematically the ones with less funding per applicant — which relocates the question rather than answering it. If a field is underfunded because women enter it, disaggregating by department conditions the bias away instead of measuring it.
The story is almost always deployed to prove the opposite of what the paper says.
Used as “so there was no bias after all”, it inverts the authors' own caution. The finding is that the aggregate measure was mislocated, not that nothing was there. This is a case where the correct technical observation reliably produces the wrong public conclusion.
Evidence that departmental funding and applicant composition were independent would make the simple reading defensible.