Start with Science Fictions. Stuart Ritchie organises the failure modes as fraud, bias, negligence and hype, and is clear about which of them is the big one — it is not fraud. If you read a single book on this subject, read that one. Then The seven deadly sins of psychology, the account from inside the field that broke first, by the researcher who did more than anyone to establish registered reports as a journal format: more technical than Ritchie and much more concrete about the reforms.
A scoping note before the stages, because this page and one other cover adjacent ground. This list is about method and statistics — how false findings get generated, published and defended when nobody has decided to cheat. The forensically documented fraud cases, the fabricated datasets and the retracted careers, are the subject of a separate list at /subjects/scientific-fraud-and-the-replication-crisis. John Carreyrou's Bad Blood, for instance, is deliberately not here. The two problems live in the same ecosystem, but the methods problem is both larger and harder, precisely because it needs no villain.
What went wrong
Rigor Mortis moves the same story into biomedicine, where the consequences are failed drug trials and wasted patient effort. Contaminated cell lines, underpowered animal studies and unusable antibodies are a different mechanism producing the same result, which is the clearest evidence that this is not a psychology story.
What all three describe is a machine rather than a conspiracy. Researchers made dozens of small analytic choices after seeing their data — which outliers to exclude, which covariates to include, when to stop collecting — each defensible on its own, and the combination made a significant result nearly inevitable. Journals published the positives. Nobody was funded to check. Several findings taught for years in introductory psychology courses have since failed to replicate in large multi-laboratory attempts, and researchers continue to argue about what the failures mean; the honest summary is that specific effects are disputed, not that a whole discipline has been voided.
The statistical machinery that produced it
How to Lie with Statistics is seventy years old, 142 pages, and still the fastest way to acquire suspicion of a chart. Statistics Done Wrong is the core text of this stage: 152 pages cataloguing the errors that actually appear in published papers — underpowering, pseudoreplication, multiple comparisons, truth inflation — written for researchers rather than statisticians. It is the book that connects the stories above to a mechanism. The Art of Statistics is the constructive counterpart, on how to reason from data properly, and it lands better after Reinhart has shown you what going wrong looks like.
Clinical research, where the stakes are highest
Bad Science is the general primer: how to read a claim about a treatment, what a systematic review is for, why anecdote and mechanism are not evidence. Bad Pharma: How Medicine is Broken, and How We Can Fix it assumes it, and is the strongest single argument on this path that publication bias is not a statistical curiosity but a mechanism with a body count — trials run and never published, outcomes swapped after the fact, comparators chosen to lose. Our catalogue record carries the wrong year for it; the book is from 2012.
Testing treatments is the shortest and most practical thing here, written for patients on what makes evidence about a treatment trustworthy, and it is the best answer to what a non-scientist should do with all this. Ending Medical Reversal documents practices adopted on surrogate endpoints and observational data and then reversed when someone finally ran the trial — the clearest demonstration that low evidential standards cost patients rather than merely credibility.
Reading and doing better science
Calling Bullshit teaches how to spot quantitative nonsense without checking the mathematics: selection effects, base rates, misleading visualisation, the difference between a big number and an important one. Everything is obvious explains why social-scientific findings feel obvious in hindsight, and why that feeling let implausible results through for decades.
The last two are harder and optional. The Book of Why makes the causal-inference case — that many failed findings are causal claims dressed as associations — and is occasionally polemical about Judea Pearl's own contribution. Statistical Rethinking is not a read-through at all but a full Bayesian statistics course, included because the honest conclusion of everything above is that the fix requires analysing data differently, and this is the book that teaches it. The staged version, with study plans and questions, is at /paths/pt_ai_the-replication-crisis-in-science.