The Replication Crisis: The Best Books on Why Studies Fail to Replicate, in Order
Around 2011 a set of results that everybody had been taught stopped reproducing, and the reason turned out to be structural rather than personal. Ordinary, well-intentioned researchers were making dozens of small analytic choices after seeing their data — which outliers to exclude, which covariates to include, when to stop collecting — and each choice was defensible in isolation while the combination made a statistically significant result almost inevitable. Journals published the positives and not the negatives; nobody was funded to check. That is the story this path tells: the mechanics of how false findings are generated, tested and published, and what the reforms — preregistration, registered reports, larger samples, open data — are actually meant to fix. A scoping note: outright fraud, the fabricated datasets and the retracted careers, is the subject of a separate reading path. It gets mentioned here because it lives in the same ecosystem, but the argument below is about method and statistics, which is both the larger problem and the harder one, precisely because it needs no villain.
What Went Wrong
IntermediateGet the anatomy of the crisis from three insider accounts covering psychology and biomedicine. By the end you should be able to define p-hacking, HARKing, publication bias, the garden of forking paths and the file-drawer problem in your own words, and explain why each of them produces false positives without anyone deciding to cheat. Read all three before the statistics stage; they supply the motivating examples for everything that follows.
▸ Study plan for this stage
Pace: Five to six weeks for 941 pages, all three books written for a general reader and none requiring statistics you do not already have. Ritchie's Science Fictions is 361 pages and is the right first book - a psychologist's tour of the failure modes organised as fraud, bias, negligence and hype, and unu
- P-hacking: making analytic choices after seeing the data - which outliers to drop, which covariates to include, when to stop collecting - each defensible alone
- The garden of forking paths, and why a researcher who made only one analytic decision can still have been running many implicit tests
- HARKing: hypothesising after the results are known, and why it converts an exploratory finding into a confirmatory-looking one
- Publication bias and the file-drawer problem: journals publish positives, so the literature is a biased sample of the studies that were run
- Why none of the above requires anyone to decide to cheat, which is the structural claim the whole path rests on
- The distinct biomedical mechanisms Harris documents - contaminated cell lines, underpowered animal studies, unusable antibodies - producing the same irreproducibility by a different route
- Ritchie's four-way split into fraud, bias, negligence and hype, and his judgement about which does the most damage
- Define p-hacking, HARKing, publication bias and the file-drawer problem in your own words, without using the others in any definition.
- Explain how a well-intentioned researcher making only defensible choices can nevertheless produce a false positive almost by construction.
- Harris's mechanisms in biomedicine are quite different from psychology's. What do they have in common that makes them the same crisis?
- Ritchie says one of his four categories does far more damage than the others. Which, and what is his argument?
- Chambers is an insider proposing reforms. Which reform does he treat as most important, and what specific failure is it designed to block?
- Why is the structural account harder to fix than a fraud account, even though it is less shocking?
- Take one failed finding described in Ritchie and write out, as a numbered list, every analytic decision the original researchers would have had to make. Then count them. The number is usually larger than people expect, and it is the concrete form of the forking-paths argument.
- For three studies mentioned across the three books, write down what result would have been publishable and what result would have gone in a drawer. Doing this repeatedly makes publication bias visible as a property of the literature rather than of any individual.
- Read Chambers's account of registered reports and write down, in order, what a researcher has to commit to and when. Then say which of the failure modes from Ritchie each commitment blocks. That mapping is the reform argument in one page.
- Take Harris's cell-line and antibody problems and ask what the psychological equivalent would be - a measurement that everyone uses and nobody has validated. Name one, from Ritchie or from your own field. The parallel is the point of having biomedicine on this path.
Next up: You now have the failure modes as stories; the next stage supplies the statistical machinery that explains exactly why each of them produces a false result.

The best single overview: a psychologist's tour of the failure modes, organised as fraud, bias, negligence and hype, and clear about which of them is the big one. The right first book on the path, and the one to read if you read only one.

The account from inside the field that broke first, by the researcher who did more than anyone to establish registered reports as a journal format. More technical than Ritchie and more concrete about the reforms, which is why it comes second.

The same problem in biomedicine, where the consequences are measured in failed drug trials and wasted patient effort. Essential for showing that this is not a psychology story: contaminated cell lines, underpowered animal studies and unusable antibodies are a different mechanism producing the same result.
The Statistical Machinery That Produced It
IntermediateUnderstand the statistics well enough to see how the errors happen, without needing to be a statistician. By the end you should be able to say what a p-value does and does not mean, why a study can be significant and still almost certainly false when it is underpowered and tests an unlikely hypothesis, why multiple comparisons need correction, and why the effect size and confidence interval matter more than the significance flag. This stage assumes no maths beyond arithmetic; the two general sta
▸ Study plan for this stage
Pace: Six to eight weeks for 742 pages, and this stage assumes no mathematics beyond arithmetic. Huff and Geis's How to Lie with Statistics is only 142 pages and is decades old; read it in an evening or two as calibration, because almost every modern example in this stage is a more sophisticated version o
- What a p-value is - the probability of data at least this extreme given the null - and the four or five things people routinely mistake it for
- Statistical power, and why an underpowered study that reaches significance has probably overestimated the effect
- Truth inflation: the winner's-curse property by which the first published estimate of an effect is systematically too large
- Why a significant result testing an a priori unlikely hypothesis is still probably false, which is the base-rate argument in its statistical form
- Multiple comparisons and why correction is required, including the implicit comparisons a researcher never counted
- Pseudoreplication: treating dependent observations as independent, and how it silently inflates a sample size
- Effect size and confidence interval as the quantities that carry the meaning, with the significance flag as the least informative thing in a results section
- Spiegelhalter's constructive frame: what good inference from data actually consists of, rather than a list of things not to do
- State what a p-value means. Then state four things it does not mean, each of which you have probably seen asserted.
- Why does low power make a significant result less rather than more trustworthy? Walk through the reasoning with a concrete effect.
- A study reports a large effect and a p just under the threshold, with a small sample. What do you expect a well-powered replication to find, and why?
- How can a researcher who ran one statistical test still have a multiple-comparisons problem?
- Why do the effect size and the confidence interval carry more information than the significance flag, and what would a results section look like if it were written around them?
- Pick one error from Reinhart and find the corresponding story from stage one. Which of Ritchie's or Harris's cases does it explain?
- Work through Reinhart's catalogue and, for each error, write down the sentence in a paper that would reveal it. That list is a reading instrument, and it is what makes the final stage's paper-reading exercises possible.
- Reproduce one of Reinhart's own worked demonstrations with his own numbers, then change one input - sample size, or the assumed base rate - and predict the effect before recalculating. Truth inflation in particular only becomes intuitive when you have watched a number move.
- Take one chart from Huff and one from a current newspaper and annotate both with the same list of distortions. Seventy years apart, the techniques are the same, and confirming that yourself is the calibration the book is here for.
- Using Spiegelhalter, rewrite the conclusion of a paper you have read so that it reports effect size and interval rather than significance. Then ask whether the paper's headline claim survives the rewrite.
- Return to a finding from stage one and write down which specific statistical mechanism produced it - not 'p-hacking' as a label, but the actual chain from analytic freedom to inflated estimate to publication.
Next up: The mechanics are the same everywhere, but the cost is not - the next stage takes them into clinical research, where a false positive changes what a doctor does.

Seventy years old, an hour long, and still the fastest way to acquire suspicion of a chart. Start here as calibration; almost every modern example in this stage is a more sophisticated version of something Huff already names.

The core text of this stage: a short, precise catalogue of the statistical errors that actually appear in published papers — underpowering, pseudoreplication, multiple comparison, truth inflation — written for researchers rather than for statisticians. This is the book that connects stage one's stories to a mechanism.

The constructive counterpart: how to reason from data properly, by a statistician who has spent a career communicating uncertainty. Read after Reinhart, so the positive account lands against a clear picture of what goes wrong.
Clinical Research, Where the Stakes Are Highest
IntermediateSee the same failures where the consequence is a treatment decision. By the end you should understand selective outcome reporting and trial non-publication, why the AllTrials campaign and trial registration exist, why a randomised trial is the design of choice and how it can still mislead, and what medical reversal is — a practice adopted on weak evidence and abandoned when someone finally ran the trial. This is the most consequential stage; the reforms described here are furthest along in medic
▸ Study plan for this stage
Pace: Ten to twelve weeks for 1,171 pages, and this is the most consequential stage on the path. Goldacre's Bad Science is 338 pages and comes first because his own later book assumes it: how to read a claim about a treatment, what a systematic review is for, and why anecdote and mechanism are not evidenc
- Selective outcome reporting: the primary outcome declared in advance versus the one eventually published, and why the swap is invisible without a registry
- Trial non-publication, and the arithmetic point that a systematic review of the published trials inherits their bias in full
- Why trial registration and the AllTrials campaign exist, and precisely what registration does and does not prevent
- Randomisation as the design that licenses a causal claim, and the specific ways a randomised trial can still mislead - comparator choice, dose, duration, surrogate endpoints, and who was excluded
- Surrogate endpoints: measuring the marker instead of the outcome patients care about, and the cases where the two moved in opposite directions
- Medical reversal as a named phenomenon - adopted on weak evidence, abandoned when someone finally ran the trial - with stents and hormone therapy as the worked examples
- Fair tests as Evans frames them, and what a patient can reasonably ask about the evidence behind a recommendation
- Why the evidential standard, not any individual's conduct, is the thing that costs patients
- Explain how a systematic review can be conducted impeccably and still reach the wrong conclusion.
- What exactly does prospective trial registration prevent, and what does it leave untouched?
- Give an example from Prasad and Cifu where a surrogate endpoint moved favourably while the outcome that mattered did not. What should have been required before adoption?
- A randomised trial is the design of choice and can still mislead. Name three ways, with an example of each from these books.
- What is medical reversal, and what does the existence of a long list of reversals imply about the current standard of evidence for adoption?
- Evans writes for patients. What are the three questions a patient could ask that would do the most to expose weak evidence behind a recommendation?
- Take a treatment you or someone close to you actually uses and follow it back: find the trials, check whether they were registered, and compare the registered primary outcome with the published one. This is the exercise the whole stage exists for, and it takes an afternoon.
- Reconstruct one of Goldacre's documented cases in Bad Pharma as a timeline - trial run, result, publication or non-publication, guideline, prescription. Seeing the whole chain once makes the phrase 'publication bias' permanently concrete.
- For two reversals in Prasad and Cifu, write down the evidence that supported adoption and the evidence that forced the reversal, and identify the specific difference between them. In most cases it is design rather than size, which is the lesson.
- Use Evans to write a one-page list of questions to take to a clinical appointment about a proposed treatment. It is short, practical, and the most directly usable output of this path.
- Apply the reading instrument you built from Reinhart in the previous stage to a single clinical paper - power, primary outcome, comparator, effect size, interval. Note how much you can now say about it without any specialist knowledge of the condition.
Next up: Diagnosis is complete; the final stage turns to competence - reading a paper well, seeing why implausible findings felt obvious, and the causal and statistical machinery that a real fix requires.

Goldacre's first book and the general primer: how to read a claim about a treatment, what a systematic review is for, and why anecdote and mechanism are not evidence. Read it before Bad Pharma, which assumes it.

The central case: trials that are run and never published, outcomes swapped after the fact, and comparators chosen to lose. The strongest single argument on this path that publication bias is not a statistical curiosity but a mechanism with a body count. Note the catalogue record carries a wrong year; the book is from 2012.

A short, free-to-read book written for patients and the public on what makes evidence about a treatment trustworthy, and on fair tests. The most practical thing here, and the best answer to the question of what a non-scientist should actually do with all this.

Prasad and Cifu on practices adopted on the basis of surrogate endpoints and observational data, then reversed when a proper trial was finally run — stents, hormone therapy, and others. The clearest demonstration that low evidential standards cost patients, not merely credibility.
Reading and Doing Better Science
IntermediateMove from diagnosis to competence. By the end you should be able to read a paper and judge whether it was preregistered, whether it was powered for the effect it reports, and whether the causal claim is supported by the design — and be able to say why observing a correlation almost never licenses the intervention people want to draw from it. This stage is harder than the rest: Pearl assumes comfort with formal argument and McElreath is a genuine statistics course, included as the endpoint for an
▸ Study plan for this stage
Pace: Four to six months for 1,583 pages, and this stage is genuinely harder than the rest - the page counts understate it. Bergstrom and West's Calling Bullshit is 336 pages and is the most immediately usable: spotting quantitative nonsense without checking the maths, through selection effects, base rate
- Reading a paper for preregistration, power and design before reading its conclusions - the three checks that do most of the work
- Selection effects as the most common source of a striking finding, and how to spot the population that was never sampled
- Base rates, and why a test's accuracy tells you almost nothing without them
- Hindsight obviousness: why a social-science result and its opposite both feel like common sense, which is Watts's central demonstration
- Why common sense is a poor referee for social science, and what has to replace it
- The distinction between association and causation formalised: what a causal diagram encodes and why 'controlling for' variables can create bias as easily as remove it
- Confounding, colliders and mediators as different objects, and why treating them alike is a design error rather than an analysis error
- Generative modelling as an alternative to significance testing, which is the discipline McElreath teaches and the honest endpoint of everything before it
- You are handed a paper reporting a surprising social-science finding. List, in order, the first five things you check and why each comes where it does.
- Give an example of a selection effect that would produce a striking correlation from a completely uninteresting underlying reality.
- Watts argues common sense cannot referee social science. State his argument and the strongest objection to it.
- Why can controlling for an additional variable introduce bias? Give the structural situation in which it does, in Pearl's terms.
- Pearl claims many failed findings are causal claims dressed as associations. Take a specific failure from stage one and say whether his diagnosis fits it.
- What would preregistering a study actually require you to specify in advance, and which of the failure modes from stage one does each specification block?
- Pick three published papers in a field you care about and run the full check on each: preregistered or not, powered for the effect reported or not, causal claim supported by the design or not. Write a paragraph on each. Doing this three times is what converts this path from reading into a skill.
- Draw the causal diagram for one finding you believe, using Pearl's notation, and identify what you would have to condition on and what you must not condition on. Most readers discover the study they trusted controlled for a collider.
- Take one result that felt obvious to you in stage one and construct the equally plausible opposite result, as Watts recommends. Then find out which one the literature actually reports. This is the fastest available demonstration of his argument.
- Apply the Calling Bullshit toolkit to a single number in the news: find its source, its denominator and its sampling frame. Bergstrom and West's claim is that this can be done without technical skill, and testing that claim on a real number is the exercise.
- If you take up McElreath, work the first three chapters' exercises with his code rather than reading them. The difference between following a Bayesian argument and building a model is exactly the difference this path has been arguing about all along.
- Write a final page: for one finding you personally believed before starting, state what evidence supported it, which of the failure modes it was exposed to, and what would now convince you either way. That page is the deliverable of the whole path.
Next up: This closes the path. The fraud side of the ecosystem - the fabricated datasets and the retracted careers - is a separate reading path, and the natural next step here is either McElreath's course worked properly or the current methodological literature in your own field.

Bergstrom and West on spotting quantitative nonsense without needing to check the maths — selection effects, base rates, misleading visualisation, the difference between a big number and an important one. The practical skill this path exists to build.

Why social-scientific findings feel obvious in hindsight and why that feeling makes bad results hard to detect. Watts's argument that common sense is a poor referee for social science explains how implausible findings were waved through for decades.

The causal-inference case: that the statistics profession spent a century refusing to formalise causation, and that many failed findings are causal claims dressed as associations. Demanding, occasionally polemical about Pearl's own contribution, and the most intellectually substantial book on the path.

The endpoint rather than a read-through: a full Bayesian statistics course built around thinking generatively about a model instead of running a significance test. Included because the honest conclusion of everything above is that the fix requires learning to analyse data differently, and this is the book that teaches it.
Discussion
Keep reading
Paths that share books, cover the same subject, or open a related topic.