Best Books on the Replication Crisis and Scientific Fraud
This path is built around fraud specifically: the people who fabricated data, how they got away with it for years, and the incentive structure that made them useful to their institutions until the moment they were not. That is a narrower subject than the replication crisis, which is mostly about honest work that does not hold up, and the two are constantly conflated — so the path opens by drawing the line and then stays on the fraud side of it. It runs from two general surveys, through four case studies chosen because each was investigated in forensic detail, to the historical and analytical literature on misconduct, and finally to the reward system that produces it. No scientific training is assumed, though the psychology and biomedicine chapters go easier if you have met a p-value.
The general shape of the problem
IntermediateDistinguish the four things routinely lumped together — fraud, bias, negligence and hype — and understand why deliberate fabrication is both the rarest and the least important of them, and still worth a whole reading path.
▸ Study plan for this stage
Pace: Three to four weeks for 699 pages, and both books read at ordinary pace — 30-40 pages an evening — with no scientific training assumed. Goldacre's Bad Science is 338 pages and is the entry point even if it looks too introductory: it is still the best training available in reading a scientific claim
- Ritchie's four categories — fraud, bias, negligence, hype — and the fact that they need different remedies
- Fabrication, falsification and plagiarism as the formal definitions of misconduct, and what falls outside them
- Regression to the mean, and how it manufactures apparent effects in before-and-after designs
- Randomisation, blinding and controls, and what each specifically protects against
- Publication bias and the file-drawer problem as structural rather than individual failures
- The press release as a distinct link in the chain where claims inflate
- Why the rarest failure — deliberate fabrication — is nonetheless worth a whole reading path
- Distinguish fraud from bias with a concrete example of each where the published paper would look identical
- Why does regression to the mean produce apparent improvement in an uncontrolled trial, and how does a control group remove it?
- Which of Ritchie's four categories does he think does the most cumulative damage, and what is his argument?
- What is publication bias, and why is it not solved by better peer review?
- At which point in the chain from experiment to headline does Goldacre locate most distortion, and does Ritchie agree?
- Take a health claim from this week's news, trace it back to the paper, and classify the distortion — if any — into one of Ritchie's four categories
- Find a study with an uncontrolled before-and-after design and estimate how much of its effect regression to the mean could account for
- Read Ritchie's fraud chapter and list the cases he mentions, then check which of them the next stage covers at length and which it does not
- Compare a university press release with the abstract of the paper it describes, and mark every claim that was strengthened
- Write a one-paragraph definition of research fraud that would correctly include the four cases in the next stage and exclude honest error
Next up: Ritchie's fraud chapter is a summary; the next stage is four cases documented in enough forensic detail to show how the fabrication was actually done.

The entry point, and still the best training in reading a scientific claim sceptically: what a controlled trial is for, how regression to the mean fools people, and how the nutritionists and the newspapers between them manufacture findings. Start here even if it looks too introductory.

The best single survey of the whole territory, and the book that makes this path's distinction for you: Ritchie separates fraud, bias, negligence and hype into four chapters and shows what a fix for each would look like. Read it second, and treat its fraud chapter as the map for the next stage.
Four frauds, investigated
IntermediateLearn what a real case looks like from the inside — how the fabrication was done, why coauthors and referees and journals did not catch it, and what finally exposed it.
▸ Study plan for this stage
Pace: Two to three months for 1,416 pages of investigative narrative, all readable at 30-40 pages an evening. Reich's Plastic fantastic is 266 pages and is the best case study on the path: the Bell Labs physicist Jan Hendrik Schon fabricated a run of breakthrough results in molecular electronics, and the
- Duplicated data as a detectable signature, and why duplication is the most common forensic tell
- Coauthorship without verification, and what a senior author's name actually certifies in practice
- Peer review's actual capabilities — it checks plausibility, not data
- Undisclosed conflict of interest as a separate offence from fabrication
- The difference between fraud that damages a literature and fraud that damages patients
- Unverifiable methods sections, and what Cahalan could not confirm about the pseudopatient study
- Institutional incentives to investigate slowly, and the role of the prestige journal
- Credentialed authority substituting for data — the failure mode common to all four cases
- How was Schon's fabrication eventually detected, and why did the detection depend on the figures rather than on the physics?
- What did Deer document about the Wakefield paper beyond its scientific errors, and which of those findings is a conflict-of-interest question rather than a fraud question?
- What specifically could Cahalan not verify about the Rosenhan study, and what does she say she found instead?
- Why did the Theranos board's eminence make the failure worse rather than better, on Carreyrou's account?
- In which of the four cases did peer review have any realistic chance of catching the problem, and why so few?
- Who first raised the alarm in each of the four cases, and what does the pattern across them suggest?
- Take the duplicated-figure problem in Reich's account and look up one of the retracted Schon papers, then see whether you can identify the reuse yourself
- Build a timeline of the Wakefield case from Deer, marking separately the scientific claims, the financial disclosures and the institutional responses
- List Cahalan's specific evidentiary gaps in the Rosenhan study, and for each note what a modern preregistration requirement would have forced into the open
- For each of the four cases, write down at what point it could first have been caught and by whom, and what would have had to be different
- Compare Reich's physics case with Carreyrou's commercial one on a single dimension — who had the incentive to check, and why did none of them
Next up: Four cases are not a pattern; the next stage is the literature about misconduct itself — how common it is, what forms it takes, and how institutions respond.

The definitive account of Jan Hendrik Schon, the Bell Labs physicist who fabricated a run of breakthrough results in molecular electronics and published them in Science and Nature. The best case study on this path because the fraud was in the data files themselves — the same curve reappearing in different experiments — and dozens of physicists looked straight at it.

By the journalist who spent years documenting Andrew Wakefield's MMR paper, including the undisclosed litigation money and the alteration of the children's medical records. Read it as the case where the harm was not to the literature but to public health. Note that our catalogue record drops the leading article from the display title.

Cahalan set out to celebrate David Rosenhan's famous On Being Sane in Insane Places study and found she could not verify most of it — missing pseudopatients, a participant whose account contradicts the paper. Included because it is the rarer and more unsettling case: a fabrication that reshaped a whole profession and went unchecked for fifty years.

Theranos, and the boundary case of this stage: not academic fraud but a commercial one built on unverifiable scientific claims, with a board of eminent people who never asked for the data. Placed last here because it shows the same failure mode — credentialed authority substituting for verification — outside the university.
Fraud as a subject: history and anatomy
IntermediateMove from individual cases to the literature about misconduct itself: how common it is, what forms it takes, how institutions respond, and why detection almost always comes from a junior colleague rather than peer review.
▸ Study plan for this stage
Pace: Two to three months for 1,128 pages, and the register shifts from narrative to analysis, so the pace drops to 20-25 pages an evening. Broad and Wade's Betrayers of the truth is 256 pages, written by two Science reporters, and is the book that opened the subject by arguing misconduct was structural r
- Goodstein's three risk factors, and why all three were present in each of the previous stage's cases
- The bad-apple versus structural explanation, and what evidence would distinguish them
- Base rates of misconduct: survey estimates, their known biases, and why the true figure is unmeasurable
- How a university misconduct investigation actually proceeds, and where the incentives point
- Whistleblowers, and the consistent finding that detection comes from a junior colleague rather than peer review
- Image manipulation as the dominant modern modality, and the tools that detect it
- Paper mills and the industrialisation of fabricated output
- Retraction as a mechanism: how slow it is, and how much the retracted work continues to be cited
- What was the core claim of Broad and Wade that so angered the scientific establishment, and how well has it held up?
- Why does Judson argue that institutional self-investigation is structurally compromised, and what alternative does he consider?
- State Goodstein's three risk factors and test them against each of the four cases from the previous stage
- Why is the base rate of research fraud so hard to estimate, and what do the anonymous survey figures actually measure?
- What makes image duplication detectable at scale in a way that fabricated numerical data is not?
- Why do retracted papers keep being cited, and what would have to change about the literature for that to stop?
- Apply Goodstein's three risk factors to a research area you know and identify where a fabrication would be least likely to be caught
- Take one historical case from Broad and Wade — Summerlin is a good target — and find a later scholarly treatment of it to see what has been revised
- Read Judson on institutional response, then find a real university misconduct policy and mark which of his criticisms it addresses and which it does not
- Learn one image-forensics technique from Chevassus-au-Louis and apply it to a set of published figures from any open-access paper
- Look up a retracted paper on a retraction database, count its citations before and after retraction, and see how many post-retraction citations acknowledge it
- Write a half-page comparison of Broad and Wade's structural claim with Goodstein's insider account — do they disagree, or are they describing the same thing from two positions?
Next up: The literature on misconduct keeps arriving at the incentive structure as the cause; the last stage is about that structure directly, and about the reforms proposed to change it.

The book that opened the subject in 1982, by two Science reporters, and the first to argue that misconduct was structural rather than a handful of bad apples. Fiercely attacked at the time and largely vindicated since; read it first here for the historical cases, from Ptolemy to Summerlin.

The considered scholarly successor to Broad and Wade, by the historian of molecular biology. Judson is the best on institutional response — how universities investigate their own, and why the incentives point toward quiet resolution — and on the Baltimore case in particular.

Short, and written from inside the machinery: Goodstein handled misconduct allegations as vice-provost at Caltech. His three risk factors for fraud — career pressure, a belief you already know the answer, and data nobody will try to reproduce — are the most useful diagnostic on this path.

The most current survey, and the strongest on the modern mechanics: image manipulation and its detection, paper mills, the retraction economy, and the sleuths who now do this work publicly. Translated from the French by Nicholas Elliott, who appears in the record alongside the author.
The incentive structure that produces it
IntermediateUnderstand why the system rewards unreliable results even when nobody is lying, and evaluate the proposed fixes — preregistration, registered reports, open data, mandatory trial publication — on their evidence.
▸ Study plan for this stage
Pace: Two to three months for 1,297 pages, and these four are argument-driven books that read at 25-30 pages an evening; the p-value material goes easier if you have met one, but nothing requires statistical training. Chambers's The seven deadly sins of psychology is 292 pages and comes first: written fro
- P-hacking and researcher degrees of freedom as a continuum with fabrication at one end
- HARKing — hypothesising after results are known — and why it invalidates the reported test
- Preregistration and registered reports, and the specific difference between them
- Statistical power, and why underpowered studies produce exaggerated effect sizes rather than merely uncertain ones
- Reagent and model validity: misidentified cell lines, and unreproducible animal models
- Trial withholding as publication bias with a commercial mechanism behind it
- Medical reversal: the pattern of a treatment adopted on weak evidence and abandoned after a proper trial
- Which reforms have evidence behind them and which are so far only plausible
- Where exactly is the line between p-hacking and fraud, and can it be drawn by looking at the paper alone?
- What does a registered report change about the incentives, and at which point in the process does it bind?
- Why does low statistical power inflate the effect sizes of the studies that do reach significance?
- How does the cell line problem Harris describes differ in kind from the psychology problems Chambers describes?
- What did AllTrials specifically demand, and how much of the published record does trial withholding affect?
- Why do Prasad and Cifu insist their reversals are not fraud, and what does that distinction cost or gain their argument?
- Take a published psychology finding and list every researcher degree of freedom that was available in its design, then estimate how many analyses could have been run
- Read one registered report and its preregistration side by side, and note what the report contains that the preregistration did not permit
- Compute the statistical power of a study you can find full details for, then compute what effect size it could detect, and compare with the effect it reported
- Search a trials registry for a drug you know, count registered trials against published ones, and see whether the gap runs in the direction Goldacre predicts
- Take three reversals from Prasad and Cifu and trace back what evidence originally justified each treatment's adoption
- Return to Goodstein's three risk factors from the previous stage and assess each of the four reforms in this stage on whether it addresses any of them
Next up: This is the end of the path: from here the natural continuations are the metascience literature itself, retraction and post-publication review platforms, and the specific reform documents — preregistration templates, reporting guidelines, open data policies — that these four books argue about.

From inside the field that broke first, by one of the architects of registered reports. Chambers is unusually concrete about the mechanisms — p-hacking, HARKing, hidden flexibility — and about which reforms have actually been adopted. The best account of a discipline attempting to repair itself.

The biomedical counterpart, and a grimmer one because the failures cost lives and money rather than citations: misidentified cell lines, unreproducible mouse models, and preclinical work that no company can replicate. Read it after Chambers to see how differently the same problem looks in a different field.

The most consequential form of publication bias: trials that produce the wrong answer get withheld, so the published record on a drug is a filtered one. Goldacre's argument here launched the AllTrials campaign. Note that our record's year is wrong metadata — the book is from 2012.

The closing book, and the one that shows what happens downstream: Prasad and Cifu catalogue established treatments that turned out not to work when finally tested properly. Not fraud at all, which is the point — it is what a field looks like after weak evidence has been acted on for a decade.
Discussion
Keep reading
Paths that share books, cover the same subject, or open a related topic.