Blog / Safety engineering and why complex systems fail

Why Complex Systems Fail: A Safety Engineering Reading Order

August 1, 2026 · 3 min read

This field looks like one subject and is really an unresolved argument. At least three incompatible accounts of why complex systems fail are in circulation: that catastrophic accidents are an inevitable property of tightly coupled systems, that they are the product of organisational drift that can be detected and arrested, and that they are control problems traceable to inadequate constraints in the design. Each has a serious literature. Whichever you read first will feel obviously correct.

So the order matters more here than in most subjects. Start with the case studies, which everyone agrees on, and only then take the theories — because once you know the accidents in detail you can test each framework against the same events rather than against its own examples.

The case studies first

To engineer is human by Henry Petroski is the gentlest possible start and still the best statement of the central idea: engineering advances by failure, and a design that has never failed teaches you very little. Inviting Disaster by James R. Chiles is a tour of industrial catastrophes written for a general reader — shallow by design, but it gives you the raw material. The logic of failure by Dietrich Dörner comes at it from cognitive psychology, using simulations to show how competent people mismanage complex, delayed-feedback systems in specific and repeatable ways.

Three theories, in the order they were argued

Normal Accidents by Charles Perrow is the origin of the pessimistic position: in systems that are both interactively complex and tightly coupled, serious accidents are a structural property, not a failure of diligence. It has been enormously influential and it is contested — the high-reliability-organisation researchers argue, with their own field evidence, that some hazardous organisations achieve very low accident rates precisely by managing those properties. Perrow's framework is a claim in an ongoing debate, not a settled result.

The Challenger launch decision by Diane Vaughan is the pivot of the whole field. Vaughan reconstructs the shuttle decision in exhausting documentary detail and finds no villains: what she finds is the normalisation of deviance, a group repeatedly accepting an anomaly until the anomaly becomes the standard. It is a long book and worth every page. Friendly Fire by Scott A. Snook does something similar for the 1994 shootdown of two US Army helicopters over Iraq, and introduces practical drift — the slow, local, individually reasonable adaptations that leave a system nothing like its design.

Human error, reframed twice

Human error by James Reason is the classical treatment and the source of the defence-in-depth model everyone has seen drawn as slices of cheese. Managing the Risks of Organizational Accidents extends it from the individual to the organisation. Reason's model has since been criticised for encouraging exactly the box-ticking it was meant to prevent, which is a fair criticism of its use rather than of the book.

The Field Guide to Understanding Human Error by Sidney Dekker is the direct challenge: that "human error" is a conclusion investigators reach rather than a cause they find, and that the useful question is why the actions made sense to the person at the time. Drift into failure generalises it into a systems account. Dekker's new-view position is well argued and is not universally accepted; some engineers read it as dissolving accountability. Read him against Reason and decide.

The engineering answer

SafeWare by Nancy Leveson is the systems-engineering treatment of software-intensive safety, and the place where the subject becomes technical rather than sociological. Engineering a safer world is her mature statement, presenting safety as a control problem and setting out the STAMP model and its analysis techniques. It is the most practically actionable book in the path and the hardest going. Read it last, when you have the case material to test it against.

Nothing here substitutes for the standards, regulators and qualified review your own domain requires. What the sequence gives you is the ability to recognise which theory a given investigation report is quietly using. The full path carries the study plan for each stage.

Follow the full ordered path here: Why Complex Systems Fail: A Safety Engineering Reading Order.

FAQ

Which theory is actually right?
The field has not settled it, and the honest answer is that they explain different things. Perrow is strongest on structural coupling, Vaughan and Dekker on how organisations rationalise anomalies, Leveson on designing constraints into a system in the first place. Practitioners tend to borrow from all three.
Is this useful outside heavy industry?
The case studies come from aviation, nuclear, chemical plants and the military, but the reasoning transfers to any system where feedback is delayed and failures are rare. Software reliability engineering borrows heavily from Dekker and Leveson in particular.

Get the books

As an Amazon Associate we earn from qualifying purchases. Some book links are affiliate links; you pay the same price and we may earn a small commission.

Follow the full reading path

Ready to learn something deeply?

Build a reading path — free

Keep reading

Explore related subjects