Start with Patrick O'Connor's Practical reliability engineering, then Charles Ebeling's An Introduction to Reliability and Maintainability Engineering. O'Connor is the readable practitioner standard and is consistently sceptical about the arithmetic — particularly about MTBF and handbook failure-rate prediction — which is a healthy attitude to acquire before the mathematics arrives. Ebeling is the course text and the one that treats maintainability as a first-class subject rather than an appendix: repair time distributions, maintenance policies, availability modelling. O'Connor supplies judgement, Ebeling supplies problem sets.
Two things to know before buying anything on this path. It assumes calculus and a first course in probability and statistics, and the system-reliability stage additionally assumes conditional probability and Markov chains. And the standard texts here are long-lived and heavily revised while our catalogue records are frequently early editions — the O'Connor record is the 1981 first edition of a book now in its sixth, co-authored with Andre Kleyner, and the difference is enormous. Buy current, on this one especially.
Failure data and the statistics of life
Robert Abernethy's The new Weibull handbook is the practitioner's manual for Weibull analysis, written as a working document rather than a textbook and full of the small-sample cases real failure data presents. Our record carries an impossible publication year — the metadata is junk, the book is not, and the fifth edition is current. Meeker and Escobar's Statistical methods for reliability data is the rigorous statistical reference that the handbooks defer to: censoring, likelihood-based inference, degradation data and accelerated testing done properly. It is the hardest book in this stage and the most durable; our record is the first edition and a heavily expanded second exists.
Tobias and Trindade's Applied reliability comes out of semiconductor reliability, where accelerated testing was pushed hardest, and is unusually good on the physics-of-failure models that justify an acceleration factor. Wayne Nelson's Accelerated Testing is the monograph on that subject specifically — models, test plans and analysis for tests that deliberately overstress a product to get answers in weeks rather than years. Dimitri Kececioglu's Reliability engineering handbook belongs on the shelf as a lookup reference rather than a read-through: exhaustive worked procedures for distribution fitting, confidence bounds and test design, useful precisely when you need the formula your software is already using.
System reliability and risk
Marvin Rausand's System reliability theory is the best single text on system-level modelling — structure functions, fault tree and event tree analysis, Markov models for repairable systems, and a genuinely good chapter on FMECA. Our record is an early edition; the current one is substantially rewritten and retitled. Modarres, Kaminskiy and Krivtsov's Reliability engineering and risk analysis is the bridge to probabilistic risk assessment, the nuclear-industry tradition that gave the field fault trees in the first place; read it after Rausand for the uncertainty and expert-judgement material he treats more briefly, and note our record is the first edition of a book now in its third.
Kumamoto and Henley's Probabilistic risk assessment and management for engineers and scientists goes further into fault tree construction and evaluation than anyone else, including the cut-set algorithms that PRA software implements — where to go when a fault tree stops being a diagram and becomes a computation. Alessandro Birolini's Reliability Engineering is the most mathematically complete of the general references and unusually strong on stochastic processes for repairable systems. That display title is shared with several other books; this is Birolini's Springer volume, now in its eighth edition against our record's fourth.
Maintainability and maintenance in practice
John Moubray's Reliability-centered maintenance changed maintenance practice across the airline, process and military industries by insisting that most components have no useful wear-out age, so scheduled overhaul is often actively harmful. Read it as the argument the rest of this stage responds to; our record is the definitive second edition, which is the one you want. Higgins and Mobley's Maintenance engineering handbook is the plant-side reference at 1,104 pages — lubrication, vibration analysis, condition monitoring, shutdown planning, and the organisational structures that make them happen. It is encyclopaedic rather than sequential; use it to answer the questions Moubray raises.
Crowe and Feinberg's Design for Reliability closes the loop, and it is the right last book because it makes the point the whole path has been building toward: reliability and maintainability are design decisions rather than test outcomes, and derating, margin and design review are where they are still cheap to make. The staged version, with a study plan per stage, is at /paths/pt_ai_reliability-and-maintainability-engineering.