Discover / Reliability and maintainability engineering / Reading path

Best Books on Reliability and Maintainability Engineering

@sciencesherpaBeginner → Intermediate
14
Books
197
Hours
4
Stages
Rate this path

Reliability engineering asks a question every other discipline avoids: not whether a design works, but how long it will keep working, and what it will cost to keep it working. The answer is part statistics — lifetime distributions, censored failure data, accelerated testing — and part engineering judgement about failure modes, redundancy and maintenance. This path runs both tracks. It assumes calculus and a first course in probability and statistics; the system-reliability stage additionally assumes comfort with conditional probability and Markov chains. One warning applies throughout: the standard texts here are long-lived and heavily revised, and the catalogue records below are frequently early editions of books now several editions on. Read the note in each entry and buy current.

1

The practitioner's grounding

Intermediate

Understand what reliability means quantitatively — the bathtub curve, MTBF and its misuse, failure rate, availability — and how a reliability programme actually runs inside an engineering organisation.

Study plan for this stage

Pace: Six to eight weeks for 932 pages, read in parallel rather than in sequence — the two books do different jobs and the combination is the point. O'Connor's Practical reliability engineering is 420 pages of readable prose and can be taken at 20-25 pages a sitting; read it for judgement, and in particul

Key concepts
  • The reliability function, failure rate and hazard rate, and the exact relationships between them
  • The bathtub curve, and O'Connor's argument about how rarely real populations exhibit all three regions
  • MTBF, what it does and does not mean, and why quoting it for a non-exponential population misleads
  • Availability, and the split into inherent, achieved and operational availability
  • Maintainability as a designed property: mean time to repair and the repair time distribution
  • Failure modes and effects analysis as an engineering rather than a statistical activity
  • How a reliability programme sits inside an organisation — specification, allocation, prediction, demonstration
You should be able to answer
  • If a component has an MTBF of 100,000 hours, what fraction survives 100,000 hours, and why is the intuitive answer wrong for anything but an exponential life?
  • What is the difference between the failure rate of a population and the hazard rate of an individual item, and when do they coincide?
  • Why is O'Connor sceptical of handbook failure-rate prediction, and what would you use instead for a design decision?
  • How do inherent, achieved and operational availability differ, and which one does a customer actually experience?
  • What does a maintainability requirement look like when written properly, and what does Ebeling say it must specify?
Practice
  • Derive the relationship between the reliability function, the probability density and the hazard rate from Ebeling's first chapters, and confirm that a constant hazard gives the exponential
  • Work Ebeling's problem sets on availability for a repairable system, computing all three availability measures for the same equipment
  • Take a product you own, do an informal FMEA on it following O'Connor's procedure, and rank the failure modes by risk priority
  • Find a published MTBF figure for a real product and write down what it would have to assume to be a meaningful statement
  • Write a one-page reliability requirement for a subsystem, specifying the mission, the environment, the acceptable failure definition and the demonstration method — all four, as O'Connor insists

Next up: O'Connor has taught you to distrust a single number for a lifetime; the next stage teaches you to fit the distribution that a single number was standing in for.

Practical reliability engineering
Patrick D. T. O'Connor · 1981 · 420 pp

The readable practitioner standard and the right first book: O'Connor is consistently sceptical about the arithmetic, particularly about MTBF and handbook failure-rate prediction, which is a healthy attitude to acquire before the mathematics arrives. Our record is the 1981 first edition; the book is now in its sixth, co-authored with Andre Kleyner, and the difference is enormous. Buy current.

An Introduction to Reliability and Maintainability Engineering
Charles E. Ebeling · 1996 · 512 pp

The course text, and the one that treats maintainability as a first-class subject rather than an appendix — repair time distributions, maintenance policies, availability modelling. Read it alongside O'Connor: he supplies judgement, Ebeling supplies problem sets. The record is the first edition; the current Waveland edition is the one assigned.

2

Failure data and the statistics of life

Intermediate

Fit and interpret lifetime distributions from real, censored failure data; run and analyse an accelerated life test; and know when a Weibull shape parameter is telling you something physical.

Study plan for this stage

Pace: Six to nine months if you work all five, which is 2,763 pages — but only two of these are read-through books and the rest are references, so plan accordingly. Abernethy's The new Weibull handbook is 252 pages written as a working manual, and it is the place to start: read it end to end at 12-15 page

Key concepts
  • Censoring — right, left and interval — and what a suspended item contributes to a likelihood
  • The Weibull distribution, and what the shape parameter says physically about infant mortality, random failure and wear-out
  • Probability plotting and median ranks, versus maximum likelihood estimation, and when each is preferable
  • Confidence bounds on a life distribution, and why small samples make them wide in a way that matters
  • The lognormal distribution and where it beats the Weibull
  • Accelerated life testing: stress-life relationships, the Arrhenius and inverse power law models, and the acceleration factor
  • Degradation data as an alternative to waiting for failures
  • Competing failure modes, and why a mixed-mode dataset produces a Weibull plot with a kink in it
You should be able to answer
  • What does a Weibull shape parameter below one tell you about a population, and what maintenance decision does it argue against?
  • Why does a suspended item still carry information, and how does it enter the likelihood?
  • When would you fit by median rank regression rather than maximum likelihood, and what does Abernethy say about small samples?
  • What physical assumption underlies the Arrhenius model, and how would you check it from test data at three temperatures?
  • How do you detect competing failure modes from a probability plot, and what do you do about them?
  • What does degradation data buy you over time-to-failure data, and what does it require you to be able to measure?
Practice
  • Take a set of twenty failure times with several suspensions and fit a Weibull two ways — by Abernethy's median rank regression on plotting paper, then by maximum likelihood — and compare both the parameters and the confidence bounds
  • Deliberately simulate a two-mode failure population, plot it, and confirm you can see the kink Abernethy describes before you know it is there
  • Work Meeker and Escobar's derivation of the censored likelihood for the Weibull and reproduce it without the book
  • Design an accelerated test to three stress levels using Nelson's test-planning chapters: state the stress-life model, the allocation of units to levels, and the extrapolation you intend, then compute how wrong the extrapolation goes if the model is misspecified
  • Reproduce one of Tobias and Trindade's semiconductor acceleration examples and identify the physical failure mechanism the model is standing in for
  • Look up a confidence-bound procedure in Kececioglu that your software already implements, and check the software's answer against the handbook's worked steps

Next up: You can now characterise the life of a component; the next stage asks the harder question of what a thousand components assembled into a system will do, which is not the product of the answers you just computed.

The new Weibull handbook
Robert B. Abernethy · 1921 · 252 pp

The practitioner's bible for Weibull analysis, written as a working manual rather than a textbook, and full of the small-sample cases real failure data actually presents. Note that our catalogue record carries an impossible publication year — the metadata is junk, the book is right; the fifth edition is the current one.

Statistical methods for reliability data
William Q. Meeker · 1998 · 680 pp

Meeker and Escobar is the rigorous statistical reference the handbooks defer to: censoring, likelihood-based inference, degradation data and accelerated testing done properly. The hardest book in this stage and the most durable. Our record is the first edition; a heavily expanded second appeared recently.

Applied reliability
Paul A. Tobias · 2009 · 600 pp

Tobias and Trindade come from semiconductor reliability, which is where accelerated testing was pushed hardest, and the book is unusually good on the physics-of-failure models that justify an acceleration factor. Read it after Meeker for the applied counterpart.

Accelerated Testing
Wayne B. Nelson · 2004 · 616 pp

The monograph on the subject: statistical models, test plans and data analysis for tests that deliberately overstress a product to get answers in weeks rather than years. Narrow and definitive; still in print as a Wiley classics reissue.

Reliability engineering handbook
Dimitri Kececioglu · 1991 · 615 pp

Included as a lookup reference rather than a read-through: exhaustive worked procedures for distribution fitting, confidence bounds and test design. Useful precisely when you need the formula someone else's software is using.

3

System reliability and risk

Beginner

Model the reliability of a whole system rather than a component — reliability block diagrams, fault trees, Markov models for repairable systems, common-cause failure — and connect it to quantitative risk assessment.

Study plan for this stage

Pace: Eight months to a year for 2,410 pages, and this is where conditional probability and Markov chains become non-negotiable — a reader who is shaky on either will follow the diagrams and misuse the arithmetic. Rausand's System reliability theory is 672 pages and is the core book of the stage: work it

Key concepts
  • Structure functions, minimal path sets and minimal cut sets
  • Reliability block diagrams, and their limits when the system is not coherent
  • Fault tree analysis: gates, top events, qualitative cut-set generation and quantitative evaluation
  • Event trees and the pairing with fault trees that constitutes a PRA
  • Markov models for repairable systems, and why they are needed the moment repair is possible
  • Redundancy, standby versus active, and the switching failures that eat the benefit
  • Common-cause failure, beta factors, and why it dominates the answer for highly redundant systems
  • Uncertainty propagation and expert judgement in a risk assessment
You should be able to answer
  • Why does a reliability block diagram give the wrong answer for a system with a shared standby, and what do you use instead?
  • How do minimal cut sets relate to the structure function, and why are the low-order cut sets the ones to look at first?
  • What does a beta-factor common-cause model assume, and why does redundancy stop buying you much once it is included?
  • For a repairable two-unit system, write the Markov state diagram and say what each transition rate is — where does the assumption of exponential repair times bite?
  • What is the difference between aleatory and epistemic uncertainty in Modarres's framing, and how is each propagated?
  • When does a fault tree stop being a diagram and start being a computation, and what makes the cut-set generation hard?
Practice
  • Write the structure function for a bridge network by hand, enumerate its minimal cut sets, and compute system reliability both exactly and by the rare-event approximation, then compare
  • Build a fault tree for a real system you understand — a domestic heating system is a good size — down to component level, and generate the minimal cut sets by Kumamoto and Henley's algorithm rather than by software
  • Solve Rausand's Markov model for a two-out-of-three repairable system for steady-state availability, then vary the repair rate and see where the sensitivity is
  • Apply a beta-factor common-cause model to a triply redundant subsystem and quantify how much of the theoretical redundancy benefit survives
  • Take one of Modarres's worked risk assessments, redo the uncertainty propagation with a different distribution on a key parameter, and see whether the conclusion moves
  • Look up a result on the availability of a repairable system in Birolini and check it against your own Markov solution

Next up: You can now compute what a system will do; the last stage is about what to do with the number — which maintenance policy the reliability model actually justifies.

System reliability theory
Marvin Rausand · 2003 · 672 pp

The best single text on system-level modelling: structure functions, fault tree and event tree analysis, Markov models for repairable systems, and a genuinely good chapter on FMECA. The core book of this stage. Our record is an early edition; the current edition, retitled around models and statistical methods, is substantially rewritten.

Reliability engineering and risk analysis
M. Modarres · 1999 · 513 pp

The bridge from component reliability to probabilistic risk assessment, the nuclear-industry tradition that gave the field fault trees in the first place. Read it after Rausand for the uncertainty and expert-judgement material Rausand treats more briefly. The record is the first edition; the third is current.

Probabilistic risk assessment and management for engineers and scientists
Hiromitsu Kumamoto · 1996 · 597 pp

Kumamoto and Henley on fault tree construction and evaluation in more depth than anyone else, including the cut-set algorithms that PRA software implements. Where to go when a fault tree stops being a diagram and starts being a computation.

Reliability Engineering
Alessandro Birolini · 2004 · 628 pp

The most mathematically complete of the general references, and unusually strong on stochastic processes for repairable systems and on statistical test design. Note that this shares a display title with several other books; this is Birolini's Springer volume, now in its eighth edition against our record's fourth.

4

Maintainability and maintenance in practice

Beginner

Turn reliability numbers into a maintenance policy: choose between run-to-failure, scheduled overhaul and condition monitoring on evidence, and design a product so that maintaining it is possible.

Study plan for this stage

Pace: Three to four months for 1,786 pages, but the arithmetic is misleading because 1,104 of those are a handbook nobody reads through. Moubray's Reliability-centered maintenance is 426 pages and is the argument the rest of the stage responds to: read it properly, end to end, at 15-20 pages a day, and tr

Key concepts
  • The RCM logic: function, functional failure, failure mode, failure effect, consequence, and the task selected from the consequence
  • Moubray's central empirical claim — that most failure patterns are not age-related — and what it implies for scheduled overhaul
  • Run-to-failure as a legitimate engineered decision rather than an admission of defeat
  • The P-F interval, and why it decides whether condition monitoring is worth doing at all
  • Condition monitoring techniques: vibration, oil analysis, thermography, and what each detects
  • Maintainability in design: access, modularity, diagnostics, derating and margin
  • Spares provisioning and the link between availability, repair time and inventory
  • Design reviews as the point where maintainability is decided
You should be able to answer
  • What are the six failure patterns Moubray reports, and what fraction of components does he claim show a wear-out region?
  • Under what conditions is run-to-failure the correct policy, and what has to be true about consequences for it to be safe?
  • What is the P-F interval, and how does it set the inspection frequency for a condition-monitored failure mode?
  • How would you decide between scheduled overhaul, on-condition maintenance and redesign for a specific failure mode using Moubray's decision logic?
  • Which design choices in Crowe and Feinberg most reduce mean time to repair, and at what cost?
  • Why does a maintenance policy chosen on cost alone tend to be wrong for safety-consequence failure modes?
Practice
  • Run a full RCM analysis on one system you know well, following Moubray's worksheets: functions, functional failures, failure modes, effects, consequences and selected tasks
  • Take a failure mode you selected for on-condition maintenance and estimate its P-F interval from evidence rather than assumption, then set the inspection frequency accordingly
  • Use the Higgins and Mobley handbook to specify a vibration monitoring programme for a rotating machine, including measurement points, frequencies and alarm criteria
  • Redesign one component of that system for maintainability using Crowe and Feinberg's criteria, and estimate the change in mean time to repair and hence in availability
  • Take the Markov availability model you built in the previous stage and re-solve it under two maintenance policies from your RCM analysis, to price the policy choice rather than argue it

Next up: This is the end of the path: from here the natural continuations are the prognostics and health management literature, the standards a certified reliability engineer is examined on, and the asset management frameworks that put maintenance policy inside a whole-life cost model.

Reliability-centered maintenance
John Moubray · 1997 · 426 pp

The book that changed maintenance practice across the airline, process and military industries by insisting that most components have no useful wear-out age, so scheduled overhaul is often actively harmful. Read it as the argument the rest of this stage responds to; our record is the definitive second edition.

Maintenance engineering handbook
Lindley R. Higgins · 2001 · 1104 pp

The practical reference for the plant side: lubrication, vibration analysis, condition monitoring, shutdown planning, and the organisational structures that make them happen. Encyclopaedic rather than sequential — use it to answer questions Moubray raises.

Design for Reliability
Dana Crowe · 2001 · 256 pp

The last book because it closes the loop: reliability and maintainability are design decisions, not test outcomes, and this is the short account of how to make them at the point where they are still cheap. Crowe and Feinberg on derating, margin, and design reviews.

Discussion

Keep reading

Paths that share books, cover the same subject, or open a related topic.

More on Traffic and transportation engineering

Best Books on Traffic and Transportation Engineering

Intermediate14books217 hrs5 stages
More on Site reliability engineering career

Site Reliability Engineering: The Best Books for an SRE Career, In Order

Beginner5books52 hrs4 stages

More on reliability and maintainability engineering