Risk
Failure Mode and Effects Analysis
A systematic, bottom-up method for asking of every element of a design or process: how could this fail, what would the effect be, how likely is it, would we catch it, and which failure modes therefore deserve preventive action first.
Also known as FMEA, Failure Modes and Effects Analysis, FMECA (with criticality analysis), Design FMEA and Process FMEA. First set out by United States armed forces (procedure MIL-P-1629); no single author in 1949; the primary source is cited in full below.
Where this is contested
No single originator: formalised as US military procedure MIL-P-1629 in 1949, developed heavily by NASA and its contractors for Apollo-era reliability work, adopted by Ford and the wider car industry from the late 1970s after the Pinto fuel-tank affair, and most recently standardised in the joint AIAG-VDA handbook of 2019, which replaced the Risk Priority Number with Action Priority.
- Format
- Scoring model
- Level
- Product · Team
- Best for
- Assess risk · Prioritise
- Decision stage
- Diagnose · Plan · Review
- Difficulty
- Intermediate
- Time to apply
- Two to five team sessions for a first pass on one process; ongoing upkeep for the life of the process.
Plate · The model
The components
Map the process or design
Decompose the subject into its elements and their functions at an agreed level of detail. The structure analysis determines everything downstream: a step that never makes it onto the map will never have its failure modes examined.
Signals of strength
The team cannot state what an element is for · Steps operators actually perform are missing from the map · Scope creeps mid-analysis as new elements keep appearing
Identify failure modes
For each function, enumerate the ways it could fail and trace each to its effects and causes. Discipline matters here: the mode is the failure of the function, the effect is what the customer experiences, the cause is what produced it.
Signals of strength
Modes, effects and causes are jumbled in one column · Only historical defects are listed, with nothing anticipated · Effects are stated internally with no reference to the end user
Score severity, occurrence and detection
Rate each failure mode on standard scales for consequence, frequency of cause, and the chance current controls catch it. The scores are ordinal judgements, useful for ranking and dangerous to treat as arithmetic quantities.
Signals of strength
Scores cluster in the safe middle of every scale · Ratings shift with who attends the meeting · Detection is scored on controls that exist only in procedure documents
Compute risk priority and act
Rank the modes, by Action Priority under AIAG-VDA 2019 or legacy RPN, and convert the ranking into resourced preventive and detective actions, then re-score once actions land. Prioritisation without action is the failure mode of FMEA itself.
Signals of strength
The FMEA ends at the number column with no action owners · High-severity modes sit ignored because detection scores flatter them · The document is untouched since the last customer audit
When it earns its keep
- You are designing or industrialising a product or process where failure carries real consequence, safety, regulatory or financial, and you want failures found on paper before they are found by customers.
- A recurring defect or field failure has shown that risks are being discovered downstream, and you need a disciplined way to work through where else the process could bite.
- A regulator, standard or customer requires documented risk analysis, as ISO 13485 and IATF 16949 environments do, and you want the document to earn its keep rather than sit in a binder.
- You are changing an established process, new equipment, new supplier, new layout, and need to assess what the change disturbs before switching it on.
And when it doesn't
- The risk is strategic or systemic rather than component-level. FMEA works element by element; for market, political or portfolio risk use scenario planning or a risk matrix.
- You need to understand a failure that has already happened. Root-cause tools such as Five Whys or a fishbone diagram dig backwards; FMEA looks forward.
- The dominant risks come from interactions and combinations of failures. FMEA considers failure modes largely one at a time; fault-tree or systems-theoretic methods handle combinatorial and emergent failure better.
- There is no capacity to act on the output. An FMEA that ends at the scoring column is expensive theatre; if no actions will be resourced, do not start.
How to run it
Before starting, gather the inputs the analysis depends on:
- A defined scope and boundary: which product, process or service, down to which level of decomposition.
- Process maps, drawings, specifications or bills of material detailed enough to enumerate elements and their functions.
- A cross-functional team including the people who operate the process, since they know the failure modes the drawings do not show.
- Agreed rating scales for severity, occurrence and detection, taken from a standard such as the AIAG-VDA handbook rather than invented in the meeting.
- History where it exists: complaints, scrap and rework data, warranty claims, near misses.
- 1
Define scope and assemble the team
Fix the boundary and the level of detail, and put operators, engineers and quality in the same room. The AIAG-VDA seven-step approach spends its first three steps on planning and structure analysis precisely because unscoped FMEAs sprawl and die.
- 2
Break the process or design into elements and functions
List each step or component and what it must do. Failure modes are defined against functions; an element whose function is vague cannot be analysed, only worried about.
- 3
Identify failure modes, effects and causes
For each function ask how it could fail to be delivered: not performed, performed wrongly, performed late, performed partially. Trace each mode to its effects on the customer or patient and back to its plausible causes.
- 4
Score severity, occurrence and detection
Rate each mode on the three scales, typically 1 to 10: how bad the effect is, how often the cause arises, and how likely current controls are to catch it before it escapes. Anchor every score in evidence and the standard scale definitions, or the numbers become mood.
- 5
Prioritise and act
Classic practice multiplied the three scores into a Risk Priority Number and worked the list from the top. The 2019 AIAG-VDA handbook replaced RPN with Action Priority, a lookup table over all thousand severity-occurrence-detection combinations that classifies each mode High, Medium or Low, ensuring high severity is never diluted by good detection scores. Either way, define preventive and detective actions with owners and dates.
- 6
Implement, re-score and keep it alive
Verify actions were taken, re-rate the affected modes, and revisit the FMEA when the process changes or field data contradicts it. A dated FMEA is a record of what the risks used to be.
Reading the result
A living register of failure modes with their effects, causes, controls and ratings, a defensible priority order across them, and a tracked set of preventive and detective actions with owners, deadlines and post-action re-scores.
- Read severity first. A severity 9 or 10 mode demands attention regardless of how rare or detectable it looks, which is exactly the judgement Action Priority hard-codes and raw RPN obscures.
- Treat identical scores as un-alike. Two modes with the same product of ratings can carry wildly different real risk; the scales are ordinal and the ranking is a starting argument, never a verdict.
- Watch the detection column for wishful thinking. Detection ratings describe controls as operated, and an FMEA is most often wrong about what the inspection step actually catches.
A worked example
A medical-device assembler runs a process FMEA on a new infusion-set line
A contract medical-device assembler in Worcester is industrialising a semi-automated line building single-use infusion sets under ISO 13485 for an OEM customer. Before process validation, a cross-functional team, production engineer, line operators, quality engineer and the sterilisation coordinator, runs a process FMEA using AIAG-VDA scales.
- Map the process or design
- The line is decomposed into fourteen steps from tube cutting to pouch sealing and labelling. The mapping itself surfaces an undocumented step: operators pre-flex the tubing before solvent bonding. It is added to the map, since an unmapped step is an unanalysed one.
- Identify failure modes
- Forty-one failure modes are logged. The heavyweights: incomplete solvent bond on the luer connector (effect: leak or disconnection during infusion), particulate left from tube cutting (effect: particulate infused to patient), pouch seal channel (effect: loss of sterility), and label mix-up between two similar set variants (effect: wrong set used clinically).
- Score severity, occurrence and detection
- Severity is scored against harm to the patient, per the customer's risk file: disconnection 9, sterility loss 9, label mix-up 9, cosmetic marks 3. Occurrence draws on trial-build data; detection is scored honestly, and the in-line leak test rates well while the end-of-line visual check for labels rates poorly, operators inspecting for a difference they rarely see.
- Compute risk priority and act
- By legacy RPN, a frequent cosmetic scuff (3 x 7 x 6 = 126) outranks the label mix-up (9 x 2 x 6 = 108). Action Priority reverses this: the mix-up's severity 9 makes it High, the scuff Medium. Actions: barcode verification camera interlocked to the pouch sealer for labels, bond-time monitoring with automatic reject on the solvent station, and a re-score once both are validated.
The read. The FMEA earns its keep twice: it forces an undocumented operator step into the controlled process, and the Action Priority logic stops a high-frequency trivial defect from outranking a rare, dangerous one, which raw RPN arithmetic had done. The honest caveat is in the detection column: two ratings rest on a visual check the team suspects is weaker than scored, so the quality engineer schedules an attribute agreement study before validation signs off.
Pitfalls
- Filling in the form to satisfy an auditor rather than analysing the process. A compliance-driven FMEA reliably contains the failure modes that already happened and none of the ones that will.
- Running it without the operators. The gap between the process as documented and as performed is where failure modes live, and only the people on the line can see it.
- Worshipping the RPN threshold. Working every mode above an arbitrary cut-off, and nothing below it, ignores severity and invites score-gaming to duck under the line; this is precisely why AIAG-VDA abandoned RPN thresholds.
- Scoring detection on paper controls. If the check is not actually performed as described, the detection rating is fiction and the priority order inherits the fiction.
- Treating the FMEA as a project deliverable rather than a living document. Processes drift, suppliers change, complaints arrive; an un-revisited FMEA quietly detaches from the process it describes.
- Analysing at the wrong altitude, either so coarse that modes hide inside steps or so fine that the team drowns before reaching the risky part.
What the critics say
The RPN arithmetic is unsound. Severity, occurrence and detection are ordinal scales, and multiplying them has no defensible mathematical meaning: different mode profiles produce identical RPNs, the three factors are weighted equally without justification, and the 1-1000 range is full of holes. Liu and colleagues review dozens of proposed alternatives driven by exactly these defects.
Liu, H.-C., Liu, L. and Liu, N. (2013) 'Risk evaluation approaches in failure mode and effects analysis: A literature review', Expert Systems with Applications, 40(2), pp. 828-838.
RPN prioritisation misleads in practice as well as theory. Bowles showed the measure is highly sensitive to small rating changes, that most of its possible values cannot occur, and that ranking by RPN can invert sensible priorities. The AIAG-VDA handbook's replacement of RPN with severity-weighted Action Priority tables in 2019 is an institutional concession to this critique.
Bowles, J. B. (2003) 'An assessment of RPN prioritization in a failure modes effects and criticality analysis', Proceedings of the Annual Reliability and Maintainability Symposium, IEEE; AIAG and VDA (2019) FMEA Handbook, 1st edn.
FMEA examines failure modes largely one at a time, so it under-detects failures arising from combinations, interactions and emergent system behaviour, and its bottom-up exhaustiveness makes it slow and dependent on the team's imagination. Systems-theoretic approaches such as STPA were developed in part to address what single-mode analysis structurally misses.
Sources and further reading
- United States Department of Defense (1949) MIL-P-1629, 'Procedures for Performing a Failure Mode, Effects and Criticality Analysis'; revised as MIL-STD-1629A (1980).
- AIAG and VDA (2019) FMEA Handbook: Design FMEA, Process FMEA, Supplemental FMEA for Monitoring and System Response. 1st edn. Southfield, MI: Automotive Industry Action Group.
- Liu, H.-C., Liu, L. and Liu, N. (2013) 'Risk evaluation approaches in failure mode and effects analysis: A literature review', Expert Systems with Applications, 40(2), pp. 828-838.
- Wikipedia: Failure mode and effects analysis (for the documented lineage from MIL-P-1629 through NASA and automotive adoption). ↗