FMEA is a disciplined way of writing down every plausible way a component or system can fail before it actually does, then ranking those failures so the worst ones get attention first. It is built as a worksheet, filled in row by row, and it produces a prioritized list of risks rather than a yes-or-no verdict. This guide explains how an FMEA is structured, how the severity, occurrence, and detection scores combine into a risk priority number, and how the finished worksheet feeds reliability decisions.
FMEA (Failure Mode and Effects Analysis) in one line: FMEA (failure mode and effects analysis) is a bottom-up, worksheet-driven method that lists each way a component or process step can fail, describes the effect of that failure, and rates it on three scales - severity, occurrence, and detection - whose product is the risk priority number (RPN). The higher the RPN, the more urgent the corrective action, making FMEA the input document that many RCM and reliability programs build on.
An FMEA starts by breaking a system into its components or a process into its steps, then working through each one systematically. For every item, the team records the potential failure modes - the specific ways it can stop performing its function - along with the effects those failures produce and the underlying causes. A bearing on a compressor, for example, might have failure modes of seizure, excessive play, and contamination, each with its own effect on the machine and its own root cause.
The bottom-up direction is the defining trait. Rather than starting from an undesired outcome and reasoning backward, FMEA starts at the component level and asks what happens if this part fails, then traces the consequences upward. That makes it thorough for cataloging single-point failures, though it is less suited to capturing combinations of failures that only matter together.
Because the worksheet is exhaustive, an FMEA is usually a team exercise. Operators, maintenance technicians, and engineers each contribute failure modes the others would miss, and the discussion itself often surfaces risks nobody had documented. The completed sheet becomes a living record that is revisited when the equipment, its duty, or its failure history changes.
Each failure mode is scored on three independent scales, conventionally from 1 to 10. Severity rates how serious the effect is if the failure happens, from a minor nuisance to a safety or environmental incident. Occurrence rates how likely the failure is to arise. Detection rates how likely current controls are to catch the failure before it causes harm, and here a high number is bad - it means the failure would probably go unnoticed.
Multiplying the three scores gives the risk priority number. Because it is a product, RPN punishes failures that are severe, frequent, and hard to detect all at once, which are exactly the ones that deserve priority. A failure that is catastrophic but almost certain to be caught early scores lower than one that is moderate but silent and common. Teams typically sort the worksheet by RPN and drive corrective actions down the list until the remaining risk is acceptable.
RPN is a ranking tool, not an absolute measure, and it has known weaknesses - very different score combinations can produce the same number, and the 1-to-10 scales are subjective. Mature programs therefore treat RPN as a starting point and also flag any high-severity failure for action regardless of its total, since a safety consequence should not be diluted by a low occurrence score. FMECA, the criticality-extended variant, adds a more formal criticality analysis on top of the basic worksheet.
FMEA rarely stands alone. In a reliability program it is the analytical input that feeds the higher-level decision process: once the worksheet identifies and ranks failure modes, RCM logic decides what to do about each one. The failure modes with detectable warning and useful lead time become candidates for condition-based monitoring, while others are handled by scheduled maintenance, redesign, or accepted run-to-failure. Without the FMEA catalog, RCM has nothing concrete to reason over.
The detection column is where field monitoring connects. When an FMEA notes that a failure mode is currently hard to detect, that is a direct prompt to add or improve instrumentation. If a compressor overheating failure would otherwise be caught only after damage, adding a monitored temperature signal lowers the detection score and reduces the RPN, changing the case for or against proactive maintenance.
In a SCADA-instrumented oil and gas facility, many of the detection controls an FMEA relies on are already the alarms and trends collected in the monitoring platform. Merobix historizes those field signals, so a failure mode the FMEA flagged as detectable-by-pressure-trend has its detecting evidence continuously recorded. The FMEA worksheet defines what to watch for; the monitoring layer is one of the controls that determines the detection rating.
FMEA catalogs failure modes and their effects and typically ranks them with a risk priority number. FMECA adds a formal criticality analysis, quantifying how critical each failure mode is, often using failure rate data. In short, FMECA is FMEA plus a dedicated criticality step, so it demands more data but gives a stronger basis for prioritization.
RPN is the product of three scores, each usually rated 1 to 10: severity of the effect, occurrence likelihood, and detection difficulty, where a high detection score means the failure is hard to catch. Multiplying them yields a number from 1 to 1000, and higher values indicate failures that should be addressed first.
No. FMEA is proactive and forward-looking - it predicts how things might fail before they do and ranks the risks. Root cause analysis is reactive, investigating a failure that has already occurred to find why it happened. They complement each other, and an RCA finding often prompts an update to the relevant FMEA worksheet.
Merobix reads your field devices into a cloud SCADA - the real thing behind these terms, live in days from any browser.