Automation Glossary • Safe Failure Fraction

What Is Safe Failure Fraction (SFF)?

Merobix Engineering • • 6 min read

Safe failure fraction is a reliability metric that describes what proportion of a subsystem's failures are either safe or dangerous but detected, as opposed to dangerous and undetected. It matters because it feeds directly into the architectural constraints that cap how high a safety integrity level a given design is allowed to claim. Even a subsystem whose calculated failure probability looks excellent can be limited to a lower integrity level if its safe failure fraction and redundancy are not good enough. Safe failure fraction is therefore one of the gatekeepers between a calculation and an achievable SIL.

Back to Blog

Safe Failure Fraction in one line: Safe failure fraction (SFF) is the fraction of a subsystem's total dangerous and safe failure rate that does not result in a dangerous undetected failure; that is, the sum of safe failures and dangerous detected failures divided by all failures. Combined with hardware fault tolerance, safe failure fraction sets the architectural constraint that caps the maximum safety integrity level a subsystem may claim, regardless of how good its calculated probability of failure looks.

How Safe Failure Fraction Is Defined

To understand safe failure fraction you first have to divide a subsystem's failures into categories. Failures are either safe, meaning they tend toward the safe state or at least do not defeat the safety function, or dangerous, meaning they could prevent the function from acting. Dangerous failures split further into detected, where diagnostics catch them, and undetected, where they stay hidden until a proof test or a real demand. The one category that truly threatens safety is the dangerous undetected failure.

Safe failure fraction is the proportion of all failures that fall outside that threatening category. It is the safe failures plus the dangerous detected failures, taken as a fraction of the total failure rate. Put simply, it answers the question: of everything that can go wrong in this subsystem, what share either lands us safe or is at least caught automatically, rather than lurking undetected? A high safe failure fraction means most failures are benign or visible; a low one means a large share of failures could quietly disable the function.

This makes safe failure fraction partly a property of the equipment and partly a property of its diagnostics. Adding effective online diagnostics moves failures from the dangerous undetected bucket into the dangerous detected bucket, which raises the safe failure fraction even though the raw failure rate has not changed. That is why good diagnostic coverage and a high safe failure fraction tend to go together, and why the two concepts are closely related in reliability engineering.

Safe Failure Fraction and Architectural Constraints

The reason safe failure fraction matters is that reliability standards do not let a design claim any safety integrity level purely on the strength of its calculated failure probability. They impose architectural constraints as a second, independent hurdle. These constraints combine safe failure fraction with hardware fault tolerance, the number of faults a subsystem can withstand and still perform its function, to set a ceiling on the claimable integrity level. A design has to clear both the probability calculation and this architectural ceiling.

The logic is a deliberate check against overconfidence in the numbers. Failure rate data always carries uncertainty, and a design that leans entirely on a favorable calculation could be fragile if that data is optimistic. By additionally requiring a certain combination of safe failure fraction and fault tolerance for each integrity level, the standards ensure the architecture has inherent robustness, not just a good spreadsheet. A subsystem with poor safe failure fraction is capped lower precisely because too many of its failures could disable the function unseen.

In practice this means safe failure fraction and hardware fault tolerance are read together from a table that maps them to a maximum integrity level. Improving one can compensate somewhat for the other: a subsystem with more redundancy can tolerate a lower safe failure fraction and still reach a given level, and vice versa. This coupling is why engineers cannot treat redundancy and safe failure fraction as separate concerns; the two jointly decide what integrity level the architecture is even allowed to claim.

Raising Safe Failure Fraction With Diagnostics and SCADA

Because dangerous detected failures count on the favorable side of safe failure fraction, anything that improves detection improves the fraction. This is the direct link to diagnostics: automatic self-tests, comparison between redundant channels, and online monitoring all convert failures that would otherwise be dangerous undetected into dangerous detected. That conversion both raises the safe failure fraction and, just as importantly, gives operators a chance to repair the fault before it can defeat a demand.

This is where continuous monitoring supports the reliability engineering directly. A SCADA platform that surfaces diagnostic alarms, compares redundant transmitters, and watches for the subtle signs of a developing fault effectively strengthens the detection that safe failure fraction rewards. A failure that the field diagnostics flag and the monitoring system escalates is no longer a hidden danger; it is a maintenance task. The architecture provides the raw diagnostics, and the monitoring layer makes sure their findings are seen and acted on quickly.

For remote and unmanned oil and gas sites, that visibility is what turns a detected failure into a repaired one within a useful timeframe. A dangerous detected failure only helps if someone acts on the detection before the next demand, and on a distant asset that depends on telemetry reaching an operations team. Cloud monitoring that reliably delivers diagnostic findings, so that detected failures are repaired rather than accumulating, helps the real system live up to the safe failure fraction its design assumed.

Frequently Asked Questions

What is the formula for safe failure fraction?

Safe failure fraction is the sum of the safe failure rate and the dangerous detected failure rate, divided by the total failure rate of safe plus dangerous failures. In words, it is the share of all failures that are either safe or at least detected, rather than dangerous and undetected. A high value means most failures are benign or caught automatically, while a low value means many failures could disable the function unseen.

How does safe failure fraction affect the achievable SIL?

Reliability standards impose architectural constraints that combine safe failure fraction with hardware fault tolerance to cap the maximum safety integrity level a subsystem may claim. A design must clear both this ceiling and its probability calculation, so a low safe failure fraction can limit the achievable level even when the calculated failure probability looks excellent. Improving redundancy can partly compensate for a lower safe failure fraction, and vice versa.

How can safe failure fraction be improved?

The most common way is to add or strengthen diagnostics, since detecting a previously undetected dangerous failure moves it into the dangerous detected category, which counts on the favorable side of the fraction. Comparison between redundant channels and effective self-tests both raise the fraction without changing the raw failure rate. Choosing equipment with inherently more benign failure modes also helps.

Safety & engineering notice. This article is general educational information, not site-specific engineering, safety, or legal advice, and it does not reflect any particular facility. Standards and regulations (for example OSHA, API, IEC, ISO, NFPA, NIST, and NERC CIP requirements) change and vary by edition, jurisdiction, and application. SCADA and remote monitoring cannot verify physical isolation, atmosphere, lockout/tagout, permit status, or a safe go/no-go decision. Qualified personnel must perform site-specific engineering, hazard analysis, and safety review, and confirm current requirements with the authority having jurisdiction, before acting.

From Definitions to a Live Dashboard

Merobix reads your field devices into a cloud SCADA - the real thing behind these terms, live in days from any browser.

Request a Free Demo +1 (903) 307-7300
More in Automation Glossary
Hardware Fault Tolerance  •  Diagnostic Coverage  •  Common Cause Failure  •  Logic Solver  •  Final Element  •  Spurious Trip  •  All Automation Glossary →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →