Automation Glossary • Hardware Fault Tolerance

What Is Hardware Fault Tolerance (HFT)?

Merobix Engineering • • 6 min read

Hardware fault tolerance is the number of faults a subsystem can suffer and still perform its safety function. It is a measure of built-in robustness through redundancy: a subsystem that keeps working after one fault has a fault tolerance of one, while one that fails after a single fault has a fault tolerance of zero. This number matters because, together with the safe failure fraction, it sets a hard ceiling on the safety integrity level a design may claim. Hardware fault tolerance is where the abstract idea of redundancy becomes a concrete constraint on what a design is allowed to achieve.

Back to Blog

Hardware Fault Tolerance in one line: Hardware fault tolerance (HFT) is the number of dangerous faults a subsystem can tolerate while still performing its safety function; a fault tolerance of one means the subsystem survives a single fault, and zero means a single fault can defeat it. Combined with the safe failure fraction, hardware fault tolerance sets the architectural constraint that caps the maximum safety integrity level a subsystem may claim, making minimum redundancy a formal requirement rather than a design preference.

What Hardware Fault Tolerance Actually Counts

Hardware fault tolerance answers a blunt question: how many faults can this subsystem absorb before it can no longer do its safety job? A single non-redundant channel has a fault tolerance of zero, because any one dangerous fault in it can defeat the function. Add a second channel arranged so that one healthy channel can still act, and the subsystem gains a fault tolerance of one, because it now takes two coincident faults to defeat it. More redundancy raises the number further.

It is important to see that fault tolerance describes robustness, not the voting behavior itself. A voting architecture like 1oo2 or 2oo3 describes how channels combine to make a decision, whereas hardware fault tolerance is the resulting count of how many faults the arrangement can survive. A 1oo2 and a 2oo3 arrangement can both provide a fault tolerance of one for the dangerous case, even though they behave very differently with respect to spurious trips. Fault tolerance is the safety-focused summary of what the redundancy buys you.

This distinction is why hardware fault tolerance is treated as its own parameter. Two designs can share the same voting label but differ in their effective fault tolerance depending on exactly how faults propagate, and two designs with different voting can share the same fault tolerance. Reliability engineering needs a clean number for how many faults are survivable, independent of the marketing name of the architecture, and hardware fault tolerance is that number.

How HFT and SFF Set the Maximum Claimable SIL

Hardware fault tolerance does not act alone. Reliability standards pair it with the safe failure fraction to form architectural constraints, an independent check that sits alongside the probability-of-failure calculation. Together, fault tolerance and safe failure fraction are read from a table that specifies the highest safety integrity level a subsystem may claim. A design has to satisfy this architectural ceiling in addition to meeting its numerical failure target, so both hurdles must be cleared.

The two parameters trade off against each other. A subsystem with a high safe failure fraction can reach a given integrity level with less fault tolerance, because most of its failures are safe or detected. A subsystem with a lower safe failure fraction needs more fault tolerance, that is, more redundancy, to reach the same level, because its failures are more likely to be dangerous and hidden. This coupling means you cannot decide redundancy in isolation; the required fault tolerance depends on how good the safe failure fraction is.

The purpose of tying integrity levels to fault tolerance is to guarantee a minimum robustness that does not rely solely on trusting failure-rate numbers. Failure data is uncertain, and a design that leans entirely on a favorable calculation could be brittle. Requiring a certain fault tolerance forces the architecture to physically survive at least a defined number of faults, giving the safety function resilience that a spreadsheet alone cannot promise. This is why minimum redundancy becomes a formal requirement for higher integrity levels, not merely good engineering taste.

Fault Tolerance, Diagnostics, and SCADA Visibility

A subtle but critical point is that fault tolerance only helps if faults are actually detected and repaired. A subsystem with a fault tolerance of one can survive its first fault, but if that first fault goes unnoticed, the subsystem is now silently running with no remaining tolerance, one fault away from defeat. The redundancy that architectural constraints demand is only as valuable as the process that finds and fixes the faults it is meant to absorb. Undetected faults quietly consume fault tolerance.

This is where diagnostics and monitoring protect the value of hardware fault tolerance. Effective online diagnostics reveal when a channel has failed, and a monitoring layer ensures that finding reaches someone who can act. A SCADA platform that compares redundant channels and raises an alarm when one diverges tells operators that the subsystem has spent one of its lives and must be restored before another fault arrives. Without that visibility, a fault-tolerant design can degrade to a fault-intolerant one without anyone knowing.

For remote and unmanned oil and gas assets, this visibility is often the only defense against silent erosion of fault tolerance. A redundant subsystem on a distant site can lose a channel and keep running, appearing healthy while its safety margin is gone. Cloud monitoring that continuously watches redundant elements, surfaces a failed channel promptly, and drives it to repair is what keeps the physical redundancy meaningful. The architecture provides the fault tolerance; monitoring makes sure it is still there when the next fault comes.

Frequently Asked Questions

What is the difference between hardware fault tolerance and voting architecture?

Voting architecture, such as 1oo2 or 2oo3, describes how redundant channels combine to make a trip decision, including how they behave for spurious trips. Hardware fault tolerance is the resulting count of how many faults the arrangement can survive while still performing its safety function. Different voting schemes can share the same fault tolerance for the dangerous case even though they behave very differently otherwise.

How do HFT and SFF work together to limit SIL?

Reliability standards read hardware fault tolerance and safe failure fraction together from a table that sets the highest safety integrity level a subsystem may claim. A design must satisfy this architectural ceiling in addition to its probability-of-failure target. The two parameters trade off, so a higher safe failure fraction allows a given level with less fault tolerance, while a lower safe failure fraction requires more redundancy.

Does hardware fault tolerance still help if faults go undetected?

Only partly, and it is a real hazard. A subsystem with a fault tolerance of one survives its first fault, but if that fault is not detected and repaired, the subsystem now runs silently with no remaining tolerance, one fault away from failure. This is why diagnostics and monitoring are essential: fault tolerance is only meaningful if the faults it absorbs are found and fixed before the next one arrives.

Safety & engineering notice. This article is general educational information, not site-specific engineering, safety, or legal advice, and it does not reflect any particular facility. Standards and regulations (for example OSHA, API, IEC, ISO, NFPA, NIST, and NERC CIP requirements) change and vary by edition, jurisdiction, and application. SCADA and remote monitoring cannot verify physical isolation, atmosphere, lockout/tagout, permit status, or a safe go/no-go decision. Qualified personnel must perform site-specific engineering, hazard analysis, and safety review, and confirm current requirements with the authority having jurisdiction, before acting.

From Definitions to a Live Dashboard

Merobix reads your field devices into a cloud SCADA - the real thing behind these terms, live in days from any browser.

Request a Free Demo +1 (903) 307-7300
More in Automation Glossary
Diagnostic Coverage  •  Common Cause Failure  •  Logic Solver  •  Final Element  •  Spurious Trip  •  Swinging Door Compression  •  All Automation Glossary →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →