Automation Glossary • MTBF

What Is MTBF?

Merobix Engineering • • 7 min read

MTBF is the number reliability engineers reach for to say, in one figure, how dependable a piece of repairable equipment is. It is widely quoted and just as widely misunderstood, so knowing exactly what it measures - and what it does not - matters.

Back to Blog

MTBF in one line: Mean time between failures (MTBF) is the average operating time between one failure and the next for a repairable asset, calculated as total operating time divided by the number of failures over that period - a core reliability metric used to compare equipment and predict availability.

How MTBF Is Calculated and What It Is Not

MTBF is total operating time divided by the number of failures. If a set of pumps accumulated 10,000 running hours and failed 5 times, the MTBF is 2,000 hours. Note that this is an average across a population and a period, not a guarantee for any single unit - a pump with a 2,000-hour MTBF will not run exactly 2,000 hours before failing; some fail sooner, some much later.

MTBF also assumes the equipment is repairable and returned to service. For non-repairable items that are simply replaced, the equivalent metric is mean time to failure (MTTF). And MTBF says nothing about how long a repair takes - that is mean time to repair (MTTR). The three are distinct: MTBF is about how often failures occur, MTTR about how long they last.

MTBF, MTTR, and Availability

MTBF and MTTR combine to give availability, one of the most useful outcomes for operations. Inherent availability is MTBF divided by the sum of MTBF and MTTR. Improving availability therefore has two levers: make failures less frequent (raise MTBF) or make repairs faster (lower MTTR). A high MTBF is undermined by slow repairs, and a low MTBF can be partly offset by very fast recovery - so both numbers deserve attention.

These metrics depend on good failure and runtime records, which is where monitoring systems come in. Runtime and downtime tracking in a SCADA platform provides the operating hours and failure events that MTBF and MTTR are computed from. A cloud SCADA system that historizes run status across many remote sites supplies exactly this raw data, so reliability metrics can be calculated from measured operating history rather than estimates - though the reliability analysis itself is a layer built on top of that data.

A Worked Availability Calculation

Put symbolic numbers on the availability formula to see how the two levers interact. Suppose a fleet of injection pumps shows an observed MTBF of 900 hours, and each repair takes 100 hours end to end. Inherent availability is 900 / (900 + 100) = 0.90, so the fleet is running 90 percent of the time it is wanted. Now spend effort on either lever. Doubling MTBF to 1,800 hours gives 1,800 / 1,900, roughly 0.947. Halving MTTR to 50 hours instead gives 900 / 950 - also roughly 0.947. The same availability gain, and the repair-side fix is often the cheaper one, because faster repairs may be a matter of spares stocking and callout logistics rather than equipment redesign.

The example also shows why chasing one metric in isolation misleads. The failure count in the MTBF denominator is exactly the event count that multiplies MTTR's cost: fewer failures means fewer repair windows, so the two levers compound rather than merely add. Which lever is cheaper at a given site is entirely site-specific - but the arithmetic makes the comparison explicit before money is spent. The repair-side lever is covered in more depth under mean time to repair.

Getting the Inputs Right

An MTBF figure is only as good as the two numbers that produce it, and both are easy to get wrong. Operating time means time the equipment was actually running or required to run - not calendar time. A pump that runs half the year on batch duty accumulates half a year of operating hours, and counting the idle months inflates MTBF flatteringly and uselessly. Runtime metering, whether from a run-status contact or drive feedback, is the honest source for the numerator.

The failure count needs an equally honest definition. Count functional failures - events where the equipment could not do its job - and exclude planned maintenance stops, operational shutdowns, and external causes such as power outages if the goal is to characterize the equipment rather than the site. Whatever definition you choose, write it down and apply it consistently; the most common way two teams get incomparable MTBF numbers from the same plant is by silently counting different things as failures.

Common Misuses of MTBF

The classic misuse is reading MTBF as service life. A quoted MTBF far longer than the design life of the equipment is not a contradiction: the figure typically assumes a constant failure rate during useful life, and says nothing about the wear-out region at the end of the bathtub curve. A unit can carry a very long MTBF and still be expected to wear out on schedule - the two statements describe different failure mechanisms.

A second misuse is comparing a manufacturer's predicted MTBF against a field-observed one as if they were the same quantity. Predictions are calculated from component reliability models under assumed conditions; field MTBF is measured under real dust, heat, vibration, and duty cycle. Both are useful - predictions for comparing candidate designs, field figures for managing the installed fleet - but a gap between them is expected, not scandalous. Treat the prediction as a ceiling set in a lab and the field number as the truth of your site.

Building MTBF Tracking on SCADA Data

A practical tracking scheme needs discipline more than tooling:

  1. Historize a run-status signal for every asset in the program, so operating hours accumulate automatically.
  2. Capture failure events with a timestamp and a cause code - even a short, fixed list of causes beats free text for later analysis.
  3. Separate failure downtime from planned downtime at the point of entry; retrofitting the distinction months later rarely works.
  4. Compute MTBF per asset class over a rolling window, not just cumulatively, so improvement or degradation shows up as a trend.
  5. Review the trend beside the maintenance program - a falling MTBF on a class of assets is the data-driven trigger for moving that class from reactive work toward preventive maintenance or, where measurements support it, predictive maintenance.

Frequently Asked Questions

What is the difference between MTBF and MTTF?

MTBF applies to repairable equipment that is fixed and returned to service, measuring average time between failures. MTTF applies to non-repairable items that are replaced rather than repaired, measuring average time until the single failure occurs.

How does MTBF relate to availability?

Inherent availability equals MTBF divided by the sum of MTBF and MTTR. Availability improves either by increasing MTBF so failures are less frequent, or by reducing MTTR so repairs are faster. Both levers matter.

Does a high MTBF guarantee a unit will not fail early?

No. MTBF is a population average over a period, not a promise for any individual unit. Some units fail well before the MTBF and some well after; the figure describes expected failure frequency across many units and hours, not a single machine's lifespan.

How many failures do I need before an MTBF figure means anything?

The confidence in an MTBF estimate is driven by the number of failures observed, not by how long you have been watching. A fleet that has failed twice gives an estimate with enormous uncertainty no matter how many running hours it has logged, while a class with dozens of recorded failures supports a usable figure and even a trend. Early in a program, report the raw counts alongside the computed MTBF so nobody over-trusts a number built on three events.

Should standby equipment accumulate MTBF hours while idle?

Only if idle time genuinely stresses the equipment the way running does, which is rarely true. The usual convention is to count operating hours for the running unit and track standby units separately, since failure mechanisms at rest, such as corrosion and seal set, differ from failure mechanisms in service. What matters most is choosing one convention, documenting it, and applying it to every unit in the comparison.

More in Maintenance & Reliability
FIT rate (failures in time)  •  MTTR (Mean Time to Repair)  •  DNP3 Control Failures  •  No-effect / no-part failures  •  MTTA  •  All Maintenance & Reliability →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →