A new design rarely arrives at its final reliability; it earns it, failure by failure, as problems surface and get corrected. That improvement over time is called reliability growth, and it is the expected pattern during development testing and the early life of a fleet, where each fix makes the next failure a little less likely. To manage this deliberately you need a way to measure whether reliability is actually improving and how fast, and the Crow-AMSAA and Duane models provide it. This page describes how reliability grows through finding and fixing failures, and how these models track cumulative failures over time to forecast whether a fleet is getting better, standing still, or getting worse.
Reliability growth (Crow-AMSAA) in one line: Reliability growth is the improvement in a system's reliability, usually measured as a rising mean time between failures, that comes from finding failures and correcting their root causes so they recur less often. The Crow-AMSAA model, closely related to the earlier Duane model, tracks cumulative failures against cumulative operating time and fits a trend that shows whether reliability is improving, flat, or degrading, and lets you forecast future MTBF. It is the standard method for monitoring test-analyze-fix programs and fielded fleets.
Reliability growth is not automatic; it is the result of a deliberate cycle usually called test-analyze-fix, or sometimes test-analyze-and-fix. A system is operated or tested until it fails, the failure is analyzed to find its underlying cause rather than just its symptom, and a corrective action is designed and implemented to remove that cause. If the fix is effective, that particular failure mode stops recurring, so the system runs longer before the next failure, and the mean time between failures rises. Repeating the cycle across many failure modes drives the steady improvement that the term reliability growth names.
Two things determine how much growth actually happens. The first is how many failures are surfaced and addressed, because a mode that never appears in testing cannot be fixed before it reaches the field. The second is the effectiveness of each corrective action: a fix that fully eliminates a failure mode contributes far more growth than one that only partly reduces it or, worse, one that is deferred and never implemented. Programs often distinguish failures that are fixed immediately from those whose fixes are held for a later design update, because the timing of corrective action changes when the growth shows up.
It helps to see reliability growth as the early, improving phase of a system's life. During development and initial fielding, infant-mortality and design weaknesses dominate, and correcting them yields rapid gains. As the easy problems are exhausted, growth naturally slows and reliability approaches a plateau set by the design's inherent limits. Recognizing which phase a system is in matters, because continued rapid growth cannot be assumed indefinitely, and a fleet that has stopped improving may simply have reached the ceiling of its current design rather than having a tracking problem.
The Duane model was the early observation that, for many systems undergoing reliability growth, plotting cumulative mean time between failures against cumulative operating time on logarithmic scales produces a roughly straight line. The slope of that line measures how fast reliability is growing: a steeper upward slope means faster improvement, a flat line means no growth, and a downward slope means reliability is actually degrading over time. This simple, visual relationship made it possible to see growth at a glance and to extrapolate the line to forecast the reliability a program would reach after further testing.
The Crow-AMSAA model, developed at the US Army Materiel Systems Analysis Activity, put this relationship on a firmer statistical footing while preserving the same essential idea. It models the cumulative number of failures as a function of cumulative time with a power-law form, whose parameters capture both the current failure intensity and the growth rate. Because it is a statistical model rather than a hand-drawn line, it supports estimating the parameters from failure data, testing whether the trend is real, and producing forecasts with a sense of their uncertainty. In practice Crow-AMSAA is the more widely used of the two today, with Duane understood as its graphical ancestor.
The parameter that draws the most attention is the growth or shape parameter, because its value tells the story directly. When it indicates that failures are arriving less and less frequently as time accumulates, the system is improving. When it indicates a constant failure intensity, reliability is flat and neither growing nor degrading. When it indicates failures arriving more frequently over time, the model is signaling deterioration, which for a fielded fleet is an early warning that wear-out, a bad batch, or a creeping problem is taking hold. Reading that parameter turns a pile of failure dates into a clear verdict on the direction the fleet is heading.
Applying Crow-AMSAA to a real fleet needs two streams of data: the cumulative operating time the fleet has accumulated and the sequence of failures against that time. Both are exactly what operational monitoring records. Runtime hours give the cumulative-time axis, and dated failure and trip events give the failure sequence, so the raw material for a growth plot is a byproduct of ordinary operation rather than something that has to be gathered specially. The main work is keeping the two consistent, so that failures are counted against the same operating time base the model plots them on.
A cloud SCADA platform such as Merobix is well suited to this because it aggregates runtime and event history across many units and dispersed sites into one record, which is what fleet-level growth tracking requires. A single machine rarely accumulates enough failures for a clean trend, but a fleet of similar assets does, and pulling their combined operating hours and failure events together lets the cumulative-failure-versus-time relationship be built for the class as a whole. That same central history makes it straightforward to keep the plot current as new failures occur, so the growth or degradation trend is always up to date rather than reconstructed after the fact.
The forecasting angle is where this becomes operationally valuable. Because the model extrapolates the trend, a fleet that is genuinely improving can have its future MTBF projected, informing spares levels and maintenance intervals, while a fleet whose growth parameter has turned toward degradation raises an early flag that a systemic problem is emerging before it shows up as a rash of failures. Surfacing that turn promptly, using the failure and runtime data a monitoring platform already holds, lets reliability engineers intervene while the trend is still a warning rather than a crisis, which is the whole point of tracking reliability growth in the first place.
The Duane model is the earlier, graphical observation that cumulative MTBF plotted against cumulative time tends to form a straight line on log scales, with its slope showing the growth rate. Crow-AMSAA takes the same idea and expresses it as a statistical power-law model, which allows formal parameter estimation, significance testing, and forecasts with uncertainty. Crow-AMSAA is the more widely used method today, and Duane is generally regarded as its graphical predecessor.
Yes, the model's growth parameter can indicate deterioration as well as improvement. When failures arrive more frequently as operating time accumulates, the parameter signals that reliability is degrading rather than growing, which for a fielded fleet is an early warning of wear-out, a bad batch, or a creeping systemic issue. This ability to distinguish improving, flat, and worsening trends from the same failure data is a large part of why the model is used for ongoing fleet monitoring.
You need the cumulative operating time the system or fleet has accumulated and the sequence of failures recorded against that time. Runtime hours provide the time axis and dated failure events provide the failure counts, both of which a control or SCADA system logs during normal operation. Aggregating this across a fleet of similar assets, rather than a single machine, gives enough failures to fit a meaningful trend and forecast future MTBF.
Merobix reads your field devices into a cloud SCADA - the real thing behind these terms, live in days from any browser.