Automation Glossary • Critical spares optimization

How Do You Optimize Critical Spares Inventory?

Merobix Engineering • • 7 min read

A critical spare is the expensive, slow-moving part that a facility rarely needs but cannot afford to be without, because when it fails the plant stops until a replacement arrives. Deciding how many of those to keep on the shelf is a genuine trade-off, not a guess: hold too few and a single failure turns into weeks of downtime while a new one is procured, hold too many and money sits idle on shelves that may never be picked. Critical spares optimization is the discipline of setting each stocking level so that the total expected cost, carrying cost plus the cost of stockout-driven downtime, is as low as it can be. This guide explains the trade-off and shows how failure rate, lead time, and a target availability drive the answer.

Back to Blog

Critical spares optimization in one line: Optimizing critical spares means choosing the stock level for each part that minimizes total expected cost, weighing the ongoing cost of holding a spare against the cost of the downtime you suffer if the part fails and none is on hand. The three inputs that drive the answer are how often the part is expected to fail, how long a replacement takes to arrive, and how much availability the asset is required to deliver. The higher the failure rate, the longer the lead time, or the more costly the downtime, the more spares it pays to stock.

The Trade-Off Between Carrying Cost and Downtime

Every critical spare sits between two opposing costs. On one side is the carrying or holding cost of keeping it: the capital tied up in the part, the warehouse space, the insurance, and any deterioration or obsolescence over the years it waits. For a large motor, a compressor rotor, or a specialized control card, that cost is real and recurring, and it grows with every additional unit stocked. On the other side is the cost of not having it when a failure occurs: the production lost, the emergency freight, and sometimes the safety or contractual penalties that come with an extended outage. Optimization is simply finding the stock level where the sum of these two is smallest.

What makes critical spares different from ordinary consumables is the shape of the demand. A fast-moving part fails often enough that classic reorder-point logic works well, because usage is steady and predictable. A critical spare fails rarely, so demand over any given lead time is usually zero and occasionally one, which means the whole decision hinges on the probability of that rare event rather than on an average consumption rate. Treating a slow-moving expensive spare with the same rules as a box of gaskets is a common way to either overstock capital-heavy items or leave a plant exposed to a long outage.

The right framing is expected cost, not certainty. You cannot know whether a particular pump seal will fail next year, but you can estimate the probability, multiply it by the consequence, and compare that expected downtime cost against the certain cost of holding an extra spare. When the expected downtime avoided by adding one more unit is worth more than the cost of carrying it, the extra unit is justified; when it is not, the stock level is already high enough. That marginal comparison is the heart of the method.

How Failure Rate, Lead Time, and Availability Set the Level

Three variables move the optimal stock level. The first is the part's failure rate, usually expressed as how many failures are expected per unit of time across the installed population. A higher failure rate, or a larger number of identical units in service, raises the chance that a spare will be needed within any window, and pushes the stocking level up. The second is procurement lead time: the elapsed time from raising an order to having the replacement fitted and running. The relevant quantity is expected demand during that lead time, because that is the window you are exposed to before a fresh order can arrive, and long lead-time items therefore demand more protection on the shelf.

The third variable is the target availability, or equivalently the acceptable stockout risk, for the asset the spare protects. A part that supports a redundant, non-critical service can tolerate a real chance of being out of stock, because a failure there is an inconvenience rather than an outage. A part whose failure stops the whole plant, or trips a safety function, is held to a much higher availability target, which translates into a lower tolerated stockout probability and so a higher stock level. In practice each critical spare is assigned a service level that reflects the consequence of its asset going down, and the stock quantity is then chosen to meet that level given the failure rate and lead time.

These three inputs interact rather than adding up independently. A part with a modest failure rate but a very long lead time can warrant the same stock as a more failure-prone part that can be replaced quickly, because both leave the plant exposed for a similar expected demand. Likewise, raising the availability target on a long-lead, failure-prone item can jump the required quantity from one to two, while relaxing it on a redundant service can justify holding none at all and accepting emergency procurement. Good optimization looks at each spare on its own inputs rather than applying a single blanket policy across the storeroom.

Feeding the Decision With Field and SCADA Data

The quality of a spares decision rests on the quality of the failure and downtime data behind it, and that is where operational monitoring earns its keep. Runtime hours, start counts, trips, and alarm history from a control or SCADA system give a far better basis for estimating a part's real failure rate than nameplate assumptions, because they reflect how the asset is actually loaded and cycled on your site. A pump that runs continuously and a spare of the identical pump that runs only on demand will fail at very different real-world rates, and only operating data reveals which is which.

A cloud SCADA platform such as Merobix helps by aggregating this history across many dispersed assets and sites into one place, so the population data needed to estimate failure rates is available without visiting each panel. Trended runtime and event counts let a reliability engineer see which parts are actually accumulating hours and cycles, and alarm and trip logs flag the near-misses and repeated faults that often precede a hard failure. That same central record makes it easy to track how long past replacements actually took to arrive, which is the lead-time input the optimization needs, rather than relying on an optimistic vendor quote.

Monitoring also shortens the exposure window that spares exist to cover. Early warning of a degrading asset, through rising vibration, temperature, or cycle counts surfaced by SCADA, gives the storeroom time to pre-position or expedite a spare before the failure is complete, effectively buying back part of the lead time. Combined with a rational stock level, that early notice lets a facility hold fewer spares for the same protection, because some of the risk is being managed by seeing failures coming rather than by carrying inventory against them.

Frequently Asked Questions

What makes a spare part critical?

A spare is treated as critical when its failure would stop production, trip a safety function, or cause a long outage, and when a replacement is expensive or slow to obtain. These parts are usually low-usage but high-consequence, such as major rotating equipment components or long-lead control hardware. Because a stockout is so costly, they justify a deliberate stocking analysis rather than the simple reorder rules used for common consumables.

How do lead time and failure rate change how many spares to keep?

Longer lead time widens the window during which you are exposed before a replacement can arrive, so it raises the number of spares needed to hold a given availability. A higher failure rate, or more identical units in service, increases the chance a spare is called for within that window, which also pushes the stock level up. The two combine into the expected demand over the lead time, which is the quantity the stocking decision is built around.

Is holding more spares always safer?

More spares reduce stockout risk but not for free, since each additional unit adds carrying cost, ties up capital, and can become obsolete before it is ever used. Beyond the point where the downtime avoided by one more spare is worth less than the cost of holding it, extra stock destroys value rather than protecting it. The goal is the level that minimizes total expected cost, which is often only one or two units even for a very critical part.

From Definitions to a Live Dashboard

Merobix reads your field devices into a cloud SCADA - the real thing behind these terms, live in days from any browser.

Request a Free Demo +1 (903) 307-7300
More in Automation Glossary
Spares quantity from failure rate  •  FIT rate (failures in time)  •  Parts count reliability prediction  •  Markov reliability model  •  Monte Carlo reliability simulation  •  Reliability growth (Crow-AMSAA)  •  All Automation Glossary →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →