Automation Glossary • Partial Site Outage

Partial Site Outage: Triage the Half That Dropped

Merobix Engineering • • 8 min read

Part of a site is still reporting normally while another part has gone dark. A partial outage is more informative than a total one, because the fact that some tags survived proves the site has power and a working path to you, so the fault sits at a boundary inside the site. This page is a symptom-first tree that reads the boundary between the live tags and the dead ones to name the shared element - a second controller, a downstream network segment, a separate power feed, or one interface - that failed.

Back to Blog

Partial Site Outage in one line: In a partial site outage, the surviving tags prove the site has power and a link to you, so the fault is at an internal boundary. Draw the line between the dead tags and the live ones and find what the dead group shares that the live group does not: a second PLC or RTU, a downstream switch or network segment, a separate power branch, or one protocol interface. The boundary of the outage points straight at the failed element.

First Checks: Draw the Boundary Exactly

The whole diagnosis begins by drawing the exact line between what dropped and what survived. List the dead tags and the live tags and look for the boundary: are the dead tags all from one controller, one network segment, one power branch, one interface, or one physical area of the site? Whatever the dead group has in common that the live group lacks is the failed element. This is the same scope-mapping move used for any all-at-once symptom, but a partial outage is easier because the surviving tags give you a working reference to contrast against.

Confirm the survivors are genuinely healthy, not merely stale. A tag that has not updated in a while can look alive on a screen while actually being frozen at its last value, which would blur the boundary. Verify the live tags are still changing and carry Good quality, so the line you draw is between truly-reporting and truly-dark rather than between fresh and stale. A clean boundary needs a reliable definition of both sides, and a stale survivor mistaken for a live one will send you looking for the wrong shared element.

Note the timing of the partial drop against the survivors. Did the dead group stop at one instant while the survivors carried on without a blip, or did the survivors also hiccup at that moment and then recover? A clean drop of one group with no disturbance to the rest points at a failure isolated to that group's own element. A moment where everything stuttered and only part recovered points at a shared event - a brief power dip, a network glitch - that some equipment rode out and some did not. The survivors' behavior at the drop time is part of the evidence.

Reading the Boundary to Name the Element

If the dead tags are exactly one controller's worth and the survivors are on a different controller at the same site, the site link and power are proven and one controller or its connection failed. The site clearly reaches you - the other controller's tags are arriving - so this is a device-or-interface fault on the dark controller: it crashed, hung, lost its local network drop, or sits in fault mode. The fix follows the single-controller path, and a controller in fault mode is a common cause. The rest of the site being fine tells the responder exactly which box to attend to.

If the boundary follows a network segment - the dead tags are all downstream of one switch or one network branch while upstream devices survive - the fault is that segment. A failed switch, a pulled cable between segments, or a downstream network problem cuts off everything past it while everything before it keeps reporting. The signature is that the boundary matches the network topology rather than the device list: devices that share a downstream path all died together, and their common ancestor in the network is the suspect. Knowing the site's network layout turns this into a quick read of which segment lost its uplink.

If the boundary follows a power branch, the dead tags share a supply the survivors do not. Sites often run instruments and controllers on separate power branches - a tripped breaker, a blown fuse, or a failed supply on one branch takes down everything on it while the other branch runs on. The tell is that the dead group maps to a power distribution boundary rather than a network or device one. And if the boundary follows a single protocol interface - the dead tags all come through one interface or driver while another interface at the same site is fine - the fault is that interface or its configuration, not the field. Reading which kind of boundary the outage follows names the class of the failed element.

Verifying the Element and Confirming Recovery

Verify the named element with a targeted test rather than acting on the boundary alone. If you suspect one controller, try to reach it directly; if you suspect a network segment, check whether the switch or branch is reachable; if you suspect a power branch, check any available power telemetry or ask a site contact to look at the relevant breaker. Each test confirms the specific element before a dispatch commits to it. The boundary tells you where to look; the targeted test confirms what you find there, and the two together justify the response.

Confirm recovery by watching exactly the dead group return while the survivors stay untouched. When the right element is restored - the controller rebooted, the switch replaced, the breaker reset - the previously dark tags come back and rejoin the survivors, and the boundary disappears. A recovery that brings back only some of the dead group, or that disturbs the survivors, means the element was misidentified or there is more than one fault. Matching the recovery to the predicted boundary is the proof the diagnosis was right.

In a cloud SCADA such as Merobix, a partial outage is easy to characterize because every tag carries its controller, interface, and site, so an engineer can sort the dead tags from the live ones and see the boundary directly rather than inferring it. Contrasting the failed group against the working group in one view is exactly the move this tree depends on, and watching the dark tags rejoin the survivors after a fix confirms the element without a second guess. The surviving half of the site is the reference that makes the failed half diagnosable.

When to Escalate

Escalate to the layer that owns the failed element once the boundary names it: controls for a dead controller or interface, network for a failed segment or switch, and electrical for a tripped branch or failed supply. Because the surviving tags prove the site is otherwise healthy, the dispatch can be narrow and specific - one box, one segment, one breaker - rather than a general site call, which is the efficiency a partial outage buys you.

Escalate with the boundary map in hand, because the single most useful thing you can give a responder is the exact line between what works and what does not, plus the element it points at. That map turns a site visit into a targeted repair and prevents the responder from re-diagnosing a whole site when three quarters of it was never affected. A partial outage is a gift of information; escalating without passing along the boundary throws that gift away.

Frequently Asked Questions

Why are some tags from a site working while others are dark?

Because the surviving tags prove the site has power and a working path to you, so the fault is at a boundary inside the site rather than a total loss. Draw the line between the dead tags and the live ones and find what the dead group shares that the survivors do not: a second controller, a downstream network segment, a separate power branch, or one protocol interface. That shared element is the failed one. A partial outage is easier to diagnose than a total one precisely because the working tags give you a healthy reference to contrast against.

The dead tags are all on one PLC and another PLC is fine. What does that mean?

It means the site link and power are healthy, because the other controller's tags are still arriving, so the fault is isolated to the dark controller or its connection. That controller has crashed, hung, lost its local network drop, or dropped into fault mode. The fix follows the single-controller path: reach it directly to confirm it is unresponsive, then power-cycle or clear the fault, on site if no remote reset exists. The rest of the site being fine tells the responder exactly which box needs attention and rules out a site-wide cause.

How do I know if a partial outage follows the network or the power?

Compare the boundary of the dead group against the site's topology. If the dead tags are all downstream of one switch or network branch while upstream devices survive, the boundary follows the network and that segment is the suspect. If the dead tags map instead to a shared power branch that the survivors do not use, the boundary follows the power and a tripped breaker or failed supply on that branch is likely. Reading whether the outage line matches the network layout or the power distribution names the class of the failed element before you dispatch.

More in OT Cybersecurity
Field Site Went Dark  •  Alarms With No Process Change  •  All Tags Flatlined at Once  •  Bad Quality Across a Site  •  OT Alert Triage  •  All OT Cybersecurity →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →