Field Site Went Dark: How to Triage It
A remote site just stopped reporting, every tag from it has gone quiet, and you have no one standing at the panel. Before you dispatch a truck or start rebooting, you want to know whether the site lost power, lost its communication path, or has a dead controller, because each one sends a different person with different parts. This page is a symptom-first triage tree that splits those three causes apart from the office, using only what the SCADA system still shows you.
Field Site Went Dark in one line: When a field site goes dark, work three questions in order: did other sites on the same path also drop (points at the network or head end, not the site), does the site still answer a ping or show carrier on its modem (comms path alive, controller suspect), and did the values freeze at plausible numbers or drop to bad quality (freeze suggests the poll stopped, bad quality suggests the device stopped). Isolate the layer before you dispatch.
First Checks You Can Do From the Office
Start with the checks that cost nothing and take seconds, because they often localize the fault before anyone leaves the building. Look at the exact time every tag from the site stopped: a clean simultaneous stop across all of them is a single common cause, which is almost always power, the comms link, or the controller itself, not a field sensor. A ragged stop where tags dropped over minutes points somewhere else entirely and moves you off this tree.
Next, check whether other sites that share the same path went dark at the same moment. If the site talks through a shared cellular carrier, a repeater, a poll engine, or a head-end router, and its neighbors dropped with it, the fault is upstream and dispatching to the site is wasted. Your own SCADA is the fastest instrument you have here: a group of sites failing together draws a map of the common element, and that map is usually more accurate than any single site alarm. If only this one site dropped, the fault is local to it and the rest of the tree applies.
Finally, read the last values the site sent before it went quiet. A slow decline in a battery-voltage or panel-voltage tag in the minutes before the stop is a power failure narrating itself. A comms-quality or signal-strength tag that degraded first points at the link. Values that were completely healthy right up to the instant they stopped, with no warning, point at an abrupt loss such as a tripped breaker, a severed line, or a controller crash. The pre-failure trend is evidence, and it is already recorded.
Split Power From Comms From Device
With the office checks done, the goal is to place the fault in exactly one of three layers, because each dispatches differently. The power layer means the site or its comms equipment lost supply: a tripped breaker, a dead battery on a solar site, a blown fuse, a failed power supply. The comms layer means the site has power but cannot reach you: a down modem, a carrier outage, a pulled Ethernet cable, an antenna problem. The device layer means power and comms are fine but the controller stopped answering: a crashed PLC or RTU, a hung program, a locked-up communications processor.
The single most useful discriminator is whether anything at the site still answers. If the site's cellular gateway or router still responds to a ping, or its modem still shows registered on the carrier, then power to that device is present and the comms path exists at least partway, which pushes the fault toward the controller behind it. If nothing at the site answers at any layer, power is the leading suspect because it takes down comms and controller together. A device that keeps dropping and recovering rather than staying dark is a different symptom with its own workflow, covered in the guide on diagnosing a cellular gateway that keeps dropping.
The last-value pattern refines the split further. When the poll simply stops but the controller is actually still alive, many SCADA systems freeze the tags at their last-received numbers and eventually mark them stale, so you see plausible frozen values that go bad-quality after a timeout. When the controller itself dies but the comms path is intact, you often get an explicit communication-failure indication or an exception response rather than a silent freeze. Reading which of those you got tells you whether to send a comms tech or a controls tech.
Confirming the Layer Before You Dispatch
Before committing a truck roll, prove the layer with one more test rather than acting on the first hypothesis. If you suspect comms, try to reach the site's gateway directly on its management path, bypassing the poll engine; reaching it proves the link and convicts the controller, failing to reach it convicts the link or power. If you suspect the controller, and the comms path is confirmed alive, a raw protocol read attempt that returns a connection but no valid response is strong evidence the device is hung. Each test is a bisection that halves the remaining possibilities.
The value of this discipline is that it changes who you send and what they carry. A power fault sends someone with a meter, spare fuses, and possibly a battery. A comms fault sends someone with a spare modem, antenna, and cable, or a call to the carrier before anyone drives. A device fault sends a controls technician who can power-cycle and reprogram the PLC. Guessing wrong means a second trip, and on a remote site a second trip can be a day lost. The triage tree exists to make the first dispatch the right one.
A cloud SCADA platform such as Merobix helps this triage because it timestamps every tag stop, records the pre-failure trend of power and signal tags, and shows at a glance whether a group of sites failed together or just one. Rather than logging into several tools to assemble the picture, an engineer sees the common-cause map and the last-good values in one place, which is exactly the evidence this tree runs on. The platform does not fix the site, but it tells the dispatcher which layer broke.
When to Escalate
Escalate to a coordinated response when multiple sites went dark together, because that is an infrastructure event - a carrier outage, a head-end failure, or a wide-area power problem - and it needs the network or carrier engaged, not a truck per site. Trying to work each site individually during a shared outage wastes the crew and delays the real fix. Treat a group failure as one incident with one root cause until proven otherwise.
Escalate a single dark site to a field dispatch once the office tree has placed the fault in a layer and a remote recovery has failed or is impossible. If the controller is confirmed hung and no remote reset exists, someone has to go. If power is confirmed lost and no remote switching is available, someone has to go. The point of the tree is to make that dispatch informed, not to avoid it. Send the right person, with the right parts, on the first trip, and hand them the pre-failure trend so they are not diagnosing from zero when they arrive.
Frequently Asked Questions
How do I know if a dark site is a power problem or a comms problem?
Check whether anything at the site still answers a ping or shows carrier on its modem. If the gateway responds, power to that device is present and the fault is more likely the controller or a partial comms issue. If nothing at the site answers at any layer, power is the leading suspect because losing supply takes down the modem and the controller together. The pre-failure trend of a battery or panel-voltage tag often settles it: a slow voltage decline before the stop is power narrating its own failure.
All my tags from one site stopped at exactly the same second. What does that mean?
A clean, simultaneous stop across every tag from a site is a single common cause, not a field sensor problem. The common element is almost always power, the communication link, or the controller itself, because those three are the only things every tag from that site shares. A ragged stop where tags dropped over minutes points at something else and moves you off the single-cause tree. The exact-second stop is actually good news because it narrows the search to three candidates.
Should I reboot the modem before I dispatch someone?
Only after the triage tree points at comms, and only if a remote reboot is available. Rebooting blindly can clear a transient fault, but it also destroys the evidence of what failed and can waste the one recovery attempt on the wrong layer. Work the office checks first: confirm whether neighbors dropped too, read the pre-failure trend, and test whether the gateway still answers. If those point at the link and a remote reboot exists, then reboot; if they point at power or the controller, a modem reboot changes nothing.
Automation services
Need help turning this into a working system?
Merobix integrates SCADA, programs Allen-Bradley and Siemens PLCs, and designs and fabricates industrial control panels.
Meeting requests are reviewed before confirmation.