Automation Glossary • Set a failover hold-down timer

How to Set a Failover Hold-Down Timer

Merobix Engineering • • 5 min read

A failover hold-down timer is the setting that decides how long a link must look bad before the gateway gives up on it, and how long a recovered link must prove itself before the gateway trusts it again. Set too short, the gateway ping-pongs on every momentary blip; set too long, a real outage leaves the site dark while the timer counts down. This procedure is for the technician tuning a dual-SIM or dual-WAN gateway who needs the timer to ignore noise but react to genuine failures.

Back to Blog

Set a failover hold-down timer in one line: To set a failover hold-down timer, choose a detection window long enough to ride out momentary blips, a switch delay that commits to the backup only after the primary is convincingly down, and a failback dwell that requires the recovered primary to stay healthy before returning. Then test both edges: a brief blip should not trigger a switch, and a sustained outage should. Lengthen the timer if the site flaps, shorten it if outages linger.

Understand the Two Directions the Timer Governs

A hold-down timer works in both directions, and confusing them is the usual reason a gateway misbehaves. On the way out, it sets how long the primary must appear failed before the gateway switches to the backup, which keeps a one-second signal dip from triggering a needless switchover. On the way back, it sets how long the recovered primary must stay healthy before the gateway fails back, which stops a primary that flickers back to life from yanking the site off a working backup prematurely. The flapping problem this solves is described in the guide to the failover hold-down timer.

The principle is hysteresis: the threshold to leave a path and the threshold to return to it should not be the same, so the gateway does not oscillate at the boundary. A link that hovers right at the switching point is exactly where a naive failover thrashes, and the hold-down times are what give the gateway the memory to commit to a decision rather than reconsider every second.

Choose Values That Fit the Site

Set the detection window against the site's normal signal behavior. A site with a clean, stable primary can afford a short window because real drops are rare and decisive; a site whose primary is known to blip needs a longer window so ordinary noise does not read as failure. There is no universal number, because it depends on how the primary link actually behaves at this location, so base it on what the survey and early trends show rather than a default. The point is to make the timer longer than the site's typical blip and shorter than an outage you would consider intolerable.

Set the failback dwell generously. Failing back the instant the primary twitches alive is a classic cause of flapping, so require the primary to hold healthy for a meaningful, sustained period before returning to it. Many sites are better off with an aggressive failover and a patient failback: switch away from a broken primary quickly, but return to it only once it has clearly recovered. If the site would rather stay put than thrash, you can even prefer to hold on the backup until a maintenance window, which relates to the broader reconnect-timing logic in the guide to reconnect backoff for telemetry clients.

Test Both Edges and Record the Setting

Prove the timer with two tests. First, cause a brief blip on the primary shorter than the detection window and confirm the gateway does not switch, because a timer that reacts to a blip is set too short and will thrash in service. Second, cause a sustained outage longer than the window and confirm the gateway does switch and reconnects to your host, because a timer that ignores a real outage is set too long and leaves the site dark. Adjust and repeat until the gateway rides out the blip and reacts to the outage.

Then test failback: restore the primary and confirm the gateway returns only after the dwell period, once and cleanly, without oscillating. If it flaps, lengthen the failback dwell. This pair of tests, one short and one long, is what separates a hold-down timer that was configured from one that was verified.

Record the detection, switch, and failback times you settled on, and note the site behavior that justified them. When a platform such as Merobix trends the site's connectivity afterward, a gateway that starts flapping despite the timer is telling you the primary's behavior changed and the setting needs revisiting, while a clean single switch during a real outage confirms the timer is doing its job. The timer is tuned once; the trend tells you when the site outgrew it.

Frequently Asked Questions

What is a good hold-down time for cellular failover?

There is no universal value, because it depends on how the primary link behaves at the specific site. Set the detection window longer than the site's typical momentary blip and shorter than an outage you would find intolerable, and set a generous failback dwell so a recovered primary must prove stable before the gateway returns to it. Base the numbers on the site's real signal behavior, then test both edges.

Should failover and failback use the same timer?

No. Using the same threshold in both directions invites oscillation at the boundary. Favor a faster failover so a broken primary does not keep the site dark, and a slower, patient failback so a primary that flickers back to life cannot yank the site off a working backup. That asymmetry, a form of hysteresis, is what stops flapping.

My gateway ignores a real outage - is the hold-down too long?

Likely yes. If a sustained outage does not trigger a switch, the detection window is longer than the outage you care about, so the site stays dark waiting for the timer. Shorten the detection window until a genuine outage triggers the switch, while keeping it long enough to ignore ordinary blips, then re-test both edges.

Sources and verification

This page references the vendor products and their official documentation published by the organizations below. Editions, product capabilities, and documentation change over time - confirm current requirements and specifications directly with the source.

Merobix is not affiliated with, endorsed by, or sponsored by these organizations; their names are used only to identify the standards and products discussed.

More in PLC, RTU, HMI & DCS
Failover hold-down  •  Commission a Pump Anti-Cycling Timer  •  Deadman timer  •  Lone worker check-in timer  •  POC Timer Mode  •  All PLC, RTU, HMI & DCS →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →