Skip to main content
JPMfg
Visitor III
July 16, 2026
Question

Reboot the failed HA node after HA cluster failover

  • July 16, 2026
  • 6 replies
  • 62 views

We’ve had several cases of memory exhaustion with different processes (node, wad, ips), and are using the “set failover-memory enable” setting to cause the Clusters to automatically fail-over when going into conserve mode. (as well as cpu-threshold)

While this is fine as it no longer causes prolonged service disruptions, it does leave the clusters in a degraded state: the failed node usually does not recovery by itself and needs to be rebooted in order to recover from the cause of the memory consumption and restore the cluster redundancy.

Is there a simple way (e.g. with automation stitches targeting only the currently active or passive node) to automatically trigger a reboot on the now passive node after such a failover event?

Or do we need a feature request to allow automatic reboot of the failed node after a failover that was triggered by an internal event (memory, processes, RIB/FIB, cpu)? We probably don’t want to auto-reboot after an external event (link failure/ping-probe fail).

6 replies

Toshi_Esumi
SuperUser
SuperUser
July 16, 2026

You’re probably looking like this:
 

However, you need to figure out and fix the cause(s) of memory exhaustion. It could be:

  • software issue
  • hardware issue
  • capacity issue

Open a ticket at TAC to start with.

 

Toshi

JPMfg
JPMfgAuthor
Visitor III
July 17, 2026

Yeah, but that does take time. Usually more time than to the next conserve-mode situation.
And both “node” and “wad” are notoriously full of memory holes, there will never be a release that fixes all of them entirely.

sjoshi
Staff
Staff
July 17, 2026

Hi ​@JPMfg 

When the device is going into high memory state it does the automatic failover. Do you still have access to the issue device after failover? If yes please collect logs before the reboot that can help TAC to identify the root cause.

System status    get system status
Memory overview    get system performance status (run 3x)
Memory details    diagnose hardware sysinfo memory
Conserve mode state    diagnose hardware sysinfo conserve
Top memory processes    diagnose sys top-mem 99
Live process monitor    diagnose sys top 1 40 (run 1 min)
Crash log    diagnose debug crashlog read
Session stats    diagnose sys session stat
Slab memory    diagnose hardware sysinfo slab
 

Thanks, Salon
JPMfg
JPMfgAuthor
Visitor III
July 17, 2026

> When the device is going into high memory state it does the automatic failover. 

Not by default. The default for the “config system ha” setting “memory-based-failover” is *disable*. Failover will not happen automatically if the primary node goes into conserve mode.

sjoshi
Staff
Staff
July 17, 2026

try to collect the mentioned logs when the issue is happening. You need to find out which daemon is actually causing issue and those logs are needed during the issue time

Thanks, Salon
Thought Leadership Security Summit. Outpace New Threats with AI - enhanced defense. Tuesday, Septmeber 15, 8:30 AM - 2:30 PM PT. The Golf Club at Newcastle, WA.
Fortinet Flag the Hack. Wednesday, August 26, 9:00 AM - 5:00 PM ET, COSM, Atlanta, GA.
Virtual event | September 2026. SASE summit. The age of autonomous trust. Register here!