6300F FPC ran out of memory after 16 days – kernel slab leak? FortiOS 7.6.6
Has anyone seen this on a 6300F or 6500F? Looking for a cleaner fix than rebooting the FPC.
Specifically wondering:
1. Is this a known bug in 7.6.x with a fix in a later build?
2. Is there any way to reclaim kernel slab memory without rebooting the FPC?
We had an incident last night where FPC1 on our 6300F started dropping packets after about 16 days of uptime. The other 5 FPCs were completely fine. Rebooted FPC1 and everything came back to normal immediately.
The log message we saw:fw_forward_handler line=788 msg="The system is in extreme-low-memory state. Drop the packet."
When we dug into it with diag hardware sysinfo memory we found the problem — SUnreclaim on FPC1 had grown to 22GB while every other FPC was sitting at around 600MB. MemFree on FPC1 was down to 2%.
At incident:
FPC1 - SUnreclaim: 22,029,000 kB :warning: - MemFree: 692,292 kB (2%)
FPC2 - SUnreclaim: 626,632 kB - MemFree: 21,902,960 kB (66%)
FPC3 - SUnreclaim: 621,352 kB - MemFree: 21,918,292 kB (66%)
FPC4 - SUnreclaim: 625,980 kB - MemFree: 21,930,484 kB (66%)
FPC5 - SUnreclaim: 620,464 kB - MemFree: 21,918,412 kB (66%)
FPC6 - SUnreclaim: 632,240 kB - MemFree: 21,928,744 kB (66%)
After FPC1 reboot:
FPC1 - SUnreclaim: 506,964 kB
FPC2 - SUnreclaim: 653,604 kB
FPC3 - SUnreclaim: 616,284 kB
FPC4 - SUnreclaim: 622,384 kB
FPC5 - SUnreclaim: 620,600 kB
FPC6 - SUnreclaim: 617,336 kB
Since SUnreclaim is non-reclaimable kernel memory this doesn't show up in process monitoring at all — we only caught it by running diag hardware sysinfo memory across all slots when the drops started.
For now we're monitoring daily with:diag hardware sysinfo memory | grep -E "Slot|SUnreclaim"
And will proactively reboot any FPC that goes over 5GB before it becomes a problem. The slot reboot is non-disruptive so it's manageable, just not ideal.
Platform: FortiGate 6300F
FortiOS: 7.6.6 build 3652 (GA)
```
