IPSEC Tunnel dropping its route
- February 19, 2016
- 3 replies
- 9437 views
Hello All,
Sorry for the long explanation - trying to answer foreseeable questions in advance!
We have a pair of FG-100Ds in a HA Active-Passive configuration, running v5.2.3. (Would like to upgrade to v5.2.5, but no one wants to approve the change!)
Connected to this are 22 small sites (couple of devices each) connected via IPSec tunnels, each with Cisco 887 (ADSL plus 3G Failover) and 4 x Phase 2 selectors.
The Fortigate has a private IP address on the WAN1 port, so all tunnels are enabled for NAT-T on both sides.
Head office has a range of 172 and 10 subnets, and each site has both a 10.x.x.0/28 and 172.x.x.0/28 subnet.
These tunnels were all configured by a third party in August 2015, and have been quite robust.
Needed to alter the Interesting Traffic selectors, after testing one site successfully, rolled out the change to each of the Cisco routers and on its Fortigate tunnel.
Then I started to notice semi-frequent dropouts in links from HQ to individual sites. Most sites drop 1 or 2 times a day, some up to 5 times a day, and some have never dropped, despite the fact configs have been checked 50 times and are identical.
When I check the Fortigate, I find the tunnels are reporting as up, but the route specific to that P2 has vanished.
VPN, Monitor, IPsec Monitor, Select the appropriate Phase 2, right click, Bring Down will resolve the issue within a second or two.
My thinking has been P1 / P2 timeouts.
P1 was correct - both sides 86,400 seconds - 1 day.
Then I discovered P2 was default on the Cisco (3,600 seconds, 4,608,000 Kbytes) and set to 43,200 on the Fortigate. Bingo!
Reconfigured all the P2s (all 88 of them) on the Fortigate to Key Lifetime "Both" Seconds 3,600 and Kilobytes 4,608,000 to match the defaults on the Cisco.
Reliability appears better, but still experiencing dropouts.
Some of the Cisco's have been rebooted, but the Fortigate has not.
DPD enabled both sides. NAT-T configured both sides. I believe MTU configuration on the Cisco is good.
Spent most of the last two weeks on this! We have site monitoring which pings every site every 30 minutes. Before the change this was stable. Now getting 30 or 40 alerts a day. Manually bringing down the tunnel whenever a site is unreachable resolves it, but sites are offline for ~15-20 minutes before I give it a kick, which is unacceptable.
If I don't intervene, it does come back by itself within an hour, again seeming to indicate timeout related.
The biggest thing I don't understand in all of this is how the tunnels were stable previously with the Phase 2 lifetimes being incorrect. I suppose since they were lower on the Cisco, it just re-established the P2, and the Fortigate said "Oh, OK!".
The fortigate hasn't been restarted in over 200 days, probably not since these tunnels were commissioned.
Screen grab of snippets of the interfaces view and routes view are attached.
Would like to create an explicit static route for each of the sites on the Fortigate pointing to their tunnel. But new route doesn't list the tunnels as a selectable interface. Setting in the CLI also fails. The routes from that screen shot are generated automatically and vanish when the issue occurs.
I feel like I need a way to completely reset each tunnel on the Fortigate (rebooting it has been tempting!) - perhaps just dropping the tunnel isn't resetting the lifetime back to 0 for all phases on both the Fortigate and Cisco?
Need help with advanced troubeshooting / monitoring CLI commands for the Fortigate. Have done most of my debug on the Cisco's thus far, and haven't found anything.
Help!
TIA
Tony
