Netgate Discussion Forum
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Search
    • Register
    • Login
    Introducing Netgate Nexus: Multi-Instance Management at Your Fingertips.

    Failover Advanced Option

    Scheduled Pinned Locked Moved Routing and Multi WAN
    6 Posts 4 Posters 2.5k Views 7 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • chpalmerC Offline
      chpalmer
      last edited by

      Curious if this would be an easy change? And if anyone else would like to see it..

      It would be nice in my opinion if an option for trigger level would include "Packet Loss and High Latency"..

      Im trying to speed up the failover just a bit without false triggering. Im a little spoiled with another commercial product out there that we use where I work. (failover happens within a second or two.)

      GatewayFailover.png

      Triggering snowflakes one by one..
      Primary- Intel(R) Pentium(R) CPU G4400 @ 3.30GHz on an M470 WG box. pfSense CE 2.8.1
      Lab Unit- Intel(R) Core(TM) i5-4590T CPU @ 2.00GHz on an M400 WG box. pfSense+

      1 Reply Last reply Reply Quote 0
      • G Offline
        giuliafw70
        last edited by

        I would be careful with an AND-style trigger here. In a real WAN problem you often see either loss or latency first, not both at the same time, so “packet loss and high latency” could actually delay failover instead of making it cleaner.

        For avoiding false triggers, I would first tune the gateway monitor target and the dpinger thresholds/intervals per gateway. A very fast 1–2 second failover is possible in some setups, but on an Internet WAN it can also flap quickly if the monitor host or upstream path has a short hiccup.

        chpalmerC 1 Reply Last reply Reply Quote 2
        • chpalmerC Offline
          chpalmer @giuliafw70
          last edited by

          @giuliafw70 Thanks!

          I only use "Packet Loss" now and it seems to bring it in around 13 seconds which isn't the end of the world.. I figured though if I had an "and" option that I could dial down to tighter numbers for packet loss and during an actual outage the latency would increase at the same time triggering a bit sooner.. but I see your point.

          During evening hours we tend to go up in latency to my first gateway to over 30ms. (its normally somewhere from 8ms to 12ms.) I do not want that by itself triggering an outage so would have to at least double that to stay away from false triggering.

          Ill keep trying to tighten it up a bit more.

          Triggering snowflakes one by one..
          Primary- Intel(R) Pentium(R) CPU G4400 @ 3.30GHz on an M470 WG box. pfSense CE 2.8.1
          Lab Unit- Intel(R) Core(TM) i5-4590T CPU @ 2.00GHz on an M400 WG box. pfSense+

          1 Reply Last reply Reply Quote 0
          • M Offline
            MalagaFirewall8
            last edited by

            One practical way to get closer without an AND trigger is to separate the two problems: use packet loss as the actual trigger, then make latency less sensitive than the evening baseline. If your WAN normally sits at 8-12 ms but sometimes reaches 30 ms, I’d avoid setting the high-latency alarm anywhere near 30; use the dpinger advanced values per gateway and leave enough headroom that the monitor target itself does not become the flap source. Also worth testing with a monitor IP inside the ISP/upstream path versus a public resolver, because that choice can change the failover time more than the trigger label.

            1 Reply Last reply Reply Quote 1
            • dennypageD Offline
              dennypage
              last edited by

              The default of 2 pings per second offers a reasonable detection time while not annoying the ping target. That said, I've also seen people operate with 20 pings per second and a 5 second time interval to achieve very fast failover detection.

              It's a trade off. You can detect outages quickly, or you can be nice to your ICMP target. I generally recommend being nice to your target unless you have specific permission from the administrator.

              1 Reply Last reply Reply Quote 1
              • chpalmerC Offline
                chpalmer
                last edited by

                Thanks everyone..

                With a little playing I am down to 10 seconds which is not terrible even if we are on a Zoom call if it happens.

                Every time I believe our ISP has to good of an uptime to worry about it someone runs over a telephone pole somewhere or a tree falls taking down a fiber route.. Lost power to the node just a few days ago and it is painfully obvious that the ISP needs to replace the battery there..

                Ill keep messing with it.

                Triggering snowflakes one by one..
                Primary- Intel(R) Pentium(R) CPU G4400 @ 3.30GHz on an M470 WG box. pfSense CE 2.8.1
                Lab Unit- Intel(R) Core(TM) i5-4590T CPU @ 2.00GHz on an M400 WG box. pfSense+

                1 Reply Last reply Reply Quote 0
                • First post
                  Last post
                Copyright 2026 Rubicon Communications LLC (Netgate). All rights reserved.