Skip to content

Dual-WAN failover

The Gateway v2.14.0 can monitor multiple uplinks and move its active egress to a usable standby when the preferred uplink loses internet connectivity. It returns to the preferred uplink after a healthy hold period.

This release provides failover, not load balancing or bonding. The weight field is accepted and stored, but it does not distribute traffic across WAN links. The previous raw-WAN splitting recipe does not describe the supported v2.14.0 workflow.

A home or small office has a primary ISP connection and a backup, such as fibre plus LTE:

  • Prefer the uplink with the lowest priority number.
  • Detect an ISP outage even if the local router still answers ping.
  • Move to a usable standby and restore the preferred path with hysteresis.
  • Keep tunnel and privacy enforcement during the transition, accepting a reconnection window rather than bypassing protection.

The WAN manager is inactive with fewer than two enabled uplinks. Existing single-uplink installations do not require a configuration change.

Install or update using the installation guide. Use Interfaces → WAN Uplinks & Failover to edit interfaces, priorities, enabled state, optional gateways and probe targets.

Start with manual mode. It reports the switch the manager would make without changing the active uplink. Equivalent configuration:

network:
wan_iface: eth0
lan_iface: eth2
wan_failover_mode: manual
wans:
- iface: eth0
priority: 100
enabled: true
- iface: eth1
priority: 200
enabled: true

Interface names are examples: select the actual uplinks on your node, never the client-facing LAN. Lower priority numbers are preferred. Each uplink needs a usable address and a resolvable next hop. A standby with carrier but no address can receive a per-interface DHCP client from the manager. Primary and standby may be on the same LAN; that does not remove shared upstream failure risks.

  1. Confirm the UI lists the intended primary and standby, their health and their order.
  2. Inspect the manager’s reasons and proposed decisions in manual mode.
  3. Confirm each uplink reaches the internet, not just its first-hop router.
  4. Keep local or out-of-band access and a backup before a disruptive test.
  5. Change wan_failover_mode to auto only when the proposed decisions match the intended topology.

An omitted mode defaults to auto; set manual explicitly for a staged rollout. The manager picks up saved uplink changes on its next tick, normally within five seconds.

All paths below are under /api/v1 and require an appropriately scoped API key. Mutations require an operator role permitted by the server.

Method Path Purpose
GET /system/wan-uplinks Configured uplinks, mode, active interface, health and manager state.
PUT /system/wan-uplinks Replace the uplink list and mode after validation.
GET /system/wan-manager Active/standby selection, reasons, order, flap pins and timings.
GET /system/wan-health Per-uplink reachability results.

Read-only example; replace the example host and use its trusted CA:

Terminal window
curl --fail --silent --show-error \
-H "Authorization: Bearer $API_KEY" \
https://gateway.example:8080/api/v1/system/wan-manager

Reachability checks use device-bound TCP connections over each uplink’s own routing table and mark. Defaults are 1.1.1.1:443 and 8.8.8.8:443; configured per-uplink probe_targets are honoured. These are deliberate external connectivity checks. The manager does not send its probes to DNS port 53.

The same transition updates the kernel default, system-default policy, policy uplink routes, NAT, tunnel binding, conntrack and DNS pins. Existing sessions may need to reconnect; this is not seamless connection migration.

Situation Behaviour
Both uplinks healthy One preferred uplink carries egress; the other is standby.
Primary loses internet The manager selects a usable standby when failure is confirmed.
Primary recovers Failback is held for 60 seconds, subject to a 30-second minimum dwell.
Repeated flapping Three flaps in ten minutes trigger flap pinning. Inspect the manager’s reason and state.
Manual mode Health and proposed decisions are reported; the manager does not move egress.
Fewer than two enabled uplinks The WAN manager is inactive.

The project’s v2.14.0 release record reports three real primary-link outages on a Raspberry Pi 4: primary declared down after 20–22 seconds, failover after 21–24 seconds, and standby verification 22 pass / 0 fail, with 9/9 leak checks. Those timings describe that setup, not a guaranteed SLA.

In the recorded transitions, protected clients remained fail-closed for about 10 seconds while a tunnel rebound; DNS through the gateway was unavailable for the same window. See release verification for evidence scope.

  • No WAN load balancing, aggregate-bandwidth promise or channel bonding in v2.14.0.
  • The wan-failover alert is an event and does not automatically resolve on failback.
  • The stall watchdog takes its uplink set at startup; a newly added uplink is watched after the next daemon restart. Schedule that restart with a recovery path.
  • Repeat failover and failback tests on your own hardware, ISP and tunnel configuration before relying on automatic operation.
  • Re-run node and leak checks after a move; an available backup path is not by itself evidence of correct policy enforcement.