Dual-WAN failover
The Gateway v2.14.0 can monitor multiple uplinks and move its active egress to a usable standby when the preferred uplink loses internet connectivity. It returns to the preferred uplink after a healthy hold period.
This release provides failover, not load balancing or bonding. The weight field is accepted and stored, but it does not distribute traffic across WAN links. The previous raw-WAN splitting recipe does not describe the supported v2.14.0 workflow.
A home or small office has a primary ISP connection and a backup, such as fibre plus LTE:
- Prefer the uplink with the lowest priority number.
- Detect an ISP outage even if the local router still answers ping.
- Move to a usable standby and restore the preferred path with hysteresis.
- Keep tunnel and privacy enforcement during the transition, accepting a reconnection window rather than bypassing protection.
The WAN manager is inactive with fewer than two enabled uplinks. Existing single-uplink installations do not require a configuration change.
Configure uplinks
Section titled “Configure uplinks”Install or update using the installation guide. Use Interfaces → WAN Uplinks & Failover to edit interfaces, priorities, enabled state, optional gateways and probe targets.
Start with manual mode. It reports the switch the manager would make without changing the active uplink. Equivalent configuration:
network: wan_iface: eth0 lan_iface: eth2 wan_failover_mode: manual wans: - iface: eth0 priority: 100 enabled: true - iface: eth1 priority: 200 enabled: trueInterface names are examples: select the actual uplinks on your node, never the client-facing LAN. Lower priority numbers are preferred. Each uplink needs a usable address and a resolvable next hop. A standby with carrier but no address can receive a per-interface DHCP client from the manager. Primary and standby may be on the same LAN; that does not remove shared upstream failure risks.
Verify before enabling automatic moves
Section titled “Verify before enabling automatic moves”- Confirm the UI lists the intended primary and standby, their health and their order.
- Inspect the manager’s reasons and proposed decisions in manual mode.
- Confirm each uplink reaches the internet, not just its first-hop router.
- Keep local or out-of-band access and a backup before a disruptive test.
- Change
wan_failover_modetoautoonly when the proposed decisions match the intended topology.
An omitted mode defaults to auto; set manual explicitly for a staged rollout. The manager picks up saved uplink changes on its next tick, normally within five seconds.
API and health
Section titled “API and health”All paths below are under /api/v1 and require an appropriately scoped API key. Mutations require an operator role permitted by the server.
| Method | Path | Purpose |
|---|---|---|
GET |
/system/wan-uplinks |
Configured uplinks, mode, active interface, health and manager state. |
PUT |
/system/wan-uplinks |
Replace the uplink list and mode after validation. |
GET |
/system/wan-manager |
Active/standby selection, reasons, order, flap pins and timings. |
GET |
/system/wan-health |
Per-uplink reachability results. |
Read-only example; replace the example host and use its trusted CA:
curl --fail --silent --show-error \ -H "Authorization: Bearer $API_KEY" \ https://gateway.example:8080/api/v1/system/wan-managerReachability checks use device-bound TCP connections over each uplink’s own routing table and mark. Defaults are 1.1.1.1:443 and 8.8.8.8:443; configured per-uplink probe_targets are honoured. These are deliberate external connectivity checks. The manager does not send its probes to DNS port 53.
What happens during a move
Section titled “What happens during a move”The same transition updates the kernel default, system-default policy, policy uplink routes, NAT, tunnel binding, conntrack and DNS pins. Existing sessions may need to reconnect; this is not seamless connection migration.
| Situation | Behaviour |
|---|---|
| Both uplinks healthy | One preferred uplink carries egress; the other is standby. |
| Primary loses internet | The manager selects a usable standby when failure is confirmed. |
| Primary recovers | Failback is held for 60 seconds, subject to a 30-second minimum dwell. |
| Repeated flapping | Three flaps in ten minutes trigger flap pinning. Inspect the manager’s reason and state. |
| Manual mode | Health and proposed decisions are reported; the manager does not move egress. |
| Fewer than two enabled uplinks | The WAN manager is inactive. |
The project’s v2.14.0 release record reports three real primary-link outages on a Raspberry Pi 4: primary declared down after 20–22 seconds, failover after 21–24 seconds, and standby verification 22 pass / 0 fail, with 9/9 leak checks. Those timings describe that setup, not a guaranteed SLA.
In the recorded transitions, protected clients remained fail-closed for about 10 seconds while a tunnel rebound; DNS through the gateway was unavailable for the same window. See release verification for evidence scope.
Limits and operational checks
Section titled “Limits and operational checks”- No WAN load balancing, aggregate-bandwidth promise or channel bonding in v2.14.0.
- The
wan-failoveralert is an event and does not automatically resolve on failback. - The stall watchdog takes its uplink set at startup; a newly added uplink is watched after the next daemon restart. Schedule that restart with a recovery path.
- Repeat failover and failback tests on your own hardware, ISP and tunnel configuration before relying on automatic operation.
- Re-run node and leak checks after a move; an available backup path is not by itself evidence of correct policy enforcement.
Related
Section titled “Related”- Release notes — v2.14.0 changes and upgrade notes.
- Performance evidence & release verification — scope of recorded results.
- Fail-closed routing — why protected traffic may stop during recovery.
- Operations playbook — routine checks and recovery planning.