Skip to content

BGP & Load Balancing

Every LoadBalancer IP in the cluster, including the Kubernetes API, is announced to the UniFi UDM over BGP. The NAS announces its own address the same way, using FRR in a container.

graph LR
    client(Client) -->|hashed flow| udm("`**UDM**
    _ASN 64513_`")
    udm -->|ECMP| k1("`**k8s-01**
    10.73.20.10`")
    udm -->|ECMP| k2("`**k8s-02**
    10.73.20.20`")
    udm -->|ECMP| k3("`**k8s-03**
    10.73.20.30`")
    udm -->|ECMP| nas("`**nas**
    10.73.1.10`")
    k1 & k2 & k3 -. "`**BGP** _ASN 64514_
    VIPs from 10.73.20.0/24`" .-> udm
    nas -. "`**BGP** _ASN 64515_
    VIP 10.73.1.10/32`" .-> udm
    client@{ shape: browser}
    nas@{ shape: lin-cyl }
ASN Speaker Announces
64513 UDM Pro Max (peer)
64514 Cilium on each node LoadBalancer IPs from 10.73.20.0/24
64515 FRR on the NAS (docker/nas/00-frr) 10.73.1.10/32

The Cilium side is configured in kubernetes/apps/kube-system/cilium/config/ (CiliumBGPClusterConfig, CiliumBGPPeerConfig, CiliumBGPAdvertisement).

Kubernetes API VIP

The Kubernetes API is fronted by the Cilium LoadBalancer Service kube-api (10.73.20.100) with externalTrafficPolicy: Local, so only nodes with a healthy apiserver announce the route. Static DNS in UniFi points k8s.internal at it.

When the CNI is down

k8s.internal depends on Cilium being healthy. If it isn't, reach the API directly at https://10.73.20.{10,20,30}:6443, and the Talos API at the same node addresses. Neither depends on the CNI. Bootstrap uses the node IP for the same reason.

UDM FRR config

UniFi accepts one FRR config upload per device (Settings → Routing Table → BGP):

router bgp 64513
  bgp router-id 10.73.0.254
  no bgp ebgp-requires-policy

  neighbor k8s peer-group
  neighbor k8s remote-as 64514

  neighbor 10.73.20.10 peer-group k8s
  neighbor 10.73.20.20 peer-group k8s
  neighbor 10.73.20.30 peer-group k8s

  neighbor nas peer-group
  neighbor nas remote-as 64515

  neighbor 10.73.1.10 peer-group nas

  address-family ipv4 unicast
    maximum-paths 3
    neighbor k8s next-hop-self
    neighbor k8s soft-reconfiguration inbound
    neighbor nas next-hop-self
    neighbor nas soft-reconfiguration inbound
  exit-address-family
exit

maximum-paths 3 enables true ECMP across the three nodes. FRR's eBGP default is a single best path.

Warning

Re-uploading the FRR config briefly bounces established BGP sessions.

ECMP flow hashing

The kernel default (fib_multipath_hash_policy=0) hashes on source and destination IP only, so a given client always lands on the same node. Policy 1 adds ports to the hash and spreads individual connections across the ECMP next-hops. This is persisted with a UDM boot script.

Verifying

On the UDM:

vtysh -c "show bgp summary"            # all sessions Established
vtysh -c "show ip bgp 10.73.20.100"    # every path tagged "multipath"
vtysh -c "show ip route"               # 10.73.20.100/32 with one path per healthy apiserver
ip route show 10.73.20.100             # one "nexthop" line per node

A single flat line in ip route means multipath is not installed in the kernel.

To check that flows are spread, run this a few times from one machine and expect the node name in the certificate SAN to vary:

openssl s_client -connect k8s.internal:6443 </dev/null 2>/dev/null \
  | openssl x509 -noout -ext subjectAltName

From a workstation: curl -k https://k8s.internal:6443/livez.