Upgrades¶
Automatic (tuppr)¶
tuppr handles Talos and Kubernetes
upgrades declaratively. The target versions are pinned in two places, and
Renovate's talos and kubernetes groups update them together in one PR each
(the Talos group also bumps talosctl in .mise/config.toml):
kubernetes/talos/version.yaml: used byjust talosand bootstrapkubernetes/apps/system-upgrade/tuppr/upgrades/: theTalosUpgradeandKubernetesUpgraderesources
Merging the PR is the upgrade. tuppr upgrades one node at a time
(rebootMode: powercycle for Talos), and only starts each step when:
- no Kopiur
SnapshotisRunning - no Kopiur
RestoreisResolvingorRestoring - the
CephClusterreportsHEALTH_OK
Nodes due for an upgrade get the tuppr.home-operations.com/outdated taint.
Ceph daemons tolerate it, so they aren't evicted early.
What normal looks like¶
A full rolling Talos upgrade takes a while and logs some alarming but harmless events:
- image pre-pull retries before a node starts
- CNPG's Barman plugin pods being evicted and rescheduled
- Envoy taking around 3 minutes to drain connections
- the Ceph OSD's
expand-bluefsinit step taking around 75 seconds after a reboot
Watch progress with:
kubectl get talosupgrade,kubernetesupgrade
kubectl -n system-upgrade logs -l app.kubernetes.io/name=tuppr -f
Manual¶
When tuppr can't be used (it's broken, or a node needs a one-off), use the
just talos recipes. Each one asks for confirmation.
just talos upgrade-node <node> # Talos, using the node's schematic image
just talos upgrade-k8s <version> # Kubernetes, across the cluster
Upgrade one node at a time, and wait for ceph status to return to
HEALTH_OK before moving to the next.
Changing the schematic¶
Adding a system extension or kernel argument changes the schematic ID, and
therefore the installer image. Edit kubernetes/talos/schematic.yaml.j2, then
run just talos upgrade-node <node> for each node to move it onto the new
image at the current Talos version.
Applying machine config changes¶
Most changes to the Talos templates apply without a reboot: