Bootstrap¶
Takes freshly installed Talos nodes all the way to a self-managing Flux
cluster, and provisions a fresh TrueNAS server with its Docker stacks.
Everything is driven by
bootstrap/mod.just.
Prerequisites¶
- mise installed and activated, and
mise installrun - A signed-in 1Password CLI (
op). Everyop://reference is resolved at apply time withop inject. - A valid
talosconfigat the repo root. The controller endpoint and node list are derived fromtalosctl config info, so nothing is hardcoded. - The UDM BGP config and
static DNS records in place.
k8s.internalonly works once Cilium is running, so bootstrap talks to the controller's node IP until then. - Nodes booted into Talos maintenance mode, for example from a
Bootimus PXE boot or an ISO from
just talos download-image <version>.
Cluster¶
graph LR
nodes --> k8s --> kubeconfig --> base --> apps
- nodes: renders each node's Talos config (templates plus 1Password
injection) and applies it with
talosctl apply-config --insecure. Nodes that are already configured are skipped. - k8s: runs
talosctl bootstrapagainst the controller, retrying until etcd reports that the cluster exists. - kubeconfig: fetches the kubeconfig and rewrites the server to the
controller's node IP, because the generated
https://k8s.internal:6443points at a Cilium VIP that doesn't exist yet. - base: waits for every apiserver to answer
/readyzand for nodes to register (they stayReady=Falseuntil the CNI is up), then applies:kustomize/: bootstrap Secrets rendered throughop inject(1Password Connect credentials and token, Cloudflare tunnel ID) plus their namespaces, so no controller deadlocks on a missing Secrethelmfile/crds.yaml: CRDs extracted from upstream charts (envoy-gateway, grafana-operator, kopiur, kube-prometheus-stack)
-
apps:
helmfile syncof the minimal release chain Flux needs before it can take over:cilium → coredns → spegel → cert-manager → external-secrets → onepassword-connect → flux-operator → flux-instanceThe kubeconfig is then fetched again, so its endpoint is back to
k8s.internal.
Once flux-instance is healthy, Flux reconciles kubernetes/ and manages the
same releases from then on.
Every stage is safe to re-run
If bootstrap fails partway, fix the issue and run just bootstrap cluster
again.
Single source of truth¶
The helmfiles define no chart versions or values of their own. Each release's
chart and version are read from the app's ocirepository.yaml, and its values
from the app's helmrelease.yaml, under kubernetes/apps/ (see
bootstrap/kubernetes/helmfile/templates/).
Bootstrap therefore installs exactly what Flux will reconcile later, and
Renovate only updates one place.
Data¶
Bootstrap restores no application data itself. As Flux deploys each app, its
PVC is populated from the latest Kopiur snapshot, and CNPG clusters recover
from Barman. Pods stay Pending until their volume is restored. See
Backups → Deploy-or-restore.
Verifying¶
kubectl get nodes # all Ready
flux get ks -A | grep -v True # nothing stuck
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph status
curl -k https://k8s.internal:6443/livez
NAS¶
This runs the Ansible playbook against the hosts in ansible/inventory.yaml.
It provisions TrueNAS and deploys doco-cd, which then reconciles the
docker/nas/ stacks. Nothing in bootstrap/ is used again until the next
provisioning.