infra/RENUMBER.md
Graham McIntire 39bedb08d1
renumber infrastructure from 10.0.15.0/24 to 10.0.16.0/22 (hosts in 10.0.19.x)
Proxmox: node1-3 → 10.0.19.101-103
Talos cp1-3 → 10.0.19.1-3, workers → 10.0.19.4-6
K8s endpoint → VIP https://10.0.19.10:6443
Ansible: prom → 10.0.19.31, db → 10.0.19.30
Talos: added VIP block to controlplane.yaml base config
Promtail: Loki URL → 10.0.19.31
Docs: all references updated, talos4 removed
2026-07-18 08:28:47 -05:00

9.2 KiB
Raw Permalink Blame History

Renumber: 10.0.15.0/24 → 10.0.16.0/22 (hosts in 10.0.19.x)

Move Proxmox nodes, Talos cluster, and support VMs onto 10.0.16.0/22 (mask 255.255.252.0, broadcast 10.0.19.255). Server range: 10.0.19.x. Delete this file when done.

Address map

Host MAC Old New
talos-cp1 BC:24:11:9B:48:92 10.0.15.1 10.0.19.1
talos-cp2 BC:24:11:62:7B:3F 10.0.15.2 10.0.19.2
talos-cp3 BC:24:11:D4:2F:ED 10.0.15.3 10.0.19.3
talos-worker1 BC:24:11:43:3F:FF 10.0.15.4 10.0.19.4
talos-worker2 BC:24:11:62:C4:8F 10.0.15.5 10.0.19.5
talos-worker3 BC:24:11:3F:8E:1A 10.0.15.6 10.0.19.6
K8s API VIP — (floats on CPs) 10.0.19.10
db.w5isp.com static 10.0.15.30 10.0.19.30
prom.w5isp.com static 10.0.15.31 10.0.19.31
node1 static 10.0.15.101 10.0.19.101
node2 static 10.0.15.102 10.0.19.102
node3 static 10.0.15.103 10.0.19.103
router/gateway 10.0.15.254 10.0.19.254

Ground rules

  • Both subnets must coexist on the same L2 for the whole migration — already satisfied: the formerly separate subnets now share one flat L2. Talos updates each etcd member's peer URL automatically when its IP changes, but only if the remaining members can still reach it — moving to an unreachable network bricks etcd (siderolabs/talos#9075, #6080).
  • Control planes strictly one at a time. The etcd member update needs quorum (2/3 up).
  • Everything through Phase 3 step 2 is reversible by putting the DHCP reservation back and rebooting the node.

Phase 0 — Backups + router prep

  1. Backups (from talos/):
    talosctl --talosconfig talosconfig -n 10.0.15.1 etcd snapshot etcd-$(date +%F).snap
    cp secrets.yaml talosconfig ~/backups/  # or wherever off-repo copies live
    
    (DB dumps intentionally skipped — operator decision 2026-07-17.)
  2. Router:
    • Add 10.0.19.254/22 to the LAN interface alongside 10.0.15.254/24.
    • New DHCP scope on 10.0.16.0/22, gateway 10.0.19.254. Keep the dynamic pool out of 10.0.19.010.0.19.127 (statics + VIP live there). 10.0.19.10 (VIP) must never be leased.
    • Do not change the Talos MAC reservations yet — that happens per-node in Phase 3.
    • Clone any firewall rules referencing 10.0.15.0/24 to also cover 10.0.16.0/22 (delete the old ones in Phase 5).
  3. Sanity: from the workstation, ping 10.0.19.254.

Phase 1 — Proxmox nodes (node1/2/3)

Guests keep running throughout; no VM reboots. Corosync quorum drops mid-change — expected, harmless without HA.

  1. Discover what actually references the IPs, per node:
    ssh root@10.0.15.101 "grep -r '10\.0\.15' /etc/pve /etc/network/interfaces /etc/hosts; pveceph status 2>&1 | head -2"
    
    (pveceph status should say Ceph isn't configured — cluster/ceph-csi/ is stale. If Ceph is somehow alive, STOP; mon renumbering is its own runbook.)
  2. If /etc/pve/corosync.conf ring0_addrs are hostnames: no corosync edit needed — only /etc/hosts. If they're IPs: on ONE node while still quorate, edit /etc/pve/corosync.conf — all three ring0_addrs to 10.0.19.x and bump config_version.
  3. Then on each node in turn (node1 → node2 → node3):
    # /etc/network/interfaces: address 10.0.19.10X/22, gateway 10.0.19.254
    # /etc/hosts: own + peer entries to 10.0.19.x
    ifreload -a
    systemctl restart corosync pve-cluster
    
  4. After all three: pvecm status shows 3/3 quorate on new IPs; Proxmox web UI reachable at https://10.0.19.101:8006.
    • If pmxcfs wedges (can't write /etc/pve): systemctl stop pve-cluster corosync; pmxcfs -l to fix corosync.conf locally, then restart both.
  5. Update wherever node1/node2/node3 resolve on the workstation (/etc/hosts or ~/.ssh/config) — ansible inventory uses ansible_host=node1 by name.

Phase 2 — db + prom VMs

Use the Proxmox console (not SSH — you're cutting the branch you sit on).

  1. On each VM, update the static netconfig (netplan or /etc/network/interfaces): 10.0.19.30/22 and 10.0.19.31/22, gateway 10.0.19.254. Restart networking.
  2. Verify SSH: ssh ansible@10.0.19.31 (prom's ufw allows are 10.0.0.0/16-based, so no lockout exposure).
  3. Loki ingestion from the cluster (promtail → 10.0.15.31:3100) is now broken until Phase 4's configmap push — known gap.

Phase 3 — Talos cluster

Run from talos/. Cluster endpoint is still https://10.0.15.1:6443 and stays valid until step 4.

1. Workers, one at a time (worker1 → worker2 → worker3):

# a) On router: change THIS node's reservation to its 10.0.19.x address
talosctl --talosconfig talosconfig -n 10.0.15.4 reboot
# b) wait, then verify:
talosctl --talosconfig talosconfig -n 10.0.19.4 version
kubectl get nodes -o wide        # INTERNAL-IP updated, node Ready

2. Add the VIP to the control planes (still on old IPs). Create patches/vip.yaml:

machine:
  network:
    interfaces:
      - deviceSelector:
          physical: true
        dhcp: true
        vip:
          ip: 10.0.19.10
for ip in 10.0.15.1 10.0.15.2 10.0.15.3; do
  talosctl --talosconfig talosconfig -n $ip patch machineconfig --patch @patches/vip.yaml
done
curl -sk https://10.0.19.10:6443/version   # any TLS response = VIP is up

3. Flip the cluster endpoint to the VIP on ALL 6 nodes. Create patches/endpoint.yaml:

cluster:
  controlPlane:
    endpoint: https://10.0.19.10:6443
for ip in 10.0.15.1 10.0.15.2 10.0.15.3 10.0.19.4 10.0.19.5 10.0.19.6; do
  talosctl --talosconfig talosconfig -n $ip patch machineconfig --patch @patches/endpoint.yaml
done
talosctl --talosconfig talosconfig kubeconfig -n 10.0.15.1 --force   # kubeconfig now points at the VIP
kubectl get nodes    # all Ready via https://10.0.19.10:6443

4. Control planes, strictly one at a time (cp2 → cp3 → cp1):

# a) On router: change THIS node's reservation
talosctl --talosconfig talosconfig -n 10.0.15.2 reboot
# b) wait, then verify BEFORE touching the next one:
talosctl --talosconfig talosconfig -n 10.0.19.2 etcd members
#    -> this member's peer URL shows 10.0.19.2, 3/3 members listed
talosctl --talosconfig talosconfig -n 10.0.19.2 health

If a member gets stuck on the old peer URL: quorum check first (etcd members from a healthy node), then reboot the stuck node once more before anything drastic. Worst case: restore the Phase 0 snapshot.

5. Client configs:

talosctl --talosconfig talosconfig config endpoint 10.0.19.1 10.0.19.2 10.0.19.3
# (real node IPs — never the VIP for the Talos API)
talosctl --talosconfig talosconfig -n 10.0.19.1 health

Phase 4 — Repo + config sweep

Edit, commit, push (promtail change must reach codeberg for Argo):

  • talos/controlplane.yaml:45 + talos/worker.yaml:44 — endpoint https://10.0.19.10:6443; add the machine.network.interfaces VIP block to controlplane.yaml so rebuilds match reality
  • ansible/hosts — prom ansible_host=10.0.19.31
  • ansible/host_vars/prom.w5isp.com.yml — ansible_host + subnet comment
  • ansible/host_vars/db.w5isp.com.ymlansible_host: 10.0.19.30
  • ansible/roles/prometheus/defaults/main.ymlprometheus_k8s_api_server: "https://10.0.19.10:6443" (VIP is in the apiserver cert SANs) + commented example targets
  • ansible/roles/grafana/defaults/main.yml — root_url 10.0.19.31
  • home/cluster/promtail/configmap.yaml — Loki URL http://10.0.19.31:3100
  • Docs: CLAUDE.md, talos/CLAUDE.md, talos/README.md, cluster/README.md, ansible/roles/prometheus/README.md, home/cluster/starlink-exporter/deployment.yaml comment; cluster/ceph-csi/ is dead — delete or mark stale
  • Then: ansible-playbook -l prom.w5isp.com playbook.yml and confirm Argo synced promtail (kubectl -n promtail rollout status ds/promtail)

Phase 5 — Retire the old subnet

Only after rg "10\.0\.15\." ~/dev/infra returns nothing but history/docs:

  • Router: remove 10.0.15.254/24, old DHCP scope, old firewall rules.
  • If a Tailscale subnet router advertises 10.0.15.0/24, re-advertise 10.0.16.0/22 and approve in the admin console.
  • Hosts elsewhere in the /22 (e.g. birdnet 10.0.17.29) that still carry /24 masks should move to /22 to avoid hairpinning via the router.

Final verification

talosctl --talosconfig talosconfig -n 10.0.19.1 health
kubectl get nodes -o wide && kubectl get pods -A | grep -v Running
curl -sk https://10.0.19.10:6443/version
ssh root@10.0.19.101 pvecm status
open http://10.0.19.31:9090/targets   # prometheus targets green
# https://aprs.me and the other cloudflared-fronted apps load

References: talos#6080 — etcd peer URLs update on IP change, talos#9075 — what happens when old/new subnets can't reach each other, Talos multihoming/advertisedSubnets docs