# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Repository Overview This is a multi-environment infrastructure-as-code repository managing: - **Ansible**: Configuration management for servers across multiple environments (VNTX, home, app servers) - **OpenTofu**: Proxmox/Talos VM provisioning and Porkbun domain NS management - **Home Lab**: Server configurations with Tailscale integration ## Common Commands ### Proxmox VM Provisioning BIND9 and future infrastructure VMs are provisioned using the main Ansible playbook: ```bash cd ansible source .envrc # Load Proxmox credentials from .envrc # Provision and configure BIND9 nameserver at 204.110.191.222 ansible-playbook playbook.yml --tags provision,bind9 # VM settings (aligned with terraform/variables.tf): # - Node: vm2-380 (var.proxmox_host) # - Template: debian-13-cloudinit (VMID 9000) # - Storage: local-lvm (for BIND9 and general infrastructure VMs) # - Network: vmbr0 # # Note: Talos K3s nodes use different template and Longhorn for persistent storage ``` ### Ansible Operations ```bash # Run main playbook cd ansible && ./run.sh # Bootstrap new host make bootstrap-init HOSTNAME=hostname ANSIBLE_HOST=ip_address # Run specific playbook ansible-playbook general.yml # Run with specific tags ansible-playbook -t caddy general.yml # Gather facts from all hosts make facts ``` ### OpenTofu Operations **IMPORTANT**: Use OpenTofu (tofu) instead of terraform for all operations. ```bash cd tofu tofu init tofu plan tofu apply # Format terraform files before committing tofu fmt ``` ### Kubernetes Operations (K3s cluster at 204.110.191.2) ```bash # Set kubeconfig export KUBECONFIG=/Users/graham/dev/infra/home/ansible/kubeconfig # Deploy APRS.me application cd home/cluster/aprs ./deploy.sh # Check deployment status kubectl get pods -n aprs kubectl get svc -n aprs ``` ### Talos Operations (Home Cluster) ```bash # Change to talos directory cd /Users/graham/dev/infra/talos # Check cluster health talosctl --talosconfig talosconfig health # View cluster nodes kubectl get nodes # View all pods kubectl get pods -A # Check talos services talosctl --talosconfig talosconfig service # View logs talosctl --talosconfig talosconfig logs -f kubelet ``` ## Architecture ### Directory Structure - `ansible/` - Ansible playbooks and roles for server configuration - `roles/` - Reusable roles (base, caddy, docker, resolvers, etc.) - `group_vars/` - Group-specific variables - `host_vars/` - Host-specific variables - Main playbooks: `playbook.yml`, `general.yml`, `bootstrap.yml` - `tofu/` - OpenTofu infrastructure management - Proxmox/Talos VM provisioning (main.tf, talos_vms.tf) - Porkbun domain NS management (dns.tf) - Provider: `marcfrederick/porkbun ~> 1.3` - Credentials via `PORKBUN_API_KEY` and `PORKBUN_SECRET_API_KEY` env vars - Manages NS records for all Porkbun domains (as393837, cloudflare, or porkbun default NS) - Backend: Local state (committed to repo) - `home/` - Home lab configurations - `cluster/` - Kubernetes application deployments - `aprs/` - APRS.me deployment with PostgreSQL/PostGIS and Redis - `wavelog/` - Wavelog amateur radio logging application ### Key Infrastructure Services **Monitoring**: Icinga2 with MariaDB backend **DNS**: Unbound resolvers, PowerDNS for reverse zones **VPN**: Tailscale integration across environments **Web Proxy**: Caddy for reverse proxy and SSL termination ### Security Notes - Ansible user (UID 10001) with passwordless sudo - SSH keys fetched from GitHub for authentication - Tailscale requires `TAILSCALE_KEY` environment variable - OpenTofu credentials should use environment variables (PORKBUN_API_KEY, PORKBUN_SECRET_API_KEY) ## Development Workflow ### Adding New Hosts 1. Add host to `ansible/hosts` inventory 2. Create host_vars file if needed 3. Run bootstrap: `make bootstrap-init HOSTNAME=x ANSIBLE_HOST=y` 4. Apply configuration: `ansible-playbook -l hostname general.yml` ### DNS Zone Changes (BIND9) 1. Edit zone YAML in `ansible/group_vars/bind9_servers/.yml` 2. Run ansible playbook to deploy zone changes ### Domain NS Changes (Porkbun Registrar) 1. Edit `tofu/dns.tf` — update the domain's NS in the `local.domains` map 2. Run `tofu fmt` to format 3. Review with `tofu plan` 4. Apply with `tofu apply` ## Testing ### Ansible ```bash # Syntax check ansible-playbook --syntax-check playbook.yml # Dry run ansible-playbook --check playbook.yml # Test on specific host ansible-playbook -l hostname playbook.yml ``` ### Terraform ```bash tofu validate tofu plan ``` ## Recent Infrastructure Updates ### Talos Home Cluster (10.0.15.1-6) - **Purpose**: Primary home lab Kubernetes cluster - **Cluster Name**: home-cluster - **Kubernetes Version**: v1.35.0 - **Control Plane Nodes**: 3 (10.0.15.1-3) - **Worker Nodes**: 3 (10.0.15.4-6) - **API Endpoint**: https://10.0.15.1:6443 - **Proxmox Cluster**: "home" (node1, node2, node3 at 10.0.15.101-103) - **VM Disks**: Proxmox Ceph for VM boot disks - **Persistent Storage**: Longhorn on passthrough NVMe drives (`/dev/sdb` on each worker, mounted at `/var/mnt/longhorn`) - Worker1: 2TB, Worker2: 2TB, Worker3: 1TB - `longhorn` StorageClass is the default - `longhorn-system` namespace requires privileged PodSecurity policy - **CNI**: Flannel - **Talos Extensions** (workers only): - `iscsi-tools` v0.2.0 — required by Longhorn - `util-linux-tools` 2.41.2 — required by Longhorn - Installer image: `factory.talos.dev/installer/613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245:v1.12.6` - Control plane nodes use the stock `ghcr.io/siderolabs/installer:v1.12.6` - **Deployment**: Managed via OpenTofu in `/tofu/` - **Configuration**: `/talos/` directory - `talosconfig` - Talosctl client config - `secrets.yaml` - Cluster secrets - `controlplane.yaml` / `worker.yaml` - Base node configs - `configs/talos-worker{1,2,3}.yaml` - Per-node worker configs (with disk/extension settings) - `configs/talos-cp{1,2,3}.yaml` - Per-node control plane configs - **Applications**: `/cluster/` directory contains Kubernetes manifests - **Documentation**: See `/talos/README.md` and `/cluster/README.md` for detailed operations - **Config Notes**: - `install.extraKernelArgs` and `install.grubUseUKICmdline` cannot be used together - Worker disks with existing data (e.g. old Ceph LVM) must be wiped before Talos will partition them: `talosctl reset --user-disks-to-wipe /dev/sdb --wipe-mode user-disks --reboot` ### K3s Cluster at 204.110.191.2 - **Purpose**: Home lab Kubernetes cluster - **Node1 IPs**: - Public: 204.110.191.2 - Tailscale: 100.102.235.20 - Internal: 10.0.101.21 - **API Access**: Currently only available on 10.0.101.21:6443 (not on Tailscale) - **Context**: Use `kubectl config use-context homelab` to access - **Storage**: Longhorn for persistent volumes - **Applications**: - **APRS.me**: Amateur radio APRS tracking application (https://aprs.me) - PostgreSQL 17 with PostGIS (100GB Longhorn PVC) - Redis for caching - 2-replica StatefulSet for high availability - Erlang clustering enabled - LoadBalancer service assigned: 10.0.101.32 - **Note**: ghcr.io/aprsme/aprs.me:latest requires authentication - Deploy with: `cd home/cluster/aprs && ./deploy.sh` - **Wavelog**: Amateur radio logging application (https://log.w5isp.com) - MariaDB 11 database (10GB storage) - 20GB storage for configs, backups, uploads, QSL cards - PHP application with custom limits (50MB uploads, 256MB memory) - Deploy with: `cd home/cluster/wavelog && ./deploy.sh` ### New Server: w5isp.w5isp.com (204.110.191.200) - **Purpose**: Single-node server running Debian 13 - **Features**: - Web services with automatic SSL - Single public IP handles all services ### DNS Configuration - **Photos**: photos.w5isp.com → 204.110.191.212 (Caddy) → 100.107.11.77:2283 - **Mailcow**: mcintire.me mail records → mail.w5isp.com (sync.w5isp.com at 204.110.191.216) - Full mail setup with MX, SPF, DKIM, DMARC records - Autodiscover/Autoconfig for mail clients - Backup MX: mail.nsnw.ca (priority 20) - **Note**: sync.w5isp.com is the Mailcow server (not using Caddy proxy) ## Code Guidelines ### Communication Practices - Do not ever reply with "you're right". Just fix the issue - Never say you're right, just accept it and move on ### Notes for Claude - As you learn new things, keep claude.md updated ## Clusters - **Talos Home Cluster**: 10.0.15.1-6 (primary home lab cluster) - **K3s Cluster**: 204.110.191.2 (legacy, see above for details) ## Application Deployment Patterns ### Standard App Deployment Process 1. **Check for existing manifests** in `/home/cluster/[app-name]/` 2. **Set namespace security policy**: `kubectl label namespace [app] pod-security.kubernetes.io/enforce=privileged --overwrite` 3. **Use ClusterIP services** (internal only) with Traefik for external access 4. **Deploy via existing deploy.sh** scripts when available 5. **Use Longhorn storage** for persistent volumes 6. **Configure Traefik IngressRoute** for SSL and routing ### Deployed Applications - **APRS.me**: https://aprs.me (PostgreSQL + Redis, 2-replica StatefulSet) - **Node-RED**: https://nodered.w5isp.com (automation platform) - **n8n**: https://n8n.w5isp.com (workflow automation, PostgreSQL backend) - **Wavelog**: https://log.w5isp.com (amateur radio logging, MariaDB backend) ### Common Issues & Solutions - **PodSecurity violations**: Set namespace to privileged security policy - **Database connectivity**: Use short service names (`postgres` not `postgres.namespace.svc.cluster.local`) - **Init container permissions**: Required for volume ownership fixes - **SSL certificates**: cert-manager with Let's Encrypt, handled automatically via Traefik - **Service DNS**: Services accessible via `service-name` within same namespace ### Traefik Configuration - **Host-based routing**: Multiple apps on same IP (204.110.191.2) - **Automatic SSL**: cert-manager integration with Let's Encrypt - **HTTP to HTTPS redirect**: Configured in IngressRoute manifests - **ClusterIP backend**: All apps use ClusterIP, Traefik provides external access