infra/CLAUDE.md

9.1 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Repository Overview

This is a multi-environment infrastructure-as-code repository managing:

  • Ansible: Configuration management for servers across multiple environments (VNTX, home, app servers)
  • OpenTofu: Proxmox/Talos VM provisioning and Porkbun domain NS management
  • Home Lab: Server configurations with Tailscale integration

Common Commands

Proxmox VM Provisioning

BIND9 and future infrastructure VMs are provisioned using the main Ansible playbook:

cd ansible
source .envrc  # Load Proxmox credentials from .envrc

# Provision and configure BIND9 nameserver at 204.110.191.222
ansible-playbook playbook.yml --tags provision,bind9

# VM settings (aligned with terraform/variables.tf):
# - Node: vm2-380 (var.proxmox_host)
# - Template: debian-13-cloudinit (VMID 9000)
# - Storage: local-lvm (for BIND9 and general infrastructure VMs)
# - Network: vmbr0
#
# Note: Talos K3s nodes use different template and Longhorn for persistent storage

Ansible Operations

# Run main playbook
cd ansible && ./run.sh

# Bootstrap new host
make bootstrap-init HOSTNAME=hostname ANSIBLE_HOST=ip_address

# Run specific playbook
ansible-playbook general.yml

# Run with specific tags
ansible-playbook -t caddy general.yml

# Gather facts from all hosts
make facts

OpenTofu Operations

IMPORTANT: Use OpenTofu (tofu) instead of terraform for all operations.

cd tofu
tofu init
tofu plan
tofu apply

# Format terraform files before committing
tofu fmt

Kubernetes Operations (K3s cluster at 204.110.191.2)

# Set kubeconfig
export KUBECONFIG=/Users/graham/dev/infra/home/ansible/kubeconfig

# Deploy APRS.me application
cd home/cluster/aprs
./deploy.sh

# Check deployment status
kubectl get pods -n aprs
kubectl get svc -n aprs

Talos Operations (Home Cluster)

# Change to talos directory
cd /Users/graham/dev/infra/talos

# Check cluster health
talosctl --talosconfig talosconfig health

# View cluster nodes
kubectl get nodes

# View all pods
kubectl get pods -A

# Check talos services
talosctl --talosconfig talosconfig service

# View logs
talosctl --talosconfig talosconfig logs -f kubelet

Architecture

Directory Structure

  • ansible/ - Ansible playbooks and roles for server configuration

    • roles/ - Reusable roles (base, caddy, docker, resolvers, etc.)
    • group_vars/ - Group-specific variables
    • host_vars/ - Host-specific variables
    • Main playbooks: playbook.yml, general.yml, bootstrap.yml
  • tofu/ - OpenTofu infrastructure management

    • Proxmox/Talos VM provisioning (main.tf, talos_vms.tf)
    • Porkbun domain NS management (dns.tf)
      • Provider: marcfrederick/porkbun ~> 1.3
      • Credentials via PORKBUN_API_KEY and PORKBUN_SECRET_API_KEY env vars
      • Manages NS records for all Porkbun domains (as393837, cloudflare, or porkbun default NS)
    • Backend: Local state (committed to repo)
  • home/ - Home lab configurations

    • cluster/ - Kubernetes application deployments
      • aprs/ - APRS.me deployment with PostgreSQL/PostGIS and Redis
      • wavelog/ - Wavelog amateur radio logging application

Key Infrastructure Services

Monitoring: Icinga2 with MariaDB backend DNS: Unbound resolvers, PowerDNS for reverse zones VPN: Tailscale integration across environments Web Proxy: Caddy for reverse proxy and SSL termination

Security Notes

  • Ansible user (UID 10001) with passwordless sudo
  • SSH keys fetched from GitHub for authentication
  • Tailscale requires TAILSCALE_KEY environment variable
  • OpenTofu credentials should use environment variables (PORKBUN_API_KEY, PORKBUN_SECRET_API_KEY)

Development Workflow

Adding New Hosts

  1. Add host to ansible/hosts inventory
  2. Create host_vars file if needed
  3. Run bootstrap: make bootstrap-init HOSTNAME=x ANSIBLE_HOST=y
  4. Apply configuration: ansible-playbook -l hostname general.yml

DNS Zone Changes (BIND9)

  1. Edit zone YAML in ansible/group_vars/bind9_servers/<domain>.yml
  2. Run ansible playbook to deploy zone changes

Domain NS Changes (Porkbun Registrar)

  1. Edit tofu/dns.tf — update the domain's NS in the local.domains map
  2. Run tofu fmt to format
  3. Review with tofu plan
  4. Apply with tofu apply

Testing

Ansible

# Syntax check
ansible-playbook --syntax-check playbook.yml

# Dry run
ansible-playbook --check playbook.yml

# Test on specific host
ansible-playbook -l hostname playbook.yml

Terraform

tofu validate
tofu plan

Recent Infrastructure Updates

Talos Home Cluster (10.0.15.1-6)

  • Purpose: Primary home lab Kubernetes cluster
  • Cluster Name: home-cluster
  • Kubernetes Version: v1.35.0
  • Control Plane Nodes: 3 (10.0.15.1-3)
  • Worker Nodes: 3 (10.0.15.4-6)
  • API Endpoint: https://10.0.15.1:6443
  • Proxmox Cluster: "home" (node1, node2, node3 at 10.0.15.101-103)
  • Storage: Proxmox Ceph for VM disks
  • CNI: Flannel
  • Deployment: Managed via OpenTofu in /tofu/
  • Configuration: /talos/ directory
    • talosconfig - Talosctl client config
    • secrets.yaml - Cluster secrets
    • controlplane.yaml / worker.yaml - Node configs
  • Applications: /cluster/ directory contains Kubernetes manifests
    • ceph-csi/ - Ceph CSI RBD driver for persistent storage (default StorageClass)
  • Documentation: See /talos/README.md and /cluster/README.md for detailed operations

K3s Cluster at 204.110.191.2

  • Purpose: Home lab Kubernetes cluster
  • Node1 IPs:
    • Public: 204.110.191.2
    • Tailscale: 100.102.235.20
    • Internal: 10.0.101.21
  • API Access: Currently only available on 10.0.101.21:6443 (not on Tailscale)
  • Context: Use kubectl config use-context homelab to access
  • Storage: Longhorn for persistent volumes
  • Applications:
    • APRS.me: Amateur radio APRS tracking application (https://aprs.me)
      • PostgreSQL 17 with PostGIS (100GB Longhorn PVC)
      • Redis for caching
      • 2-replica StatefulSet for high availability
      • Erlang clustering enabled
      • LoadBalancer service assigned: 10.0.101.32
      • Note: ghcr.io/aprsme/aprs.me:latest requires authentication
      • Deploy with: cd home/cluster/aprs && ./deploy.sh
    • Wavelog: Amateur radio logging application (https://log.w5isp.com)
      • MariaDB 11 database (10GB storage)
      • 20GB storage for configs, backups, uploads, QSL cards
      • PHP application with custom limits (50MB uploads, 256MB memory)
      • Deploy with: cd home/cluster/wavelog && ./deploy.sh

New Server: w5isp.w5isp.com (204.110.191.200)

  • Purpose: Single-node server running Debian 13
  • Features:
    • Web services with automatic SSL
    • Single public IP handles all services

DNS Configuration

  • Photos: photos.w5isp.com → 204.110.191.212 (Caddy) → 100.107.11.77:2283
  • Mailcow: mcintire.me mail records → mail.w5isp.com (sync.w5isp.com at 204.110.191.216)
    • Full mail setup with MX, SPF, DKIM, DMARC records
    • Autodiscover/Autoconfig for mail clients
    • Backup MX: mail.nsnw.ca (priority 20)
    • Note: sync.w5isp.com is the Mailcow server (not using Caddy proxy)

Code Guidelines

Communication Practices

  • Do not ever reply with "you're right". Just fix the issue
  • Never say you're right, just accept it and move on

Notes for Claude

  • As you learn new things, keep claude.md updated

Clusters

  • Talos Home Cluster: 10.0.15.1-6 (primary home lab cluster)
  • K3s Cluster: 204.110.191.2 (legacy, see above for details)

Application Deployment Patterns

Standard App Deployment Process

  1. Check for existing manifests in /home/cluster/[app-name]/
  2. Set namespace security policy: kubectl label namespace [app] pod-security.kubernetes.io/enforce=privileged --overwrite
  3. Use ClusterIP services (internal only) with Traefik for external access
  4. Deploy via existing deploy.sh scripts when available
  5. Use Longhorn storage for persistent volumes
  6. Configure Traefik IngressRoute for SSL and routing

Deployed Applications

Common Issues & Solutions

  • PodSecurity violations: Set namespace to privileged security policy
  • Database connectivity: Use short service names (postgres not postgres.namespace.svc.cluster.local)
  • Init container permissions: Required for volume ownership fixes
  • SSL certificates: cert-manager with Let's Encrypt, handled automatically via Traefik
  • Service DNS: Services accessible via service-name within same namespace

Traefik Configuration

  • Host-based routing: Multiple apps on same IP (204.110.191.2)
  • Automatic SSL: cert-manager integration with Let's Encrypt
  • HTTP to HTTPS redirect: Configured in IngressRoute manifests
  • ClusterIP backend: All apps use ClusterIP, Traefik provides external access