Previously Valkey ran as a sidecar container in the main deployment,
causing it to restart every time the web app deployed. This resulted
in connection errors and cache loss during deployments.
Changes:
- Created valkey-statefulset.yaml with StatefulSet resource (1 replica)
- Created valkey-service.yaml with headless service
- Removed Valkey sidecar container from deployment.yaml
- Updated REDIS_HOST from 127.0.0.1 to valkey-0.valkey.towerops.svc.cluster.local
- Added new resources to kustomization.yaml
Benefits:
- Valkey persists across web app deployments
- No connection errors during rolling updates
- Cache data preserved
- Independent scaling and resource management
Prevent WebSocket disconnections during deployments by:
1. Rolling Update Strategy:
- maxSurge: 1 (allow 1 extra pod during rollout)
- maxUnavailable: 0 (keep all pods running)
- minReadySeconds: 10 (wait before continuing rollout)
2. Graceful Shutdown:
- terminationGracePeriodSeconds: 30 (allow Phoenix to drain connections)
3. PodDisruptionBudget:
- minAvailable: 1 (ensure at least 1 pod always available)
- Prevents all pods from being terminated simultaneously
4. Improved Health Checks:
- Faster readiness probe (5s initial, 3s period)
- More aggressive success/failure thresholds
- New pods marked ready faster
This ensures new pods are fully ready and serving traffic before old
pods are terminated, maintaining WebSocket connections and preventing
'something went wrong' errors during deployments.
Secrets are now managed directly in the cluster rather than generated
from .envrc files. This fixes FluxCD reconciliation errors since .envrc
is gitignored and cannot be used in GitOps workflows.
All secrets have been backed up to 1Password for disaster recovery.