# SRE/DevOps Best Practices Reference > Practical patterns, real configs, and battle-tested commands. Not academic theory. --- ## Table of Contents 1. [Reliability Engineering](#1-reliability-engineering) 2. [Observability](#2-observability) 3. [Infrastructure as Code](#3-infrastructure-as-code) 4. [Kubernetes](#4-kubernetes) 5. [CI/CD](#5-cicd) 6. [Networking & Security](#6-networking--security) 7. [Database Operations](#7-database-operations) 8. [On-Call & Alerting](#8-on-call--alerting) --- ## 1. Reliability Engineering ### SLIs, SLOs, and SLAs **SLI (Service Level Indicator)** — A quantitative measure of some aspect of service. **SLO (Service Level Objective)** — A target value or range for an SLI. **SLA (Service Level Agreement)** — A contract with consequences if the SLO is missed. Rule of thumb: SLA threshold < SLO threshold (always have internal margin). #### Common SLIs | Service Type | SLI | How to Measure | |---|---|---| | HTTP API | Availability | `successful_requests / total_requests` | | HTTP API | Latency | `requests < 300ms / total_requests` | | Data pipeline | Freshness | `time_since_last_successful_run` | | Storage | Durability | `intact_objects / total_objects` | | Batch job | Throughput | `records_processed / expected_records` | #### Defining SLOs — Concrete Example ```yaml # slo.yaml — Structured SLO definition service: towerops-api slos: - name: availability sli: "ratio of HTTP requests with status < 500" target: 99.9% # 43.8 min/month downtime budget window: 30d rolling - name: latency-p99 sli: "99th percentile response time" target: 500ms window: 30d rolling - name: latency-p50 sli: "50th percentile response time" target: 100ms window: 30d rolling ``` #### Error Budgets Error budget = `1 - SLO`. For a 99.9% availability SLO over 30 days: ``` Budget = 0.1% × 30 days × 24 hours × 60 minutes = 43.2 minutes of downtime ``` **Error budget policy:** | Budget Remaining | Action | |---|---| | > 50% | Ship features freely | | 20–50% | Require extra review, feature flag new code | | 5–20% | Freeze non-critical changes, focus on reliability | | < 5% | Full freeze — only reliability work ships | PromQL to track error budget consumption: ```promql # Error budget remaining (0 = exhausted, 1 = full) 1 - ( ( sum(rate(http_requests_total{status=~"5.."}[30d])) / sum(rate(http_requests_total[30d])) ) / (1 - 0.999) ) ``` ### Toil Reduction **Toil** = manual, repetitive, automatable, tactical, devoid of long-term value, scales linearly with service growth. #### Identifying Toil Track it explicitly. Every on-call rotation, log toil: ```markdown ## Toil Log — Week of 2026-03-09 | Task | Time Spent | Frequency | Automatable? | Ticket | |---|---|---|---|---| | Restart stuck workers | 15 min | 3x/week | Yes — add health check + auto-restart | TOIL-042 | | Manually rotate certs | 30 min | Monthly | Yes — cert-manager | TOIL-043 | | Resize PVCs | 20 min | 2x/month | Yes — auto-expand policy | TOIL-044 | ``` **Target: Keep toil below 50% of on-call time.** If it's above that, you're not doing engineering — you're doing ops. #### Automation Priority Matrix ``` High frequency + Easy to automate → DO IT NOW High frequency + Hard to automate → Invest time, high ROI Low frequency + Easy to automate → Quick win, do it Low frequency + Hard to automate → Probably not worth it ``` ### Incident Response #### Severity Levels | Severity | Criteria | Response Time | Examples | |---|---|---|---| | **SEV-1** | Complete outage, data loss risk, security breach | 15 min | API down, DB corruption, credentials leaked | | **SEV-2** | Major feature degraded, significant user impact | 30 min | Auth broken, payments failing, 50%+ error rate | | **SEV-3** | Minor feature degraded, workaround exists | 4 hours | Slow queries, non-critical service down | | **SEV-4** | Cosmetic, minimal impact | Next business day | Dashboard glitch, log noise | #### Incident Response Process ``` 1. DETECT → Alert fires or user reports 2. TRIAGE → Assign severity, page on-call if needed 3. MITIGATE → Stop the bleeding (rollback, scale, redirect) 4. DIAGNOSE → Find root cause (but mitigation comes first!) 5. RESOLVE → Fix the underlying issue 6. FOLLOWUP → Postmortem within 48 hours ``` **Key principle: Mitigate first, diagnose second.** Rolling back a deploy takes 2 minutes. Finding the bug in the deploy might take 2 hours. Your users don't care about root cause during an outage. #### Incident Commander Checklist ```markdown ## Incident Commander Checklist - [ ] Declare incident and severity in #incidents channel - [ ] Assign roles: IC (you), Communications Lead, Ops Lead - [ ] Start incident doc (copy from template) - [ ] Set up war room (video call link in channel) - [ ] Post status update every 15 minutes (SEV-1) or 30 minutes (SEV-2) - [ ] Communicate to stakeholders via status page - [ ] When resolved: announce resolution, schedule postmortem ``` #### Runbook Template See [Section 8](#runbook-template) for the full runbook template. ### Capacity Planning #### The Four Signals to Monitor ```promql # 1. Utilization — How full is the resource? avg(node_cpu_seconds_total{mode!="idle"}) by (instance) sum(container_memory_working_set_bytes) / sum(machine_memory_bytes) # 2. Saturation — How queued/overloaded is it? rate(node_disk_io_time_weighted_seconds_total[5m]) avg_over_time(node_load15[1h]) / count(node_cpu_seconds_total{mode="idle"}) by (instance) # 3. Rate of growth predict_linear(node_filesystem_avail_bytes[7d], 30*24*3600) # 30-day projection # 4. Headroom — How much buffer remains? (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes ``` #### Capacity Planning Process ```bash # Export resource usage trends for capacity review # Run monthly, review in capacity planning meeting # CPU: average and peak over past 30 days curl -s "http://prometheus:9090/api/v1/query?query=\ quantile_over_time(0.95, instance:node_cpu_utilisation:rate5m[30d])" | jq '.data.result' # Memory: project when we hit 80% curl -s "http://prometheus:9090/api/v1/query?query=\ predict_linear(node_memory_MemAvailable_bytes[30d], 90*24*3600)" | jq '.data.result' # Disk: days until full curl -s "http://prometheus:9090/api/v1/query?query=\ node_filesystem_avail_bytes / ignoring(mountpoint) \ rate(node_filesystem_avail_bytes[7d]) / -86400" | jq '.data.result' ``` **Rule of thumb:** Plan for 2x current peak. If you're above 60% sustained utilization, start provisioning. --- ## 2. Observability The three pillars: **Metrics**, **Logs**, **Traces**. Each answers different questions: - **Metrics** → "What is broken?" (symptoms) - **Logs** → "Why is it broken?" (details) - **Traces** → "Where is it broken?" (request path) ### Prometheus + Grafana #### Prometheus Configuration ```yaml # prometheus.yml global: scrape_interval: 15s evaluation_interval: 15s external_labels: cluster: production environment: prod rule_files: - "/etc/prometheus/rules/*.yml" alerting: alertmanagers: - static_configs: - targets: ["alertmanager:9093"] scrape_configs: - job_name: "kubernetes-pods" kubernetes_sd_configs: - role: pod relabel_configs: # Only scrape pods with annotation prometheus.io/scrape: "true" - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] action: keep regex: true # Use custom port if annotated - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port] action: replace target_label: __address__ regex: (.+) replacement: ${1}:${2} # Use custom path if annotated - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path] action: replace target_label: __metrics_path__ regex: (.+) - job_name: "node-exporter" static_configs: - targets: ["node-exporter:9100"] - job_name: "postgres" static_configs: - targets: ["postgres-exporter:9187"] ``` #### Essential PromQL Queries ```promql # --- HTTP Request Rate --- # Requests per second by status code sum(rate(http_requests_total[5m])) by (status_code) # Error rate (5xx / total) sum(rate(http_requests_total{status_code=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) # --- Latency --- # P50/P95/P99 from histogram histogram_quantile(0.50, sum(rate(http_request_duration_seconds_bucket[5m])) by (le)) histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le)) histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le)) # Apdex score (satisfied < 0.3s, tolerating < 1.2s) ( sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m])) + sum(rate(http_request_duration_seconds_bucket{le="1.2"}[5m])) ) / 2 / sum(rate(http_request_duration_seconds_count[5m])) # --- Saturation --- # CPU saturation (load average > CPU count) node_load1 / count without(cpu, mode) (node_cpu_seconds_total{mode="idle"}) # Memory pressure 1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) # Disk I/O saturation rate(node_disk_io_time_weighted_seconds_total[5m]) # --- Kubernetes-specific --- # Pod restart rate (crash loops) rate(kube_pod_container_status_restarts_total[15m]) > 0 # Container OOMKills kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} # HPA at max replicas (can't scale further) kube_horizontalpodautoscaler_status_current_replicas == kube_horizontalpodautoscaler_spec_max_replicas # Pending pods (scheduling issues) kube_pod_status_phase{phase="Pending"} > 0 # PVC usage approaching limits kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes > 0.85 # --- PostgreSQL --- # Active connections vs max pg_stat_activity_count / pg_settings_max_connections # Slow queries (from pg_stat_statements exporter) topk(10, pg_stat_statements_mean_time_seconds) # Replication lag pg_replication_lag_seconds > 30 # Cache hit ratio (should be > 99%) pg_stat_database_blks_hit / (pg_stat_database_blks_hit + pg_stat_database_blks_read) ``` #### Alert Rules ```yaml # alerts.yml groups: - name: slo-alerts rules: # Multi-window multi-burn-rate alert (Google SRE book pattern) # Fast burn: 14.4x burn rate over 1h (uses 2% of budget in 1h) - alert: HighErrorBudgetBurn_Fast expr: | ( sum(rate(http_requests_total{status_code=~"5.."}[1h])) / sum(rate(http_requests_total[1h])) ) > (14.4 * 0.001) and ( sum(rate(http_requests_total{status_code=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) ) > (14.4 * 0.001) for: 2m labels: severity: critical annotations: summary: "High error budget burn rate" description: "Burning error budget at 14.4x. {{ $value | humanizePercentage }} error rate over 1h." runbook_url: "https://wiki.internal/runbooks/high-error-rate" # Slow burn: 1x burn rate over 3d (steady degradation) - alert: HighErrorBudgetBurn_Slow expr: | ( sum(rate(http_requests_total{status_code=~"5.."}[3d])) / sum(rate(http_requests_total[3d])) ) > (1 * 0.001) and ( sum(rate(http_requests_total{status_code=~"5.."}[6h])) / sum(rate(http_requests_total[6h])) ) > (1 * 0.001) for: 1h labels: severity: warning annotations: summary: "Slow error budget burn detected" description: "Steady error rate consuming budget. Review recent changes." - name: infrastructure rules: - alert: NodeHighCPU expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85 for: 15m labels: severity: warning annotations: summary: "High CPU on {{ $labels.instance }}" description: "CPU usage above 85% for 15 minutes. Current: {{ $value }}%" - alert: NodeHighMemory expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90 for: 10m labels: severity: warning annotations: summary: "High memory on {{ $labels.instance }}" description: "Memory usage above 90%. Current: {{ $value }}%" - alert: DiskWillFillIn24h expr: predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|fuse.lxcfs"}[6h], 24*3600) < 0 for: 30m labels: severity: warning annotations: summary: "Disk will fill within 24h on {{ $labels.instance }}" description: "Filesystem {{ $labels.mountpoint }} projected to run out of space." - alert: PodCrashLooping expr: rate(kube_pod_container_status_restarts_total[15m]) * 60 * 15 > 3 for: 5m labels: severity: warning annotations: summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} crash looping" description: "Container {{ $labels.container }} restarted {{ $value }} times in 15m." - name: postgres rules: - alert: PostgresHighConnections expr: pg_stat_activity_count / pg_settings_max_connections > 0.8 for: 5m labels: severity: warning annotations: summary: "PostgreSQL connections at {{ $value | humanizePercentage }}" - alert: PostgresReplicationLag expr: pg_replication_lag_seconds > 60 for: 5m labels: severity: critical annotations: summary: "Replication lag {{ $value }}s on {{ $labels.instance }}" - alert: PostgresLowCacheHitRatio expr: | pg_stat_database_blks_hit{datname!~"template.*"} / (pg_stat_database_blks_hit{datname!~"template.*"} + pg_stat_database_blks_read{datname!~"template.*"}) < 0.99 for: 15m labels: severity: warning annotations: summary: "Low cache hit ratio on {{ $labels.datname }}: {{ $value }}" ``` #### Alertmanager Configuration ```yaml # alertmanager.yml global: resolve_timeout: 5m slack_api_url: "https://hooks.slack.com/services/T00/B00/XXXX" route: receiver: "default" group_by: ["alertname", "namespace"] group_wait: 30s # Wait to batch related alerts group_interval: 5m # Wait before sending updated group repeat_interval: 4h # Re-notify interval routes: - match: severity: critical receiver: "pagerduty-critical" group_wait: 10s repeat_interval: 1h - match: severity: warning receiver: "slack-warnings" repeat_interval: 8h receivers: - name: "default" slack_configs: - channel: "#alerts" title: '{{ .GroupLabels.alertname }}' text: >- {{ range .Alerts }} *{{ .Labels.severity | toUpper }}*: {{ .Annotations.summary }} {{ .Annotations.description }} {{ end }} - name: "pagerduty-critical" pagerduty_configs: - service_key: "" description: '{{ .GroupLabels.alertname }}: {{ (index .Alerts 0).Annotations.summary }}' severity: critical - name: "slack-warnings" slack_configs: - channel: "#alerts-warnings" title: '⚠️ {{ .GroupLabels.alertname }}' text: '{{ (index .Alerts 0).Annotations.description }}' inhibit_rules: # Don't alert on warning if critical is already firing - source_match: severity: critical target_match: severity: warning equal: ["alertname", "namespace"] ``` ### Structured Logging **Stop logging unstructured strings.** Use JSON with consistent fields: ```json { "timestamp": "2026-03-13T21:27:00.000Z", "level": "error", "message": "Failed to process payment", "service": "billing-api", "trace_id": "abc123def456", "span_id": "span789", "request_id": "req-uuid-here", "user_id": "usr_12345", "error": { "type": "PaymentGatewayTimeout", "message": "Connection timed out after 30s", "stack": "..." }, "context": { "payment_amount": 49.99, "currency": "USD", "gateway": "stripe", "retry_count": 2 } } ``` #### Required Log Fields Every log line should have at minimum: ``` timestamp — ISO 8601 (with timezone) level — debug, info, warn, error, fatal message — Human-readable description service — Which service produced this trace_id — For correlation across services ``` #### Elixir Structured Logging Example ```elixir # config/config.exs config :logger, :default_formatter, format: {LoggerJSON.Formatters.Basic, :format}, metadata: [:request_id, :trace_id, :span_id, :user_id] # In application code Logger.error("Payment failed", user_id: user.id, payment_id: payment.id, error: inspect(reason), gateway: "stripe", amount: amount ) ``` ### OpenTelemetry Tracing #### Elixir Setup ```elixir # mix.exs defp deps do [ {:opentelemetry, "~> 1.4"}, {:opentelemetry_exporter, "~> 1.7"}, {:opentelemetry_phoenix, "~> 1.2"}, {:opentelemetry_ecto, "~> 1.2"}, ] end # config/runtime.exs config :opentelemetry, span_processor: :batch, exporter: {:opentelemetry_exporter, %{ protocol: :grpc, endpoints: [System.get_env("OTEL_EXPORTER_OTLP_ENDPOINT", "http://localhost:4317")] }} config :opentelemetry, :resource, service: [name: "towerops-api", version: Mix.Project.config()[:version]] # application.ex — instrument Phoenix and Ecto def start(_type, _args) do OpentelemetryPhoenix.setup(adapter: :cowboy2) OpentelemetryEcto.setup([:towerops, :repo]) # ... end ``` #### Custom Spans ```elixir require OpenTelemetry.Tracer, as: Tracer def process_device_command(device_id, command) do Tracer.with_span "device.process_command", %{ attributes: [ {"device.id", device_id}, {"command.type", command.type} ] } do with {:ok, device} <- lookup_device(device_id), {:ok, result} <- send_command(device, command) do Tracer.set_attribute("command.result", "success") {:ok, result} else {:error, reason} -> Tracer.set_status(:error, inspect(reason)) Tracer.set_attribute("command.result", "failure") {:error, reason} end end end ``` ### Log Aggregation with Loki #### Promtail Configuration (ships logs to Loki) ```yaml # promtail-config.yml server: http_listen_port: 9080 positions: filename: /tmp/positions.yaml clients: - url: http://loki:3100/loki/api/v1/push scrape_configs: - job_name: kubernetes-pods kubernetes_sd_configs: - role: pod pipeline_stages: # Parse JSON logs - json: expressions: level: level message: message trace_id: trace_id # Set log level as label - labels: level: # Drop debug logs in production - match: selector: '{level="debug"}' action: drop # Extract timestamp from log line - timestamp: source: timestamp format: RFC3339Nano relabel_configs: - source_labels: [__meta_kubernetes_pod_label_app] target_label: app - source_labels: [__meta_kubernetes_namespace] target_label: namespace ``` #### LogQL Query Examples ```logql # All errors from billing service in last hour {app="billing-api"} |= "error" | json | level="error" # Errors with specific trace ID (incident investigation) {namespace="production"} | json | trace_id="abc123def456" # Top 10 error messages sum by (message) (count_over_time({app="towerops-api"} | json | level="error" [1h])) # Latency from logs (if you log request duration) {app="towerops-api"} | json | duration > 1s # Rate of errors per service sum(rate({namespace="production"} | json | level="error" [5m])) by (app) ``` --- ## 3. Infrastructure as Code ### Terraform Patterns #### Project Structure ``` terraform/ ├── modules/ │ ├── vpc/ │ │ ├── main.tf │ │ ├── variables.tf │ │ ├── outputs.tf │ │ └── versions.tf │ ├── k8s-cluster/ │ ├── database/ │ └── cdn/ ├── environments/ │ ├── prod/ │ │ ├── main.tf # Composes modules │ │ ├── terraform.tfvars │ │ └── backend.tf │ ├── staging/ │ └── dev/ ├── global/ # Shared resources (DNS, IAM) │ ├── main.tf │ └── backend.tf └── scripts/ └── plan-all.sh ``` #### Module Pattern — VPC Example ```hcl # modules/vpc/variables.tf variable "name" { type = string description = "VPC name" } variable "cidr" { type = string default = "10.0.0.0/16" } variable "availability_zones" { type = list(string) description = "AZs for subnet distribution" } variable "enable_nat_gateway" { type = bool default = true } # modules/vpc/main.tf resource "aws_vpc" "main" { cidr_block = var.cidr enable_dns_hostnames = true enable_dns_support = true tags = { Name = var.name ManagedBy = "terraform" Environment = var.name } } resource "aws_subnet" "public" { count = length(var.availability_zones) vpc_id = aws_vpc.main.id cidr_block = cidrsubnet(var.cidr, 8, count.index) availability_zone = var.availability_zones[count.index] map_public_ip_on_launch = true tags = { Name = "${var.name}-public-${var.availability_zones[count.index]}" Tier = "public" } } resource "aws_subnet" "private" { count = length(var.availability_zones) vpc_id = aws_vpc.main.id cidr_block = cidrsubnet(var.cidr, 8, count.index + length(var.availability_zones)) availability_zone = var.availability_zones[count.index] tags = { Name = "${var.name}-private-${var.availability_zones[count.index]}" Tier = "private" } } # modules/vpc/outputs.tf output "vpc_id" { value = aws_vpc.main.id } output "public_subnet_ids" { value = aws_subnet.public[*].id } output "private_subnet_ids" { value = aws_subnet.private[*].id } ``` #### Environment Composition ```hcl # environments/prod/main.tf module "vpc" { source = "../../modules/vpc" name = "prod" cidr = "10.0.0.0/16" availability_zones = ["us-east-1a", "us-east-1b", "us-east-1c"] enable_nat_gateway = true } module "k8s" { source = "../../modules/k8s-cluster" cluster_name = "prod" vpc_id = module.vpc.vpc_id subnet_ids = module.vpc.private_subnet_ids node_count = 5 instance_type = "m6i.xlarge" k8s_version = "1.29" } module "database" { source = "../../modules/database" name = "prod" vpc_id = module.vpc.vpc_id subnet_ids = module.vpc.private_subnet_ids instance_class = "db.r6g.xlarge" engine_version = "16.2" multi_az = true backup_retention = 30 } ``` #### Remote State Configuration ```hcl # environments/prod/backend.tf terraform { backend "s3" { bucket = "towerops-terraform-state" key = "prod/terraform.tfstate" region = "us-east-1" dynamodb_table = "terraform-locks" encrypt = true } } # Reading state from another workspace data "terraform_remote_state" "global" { backend = "s3" config = { bucket = "towerops-terraform-state" key = "global/terraform.tfstate" region = "us-east-1" } } # Use: data.terraform_remote_state.global.outputs.route53_zone_id ``` #### Terraform Best Practices — Quick Reference ```bash # Always plan before apply terraform plan -out=plan.tfplan terraform apply plan.tfplan # Import existing resources terraform import aws_instance.web i-1234567890abcdef0 # Target specific resources (use sparingly) terraform plan -target=module.database # Move state (rename without destroy/recreate) terraform state mv aws_instance.old aws_instance.new # Detect drift terraform plan -detailed-exitcode # Exit 2 = drift detected # Workspace management (alternative to directory-per-env) terraform workspace new staging terraform workspace select staging terraform workspace list ``` **State locking is non-negotiable.** Always use DynamoDB (AWS) or equivalent. Two engineers running `terraform apply` simultaneously = destroyed infrastructure. ### NixOS for Reproducible Servers #### NixOS Server Configuration ```nix # /etc/nixos/configuration.nix { config, pkgs, ... }: { imports = [ ./hardware-configuration.nix ./services/prometheus.nix ./services/postgres.nix ./networking.nix ]; # System basics system.stateVersion = "24.11"; time.timeZone = "America/Chicago"; # Automatic updates with rollback capability system.autoUpgrade = { enable = true; allowReboot = false; # Manual reboot approval channel = "https://nixos.org/channels/nixos-24.11"; }; # Minimal packages environment.systemPackages = with pkgs; [ vim git htop curl jq ripgrep ]; # Firewall networking.firewall = { enable = true; allowedTCPPorts = [ 22 80 443 ]; allowedUDPPorts = [ ]; }; # SSH hardening services.openssh = { enable = true; settings = { PasswordAuthentication = false; PermitRootLogin = "no"; KbdInteractiveAuthentication = false; }; extraConfig = '' AllowUsers deploy MaxAuthTries 3 ''; }; # Automatic garbage collection nix.gc = { automatic = true; dates = "weekly"; options = "--delete-older-than 30d"; }; # Users users.users.deploy = { isNormalUser = true; extraGroups = [ "wheel" "docker" ]; openssh.authorizedKeys.keys = [ "ssh-ed25519 AAAA... deploy@towerops" ]; }; # Security security.sudo.wheelNeedsPassword = true; } ``` #### NixOS Service Module Example ```nix # services/prometheus.nix { config, pkgs, ... }: { services.prometheus = { enable = true; port = 9090; retentionTime = "90d"; globalConfig = { scrape_interval = "15s"; evaluation_interval = "15s"; }; scrapeConfigs = [ { job_name = "node"; static_configs = [{ targets = [ "localhost:9100" ]; }]; } { job_name = "postgres"; static_configs = [{ targets = [ "localhost:9187" ]; }]; } ]; rules = [ (builtins.toJSON { groups = [{ name = "system"; rules = [{ alert = "HighCPU"; expr = "100 - (avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) * 100) > 85"; for = "15m"; labels.severity = "warning"; }]; }]; }) ]; }; services.prometheus.exporters.node = { enable = true; port = 9100; enabledCollectors = [ "systemd" "processes" ]; }; services.grafana = { enable = true; settings = { server = { http_port = 3000; domain = "grafana.towerops.net"; }; security = { admin_password = "$__file{/run/secrets/grafana-admin-password}"; }; }; }; } ``` #### NixOS Deployment ```bash # Build without switching (test first) nixos-rebuild build --flake .#my-server # Switch to new configuration nixos-rebuild switch --flake .#my-server # Rollback to previous generation nixos-rebuild switch --rollback # List generations (your deployment history) nix-env --list-generations --profile /nix/var/nix/profiles/system # Remote deployment nixos-rebuild switch --flake .#prod-server \ --target-host deploy@prod.towerops.net \ --use-remote-sudo ``` **Key NixOS advantage:** Every deployment is atomic and rollbackable. If generation 42 breaks, `nixos-rebuild switch --rollback` takes you back to generation 41 in seconds. No Ansible "undo" scripts needed. --- ## 4. Kubernetes ### Deployment with Best Practices ```yaml # deployment.yaml apiVersion: apps/v1 kind: Deployment metadata: name: towerops-api namespace: production labels: app: towerops-api version: v1.4.2 spec: replicas: 3 revisionHistoryLimit: 5 # Keep 5 old ReplicaSets for rollback selector: matchLabels: app: towerops-api strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 # 1 extra pod during rollout maxUnavailable: 0 # Never drop below desired count template: metadata: labels: app: towerops-api version: v1.4.2 annotations: prometheus.io/scrape: "true" prometheus.io/port: "4000" prometheus.io/path: "/metrics" spec: serviceAccountName: towerops-api terminationGracePeriodSeconds: 60 # Don't schedule all replicas on the same node topologySpreadConstraints: - maxSkew: 1 topologyKey: kubernetes.io/hostname whenUnsatisfiable: DoNotSchedule labelSelector: matchLabels: app: towerops-api containers: - name: towerops-api image: ghcr.io/towerops/api:v1.4.2 # Always pin exact tags ports: - containerPort: 4000 protocol: TCP # Resource limits — ALWAYS set these resources: requests: cpu: 250m # Guaranteed minimum memory: 256Mi limits: cpu: "1" # Hard ceiling memory: 512Mi # OOMKilled if exceeded # Liveness: is the process alive? Restart if not. livenessProbe: httpGet: path: /healthz port: 4000 initialDelaySeconds: 15 periodSeconds: 10 timeoutSeconds: 3 failureThreshold: 3 # Readiness: can it serve traffic? Remove from LB if not. readinessProbe: httpGet: path: /readyz port: 4000 initialDelaySeconds: 5 periodSeconds: 5 timeoutSeconds: 3 failureThreshold: 2 # Startup: for slow-starting apps. Liveness/readiness wait until this passes. startupProbe: httpGet: path: /healthz port: 4000 failureThreshold: 30 # 30 × 2s = 60s to start periodSeconds: 2 # Graceful shutdown lifecycle: preStop: exec: command: ["/bin/sh", "-c", "sleep 5"] # Let LB drain connections env: - name: MIX_ENV value: "prod" - name: PHX_HOST value: "api.towerops.net" - name: PORT value: "4000" - name: DATABASE_URL valueFrom: secretKeyRef: name: towerops-secrets key: database-url - name: SECRET_KEY_BASE valueFrom: secretKeyRef: name: towerops-secrets key: secret-key-base volumeMounts: - name: tmp mountPath: /tmp volumes: - name: tmp emptyDir: sizeLimit: 100Mi # Security context securityContext: runAsNonRoot: true runAsUser: 1000 fsGroup: 1000 seccompProfile: type: RuntimeDefault ``` ### Health Check Implementation ```elixir # In your Phoenix router — separate liveness and readiness scope "/", TowerOpsWeb do # Liveness: "is the BEAM VM running?" get "/healthz", HealthController, :liveness # Readiness: "can we serve traffic?" # Checks DB connection, required services, etc. get "/readyz", HealthController, :readiness end # lib/towerops_web/controllers/health_controller.ex defmodule TowerOpsWeb.HealthController do use TowerOpsWeb, :controller def liveness(conn, _params) do conn |> put_status(200) |> json(%{status: "ok"}) end def readiness(conn, _params) do checks = %{ database: check_database(), redis: check_redis() } if Enum.all?(Map.values(checks), &(&1 == :ok)) do conn |> put_status(200) |> json(%{status: "ready", checks: checks}) else conn |> put_status(503) |> json(%{status: "not_ready", checks: checks}) end end defp check_database do case Ecto.Adapters.SQL.query(TowerOps.Repo, "SELECT 1", []) do {:ok, _} -> :ok _ -> :error end end defp check_redis do case Redix.command(:redix, ["PING"]) do {:ok, "PONG"} -> :ok _ -> :error end end end ``` ### Service & Ingress ```yaml # service.yaml apiVersion: v1 kind: Service metadata: name: towerops-api namespace: production spec: type: ClusterIP selector: app: towerops-api ports: - port: 80 targetPort: 4000 protocol: TCP --- # ingress.yaml apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: towerops-api namespace: production annotations: cert-manager.io/cluster-issuer: letsencrypt-prod nginx.ingress.kubernetes.io/rate-limit: "100" nginx.ingress.kubernetes.io/rate-limit-window: "1m" nginx.ingress.kubernetes.io/proxy-body-size: "10m" nginx.ingress.kubernetes.io/ssl-redirect: "true" nginx.ingress.kubernetes.io/force-ssl-redirect: "true" spec: ingressClassName: nginx tls: - hosts: - api.towerops.net secretName: towerops-api-tls rules: - host: api.towerops.net http: paths: - path: / pathType: Prefix backend: service: name: towerops-api port: number: 80 ``` ### Horizontal Pod Autoscaler ```yaml # hpa.yaml apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: towerops-api namespace: production spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: towerops-api minReplicas: 3 maxReplicas: 20 behavior: scaleUp: stabilizationWindowSeconds: 60 # Prevent flapping policies: - type: Pods value: 4 # Add max 4 pods at a time periodSeconds: 60 scaleDown: stabilizationWindowSeconds: 300 # Wait 5 min before scaling down policies: - type: Percent value: 25 # Remove max 25% at a time periodSeconds: 60 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70 - type: Resource resource: name: memory target: type: Utilization averageUtilization: 80 # Custom metric: requests per second per pod - type: Pods pods: metric: name: http_requests_per_second target: type: AverageValue averageValue: "100" ``` ### Pod Disruption Budget ```yaml # pdb.yaml — Ensure minimum availability during voluntary disruptions apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: towerops-api namespace: production spec: # At least 2 pods must always be running minAvailable: 2 # OR: at most 1 pod can be disrupted at a time # maxUnavailable: 1 selector: matchLabels: app: towerops-api ``` ### Helm vs Kustomize #### When to Use Each | Criteria | Helm | Kustomize | |---|---|---| | Third-party charts | ✅ Great | ❌ Not designed for this | | Environment overlays | Works (values files) | ✅ Built for this | | Templating | Full Go templates | Patches only (no templates) | | Complexity | Higher | Lower | | Learning curve | Steeper | Gentler | | Debugging | `helm template` + squint | WYSIWYG patches | **Recommendation:** Use Helm for third-party dependencies (nginx-ingress, cert-manager, prometheus). Use Kustomize for your own application manifests. #### Kustomize Example ``` k8s/ ├── base/ │ ├── kustomization.yaml │ ├── deployment.yaml │ ├── service.yaml │ └── ingress.yaml ├── overlays/ │ ├── staging/ │ │ ├── kustomization.yaml │ │ ├── replicas-patch.yaml │ │ └── env-patch.yaml │ └── production/ │ ├── kustomization.yaml │ ├── replicas-patch.yaml │ └── hpa.yaml ``` ```yaml # base/kustomization.yaml apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization resources: - deployment.yaml - service.yaml - ingress.yaml commonLabels: app: towerops-api # overlays/production/kustomization.yaml apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization namespace: production resources: - ../../base - hpa.yaml patches: - path: replicas-patch.yaml images: - name: ghcr.io/towerops/api newTag: v1.4.2 # overlays/production/replicas-patch.yaml apiVersion: apps/v1 kind: Deployment metadata: name: towerops-api spec: replicas: 3 ``` ```bash # Preview what will be applied kubectl kustomize overlays/production # Apply kubectl apply -k overlays/production # Diff against live cluster kubectl diff -k overlays/production ``` ### Kubernetes Operations — Quick Reference ```bash # --- Deployments --- kubectl rollout status deployment/towerops-api -n production kubectl rollout history deployment/towerops-api -n production kubectl rollout undo deployment/towerops-api -n production # Rollback to previous kubectl rollout undo deployment/towerops-api -n production --to-revision=3 # Specific revision # --- Debugging --- kubectl get pods -n production -o wide # Pod status + node kubectl describe pod -n production # Events, conditions kubectl logs -n production --tail=100 -f # Stream logs kubectl logs -n production -c --previous # Crashed container logs kubectl top pods -n production --sort-by=memory # Resource usage kubectl get events -n production --sort-by=.lastTimestamp # Recent events # --- Emergency --- kubectl scale deployment/towerops-api -n production --replicas=10 # Scale manually kubectl cordon # Prevent scheduling kubectl drain --ignore-daemonsets --delete-emptydir-data # Evacuate node ``` --- ## 5. CI/CD ### Pipeline Design Principles 1. **Fast feedback first** — Lint and unit tests run in < 2 minutes 2. **Fail fast** — Don't run slow integration tests if linting fails 3. **Immutable artifacts** — Build once, promote to each environment 4. **No manual steps** — If a human has to click something, automate it #### GitHub Actions Pipeline ```yaml # .github/workflows/ci.yml name: CI/CD on: push: branches: [main] pull_request: branches: [main] env: REGISTRY: ghcr.io IMAGE_NAME: ${{ github.repository }} jobs: lint-and-test: runs-on: ubuntu-latest services: postgres: image: postgres:16 env: POSTGRES_PASSWORD: test POSTGRES_DB: towerops_test ports: ["5432:5432"] options: >- --health-cmd pg_isready --health-interval 10s --health-timeout 5s --health-retries 5 steps: - uses: actions/checkout@v4 - uses: erlef/setup-beam@v1 with: elixir-version: "1.19" otp-version: "28" - uses: actions/cache@v4 with: path: | deps _build key: ${{ runner.os }}-mix-${{ hashFiles('mix.lock') }} - run: mix deps.get - run: mix format --check-formatted - run: mix compile --warnings-as-errors - run: mix credo --strict - run: mix dialyzer - run: mix test --cover env: DATABASE_URL: postgres://postgres:test@localhost/towerops_test build-and-push: needs: lint-and-test if: github.ref == 'refs/heads/main' runs-on: ubuntu-latest permissions: contents: read packages: write outputs: image-tag: ${{ steps.meta.outputs.version }} steps: - uses: actions/checkout@v4 - uses: docker/login-action@v3 with: registry: ${{ env.REGISTRY }} username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - id: meta uses: docker/metadata-action@v5 with: images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }} tags: | type=sha,prefix= type=semver,pattern={{version}} - uses: docker/build-push-action@v6 with: context: . push: true tags: ${{ steps.meta.outputs.tags }} cache-from: type=gha cache-to: type=gha,mode=max deploy-staging: needs: build-and-push runs-on: ubuntu-latest environment: staging steps: - uses: actions/checkout@v4 - name: Update image tag in kustomization run: | cd k8s/overlays/staging kustomize edit set image \ ghcr.io/towerops/api=ghcr.io/towerops/api:${{ needs.build-and-push.outputs.image-tag }} - name: Apply to staging run: kubectl apply -k k8s/overlays/staging env: KUBECONFIG: ${{ secrets.STAGING_KUBECONFIG }} - name: Wait for rollout run: kubectl rollout status deployment/towerops-api -n staging --timeout=300s env: KUBECONFIG: ${{ secrets.STAGING_KUBECONFIG }} - name: Smoke test run: | for i in {1..10}; do status=$(curl -s -o /dev/null -w '%{http_code}' https://staging-api.towerops.net/healthz) if [ "$status" = "200" ]; then echo "✅ Healthy"; exit 0; fi sleep 5 done echo "❌ Smoke test failed"; exit 1 deploy-production: needs: [build-and-push, deploy-staging] runs-on: ubuntu-latest environment: production # Requires manual approval in GitHub steps: - uses: actions/checkout@v4 - name: Update image tag run: | cd k8s/overlays/production kustomize edit set image \ ghcr.io/towerops/api=ghcr.io/towerops/api:${{ needs.build-and-push.outputs.image-tag }} - name: Apply to production run: kubectl apply -k k8s/overlays/production env: KUBECONFIG: ${{ secrets.PROD_KUBECONFIG }} - name: Wait for rollout run: kubectl rollout status deployment/towerops-api -n production --timeout=600s env: KUBECONFIG: ${{ secrets.PROD_KUBECONFIG }} ``` ### GitOps with Flux ```yaml # flux-system/gotk-components.yaml — installed via: # flux bootstrap github \ # --owner=towerops \ # --repository=fleet \ # --branch=main \ # --path=clusters/production # clusters/production/towerops-api.yaml --- apiVersion: source.toolkit.fluxcd.io/v1 kind: GitRepository metadata: name: towerops-api namespace: flux-system spec: interval: 1m url: https://github.com/towerops/towerops ref: branch: main secretRef: name: github-token --- apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: towerops-api namespace: flux-system spec: interval: 5m path: ./k8s/overlays/production prune: true # Delete resources removed from git sourceRef: kind: GitRepository name: towerops-api healthChecks: - apiVersion: apps/v1 kind: Deployment name: towerops-api namespace: production timeout: 10m ``` ### GitOps with ArgoCD ```yaml # argocd/applications/towerops-api.yaml apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: towerops-api namespace: argocd finalizers: - resources-finalizer.argocd.argoproj.io spec: project: default source: repoURL: https://github.com/towerops/towerops targetRevision: main path: k8s/overlays/production destination: server: https://kubernetes.default.svc namespace: production syncPolicy: automated: prune: true selfHeal: true # Revert manual kubectl changes syncOptions: - CreateNamespace=true retry: limit: 3 backoff: duration: 5s maxDuration: 3m factor: 2 ``` ### Secrets Management #### SOPS + Age (Simple, Git-friendly) ```bash # Install # nix-env -iA nixpkgs.sops nixpkgs.age # Generate age key age-keygen -o ~/.config/sops/age/keys.txt # .sops.yaml — tells SOPS which keys to use cat > .sops.yaml << 'EOF' creation_rules: - path_regex: secrets/.*\.yaml$ age: >- age1ql3z7hjy54pw3hyww5ayyfg7zqgvc7w3j2elw8zmrj2kg5sfn9aqmcac8p EOF # Create encrypted secret sops secrets/production.yaml # Opens $EDITOR — write plaintext YAML, SOPS encrypts on save # Example encrypted file (values encrypted, keys readable) cat secrets/production.yaml # database_url: ENC[AES256_GCM,data:abc123...,iv:...,tag:...,type:str] # secret_key_base: ENC[AES256_GCM,data:def456...,iv:...,tag:...,type:str] # Decrypt and use sops -d secrets/production.yaml # Use with kubectl sops -d secrets/production.yaml | kubectl create secret generic towerops-secrets \ --from-file=database-url=/dev/stdin \ --dry-run=client -o yaml | kubectl apply -f - ``` #### HashiCorp Vault ```bash # --- Vault Setup --- # Enable KV secrets engine vault secrets enable -path=secret kv-v2 # Store a secret vault kv put secret/towerops/production \ database_url="postgres://user:pass@db:5432/towerops" \ secret_key_base="$(openssl rand -hex 64)" # Read a secret vault kv get -field=database_url secret/towerops/production # --- Kubernetes Auth --- vault auth enable kubernetes vault write auth/kubernetes/config \ kubernetes_host="https://kubernetes.default.svc" # Policy: towerops-api can only read its own secrets vault policy write towerops-api - < interval '5 minutes';") if [ "$LONG_TX" -gt 0 ]; then echo "⚠️ $LONG_TX long-running transactions detected. Proceeding with caution." fi echo "=== Running migrations ===" # For Elixir/Ecto: mix ecto.migrate # Verify migration status echo "=== Post-migration status ===" mix ecto.migrations echo "✅ Migration complete" ``` **Golden rule for zero-downtime migrations:** Never remove a column or rename a column in a single deploy. Always use expand-contract: ``` Deploy 1: Add new column (nullable or with default) Deploy 2: Backfill data, update code to write to both old and new Deploy 3: Switch reads to new column Deploy 4: Stop writing to old column Deploy 5: Drop old column ``` --- ## 6. Networking & Security ### TLS / mTLS #### cert-manager with Let's Encrypt ```yaml # cert-manager ClusterIssuer apiVersion: cert-manager.io/v1 kind: ClusterIssuer metadata: name: letsencrypt-prod spec: acme: server: https://acme-v02.api.letsencrypt.org/directory email: ops@towerops.net privateKeySecretRef: name: letsencrypt-prod-key solvers: - http01: ingress: class: nginx # Or DNS challenge for wildcard certs: - dns01: cloudflare: email: ops@towerops.net apiTokenSecretRef: name: cloudflare-api-token key: api-token selector: dnsZones: - towerops.net --- # Certificate resource (or use ingress annotation instead) apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: towerops-wildcard namespace: production spec: secretName: towerops-wildcard-tls issuerRef: name: letsencrypt-prod kind: ClusterIssuer dnsNames: - "*.towerops.net" - towerops.net ``` #### mTLS Between Services (Istio / Linkerd) ```yaml # Istio PeerAuthentication — enforce mTLS for all services in namespace apiVersion: security.istio.io/v1 kind: PeerAuthentication metadata: name: default namespace: production spec: mtls: mode: STRICT # All traffic must be mTLS --- # Or with Linkerd (simpler): # Just annotate the namespace apiVersion: v1 kind: Namespace metadata: name: production annotations: linkerd.io/inject: enabled # Auto-injects sidecar proxy with mTLS ``` ### Zero-Trust Networking **Principle: Never trust the network. Always verify identity.** ```yaml # Kubernetes NetworkPolicy — default deny all apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: default-deny-all namespace: production spec: podSelector: {} policyTypes: - Ingress - Egress --- # Allow specific traffic apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: towerops-api-policy namespace: production spec: podSelector: matchLabels: app: towerops-api policyTypes: - Ingress - Egress ingress: # Allow from ingress controller only - from: - namespaceSelector: matchLabels: name: ingress-nginx ports: - port: 4000 protocol: TCP egress: # Allow to PostgreSQL - to: - podSelector: matchLabels: app: postgresql ports: - port: 5432 protocol: TCP # Allow to Redis - to: - podSelector: matchLabels: app: redis ports: - port: 6379 protocol: TCP # Allow DNS - to: - namespaceSelector: {} ports: - port: 53 protocol: UDP - port: 53 protocol: TCP ``` ### RBAC #### Kubernetes RBAC ```yaml # Least-privilege service account for CI/CD deployments apiVersion: v1 kind: ServiceAccount metadata: name: ci-deployer namespace: production --- apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: deployer namespace: production rules: - apiGroups: ["apps"] resources: ["deployments"] verbs: ["get", "list", "patch", "update"] - apiGroups: [""] resources: ["configmaps", "secrets"] verbs: ["get", "list", "create", "update", "patch"] - apiGroups: [""] resources: ["pods"] verbs: ["get", "list", "watch"] - apiGroups: [""] resources: ["pods/log"] verbs: ["get"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: name: ci-deployer-binding namespace: production subjects: - kind: ServiceAccount name: ci-deployer namespace: production roleRef: kind: Role name: deployer apiGroup: rbac.authorization.k8s.io ``` ### DNS Patterns ```bash # DNS record types and when to use them # A — Direct IP mapping (static servers) # CNAME — Alias to another domain (CDN, load balancer) # AAAA — IPv6 address # MX — Mail routing # TXT — Verification, SPF, DKIM, DMARC # SRV — Service discovery (rare outside K8s) # CAA — Certificate Authority Authorization (security) # Example zone file for towerops.net # Always set CAA to restrict who can issue certs towerops.net. CAA 0 issue "letsencrypt.org" towerops.net. CAA 0 issuewild "letsencrypt.org" # Primary domains towerops.net. A 203.0.113.10 www.towerops.net. CNAME towerops.net. api.towerops.net. CNAME lb.towerops.net. # Points to load balancer # Low TTL for things that might change during incidents api.towerops.net. 300 CNAME lb.towerops.net. # 5 min TTL # High TTL for stable records towerops.net. 86400 A 203.0.113.10 # 24h TTL # Email security towerops.net. TXT "v=spf1 include:_spf.google.com ~all" _dmarc.towerops.net. TXT "v=DMARC1; p=reject; rua=mailto:dmarc@towerops.net" ``` ### CDN Patterns ```bash # Cloudflare as CDN + WAF + DDoS protection # Key headers to set: # Cache static assets aggressively Cache-Control: public, max-age=31536000, immutable # For hashed assets (/app-abc123.js) # Don't cache API responses Cache-Control: private, no-store # Don't cache HTML (or use short TTL) Cache-Control: public, max-age=300, s-maxage=60 # Browser: 5min, CDN: 1min # Purge cache (Cloudflare API) curl -X POST "https://api.cloudflare.com/client/v4/zones/${ZONE_ID}/purge_cache" \ -H "Authorization: Bearer ${CF_API_TOKEN}" \ -H "Content-Type: application/json" \ --data '{"purge_everything": true}' # Purge specific URLs curl -X POST "https://api.cloudflare.com/client/v4/zones/${ZONE_ID}/purge_cache" \ -H "Authorization: Bearer ${CF_API_TOKEN}" \ -H "Content-Type: application/json" \ --data '{"files": ["https://towerops.net/styles.css"]}' ``` --- ## 7. Database Operations ### PostgreSQL Backup — WAL Archiving #### Continuous Archiving Setup ```bash # postgresql.conf — WAL archiving settings wal_level = replica # Required for archiving archive_mode = on archive_command = 'pgbackrest --stanza=main archive-push %p' # Or with plain file copy: # archive_command = 'test ! -f /backup/wal/%f && cp %p /backup/wal/%f' max_wal_senders = 3 # For streaming replication wal_keep_size = 1GB # Keep WAL segments for slow replicas ``` #### pgBackRest Configuration ```ini # /etc/pgbackrest/pgbackrest.conf [global] repo1-path=/backup/pgbackrest repo1-retention-full=4 # Keep 4 full backups repo1-retention-diff=14 # Keep 14 differential backups repo1-cipher-type=aes-256-cbc repo1-cipher-pass= compress-type=zst compress-level=6 process-max=4 # Parallel backup processes # S3 storage (alternative to local) # repo1-type=s3 # repo1-s3-bucket=towerops-backups # repo1-s3-region=us-east-1 # repo1-s3-endpoint=s3.amazonaws.com [main] pg1-path=/var/lib/postgresql/16/main pg1-port=5432 ``` ```bash # Initialize backup stanza pgbackrest --stanza=main stanza-create # Full backup (weekly cron) pgbackrest --stanza=main --type=full backup # Differential backup (daily cron) pgbackrest --stanza=main --type=diff backup # Verify backups pgbackrest --stanza=main check pgbackrest --stanza=main info # Point-in-time recovery (PITR) # Restore to specific timestamp pgbackrest --stanza=main --type=time \ --target="2026-03-13 15:30:00-05" \ --target-action=promote \ restore # Crontab 0 2 * * 0 pgbackrest --stanza=main --type=full backup # Weekly full at 2 AM Sunday 0 2 * * 1-6 pgbackrest --stanza=main --type=diff backup # Daily diff at 2 AM Mon-Sat ``` #### Backup Verification (Test Your Restores!) ```bash #!/bin/bash # scripts/verify-backup.sh — Run weekly, alert on failure set -euo pipefail RESTORE_DIR="/tmp/pg-restore-test" rm -rf "$RESTORE_DIR" echo "Restoring latest backup to temp directory..." pgbackrest --stanza=main --pg1-path="$RESTORE_DIR" restore echo "Starting temporary PostgreSQL instance..." pg_ctl -D "$RESTORE_DIR" -o "-p 5433" -l /tmp/pg-restore.log start echo "Running verification queries..." psql -p 5433 -d towerops -c "SELECT count(*) FROM users;" || { echo "❌ Backup verification FAILED" pg_ctl -D "$RESTORE_DIR" stop exit 1 } echo "✅ Backup verification passed" pg_ctl -D "$RESTORE_DIR" stop rm -rf "$RESTORE_DIR" ``` ### PgBouncer — Connection Pooling ```ini # pgbouncer.ini [databases] towerops = host=127.0.0.1 port=5432 dbname=towerops [pgbouncer] listen_addr = 0.0.0.0 listen_port = 6432 auth_type = scram-sha-256 auth_file = /etc/pgbouncer/userlist.txt # Pool settings pool_mode = transaction # Best for web apps (release conn after each tx) default_pool_size = 25 # Connections per user/database pair max_client_conn = 1000 # Max incoming connections min_pool_size = 5 # Keep warm connections reserve_pool_size = 5 # Emergency overflow reserve_pool_timeout = 3 # Wait before using reserve pool # Timeouts server_idle_timeout = 600 # Close idle server connections after 10 min client_idle_timeout = 0 # No client idle timeout query_timeout = 30 # Kill queries running > 30s client_login_timeout = 60 # Logging log_connections = 0 log_disconnections = 0 log_pooler_errors = 1 stats_period = 60 # TLS client_tls_sslmode = prefer client_tls_key_file = /etc/pgbouncer/server.key client_tls_cert_file = /etc/pgbouncer/server.crt ``` ```bash # Connect to PgBouncer admin console psql -p 6432 -U pgbouncer pgbouncer # Useful admin commands: SHOW POOLS; # Pool status per database/user SHOW CLIENTS; # Connected clients SHOW SERVERS; # Backend connections SHOW STATS; # Request stats SHOW CONFIG; # Current config # Graceful restart (finishes active transactions) RELOAD; # Reload config PAUSE; # Pause new queries, finish active RESUME; # Resume after pause ``` ### Zero-Downtime Migrations #### Safe Migration Patterns ```sql -- ❌ DANGEROUS: Locks entire table for write ALTER TABLE devices ADD COLUMN firmware_version TEXT NOT NULL DEFAULT 'unknown'; -- ✅ SAFE: Add nullable column (instant, no lock) ALTER TABLE devices ADD COLUMN firmware_version TEXT; -- ✅ SAFE: Add default in separate step (PG 11+ is instant for defaults) ALTER TABLE devices ALTER COLUMN firmware_version SET DEFAULT 'unknown'; -- ✅ SAFE: Backfill in batches, not all at once DO $$ DECLARE batch_size INT := 10000; rows_updated INT; BEGIN LOOP UPDATE devices SET firmware_version = 'unknown' WHERE id IN ( SELECT id FROM devices WHERE firmware_version IS NULL LIMIT batch_size FOR UPDATE SKIP LOCKED ); GET DIAGNOSTICS rows_updated = ROW_COUNT; EXIT WHEN rows_updated = 0; RAISE NOTICE 'Updated % rows', rows_updated; PERFORM pg_sleep(0.1); -- Brief pause to reduce load COMMIT; END LOOP; END $$; -- ✅ SAFE: Add NOT NULL constraint after backfill ALTER TABLE devices ALTER COLUMN firmware_version SET NOT NULL; -- ❌ DANGEROUS: CREATE INDEX locks writes CREATE INDEX idx_devices_firmware ON devices(firmware_version); -- ✅ SAFE: CONCURRENTLY doesn't block writes (takes longer) CREATE INDEX CONCURRENTLY idx_devices_firmware ON devices(firmware_version); -- ❌ DANGEROUS: Renaming column breaks running code ALTER TABLE devices RENAME COLUMN name TO display_name; -- ✅ SAFE: Use a view or add new column, migrate, drop old -- Step 1: Add new column ALTER TABLE devices ADD COLUMN display_name TEXT; -- Step 2: Backfill (app writes to both during transition) UPDATE devices SET display_name = name WHERE display_name IS NULL; -- Step 3: (Next deploy) Switch reads to display_name -- Step 4: (Next deploy) Stop writing to name -- Step 5: (Next deploy) Drop old column ALTER TABLE devices DROP COLUMN name; ``` #### Lock Timeout Safety Net ```sql -- Always set lock_timeout for DDL in production SET lock_timeout = '5s'; -- If we can't get the lock in 5s, fail rather than queue behind long transactions BEGIN; SET lock_timeout = '5s'; ALTER TABLE devices ADD COLUMN new_col TEXT; COMMIT; ``` ### PostgreSQL Monitoring Queries ```sql -- === Connection Health === -- Current connections by state SELECT state, count(*) FROM pg_stat_activity WHERE pid != pg_backend_pid() GROUP BY state ORDER BY count DESC; -- Long-running queries (potential problems) SELECT pid, now() - pg_stat_activity.query_start AS duration, query, state, wait_event_type, wait_event FROM pg_stat_activity WHERE (now() - pg_stat_activity.query_start) > interval '5 minutes' AND state != 'idle' ORDER BY duration DESC; -- Kill a stuck query SELECT pg_cancel_backend(pid); -- Graceful (cancels query) SELECT pg_terminate_backend(pid); -- Force (kills connection) -- === Performance === -- Table bloat (dead tuples needing VACUUM) SELECT schemaname, relname, n_live_tup, n_dead_tup, round(100.0 * n_dead_tup / NULLIF(n_live_tup + n_dead_tup, 0), 1) AS dead_pct, last_vacuum, last_autovacuum FROM pg_stat_user_tables WHERE n_dead_tup > 1000 ORDER BY n_dead_tup DESC LIMIT 20; -- Index usage (find unused indexes to drop) SELECT schemaname, tablename, indexname, idx_scan AS times_used, pg_size_pretty(pg_relation_size(indexrelid)) AS index_size FROM pg_stat_user_indexes JOIN pg_index USING (indexrelid) WHERE idx_scan = 0 AND NOT indisunique -- Don't drop unique constraints ORDER BY pg_relation_size(indexrelid) DESC; -- Missing indexes (sequential scans on large tables) SELECT relname, seq_scan, seq_tup_read, idx_scan, idx_tup_fetch, seq_tup_read / NULLIF(seq_scan, 0) AS avg_rows_per_seq_scan FROM pg_stat_user_tables WHERE seq_scan > 100 AND seq_tup_read / NULLIF(seq_scan, 0) > 10000 ORDER BY seq_tup_read DESC; -- Slowest queries (requires pg_stat_statements extension) SELECT query, calls, round(total_exec_time::numeric, 2) AS total_time_ms, round(mean_exec_time::numeric, 2) AS avg_time_ms, round((100 * total_exec_time / sum(total_exec_time) OVER ())::numeric, 2) AS pct_total FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 20; -- Cache hit ratio (should be > 99%) SELECT sum(heap_blks_read) AS heap_read, sum(heap_blks_hit) AS heap_hit, round(sum(heap_blks_hit) / NULLIF(sum(heap_blks_hit) + sum(heap_blks_read), 0)::numeric * 100, 2) AS hit_ratio FROM pg_statio_user_tables; -- === Replication === -- Replication status (on primary) SELECT client_addr, state, sent_lsn, write_lsn, flush_lsn, replay_lsn, (sent_lsn - replay_lsn) AS replication_lag FROM pg_stat_replication; -- Replication lag in seconds (on replica) SELECT CASE WHEN pg_last_wal_receive_lsn() = pg_last_wal_replay_lsn() THEN 0 ELSE EXTRACT(EPOCH FROM now() - pg_last_xact_replay_timestamp()) END AS replication_lag_seconds; -- === Locks === -- Blocked queries (waiting for locks) SELECT blocked_locks.pid AS blocked_pid, blocked_activity.usename AS blocked_user, blocking_locks.pid AS blocking_pid, blocking_activity.usename AS blocking_user, blocked_activity.query AS blocked_statement, blocking_activity.query AS blocking_statement FROM pg_catalog.pg_locks blocked_locks JOIN pg_catalog.pg_stat_activity blocked_activity ON blocked_activity.pid = blocked_locks.pid JOIN pg_catalog.pg_locks blocking_locks ON blocking_locks.locktype = blocked_locks.locktype AND blocking_locks.relation = blocked_locks.relation AND blocking_locks.pid != blocked_locks.pid JOIN pg_catalog.pg_stat_activity blocking_activity ON blocking_activity.pid = blocking_locks.pid WHERE NOT blocked_locks.granted; -- === Disk Usage === -- Largest tables SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) AS total_size, pg_size_pretty(pg_relation_size(relid)) AS table_size, pg_size_pretty(pg_total_relation_size(relid) - pg_relation_size(relid)) AS index_size FROM pg_catalog.pg_statio_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 20; ``` --- ## 8. On-Call & Alerting ### Alert Design Principles **Every alert must be:** 1. **Actionable** — Someone can do something about it right now 2. **Urgent** — It needs attention within the response SLA 3. **Real** — It indicates an actual problem (< 5% false positive rate) 4. **Unique** — Not duplicated by another alert **Bad alerts → alert fatigue → ignored alerts → outages.** #### Alert Anti-Patterns ```yaml # ❌ BAD: Too sensitive, will flap - alert: HighCPU expr: node_cpu_usage > 80 for: 1m # 1 minute is too short # ✅ GOOD: Sustained high CPU - alert: HighCPU expr: node_cpu_usage > 85 for: 15m # Must sustain for 15 min # ❌ BAD: Not actionable - alert: DiskUsageAbove50Percent expr: disk_usage > 50 # What am I supposed to do at 50%? # ✅ GOOD: Predictive and actionable - alert: DiskWillFillIn24h expr: predict_linear(disk_free[6h], 24*3600) < 0 # Clear action: expand disk or clean up annotations: runbook_url: "https://wiki/runbooks/disk-full" # ❌ BAD: Causes alert storm (one per pod) - alert: PodHighMemory expr: container_memory_usage > 100Mi # ✅ GOOD: Aggregate and threshold appropriately - alert: NamespaceHighMemory expr: | sum(container_memory_usage_bytes{namespace="production"}) / sum(kube_resourcequota{resource="limits.memory", namespace="production"}) > 0.85 ``` ### Escalation Policy ``` ┌─────────────────────────────────────────────────────────┐ │ Time │ Who Gets Paged │ Channel │ ├─────────────────────────────────────────────────────────┤ │ 0 min │ Primary on-call │ PagerDuty + SMS │ │ 10 min │ Secondary on-call │ PagerDuty + SMS │ │ 20 min │ Engineering manager │ Phone call │ │ 30 min │ VP Engineering │ Phone call │ │ 45 min │ CTO │ Phone call │ └─────────────────────────────────────────────────────────┘ SEV-1: Start at top, escalate every 10 min if not acknowledged SEV-2: Start at primary on-call, escalate after 30 min SEV-3: Slack notification only, addressed next business day SEV-4: Ticket, addressed during sprint planning ``` ### Runbook Template ````markdown # Runbook: [Alert Name] ## Alert - **Name:** HighErrorBudgetBurn_Fast - **Severity:** Critical - **Dashboard:** [Link to Grafana dashboard] - **SLO:** 99.9% availability, 30d rolling ## What This Means The API error rate is consuming our error budget at 14.4x the sustainable rate. At this burn rate, we'll exhaust the entire monthly budget within ~2 hours. ## Impact - Users may see 5xx errors on API calls - Dependent services (mobile app, dashboard) will be degraded ## Diagnosis Steps 1. **Check deployment history:** ```bash kubectl rollout history deployment/towerops-api -n production # Was there a recent deploy? If so, rollback is likely the fix. ``` 2. **Check error logs:** ```bash kubectl logs -l app=towerops-api -n production --tail=50 | grep -i error ``` 3. **Check dependencies:** ```bash # Database psql "$DATABASE_URL" -c "SELECT state, count(*) FROM pg_stat_activity GROUP BY state;" # Redis redis-cli -h redis ping # External APIs curl -s -o /dev/null -w '%{http_code}' https://api.stripe.com/v1/charges -H "Authorization: Bearer $STRIPE_KEY" ``` 4. **Check resource usage:** ```bash kubectl top pods -n production -l app=towerops-api ``` ## Mitigation ### If caused by a recent deployment: ```bash kubectl rollout undo deployment/towerops-api -n production kubectl rollout status deployment/towerops-api -n production ``` ### If caused by traffic spike: ```bash kubectl scale deployment/towerops-api -n production --replicas=10 ``` ### If caused by database issues: ```bash # Kill long-running queries psql "$DATABASE_URL" -c "SELECT pg_cancel_backend(pid) FROM pg_stat_activity WHERE state = 'active' AND now() - query_start > interval '1 minute';" ``` ### If caused by external dependency: - Enable circuit breaker / fallback mode - Update status page ## Resolution Verification ```bash # Verify error rate is back to normal curl -s "http://prometheus:9090/api/v1/query?query=sum(rate(http_requests_total{status_code=~\"5..\"}[5m]))/sum(rate(http_requests_total[5m]))" # Verify pods are healthy kubectl get pods -n production -l app=towerops-api ``` ## Escalation If not resolved within 30 minutes, escalate to the team lead. ## Related - [Architecture diagram](link) - [Previous incidents](link) - [Service dependencies](link) ```` ### Blameless Postmortem Template ````markdown # Postmortem: [Incident Title] **Date:** 2026-03-13 **Duration:** 14:30 – 15:15 CDT (45 minutes) **Severity:** SEV-2 **Authors:** [Names of postmortem authors] **Status:** Action items in progress ## Summary Brief, 2-3 sentence description. What happened, who was impacted, how was it resolved. > A misconfigured database migration caused connection exhaustion in the PostgreSQL > connection pool, resulting in 503 errors for 40% of API requests over 45 minutes. > Resolved by rolling back the migration and restarting PgBouncer. ## Impact - **Users affected:** ~2,400 (40% of active users) - **Revenue impact:** ~$3,200 in failed transactions - **Error budget consumed:** 8.3% of monthly budget - **Support tickets:** 47 ## Timeline (all times CDT) | Time | Event | |------|-------| | 14:28 | Deploy #847 rolls out (includes DB migration) | | 14:30 | `HighErrorBudgetBurn_Fast` alert fires | | 14:32 | On-call engineer (Alice) acknowledges alert | | 14:35 | Alice identifies high 503 rate, checks recent deploys | | 14:38 | Correlation with deploy #847 confirmed | | 14:40 | Alice initiates rollback of deploy #847 | | 14:42 | Rollback complete, but errors persist | | 14:45 | Bob (secondary) joins, identifies PgBouncer pool exhaustion | | 14:48 | PgBouncer restarted, connection pool drains | | 14:52 | Error rate drops to normal | | 15:00 | Monitoring confirms sustained recovery | | 15:15 | Incident declared resolved | ## Root Cause The migration added a `NOT NULL` constraint to a heavily-used column, which acquired an `ACCESS EXCLUSIVE` lock on the `devices` table. This blocked all queries touching the table, causing connection pool exhaustion as requests queued up waiting for the lock. The migration was tested in staging, but staging has ~1% of production data and the lock was released in < 1 second. In production with 2.3M rows, the lock validation took > 5 minutes. ## Contributing Factors 1. No lock timeout set for DDL operations in migration scripts 2. Staging data volume doesn't reflect production 3. Migration ran inside the application deploy, not as a separate step 4. PgBouncer didn't release connections after the migration was rolled back (stale pool state) ## What Went Well - Alert fired within 2 minutes of impact starting - On-call responded in under 5 minutes - Rollback procedure worked as documented - Team communicated clearly in #incidents channel ## What Went Wrong - Migration wasn't tested against production-like data volume - Rollback didn't fully resolve the issue (PgBouncer needed manual intervention) - Status page wasn't updated until 15 minutes into the incident ## Action Items | Action | Owner | Priority | Ticket | Due | |--------|-------|----------|--------|-----| | Add `SET lock_timeout = '5s'` to all migration scripts | Alice | P1 | OPS-234 | 2026-03-17 | | Create production-sized staging dataset (anonymized) | Bob | P2 | OPS-235 | 2026-03-27 | | Separate migration step from application deploy in CI/CD | Carol | P2 | OPS-236 | 2026-03-27 | | Add PgBouncer pool drain to rollback runbook | Alice | P3 | OPS-237 | 2026-03-20 | | Automate status page updates on SEV-1/2 alerts | Dave | P3 | OPS-238 | 2026-04-03 | ## Lessons Learned 1. **DDL operations on large tables need special handling.** Always use `SET lock_timeout`, always test against production-scale data, always run migrations separately from deploys. 2. **Connection poolers have state.** Rolling back the root cause doesn't always clear downstream effects. Include dependent service restarts in rollback procedures. 3. **Staging ≠ Production.** When the difference in data volume is 100x, some issues will only manifest in production. Consider periodic production data snapshots for staging. --- *This postmortem is blameless. It focuses on systemic improvements, not individual fault. The goal is to prevent recurrence, not assign blame.* ```` --- ## Quick Reference Cards ### The RED Method (Request-driven services) ``` R — Rate: Request throughput (req/s) E — Errors: Failed requests (count or rate) D — Duration: Request latency (histograms) ``` ### The USE Method (Infrastructure resources) ``` U — Utilization: % time resource is busy S — Saturation: Queue depth / backlog E — Errors: Error events on the resource ``` ### The Four Golden Signals (Google SRE) ``` 1. Latency — Time to serve a request 2. Traffic — Demand on the system (req/s) 3. Errors — Rate of failed requests 4. Saturation — How "full" the system is ``` ### Emergency Commands ```bash # === Kubernetes === kubectl rollout undo deployment/NAME -n NAMESPACE # Rollback deploy kubectl scale deployment/NAME --replicas=N -n NAMESPACE # Scale manually kubectl delete pod NAME -n NAMESPACE # Force restart pod kubectl cordon NODE # Stop scheduling kubectl drain NODE --ignore-daemonsets # Evacuate node # === PostgreSQL === # Kill all active queries on a table SELECT pg_cancel_backend(pid) FROM pg_stat_activity WHERE query LIKE '%tablename%' AND state = 'active'; # Kill all connections to a database SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE datname = 'dbname' AND pid <> pg_backend_pid(); # Emergency VACUUM (when autovacuum is too slow) VACUUM (VERBOSE, PARALLEL 4) tablename; # === System === # Find what's eating disk du -h --max-depth=2 / | sort -rh | head -20 # Find what's eating CPU ps aux --sort=-%cpu | head -10 # Find what's eating memory ps aux --sort=-%mem | head -10 # Network: who's connected ss -tunlp # Listening ports ss -tun state established # Active connections ``` --- *Last updated: 2026-03-13* *Maintained by: TowerOps SRE Team*