Two concurrent forecast tasks plus the NFS score-file write cache (23 files x ~2 MB) plus the analysis step's wgrib2 peak repeatedly crossed the 3 Gi cgroup limit. Single-lane parallelism per pod keeps steady RSS under 2 Gi; with 2 replicas that still gives 2 concurrent tasks cluster-wide. PROP_GRID_RS_PG_CONNS follows the parallelism+2 formula, dropping from 4 to 3. |
||
|---|---|---|
| .. | ||
| deployment-backfill.yaml | ||
| deployment-grid-rs.yaml | ||
| deployment-hrrr-point-rs.yaml | ||
| deployment.yaml | ||
| flux.yaml | ||
| kustomization.yaml | ||
| metrics-service.yaml | ||
| namespace.yaml | ||
| rbac.yaml | ||
| service.yaml | ||