fix(grid-rs): bump memory limit to 2 Gi

1 Gi OOM'd during a real chain step. Each forecast hour loads both a
surface and pressure-level HRRR GRIB2 (~40 MB combined), decodes via
wgrib2 subprocess (which allocates its own buffers), scores 92k grid
points × 23 bands (~2.1M intermediates), and writes 23 score files.
2 Gi absorbs the spike; still ~3× smaller than the Elixir-per-pod
budget we're replacing.
This commit is contained in:
Graham McIntire 2026-04-19 16:37:05 -05:00
parent 25c3d4f688
commit ca33b8331b
No known key found for this signature in database
GPG key ID: F4ABF488E6029E59

View file

@ -71,14 +71,14 @@ spec:
resources:
requests:
cpu: 100m
memory: 256Mi
memory: 512Mi
limits:
cpu: "4"
# 512 Mi OOM-killed the real binary at startup. 1 Gi gives
# room for sqlx pool + reqwest pool + tokio runtime +
# tracing buffers, with headroom vs the steady-state
# target (~500 Mi during a chain step).
memory: 1Gi
# OOM at 1 Gi during active chain work (92k points × 23
# bands ≈ 2.1M scores, plus decoded GRIB2 buffers, plus
# wgrib2 subprocess memory). 2 Gi leaves headroom; we're
# still an order of magnitude below the 6 Gi Elixir budget.
memory: 2Gi
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true