Per-pod IEM rate limiter was 700ms (~1.4 req/sec); four pods put the
fleet at ~5.7 req/sec, above the observed 429 threshold. Retry storm
(Req exponential backoff up to 32s) piled concurrent sleeps on top of
live workers and tripped Bandit's acceptor + Oban's leader heartbeat.
- Raise default interval to 1500ms (~2.7 req/sec fleet).
- Hot pod /live timeout 3s → 5s (acceptor stalls under retry sleeps).
- prop-backfill exec probe period 30s → 120s, timeout 15s → 30s; the
release CLI round-trips through distributed Erlang and 15s wasn't
enough under load, restart-looping every ~10 min.
Two independent wins the latest prod telemetry pointed at:
1. Iowa Environmental Mesonet rate limiter. IEM throttles per source
IP across all in-flight requests, so dropping the `:weather` Oban
queue to 1/pod wasn't enough — workers were still issuing back-to-
back requests inside each job and drawing a steady stream of HTTP
429s (1,396 retryable on the queue). A GenServer-based token bucket
serialises acquire() with a 700ms min gap per pod; three pods give
~4 req/sec cluster-wide, well under the observed 429 threshold.
Wrapped around every IemClient fetch.
2. Parallel HRRR surface + pressure grid fetch. The two products live
in separate wrfsfcf / wrfprsf GRIB files, so their fetch → wgrib2
→ decode pipelines are independent; running them sequentially was
adding ~30s per forecast hour for no reason. Task.async the pair,
merge at the end. Halves per-fh wall time and should unblock some
of the missing-hour cases we saw on the 14:00Z chain.