Three fixes after hrrr-point-rs restarts on the first live drain:
1. Type mismatch (hard error): the `profile` column on `hrrr_profiles`
is `jsonb[]` (one element per pressure level), but the Rust worker
was binding a single jsonb array-of-objects cast as `::jsonb`. Every
insert failed with `column "profile" is of type jsonb[] but
expression is of type jsonb`. Switched to
`Vec<sqlx::types::Json<Value>>`, dropped the explicit `::jsonb`
cast, and enabled sqlx's `json` feature so the encoder maps the
array into `_jsonb` correctly.
2. Silent zero-inserts: four consecutive 2019-09-22 batches completed
with `profiles_inserted: 0` and no diagnostic. Most likely the
upstream archive doesn't keep cycles that old, but without a log
it looks identical to a snap-mismatch bug. Added a WARN when both
surface and pressure grids come back empty so fetch-miss vs.
projection-miss is distinguishable.
3. OOMKilled: the pod took four OOM restarts in ten minutes at 1 Gi.
A single CONUS decode holds the ~40 MB blob, the ~200 MB wgrib2
working set, AND the full 92k-cell merged map until all requested
points drain. 1 Gi has no headroom. Bumped to 2 Gi, matching the
rest of the per-container budgets in this namespace.
Phase 3 Stream C Rust side. Completes the HrrrFetchWorker port.
Pipeline:
- db::claim_next_hrrr_task — FOR UPDATE SKIP LOCKED on hrrr_fetch_tasks,
newest valid_time first. Accepts the points JSONB directly.
- hrrr_points::process_batch — fetch surface + pressure GRIB2 once
per task (tokio::try_join), decode via the existing wgrib2 plumbing,
then for each requested point pull the cell and UPSERT INTO
hrrr_profiles (conflict on lat/lon/valid_time).
- db::complete_hrrr_task / fail_hrrr_task — status transitions; Elixir
backfill re-enqueues failed rows on next /30-min scan.
Shipping pieces:
- new bin src/bin/hrrr_point_worker.rs
- new module src/hrrr_points.rs (process_batch, upsert_profile)
- new Cargo [[bin]] entry; Dockerfile builds both binaries in one stage
and ships them in the runtime image so a single CI pipeline covers
the whole cluster
- k8s/deployment-hrrr-point-rs.yaml (1 replica, 1 Gi limit, anti-affinity
against prop-grid-rs so chain + point work don't fight for wgrib2
slots). Uses the same image; command: override picks the right binary.
- kustomization.yaml: include the new deployment so flux applies it
- deployment-grid-rs.yaml: bump readiness initialDelaySeconds 3→15 +
failureThreshold 3→6 so a slow DB connect during startup can't race
the first probe
119 Rust tests green.