prop/lib/microwaveprop/propagation/grid_task_enqueuer.ex
Graham McIntire b80878056d
feat(grid-rs): Rust worker for HRRR f01..f18 propagation chain
Extracts the memory-hostile HRRR fetch → decode → score pipeline for
forecast hours 1–18 into a separate Rust service (`prop-grid-rs`).
Elixir retains f00 with its native-duct + NEXRAD + commercial-link
enrichment; Rust handles the other 18 steps per hourly chain.

Hand-off via a new `grid_tasks` table (`FOR UPDATE SKIP LOCKED` claim
from Rust, Postgrex `NOTIFY propagation_ready` back to Elixir). Rust
writes the same on-disk score-grid format Elixir already uses, to the
same `/data/scores` NFS tree. Phase 1 ships in shadow mode with
PROP_SCORES_DIR=/data/scores_shadow.

Rust crate layout at `rust/prop_grid_rs/` (1:1 module parity with the
Elixir source it ports):
  - grid, region, band_config, scores_file, sounding_params
  - scorer: all 10 factors + composite. Matches the Elixir scorer
    byte-for-byte across 115 golden-fixture samples (5 scenarios
    × 23 bands).
  - decoder: wgrib2 subprocess + Fortran-record lola-binary parser
  - fetcher: HRRR URL derivation + idx cache (1h TTL) + 8-way
    parallel byte-range downloads with 429/5xx retry
  - pipeline: end-to-end chain step, f00 rejected at the boundary
  - db: sqlx grid_tasks claim/complete + propagation_ready NOTIFY
  - bin/worker: tokio main loop, JSON logs, SIGTERM-safe

91 Rust tests + clippy -D warnings clean. 2,158 Elixir tests green.

Elixir additions:
  - `GridTaskEnqueuer` seeds fh=1..18 rows from
    `PropagationGridWorker.seed_chain/0`
  - `NotifyListener` LISTEN → warm `ScoreCache` → PubSub fan-out
  - `ShadowComparator` diffs prod vs shadow `.ntms` bodies
  - `mix rust.golden` task writes the Rust-side golden fixture

k8s: new `deployment-grid-rs.yaml` pinned to talos5 (32 GB NUC) via
`prop-grid-rs=primary` nodeSelector + `workload=grid-rs:NoSchedule`
toleration, 512 Mi limit, sharing the existing NFS `/data` mount.

The plan document is at `plans/vivid-hatching-quail.md` (local to my
workstation); phases 2 (cutover) and 3 (talos5 concurrency tuning)
follow after 72h of shadow-mode parity.
2026-04-19 15:42:49 -05:00

60 lines
1.7 KiB
Elixir

defmodule Microwaveprop.Propagation.GridTaskEnqueuer do
@moduledoc """
Seeds `grid_tasks` rows for the Rust `prop-grid-rs` worker.
Called from `PropagationGridWorker.seed_chain/0` alongside the existing
Elixir f00..f18 Oban fan-out. Rust only claims rows with
`forecast_hour > 0`; Elixir still owns the f00 analysis-hour chain
because of native-duct + NEXRAD + commercial-link enrichment.
Inserts are idempotent via the `(run_time, forecast_hour)` unique
index — re-seeding the same cycle is a no-op.
"""
alias Microwaveprop.Repo
require Logger
@max_forecast_hour 18
@spec seed(DateTime.t()) :: {:ok, non_neg_integer()} | {:error, term()}
def seed(%DateTime{} = run_time) do
run_time = DateTime.truncate(run_time, :second)
now = DateTime.truncate(DateTime.utc_now(), :microsecond)
rows =
for fh <- 1..@max_forecast_hour do
%{
id: Ecto.UUID.bingenerate(),
run_time: run_time,
forecast_hour: fh,
valid_time: DateTime.add(run_time, fh * 3600, :second),
status: "queued",
attempt: 0,
claimed_at: nil,
completed_at: nil,
error: nil,
inserted_at: now,
updated_at: now
}
end
{count, _} =
Repo.insert_all("grid_tasks", rows,
on_conflict: :nothing,
conflict_target: [:run_time, :forecast_hour]
)
if count > 0 do
Logger.info("GridTaskEnqueuer: seeded #{count} grid_tasks for run_time=#{run_time}")
else
Logger.info("GridTaskEnqueuer: run_time=#{run_time} already seeded")
end
{:ok, count}
rescue
e ->
Logger.error("GridTaskEnqueuer: seed failed: #{inspect(e)}")
{:error, e}
end
end