prop/k8s/deployment-grid-rs.yaml
Graham McIntire 63f25a9612
Some checks failed
Build prop-grid-rs / Test, build, push (push) Successful in 5m54s
Build and Push / Build and Push Docker Image (push) Failing after 4m44s
perf(grid-rs): dense grid, fused scoring pass, and columnar .pgrid profiles
Reworks the post-fetch half of the propagation pipeline. Fetch and GRIB2
decode were already cheap — measured against a live HRRR cycle, all 39
pressure messages decode via `wgrib2 -lola` in 0.29 s and the 31 MB
byte-range fetch takes ~3 s — so nothing here touches the decoder. All the
cost was downstream.

Also fixes a broken NOTIFY that made every chain step run up to 5 times.

pg_notify
  `NOTIFY propagation_ready, $1` is a Postgres syntax error: NOTIFY is a
  utility statement whose payload must be a literal, so a bind raises
  42601. It shared a transaction with the `status='done'` UPDATE, so every
  successful step rolled back, stayed 'running', and was requeued by
  reclaim_stale_running up to @max_reclaim_attempts times. Elixir's
  NotifyListener never fired either, so ScoreCache warm and the
  "propagation:updated" fan-out were dead.

FieldGrid
  A decoded grid was HashMap<(i32,i32), HashMap<Arc<str>, f32>> — a dense
  rectangular grid stored as ~95k nested hash maps, costing ~4.6M inserts
  on decode, ~3.7M on merge and ~14M lookups across three derivation
  passes. wgrib2 -lola already emits one dense row-major f32 block per
  message, so keep it: dense per-message planes, names hashed once per
  grid into plane ids, NaN as the missing sentinel. This is what forced
  PROP_GRID_RS_PARALLELISM=1 under a 3Gi limit.

Fused pass
  Three 95k-cell derivation passes plus 23 band-major scoring passes over
  a staged Vec<(f64,f64,Conditions,BandInvariants)> (~19MB re-streamed 23
  times) collapse into one pass: levels extracted once per cell, all 23
  bands scored while the cell is hot, scores accumulated cell-major so
  rayon chunks own disjoint slices. Scores land straight in the dense
  score-file body — no ScorePoint scatter.

.pgrid
  The profile artifact was an rmpv tree plus gzip -9, written 30x an hour,
  and ProfilesFile.read_point/3 gunzipped and unpacked the entire 95k-cell
  file to return one cell on every map click and Skew-T load. Replaced
  with a dense cell-major f32 record array carrying a self-describing
  field table. Elixir reads it via :file.pread; .mp.gz and .etf.gz remain
  readable so files written before this drain out of the 48h window.

  Measured on a full CONUS grid (95,073 cells x 48 planes x 23 bands):
    derive + score + build artifacts   0.022 s
    profile write   3.957 s -> 0.006 s (22.0 MB -> 22.4 MB on disk)
    single-cell read   whole-file decode -> 0.5 us
    23 score files     0.003 s

Also
  - hrrr_points: batched UNNEST upsert replacing one awaited INSERT per
    point. Keeps ON CONFLICT DO UPDATE — the PSKR sampler's two-pass loop
    depends on it.
  - fetcher: real semaphore capping in-flight ranges at
    MAX_PARALLEL_RANGES, which the comment claimed but the code did not do
    (it spawned all 27 while the connection pool was sized for 8).
  - metrics: per-stage histogram. Only chain-step and decode durations
    were instrumented, which is why the write cost stayed invisible.
  - profiles_file: parse_valid_time anchors on the known extension set, so
    sibling-suffixed names like <iso>.hrdps.prop no longer parse as
    <iso>.hrdps and vanish from prune and list operations.
  - PROP_GRID_RS_PARALLELISM 1 -> 3. Memory limit held at 3Gi until RSS is
    observed at the new parallelism.
  - cargo fmt over the crate; worker.rs, hrdps_fetcher.rs and nexrad.rs
    were already unformatted at HEAD and the pre-commit hook gates on it.

HRDPS still runs at 0.5 degrees. wgrib2 -lola scales linearly in output
points on rotated lat/lon (12.5 s wall, 202 s CPU for one message at
0.125 degrees) because it has no inverse projection for those grids; a raw
native dump is 0.32 s. The fix is decode-once plus a closed-form
rotated-pole index, left for a follow-up.
2026-08-01 08:23:36 -05:00

163 lines
6.9 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

apiVersion: apps/v1
kind: Deployment
metadata:
name: prop-grid-rs
namespace: prop
# Phase 2 cutover: Rust writes directly to /data/scores. Elixir only
# runs the f00 analysis-hour step (which carries enrichment Rust
# doesn't own yet — native-level duct merge, NEXRAD composite,
# commercial-link degradation, ProfilesFile write). f01..f18 land
# here from the Rust worker.
spec:
# Replicas claim from grid_tasks via FOR UPDATE SKIP LOCKED with no
# coordination; anti-affinity spreads them across distinct hosts so
# any single node going down leaves the hourly chain with other
# workers still draining.
# Replica count is HPA-managed (hpa.yaml, min 1 / max 4).
minReadySeconds: 5
# Zero-downtime rollout for the propagation chain. With HPA min=1
# and the previous maxSurge:0/maxUnavailable:1 config, every deploy
# tore the single replica down before bringing the new one up,
# leaving a multi-minute window where the :05 hourly cron had no
# worker to claim tasks. Surge-first means the new pod is ready
# before the old one drains; the two briefly overlap and race for
# `grid_tasks` rows via FOR UPDATE SKIP LOCKED, which is safe.
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app: prop-grid-rs
template:
metadata:
labels:
app: prop-grid-rs
tier: grid-rs
annotations:
# Scraped by the Prometheus at 10.0.15.31 via the
# `kubernetes-pods` job in prometheus.yml.j2 (apiserver proxy).
# Metrics endpoint is on container port 9100 (see METRICS_ADDR
# env below). `prometheus.io/job` makes the dashboard's
# `up{job="prop-grid-rs"}` selector work without per-pod
# relabel.
prometheus.io/scrape: "true"
prometheus.io/port: "9100"
prometheus.io/path: "/metrics"
prometheus.io/job: "prop-grid-rs"
spec:
# No node pinning — scheduler picks any host. Anti-affinity was
# removed so HPA can stack multiple replicas on the same node if
# that's where the scheduler wants them; grid-rs is CPU-bound but
# not memory-bound so colocating a couple of replicas is fine.
# Wide window for an in-flight chain step (typically 13 min) to
# finish before SIGKILL. The worker handles SIGTERM by stopping
# task pickup; releasing the current claim is the cron's job
# (visibility timeout reclaims abandoned `grid_tasks` rows).
terminationGracePeriodSeconds: 300
imagePullSecrets:
- name: forgejo-registry
securityContext:
runAsUser: 65534
runAsNonRoot: true
fsGroup: 65534
seccompProfile:
type: RuntimeDefault
containers:
- name: prop-grid-rs
# Image built by a new CI pipeline from rust/prop_grid_rs/.
# Multi-arch (amd64 + arm64) — reuses the repo's existing
# buildx wiring.
image: git.mcintire.me/graham/prop-grid-rs:main-1781279001-e879f29
imagePullPolicy: IfNotPresent
env:
# Elixir secrets bundle carries DATABASE_URL already. Rust
# reads it directly (sqlx connects with the same URL).
- name: HRRR_FALLBACK_BASE_URL
value: "https://noaa-hrrr-bdp-pds.s3.amazonaws.com"
- name: PROP_SCORES_DIR
value: "/data/scores"
# Shared idx-file cache on NFS. All grid-rs replicas read/write
# through the same directory; idx bodies are immutable per HRRR
# run, so cross-replica hits eliminate redundant NOAA S3 calls.
- name: HRRR_IDX_CACHE_DIR
value: "/data/hrrr_idx"
- name: RUST_LOG
value: "info"
# Was "1": a decoded grid used to be ~95 k nested HashMaps
# (~200-400 MB per in-flight task), so two concurrent tasks
# plus the wgrib2 working set crossed the 3 Gi limit. The
# dense `FieldGrid` representation stores the same data as
# ~48 f32 planes (~18 MB), and the profile artifact is now a
# flat f32 write instead of an rmpv tree + gzip -9, so the
# per-task footprint is roughly an order of magnitude lower.
#
# Raising to 3 first and holding the memory limit; drop the
# limit only after observing steady-state RSS at this
# parallelism (one knob per deploy).
- name: PROP_GRID_RS_PARALLELISM
value: "3"
# Pool size = parallelism + 2 (NOTIFY + retry slack). Kept in
# sync with PROP_GRID_RS_PARALLELISM when either changes.
- name: PROP_GRID_RS_PG_CONNS
value: "5"
- name: METRICS_ADDR
value: "0.0.0.0:9100"
envFrom:
- secretRef:
name: prop-secrets
ports:
- name: metrics
containerPort: 9100
protocol: TCP
# Liveness is "the process is still running"; let the kubelet
# restart us if we crash. Claiming work from grid_tasks is the
# implicit readiness signal — an idle pod is still healthy
# (the queue is simply empty). A minimal readiness check on
# /health confirms the metrics listener is up, which doubles
# as evidence the tokio runtime is alive.
readinessProbe:
httpGet:
path: /health
port: 9100
# Generous initialDelaySeconds so a slow DB connect during
# startup (Postgres acquire_timeout is 10s) can't race the
# first probe. failureThreshold=6 keeps the pod from
# flapping if a single scrape times out during a long
# chain step.
initialDelaySeconds: 15
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 6
resources:
requests:
cpu: 100m
memory: 768Mi
limits:
cpu: "4"
# Analysis tasks (f00) run the native-level HRRR decode
# (~530 MB GRIB2 + wgrib2 working set), which is now the
# dominant term — the per-task grid footprint dropped from
# ~200-400 MB of nested HashMaps to ~18 MB of dense f32
# planes. Holding 3 Gi for one deploy while parallelism
# goes 1 -> 3; once `container_memory_working_set_bytes`
# is confirmed steady, this can come down toward 1.5 Gi.
memory: 3Gi
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
runAsUser: 65534
capabilities:
drop:
- ALL
seccompProfile:
type: RuntimeDefault
volumeMounts:
- name: data
mountPath: /data
volumes:
- name: data
nfs:
server: 10.0.19.103
path: /data