Reworks the post-fetch half of the propagation pipeline. Fetch and GRIB2
decode were already cheap — measured against a live HRRR cycle, all 39
pressure messages decode via `wgrib2 -lola` in 0.29 s and the 31 MB
byte-range fetch takes ~3 s — so nothing here touches the decoder. All the
cost was downstream.
Also fixes a broken NOTIFY that made every chain step run up to 5 times.
pg_notify
`NOTIFY propagation_ready, $1` is a Postgres syntax error: NOTIFY is a
utility statement whose payload must be a literal, so a bind raises
42601. It shared a transaction with the `status='done'` UPDATE, so every
successful step rolled back, stayed 'running', and was requeued by
reclaim_stale_running up to @max_reclaim_attempts times. Elixir's
NotifyListener never fired either, so ScoreCache warm and the
"propagation:updated" fan-out were dead.
FieldGrid
A decoded grid was HashMap<(i32,i32), HashMap<Arc<str>, f32>> — a dense
rectangular grid stored as ~95k nested hash maps, costing ~4.6M inserts
on decode, ~3.7M on merge and ~14M lookups across three derivation
passes. wgrib2 -lola already emits one dense row-major f32 block per
message, so keep it: dense per-message planes, names hashed once per
grid into plane ids, NaN as the missing sentinel. This is what forced
PROP_GRID_RS_PARALLELISM=1 under a 3Gi limit.
Fused pass
Three 95k-cell derivation passes plus 23 band-major scoring passes over
a staged Vec<(f64,f64,Conditions,BandInvariants)> (~19MB re-streamed 23
times) collapse into one pass: levels extracted once per cell, all 23
bands scored while the cell is hot, scores accumulated cell-major so
rayon chunks own disjoint slices. Scores land straight in the dense
score-file body — no ScorePoint scatter.
.pgrid
The profile artifact was an rmpv tree plus gzip -9, written 30x an hour,
and ProfilesFile.read_point/3 gunzipped and unpacked the entire 95k-cell
file to return one cell on every map click and Skew-T load. Replaced
with a dense cell-major f32 record array carrying a self-describing
field table. Elixir reads it via :file.pread; .mp.gz and .etf.gz remain
readable so files written before this drain out of the 48h window.
Measured on a full CONUS grid (95,073 cells x 48 planes x 23 bands):
derive + score + build artifacts 0.022 s
profile write 3.957 s -> 0.006 s (22.0 MB -> 22.4 MB on disk)
single-cell read whole-file decode -> 0.5 us
23 score files 0.003 s
Also
- hrrr_points: batched UNNEST upsert replacing one awaited INSERT per
point. Keeps ON CONFLICT DO UPDATE — the PSKR sampler's two-pass loop
depends on it.
- fetcher: real semaphore capping in-flight ranges at
MAX_PARALLEL_RANGES, which the comment claimed but the code did not do
(it spawned all 27 while the connection pool was sized for 8).
- metrics: per-stage histogram. Only chain-step and decode durations
were instrumented, which is why the write cost stayed invisible.
- profiles_file: parse_valid_time anchors on the known extension set, so
sibling-suffixed names like <iso>.hrdps.prop no longer parse as
<iso>.hrdps and vanish from prune and list operations.
- PROP_GRID_RS_PARALLELISM 1 -> 3. Memory limit held at 3Gi until RSS is
observed at the new parallelism.
- cargo fmt over the crate; worker.rs, hrdps_fetcher.rs and nexrad.rs
were already unformatted at HEAD and the pre-commit hook gates on it.
HRDPS still runs at 0.5 degrees. wgrib2 -lola scales linearly in output
points on rotated lat/lon (12.5 s wall, 202 s CPU for one message at
0.125 degrees) because it has no inverse projection for those grids; a raw
native dump is 0.32 s. The fix is decode-once plus a closed-form
rotated-pole index, left for a follow-up.
163 lines
6.9 KiB
YAML
163 lines
6.9 KiB
YAML
apiVersion: apps/v1
|
||
kind: Deployment
|
||
metadata:
|
||
name: prop-grid-rs
|
||
namespace: prop
|
||
# Phase 2 cutover: Rust writes directly to /data/scores. Elixir only
|
||
# runs the f00 analysis-hour step (which carries enrichment Rust
|
||
# doesn't own yet — native-level duct merge, NEXRAD composite,
|
||
# commercial-link degradation, ProfilesFile write). f01..f18 land
|
||
# here from the Rust worker.
|
||
spec:
|
||
# Replicas claim from grid_tasks via FOR UPDATE SKIP LOCKED with no
|
||
# coordination; anti-affinity spreads them across distinct hosts so
|
||
# any single node going down leaves the hourly chain with other
|
||
# workers still draining.
|
||
# Replica count is HPA-managed (hpa.yaml, min 1 / max 4).
|
||
minReadySeconds: 5
|
||
# Zero-downtime rollout for the propagation chain. With HPA min=1
|
||
# and the previous maxSurge:0/maxUnavailable:1 config, every deploy
|
||
# tore the single replica down before bringing the new one up,
|
||
# leaving a multi-minute window where the :05 hourly cron had no
|
||
# worker to claim tasks. Surge-first means the new pod is ready
|
||
# before the old one drains; the two briefly overlap and race for
|
||
# `grid_tasks` rows via FOR UPDATE SKIP LOCKED, which is safe.
|
||
strategy:
|
||
type: RollingUpdate
|
||
rollingUpdate:
|
||
maxSurge: 1
|
||
maxUnavailable: 0
|
||
selector:
|
||
matchLabels:
|
||
app: prop-grid-rs
|
||
template:
|
||
metadata:
|
||
labels:
|
||
app: prop-grid-rs
|
||
tier: grid-rs
|
||
annotations:
|
||
# Scraped by the Prometheus at 10.0.15.31 via the
|
||
# `kubernetes-pods` job in prometheus.yml.j2 (apiserver proxy).
|
||
# Metrics endpoint is on container port 9100 (see METRICS_ADDR
|
||
# env below). `prometheus.io/job` makes the dashboard's
|
||
# `up{job="prop-grid-rs"}` selector work without per-pod
|
||
# relabel.
|
||
prometheus.io/scrape: "true"
|
||
prometheus.io/port: "9100"
|
||
prometheus.io/path: "/metrics"
|
||
prometheus.io/job: "prop-grid-rs"
|
||
spec:
|
||
# No node pinning — scheduler picks any host. Anti-affinity was
|
||
# removed so HPA can stack multiple replicas on the same node if
|
||
# that's where the scheduler wants them; grid-rs is CPU-bound but
|
||
# not memory-bound so colocating a couple of replicas is fine.
|
||
# Wide window for an in-flight chain step (typically 1–3 min) to
|
||
# finish before SIGKILL. The worker handles SIGTERM by stopping
|
||
# task pickup; releasing the current claim is the cron's job
|
||
# (visibility timeout reclaims abandoned `grid_tasks` rows).
|
||
terminationGracePeriodSeconds: 300
|
||
imagePullSecrets:
|
||
- name: forgejo-registry
|
||
securityContext:
|
||
runAsUser: 65534
|
||
runAsNonRoot: true
|
||
fsGroup: 65534
|
||
seccompProfile:
|
||
type: RuntimeDefault
|
||
containers:
|
||
- name: prop-grid-rs
|
||
# Image built by a new CI pipeline from rust/prop_grid_rs/.
|
||
# Multi-arch (amd64 + arm64) — reuses the repo's existing
|
||
# buildx wiring.
|
||
image: git.mcintire.me/graham/prop-grid-rs:main-1781279001-e879f29
|
||
imagePullPolicy: IfNotPresent
|
||
env:
|
||
# Elixir secrets bundle carries DATABASE_URL already. Rust
|
||
# reads it directly (sqlx connects with the same URL).
|
||
- name: HRRR_FALLBACK_BASE_URL
|
||
value: "https://noaa-hrrr-bdp-pds.s3.amazonaws.com"
|
||
- name: PROP_SCORES_DIR
|
||
value: "/data/scores"
|
||
# Shared idx-file cache on NFS. All grid-rs replicas read/write
|
||
# through the same directory; idx bodies are immutable per HRRR
|
||
# run, so cross-replica hits eliminate redundant NOAA S3 calls.
|
||
- name: HRRR_IDX_CACHE_DIR
|
||
value: "/data/hrrr_idx"
|
||
- name: RUST_LOG
|
||
value: "info"
|
||
# Was "1": a decoded grid used to be ~95 k nested HashMaps
|
||
# (~200-400 MB per in-flight task), so two concurrent tasks
|
||
# plus the wgrib2 working set crossed the 3 Gi limit. The
|
||
# dense `FieldGrid` representation stores the same data as
|
||
# ~48 f32 planes (~18 MB), and the profile artifact is now a
|
||
# flat f32 write instead of an rmpv tree + gzip -9, so the
|
||
# per-task footprint is roughly an order of magnitude lower.
|
||
#
|
||
# Raising to 3 first and holding the memory limit; drop the
|
||
# limit only after observing steady-state RSS at this
|
||
# parallelism (one knob per deploy).
|
||
- name: PROP_GRID_RS_PARALLELISM
|
||
value: "3"
|
||
# Pool size = parallelism + 2 (NOTIFY + retry slack). Kept in
|
||
# sync with PROP_GRID_RS_PARALLELISM when either changes.
|
||
- name: PROP_GRID_RS_PG_CONNS
|
||
value: "5"
|
||
- name: METRICS_ADDR
|
||
value: "0.0.0.0:9100"
|
||
envFrom:
|
||
- secretRef:
|
||
name: prop-secrets
|
||
ports:
|
||
- name: metrics
|
||
containerPort: 9100
|
||
protocol: TCP
|
||
# Liveness is "the process is still running"; let the kubelet
|
||
# restart us if we crash. Claiming work from grid_tasks is the
|
||
# implicit readiness signal — an idle pod is still healthy
|
||
# (the queue is simply empty). A minimal readiness check on
|
||
# /health confirms the metrics listener is up, which doubles
|
||
# as evidence the tokio runtime is alive.
|
||
readinessProbe:
|
||
httpGet:
|
||
path: /health
|
||
port: 9100
|
||
# Generous initialDelaySeconds so a slow DB connect during
|
||
# startup (Postgres acquire_timeout is 10s) can't race the
|
||
# first probe. failureThreshold=6 keeps the pod from
|
||
# flapping if a single scrape times out during a long
|
||
# chain step.
|
||
initialDelaySeconds: 15
|
||
periodSeconds: 10
|
||
timeoutSeconds: 3
|
||
failureThreshold: 6
|
||
resources:
|
||
requests:
|
||
cpu: 100m
|
||
memory: 768Mi
|
||
limits:
|
||
cpu: "4"
|
||
# Analysis tasks (f00) run the native-level HRRR decode
|
||
# (~530 MB GRIB2 + wgrib2 working set), which is now the
|
||
# dominant term — the per-task grid footprint dropped from
|
||
# ~200-400 MB of nested HashMaps to ~18 MB of dense f32
|
||
# planes. Holding 3 Gi for one deploy while parallelism
|
||
# goes 1 -> 3; once `container_memory_working_set_bytes`
|
||
# is confirmed steady, this can come down toward 1.5 Gi.
|
||
memory: 3Gi
|
||
securityContext:
|
||
allowPrivilegeEscalation: false
|
||
runAsNonRoot: true
|
||
runAsUser: 65534
|
||
capabilities:
|
||
drop:
|
||
- ALL
|
||
seccompProfile:
|
||
type: RuntimeDefault
|
||
volumeMounts:
|
||
- name: data
|
||
mountPath: /data
|
||
volumes:
|
||
- name: data
|
||
nfs:
|
||
server: 10.0.19.103
|
||
path: /data
|