Periodic 60s sweep re-warms ScoreCache from /data/scores whenever
{band, valid_time} pairs are on disk but absent in local ETS.
Covers the three silent-fail modes of the NOTIFY/LISTEN path:
dropped notifies on reconnect, missing listener, and stale NFS
reads. Each pod reconciles its own cache — no PubSub fan-out,
no coordination required. Rescue on File.Error handles the
rare window where an atomic-rename write races the reader.
docs/runbook_propagation_pipeline.md names every failure mode
in the Rust↔Elixir seam and what recovers it.
120 lines
4.8 KiB
Markdown
120 lines
4.8 KiB
Markdown
# Propagation pipeline runbook
|
||
|
||
The hourly propagation grid flows across two runtimes and a shared
|
||
NFS volume. This doc names every failure mode at that seam and says
|
||
how the system recovers.
|
||
|
||
```
|
||
prop-grid-rs (talos5) Postgres (10.0.15.30) prop pods (node1/2/3)
|
||
├─ fetch GRIB2 (NOMADS) ├─ grid_tasks (SKIP LOCKED) ├─ ScoreCache (ETS)
|
||
├─ decode ├─ LISTEN/NOTIFY channel ├─ NotifyListener
|
||
├─ score 22 bands × 92k │ "propagation_ready" ├─ ScoreCacheReconciler (60s)
|
||
├─ write /data/scores/.../.ntms │ └─ PubSub → LiveView
|
||
└─ NOTIFY propagation_ready ────┘
|
||
```
|
||
|
||
## Primary refresh path (happy case)
|
||
|
||
1. Rust completes `{band, valid_time}` compute and writes the `.ntms`
|
||
file via atomic rename.
|
||
2. Rust emits `NOTIFY propagation_ready '<iso valid_time>'`.
|
||
3. `NotifyListener` on each Elixir pod receives it and calls
|
||
`Propagation.warm_cache_and_broadcast/2` which reads the file once
|
||
and PubSub-broadcasts the payload to every peer.
|
||
4. Each pod's `ScoreCache` receives the message and writes into its
|
||
local ETS table. `/map` clients are notified via the
|
||
`"propagation:updated"` topic.
|
||
|
||
Latency budget end-to-end: < 2s.
|
||
|
||
## Failure modes
|
||
|
||
### FM1 — NOTIFY dropped on reconnect
|
||
|
||
**Symptom.** `{band, valid_time}` files land on disk but ScoreCache
|
||
stays cold until the first lazy miss.
|
||
|
||
**Why.** Postgres queues LISTEN notifications per-connection and
|
||
discards any that are unread on disconnect. A DB restart, a network
|
||
blip, or the Postgrex.Notifications connection hitting its idle
|
||
timeout all drop whatever queued notifies were in flight.
|
||
|
||
**Recovery.** `ScoreCacheReconciler` sweeps `/data/scores` every 60s
|
||
per-pod and warms missing `{band, valid_time}` pairs from disk into
|
||
local ETS. Each pod reconciles independently — no cluster messaging
|
||
required.
|
||
|
||
### FM2 — NotifyListener not running locally
|
||
|
||
**Symptom.** One pod is cold while peers are warm. `/map` load time
|
||
on that pod is ~150ms slower than average until the reconciler runs.
|
||
|
||
**Why.** The listener isn't currently part of the supervision tree
|
||
(see `lib/microwaveprop/application.ex`). When it is re-added, early
|
||
startup or Postgrex.Notifications boot failure can leave it
|
||
unsubscribed for several seconds while Rust writes land.
|
||
|
||
**Recovery.** Same as FM1 — reconciler catches up within one tick.
|
||
|
||
### FM3 — NFS stale read racing Rust write
|
||
|
||
**Symptom.** `ScoresFile.read_bounds/2` raises `File.Error` when
|
||
the reader opens the file exactly between Rust's `rename(2)` of the
|
||
`.tmp.<uniq>` file and the `fsync` landing on the NFS server.
|
||
|
||
**Why.** The rename is atomic on the same NFSv4 filesystem, but the
|
||
NFS client cache can briefly serve a stale directory listing or a
|
||
no-longer-existing path.
|
||
|
||
**Recovery.** Reconciler's `rescue File.Error` returns `:error` and
|
||
moves on; the next tick re-reads the file after NFS consistency
|
||
settles. Lazy-miss path (`Propagation.scores_at_fetch/3`) has
|
||
historically surfaced this as a LiveView error — the reconciler
|
||
warms the cache before that miss happens.
|
||
|
||
### FM4 — Cluster partition splits ScoreCache broadcast
|
||
|
||
**Symptom.** Pods on one side of a cluster partition have up-to-date
|
||
cache; pods on the other side don't, even though both can read NFS
|
||
and Postgres.
|
||
|
||
**Why.** `ScoreCache.broadcast_put/3` fans the payload out over
|
||
PubSub. A libcluster partition stops the message reaching partitioned
|
||
peers.
|
||
|
||
**Recovery.** Each pod's reconciler reads from NFS directly, so it
|
||
converges regardless of PubSub reachability. Partition does not
|
||
cause divergent scores — only transient latency until the next
|
||
sweep.
|
||
|
||
### FM5 — Rust worker dead
|
||
|
||
**Symptom.** `grid_tasks` rows with `status = 'queued'` accumulate;
|
||
no new `.ntms` files land; `/map` serves the last successful run's
|
||
scores.
|
||
|
||
**Why.** Deployment bug, Rust panic, OOM on talos5 (unlikely with
|
||
32 GB headroom), or disk-full on `/data`.
|
||
|
||
**Recovery.** Manual. `kubectl -n prop logs deploy/prop-grid-rs` for
|
||
cause. The Elixir side is read-only against the grid — it does not
|
||
regenerate scores from within the BEAM since the extraction in
|
||
`65693ed`.
|
||
|
||
## Monitoring signals
|
||
|
||
- `avg_over_time(beam_memory_total_bytes[1h])` on prop pods — steady
|
||
climb above 1 GiB means the cache isn't pruning (see
|
||
`ScoreCache.prune_older_than/1` in `NotifyListener`).
|
||
- Reconciler log line `"ScoreCacheReconciler: warmed N"` — if N is
|
||
consistently non-zero every minute, NOTIFY is dropping and
|
||
NotifyListener wants investigation.
|
||
- `grid_tasks` queue depth in Prometheus — should drain to 0 between
|
||
hourly cron ticks.
|
||
|
||
## Test coverage
|
||
|
||
- `test/microwaveprop/propagation/score_cache_reconciler_test.exs`
|
||
covers FM1 and FM3 (missing-file skip).
|
||
- Partition recovery (FM4) is untested in CI — requires multi-node
|
||
cluster and libcluster orchestration.
|