prop/docs/plans/2026-04-14-duckdb-scores-storage.md
Graham McIntire 5ec66df135
Cut propagation_scores write cost with DELETE+COPY, skip-factors, UNLOGGED
The scoring+upsert phase was ~4m40s per forecast hour and dominated
wall time. Three stacked optimizations attack it from different
angles.

replace_scores/2 is a new hot-path writer that does DELETE WHERE
valid_time = $1 followed by a plain insert_all (no ON CONFLICT
resolution). The chain worker rewrites the full (valid_time, all
bands) slice every forecast hour, so conflict detection was pure
waste. AsosAdjustmentWorker still uses upsert_scores because it
only rewrites the subset of cells near a station.

factors is now nullable. Forecast hours f01-f18 pass factors: nil
so the JSONB encode + toast write is skipped entirely — roughly
halves the data volume per run. point_detail/4 coalesces nil to
an empty map so the JS popup renders without a TypeError, and
scorer_diff only pulls the most recent valid_time that still has
factors (the f00 row).

propagation_scores is now UNLOGGED, so inserts bypass WAL entirely.
Durability tradeoff: an unclean shutdown truncates the table, but
PropagationGridWorker rebuilds it from HRRR every 3h so a lost
table is re-populated within one cron cycle.

Also adds docs/plans/2026-04-14-duckdb-scores-storage.md — a
speculative plan for a flat-file / DuckDB rewrite with explicit
trigger conditions for when to pick it up (partitioning deferred
too; revisit only if these three don't solve it).
2026-04-14 13:47:35 -05:00

208 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# DuckDB / Flat-File Propagation Scores Storage Plan
**Goal:** Replace Postgres-backed `propagation_scores` with per-`(band_mhz, valid_time)` on-disk files (DuckDB or raw binary), so the hot write path becomes a single file write instead of 475k Postgres rows through indexes + WAL.
**Status:** Speculative architectural alternative. Do not execute until the already-shipped Postgres wins (DELETE+COPY, skip-factors, UNLOGGED) have been measured in prod. If those three drop the scoring+upsert phase below ~60s per forecast hour, this rewrite is likely unnecessary.
---
## Why consider this
Each forecast-hour write to `propagation_scores` is:
- 475k rows (95k grid points × 5 bands)
- Passes through 4 indexes (PK UUID, unique tuple, valid_time, (band_mhz, valid_time))
- Even with DELETE+COPY and UNLOGGED, Postgres overhead — toast pages, page locks, vacuum state, MVCC tuple housekeeping — adds per-row cost that a flat array layout avoids entirely
- Shapes a dense numeric grid into a row-oriented representation the map never needs
The map render path also doesn't want individual rows. It wants a dense `(lat, lon) → score` array for a single `(band, valid_time)`. SQL is the wrong shape.
## Approach options
### Option A — Parquet files, read via DuckDB
One Parquet file per `(band_mhz, valid_time)`:
```
/srtm/scores/10000/2026-04-14T18-00-00Z.parquet
/srtm/scores/24000/2026-04-14T18-00-00Z.parquet
...
```
- Columns: `lat`, `lon`, `score` (plus `factors` only on f00 files)
- ~95k rows × 3 numeric columns + JSONB on f00 ≈ 1-2 MB per f00 file, ~400 KB per fh≥1 file
- Written once per forecast hour via Elixir → DuckDB WAL-less writer (`COPY ... TO 'file.parquet'`) or a Rust/Elixir native Parquet encoder
- Read on the map side via [`duckdbex`](https://hex.pm/packages/duckdbex) (Elixir NIF wrapper) or a Rust sidecar
Pros:
- Columnar compression: a single Parquet file per hour is ~300-500 KB compressed
- DuckDB can query across files: `SELECT score FROM 'scores/10000/*.parquet' WHERE lat BETWEEN ... AND ...`
- Files are portable — easy to debug, copy, inspect with `duckdb-cli`
- Pruning is `rm` — instant
Cons:
- Add a new dependency (`duckdbex` or a Rust NIF)
- Concurrent writers across multiple pods need file coordination (atomic rename, or a single writer pod)
- `factors` for f00 still needs somewhere to live — either embedded in the Parquet or a separate lookup file
### Option B — Raw binary mmap files
One file per `(band_mhz, valid_time)`, laid out as:
```
header: 16 bytes (magic + lat_min, lon_min, step, n_rows, n_cols)
scores: n_rows × n_cols × 1 byte (score 0-100)
```
- ~95k × 1 byte ≈ 95 KB per file uncompressed
- Written in Elixir via `File.write/2` on a binary blob
- Read on the map side via `:ets.new(:mmap)` / `File.read/1` cached per `(band, vt)`
Pros:
- Simplest possible layout
- Writes are ~20 ms per file
- No dependencies at all
- Matches the canvas heatmap's consumption shape exactly
Cons:
- Point-detail (single-cell lookup) needs an index by (lat, lon) — easy with the grid step but less ergonomic than SQL
- `factors` needs a separate store (another binary file, or keep it in Postgres)
- No ad-hoc query capability (can't ask "what was the score at 35.0/-97.0 two hours ago" without loading the file)
### Option C — DuckDB on persistent disk
Single `.duckdb` file under `/srtm/scores.duckdb`:
- One table with the same schema as `propagation_scores`
- DuckDB is columnar, compressed, and serializes writes via a single database file
- Writes via `INSERT INTO scores VALUES ...` batched
- Reads via direct SQL from Elixir
Pros:
- Least migration friction — keeps a SQL-ish API
- Columnar storage naturally compresses `score` columns
- DuckDB writes are dramatically faster than Postgres for bulk inserts (no WAL, no per-row overhead)
Cons:
- DuckDB is single-writer. Multi-pod clusters need to funnel writes through one pod (the propagation grid worker pod — which is exactly the worker-pod-isolation idea discussed earlier in this session)
- File corruption on crash: DuckDB has improved crash recovery but it's not Postgres-grade
- Requires `duckdbex` NIF
## Recommendation
If the already-shipped Postgres wins aren't enough, start with **Option B** (raw binary files). It's the lowest-risk, highest-impact change because:
1. The map's canvas heatmap *already* wants a dense score array — a binary file is a closer shape match than SQL rows
2. No new dependencies — `File.write/2` and `File.read/1` are enough
3. Concurrent write safety is a solved problem: write to a temp file, `File.rename/2` atomically
4. Reverting is easy: keep `propagation_scores` populated from the flat file asynchronously during a migration period
The `factors` storage question is the only thing that forces picking between Option A and Option B:
- If we're fine with **no factors for forecast hours** (already the case after the skip-factors change), factors for f00 stay in Postgres and the hot path is file-only. Option B works as-is.
- If we want factors for all hours, Option A (Parquet) is cleaner because factors can ride on the same file.
## Architectural prerequisites
Before this is worth doing:
1. **Separate worker pod**. Multi-writer coordination on a single file is painful; a single "propagation" pod that owns all writes is much simpler. The isolation idea from earlier in this session dovetails here.
2. **Shared volume or single-node**. Files need to live on a PVC shared between the worker and web pods, or the web pods need a cached-read path that pulls the file from the worker pod via HTTP. The existing `/srtm` NFS mount is already shared, which makes `/srtm/scores` free.
3. **Baseline perf measurements** from the already-shipped Postgres changes. Don't do this if the map already feels fast.
## Read-path changes
### Map render (bulk)
Current:
```elixir
Propagation.scores_at(band_mhz, valid_time, bounds)
# → list of %{lat, lon, score} maps from Postgres
```
New:
```elixir
Propagation.ScoresFile.load(band_mhz, valid_time)
# → %{lat_min, lon_min, step, rows, cols, data: <<binary>>}
```
The JS hook receives the binary directly (base64 or bytestring in an assign) and draws the canvas without an intermediate list traversal.
### Point detail (single-cell)
Current:
```elixir
Propagation.point_detail(band, lat, lon, valid_time)
# → %{score, factors, ...}
```
New:
```elixir
Propagation.ScoresFile.point(band_mhz, valid_time, lat, lon)
# → %{score}
Propagation.FactorsStore.for_point(lat, lon, valid_time)
# → %{factors} (still Postgres-backed, only f00)
```
### Scorer diff
Current: reads latest f00 rows with factors from Postgres. Unchanged — factors stay in Postgres.
### ScoreCache
Current: ETS cache of `(band, vt) → list-of-scores`. Unchanged — can wrap the file-load path the same way.
## Write-path changes
`PropagationGridWorker.run_chain_step/2` currently calls:
```elixir
scores = compute_scores(grid_data, valid_time, forecast_hour)
Propagation.replace_scores(scores, valid_time)
```
Replace with:
```elixir
scores = compute_scores(grid_data, valid_time, forecast_hour)
Propagation.ScoresFile.write!(band_mhz, valid_time, scores)
# Still write factors to Postgres for f00 only
if forecast_hour == 0, do: Propagation.replace_factors(factor_rows, valid_time)
```
Expected timings: ~10-50 ms per band per fh vs 4-5 seconds per band per fh today. The scoring CPU cost stays the same; only the storage phase collapses.
## Pruning
Current: 15-minute cron running `DELETE WHERE valid_time < now - 2h`.
New: 15-minute cron running `find /srtm/scores -mmin +120 -delete` (or Elixir equivalent).
Zero vacuum pressure, zero WAL, zero index maintenance.
## Migration path (if this gets greenlit)
1. **Commit 1**: Introduce `Propagation.ScoresFile` module. Dual-write: `replace_scores` still writes Postgres as today, but also writes the flat file. Map render still reads from Postgres. Verify file shape matches what the JS side wants.
2. **Commit 2**: Switch map-render read path to load from the flat file. Keep Postgres dual-write as a fallback. Measure wall-time improvement.
3. **Commit 3**: Stop writing to Postgres. Drop the `propagation_scores` table (or keep it empty for rollback). Point detail for forecast hours reads score from file; factors for f00 still hits a Postgres table (or a separate file).
Each commit is independently reversible.
## Risk / pushback
- **Added complexity** of a second storage system. Today: Postgres only. After: Postgres for QSOs + ASOS + weather + factors, plus a flat-file store for grid scores. The cognitive overhead is real.
- **Shared-volume assumptions** couple web and worker pods to the same PVC. Worth revisiting the separate-worker-pod idea first; if that's not on the table, the flat-file approach is harder because every web pod needs a consistent view of the files.
- **Queryability loss**. Today, ad-hoc SQL against `propagation_scores` works for debugging. After: need to boot a Python/DuckDB shell to query files.
- **Opportunity cost**. If the Postgres tuning already lands us in under 90 min per chain, this rewrite is 2-3 weeks of work for perhaps a 2× improvement on a phase that's already fast enough. The plan survives regardless — we can always pick it up if scale demands it (finer grid resolution, more forecast hours, more bands).
## When to revisit
Trigger conditions for picking this up:
- A single forecast-hour chain step still exceeds 5 min after all three Postgres wins are deployed
- A new feature requires 100m grid resolution (~8× more points → 4M rows per fh), which will blow through the Postgres wins immediately
- The Postgres pool becomes a bottleneck for other workers because `propagation` queue writes keep it saturated
Until one of those trips, the existing `propagation_scores` table + the shipped tuning is the right place to be.