+++ title = "When Postgres isn't the right choice" date = 2026-04-18 draft = true +++ I've been working on For the [Microwave Propagation](https://prop.w5isp.com) site I've been working on ┌────────────────────────────────────────┬───────────┬───────────┐ │ Phase │ Wall time │ % of fh=0 │ ├────────────────────────────────────────┼───────────┼───────────┤ │ HRRR fetch (f00) │ 27s │ 5% │ ├────────────────────────────────────────┼───────────┼───────────┤ │ Store HRRR profiles (JSONB upsert) │ 37s │ 7% │ ├────────────────────────────────────────┼───────────┼───────────┤ │ Native duct fetch (wgrib2) │ 104s │ 19% │ ├────────────────────────────────────────┼───────────┼───────────┤ │ NEXRAD fetch + merge (f00 only) │ 17s │ 3% │ ├────────────────────────────────────────┼───────────┼───────────┤ │ Commercial-link merge (f00 only) │ 95s │ 17% │ ├────────────────────────────────────────┼───────────┼───────────┤ │ Score grid + upsert propagation_scores │ 281s │ 50% │ ├────────────────────────────────────────┼───────────┼───────────┤ │ Total fh=0 │ 562s │ 100% │ └────────────────────────────────────────┴───────────┴───────────┘ The propagation map at prop.w5isp.com ingests HRRR weather every hour and re-scores the entire CONUS grid across five amateur-radio microwave bands. That is 95,000 grid points × 5 bands × 19 forecast hours = ~9 million score values per chain. For a long time those lived in a Postgres table called propagation_scores. They don't anymore, and the map went from sluggish to instantaneous. This is the story of why. The shape mismatch Each hourly run did this: DELETE FROM propagation_scores WHERE valid_time = $1; INSERT INTO propagation_scores (lat, lon, band_mhz, valid_time, score, factors) ... 475,000 rows per forecast hour, funneled through four indexes (PK UUID, uniqueness tuple, valid_time, (band_mhz, valid_time)), plus a JSONB factors blob for the analysis hour. Scoring + upsert took about 4m 40s per forecast hour. Across 19 hours that's more than an entire hourly cron cycle, so we had to back off to every-three-hours just to keep up. The problem wasn't Postgres — Postgres was doing exactly what we asked. The problem was that the shape we were asking for was wrong for how the data gets read. The map is a canvas heatmap. It wants one dense (lat, lon) → score array for a single (band, valid_time). We were storing it as half a million row tuples with MVCC bookkeeping, toast pages, and a B-tree index per column we'd never query by. Every write rematerialized metadata the reader throws away. Three stopgaps that weren't enough Before rewriting the store, we tuned what was there: 1. DELETE + INSERT … with no ON CONFLICT. The chain worker rewrites a full (valid_time, all bands) slice every hour, so conflict detection was pure waste. Skipping it shaved a chunk off. 2. Skip factors on forecast hours. factors is a ~200-byte JSONB per row describing the ten scoring components (rain, humidity, refractivity, etc.). We only show it on the analysis hour (f00) when the user clicks a cell. Passing factors: nil for f01–f18 skips a JSONB encode plus a toast write for roughly half the volume. 3. UNLOGGED table. Writes bypass the WAL entirely. On unclean shutdown the table truncates — fine, because PropagationGridWorker rebuilds it from HRRR every hour. Durability was never the point; this was a cache with extra steps. These helped, but the per-row overhead of a row-oriented store against a dense numeric grid is a ceiling you can't tune past. The phase was still the wall-clock dominator. The binary file Each (band_mhz, valid_time) grid now lands in one file at /data/scores/{band}/{iso}.ntms: magic : 4 bytes "NTMS" version : 1 byte 0x01 band_mhz : 4 bytes uint32 LE valid_time : 8 bytes int64 unix seconds LE lat_min : 4 bytes float32 LE lon_min : 4 bytes float32 LE step_deg : 4 bytes float32 LE n_rows : 2 bytes uint16 LE n_cols : 2 bytes uint16 LE scores : n_rows × n_cols bytes (uint8, 0–100; 255 = no-data) 33-byte header, dense uint8 array. A full CONUS grid serializes to ~93 KB per band per hour. A cell at (lat, lon) is at byte offset row × n_cols + col, so a point lookup is arithmetic plus a single File.read/1. No index, no query planner, no lock manager. Writes go through the temp-then-rename pattern: tmp = path <> ".tmp." <> unique_suffix() File.write!(tmp, binary, [:binary]) File.rename!(tmp, path) With Postgres off the hot path, the full f00–f18 chain dropped from ~170 min to ~45–60 min, so we ran the cron up to hourly. The 15-minute prune cron reaped expired files with rm — instant. The f00 factor breakdown One thing binary files don't handle well is the analysis-hour factor breakdown (the 10-component panel you see when you click a cell). Those factors are heterogeneous Elixir maps, not a numeric grid. The solution was a second file format, ProfilesFile: one compressed ETF file per valid_time at /data/scores/profiles/{iso}.etf.gz, storing the enriched HRRR grid keyed by {lat, lon}. On a click, we reload the relevant cell's profile and rescore on demand. Same atomic-rename write pattern, different payload — ETF where heterogeneity matters, raw binary where density does.