From aed333dbb626f008c1c4c7654c093cf64107d1cd Mon Sep 17 00:00:00 2001 From: Graham McIntire Date: Mon, 20 Apr 2026 12:45:43 -0500 Subject: [PATCH] add blog post --- content/blog/postgres1.md | 159 ++++++++++++++++++-------------------- 1 file changed, 76 insertions(+), 83 deletions(-) diff --git a/content/blog/postgres1.md b/content/blog/postgres1.md index 9990710..c3ff4dd 100644 --- a/content/blog/postgres1.md +++ b/content/blog/postgres1.md @@ -1,102 +1,95 @@ +++ title = "When Postgres isn't the right choice" -date = 2026-04-18 -draft = true +date = 2026-04-20 +draft = false +++ -I've been working on For the [Microwave Propagation](https://prop.w5isp.com) site I've been working on +I've been spending most of my recent unemployment working on the [Microwave Propagation](https://prop.w5isp.com) prediction website. - ┌────────────────────────────────────────┬───────────┬───────────┐ - │ Phase │ Wall time │ % of fh=0 │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ HRRR fetch (f00) │ 27s │ 5% │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ Store HRRR profiles (JSONB upsert) │ 37s │ 7% │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ Native duct fetch (wgrib2) │ 104s │ 19% │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ NEXRAD fetch + merge (f00 only) │ 17s │ 3% │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ Commercial-link merge (f00 only) │ 95s │ 17% │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ Score grid + upsert propagation_scores │ 281s │ 50% │ - ├────────────────────────────────────────┼───────────┼───────────┤ - │ Total fh=0 │ 562s │ 100% │ - └────────────────────────────────────────┴───────────┴───────────┘ +The propagation map ingests HRRR weather every hour and re-scores the entire CONUS grid across many +amateur radio microwave bands. That is 95,000 grid points * 23 bands * 19 forecast hours = ~41.5 +million score values per chain. For a long time those lived in a Postgres table called +propagation_scores. - The propagation map at prop.w5isp.com ingests HRRR weather every hour and re-scores the entire CONUS grid across five - amateur-radio microwave bands. That is 95,000 grid points × 5 bands × 19 forecast hours = ~9 million score values per chain. For - a long time those lived in a Postgres table called propagation_scores. They don't anymore, and the map went from sluggish to - instantaneous. This is the story of why. +## The shape mismatch - The shape mismatch +Each hourly run did this: + + DELETE FROM propagation_scores WHERE valid_time = $1; + INSERT INTO propagation_scores (lat, lon, band_mhz, valid_time, score, factors) ... - Each hourly run did this: +475,000 rows per forecast hour (at the time it was only 5 bands, now it would be 2,185,000 per +forecast hour), funneled through four indexes (PK UUID, uniqueness tuple, valid_time, (band_mhz, +valid_time)), plus a JSONB factors blob for the analysis hour. Scoring + upsert took about 4m 40s +per forecast hour. Across 19 hours that's more than an entire hourly cron cycle, so we had to back +off to once every 3 hours just to keep up. + +I was originally running the Postgres server and app on a dell r630 server I have running proxmox +on some raid5 SAS spinny drives. It quickly became obvious IOPS were the primary issue. The first +mitigation I applied was to move the database to an nvme backed server. - DELETE FROM propagation_scores WHERE valid_time = $1; - INSERT INTO propagation_scores (lat, lon, band_mhz, valid_time, score, factors) ... - - 475,000 rows per forecast hour, funneled through four indexes (PK UUID, uniqueness tuple, valid_time, (band_mhz, valid_time)), - plus a JSONB factors blob for the analysis hour. Scoring + upsert took about 4m 40s per forecast hour. Across 19 hours that's - more than an entire hourly cron cycle, so we had to back off to every-three-hours just to keep up. - - The problem wasn't Postgres — Postgres was doing exactly what we asked. The problem was that the shape we were asking for was - wrong for how the data gets read. - - The map is a canvas heatmap. It wants one dense (lat, lon) → score array for a single (band, valid_time). We were storing it as - half a million row tuples with MVCC bookkeeping, toast pages, and a B-tree index per column we'd never query by. Every write - rematerialized metadata the reader throws away. - - Three stopgaps that weren't enough - - Before rewriting the store, we tuned what was there: - - 1. DELETE + INSERT … with no ON CONFLICT. The chain worker rewrites a full (valid_time, all bands) slice every hour, so conflict - detection was pure waste. Skipping it shaved a chunk off. - 2. Skip factors on forecast hours. factors is a ~200-byte JSONB per row describing the ten scoring components (rain, humidity, - refractivity, etc.). We only show it on the analysis hour (f00) when the user clicks a cell. Passing factors: nil for f01–f18 - skips a JSONB encode plus a toast write for roughly half the volume. - 3. UNLOGGED table. Writes bypass the WAL entirely. On unclean shutdown the table truncates — fine, because PropagationGridWorker - rebuilds it from HRRR every hour. Durability was never the point; this was a cache with extra steps. - - These helped, but the per-row overhead of a row-oriented store against a dense numeric grid is a ceiling you can't tune past. The - phase was still the wall-clock dominator. - +The problem wasn't Postgres. Postgres is an amazing database when you use it for what it's good at. +The problem is trying to store ephemeral data that rotates once an hour. AUTOVACUUM couldn't keep +up even on the faster nvme, so ANALYZE never got run. It compounded every hour. - The binary file +The map is a canvas heatmap. It wants one dense (lat, lon) -> score array for a single (band, +valid_time). We were storing it as half a million row tuples with MVCC bookkeeping, toast pages, +and a B-tree index per column we'd never query by. Every write rematerialized metadata the reader +throws away. - Each (band_mhz, valid_time) grid now lands in one file at /data/scores/{band}/{iso}.ntms: +Before rewriting the store, we tuned what was there: - magic : 4 bytes "NTMS" - version : 1 byte 0x01 - band_mhz : 4 bytes uint32 LE - valid_time : 8 bytes int64 unix seconds LE - lat_min : 4 bytes float32 LE - lon_min : 4 bytes float32 LE - step_deg : 4 bytes float32 LE - n_rows : 2 bytes uint16 LE - n_cols : 2 bytes uint16 LE - scores : n_rows × n_cols bytes (uint8, 0–100; 255 = no-data) +1. DELETE + INSERT ... with no ON CONFLICT. The chain worker rewrites a full (valid_time, all bands) slice every hour, so conflict detection was pure waste. Skipping it shaved time off. +2. Skip factors on forecast hours. factors is a ~200-byte JSONB per row describing the ten scoring components (rain, humidity, refractivity, etc.). We only show it on the analysis hour (f00) when the user clicks a cell. Passing factors: nil for f01–f18 skips a JSONB encode plus a toast write for roughly half the volume. +3. UNLOGGED table. Writes bypass the WAL entirely. On unclean shutdown the table truncates — fine, because PropagationGridWorker rebuilds it from HRRR every hour. Durability was never the point, this was a cache with extra steps. - 33-byte header, dense uint8 array. A full CONUS grid serializes to ~93 KB per band per hour. A cell at (lat, lon) is at byte - offset row × n_cols + col, so a point lookup is arithmetic plus a single File.read/1. No index, no query planner, no lock - manager. +These helped, but the per row overhead of a row oriented store against a dense numeric grid is a +ceiling you can't tune past. I was still running into issues with Postgres not being able to keep +up. - Writes go through the temp-then-rename pattern: +## Enter the binary file - tmp = path <> ".tmp." <> unique_suffix() - File.write!(tmp, binary, [:binary]) - File.rename!(tmp, path) +Each (band_mhz, valid_time) grid now lands in one file at /data/scores/{band}/{iso}.ntms that's +mounted over NFS by the pods: - With Postgres off the hot path, the full f00–f18 chain dropped from ~170 min to ~45–60 min, so we ran the cron up to hourly. The - 15-minute prune cron reaped expired files with rm — instant. + magic : 4 bytes "NTMS" + version : 1 byte 0x01 + band_mhz : 4 bytes uint32 LE + valid_time : 8 bytes int64 unix seconds LE + lat_min : 4 bytes float32 LE + lon_min : 4 bytes float32 LE + step_deg : 4 bytes float32 LE + n_rows : 2 bytes uint16 LE + n_cols : 2 bytes uint16 LE + scores : n_rows * n_cols bytes (uint8, 0–100; 255 = no-data) - The f00 factor breakdown +33-byte header, dense uint8 array. A full CONUS grid serializes to ~93 KB per band per hour. A cell +at (lat, lon) is at byte offset row * n_cols + col, so a point lookup is arithmetic plus a single +File.read/1. No index, no query planner, no lock manager. - One thing binary files don't handle well is the analysis-hour factor breakdown (the 10-component panel you see when you click a - cell). Those factors are heterogeneous Elixir maps, not a numeric grid. +Writes go through the temp-then-rename pattern: - The solution was a second file format, ProfilesFile: one compressed ETF file per valid_time at - /data/scores/profiles/{iso}.etf.gz, storing the enriched HRRR grid keyed by {lat, lon}. On a click, we reload the relevant cell's - profile and rescore on demand. Same atomic-rename write pattern, different payload — ETF where heterogeneity matters, raw binary - where density does. + tmp = path <> ".tmp." <> unique_suffix() + File.write!(tmp, binary, [:binary]) + File.rename!(tmp, path) + +With Postgres off the hot path, the full f00–f18 chain dropped from ~170 min to ~45–60 min, so we +ran the cron back up to hourly. The 15-minute prune cron reaped expired files with rm which is instant. + +## The f00 factor breakdown + +One thing binary files don't handle well is the analysis-hour factor breakdown (the 10-component +panel you see when you click a cell). Those factors are heterogeneous Elixir maps, not a numeric +grid. + +The solution was a second file format, ProfilesFile: one compressed ETF file per valid_time at +/data/scores/profiles/{iso}.etf.gz, storing the enriched HRRR grid keyed by {lat, lon}. On a click, +we reload the relevant cell's profile and rescore on demand. Same atomic-rename write pattern, +different payload — ETF where heterogeneity matters, raw binary where density does. + +## Summary + +The whole propagation project has been fun to write and tweak continuously. I continue to find and +fix various bottlenecks and limitations. Each individual component in the stack is great at what it +does, but sometimes when you start throwing millions of rows per hour and then deleting them it +gets a little fussy.