ScoreCache held 437 entries × ~1.94 MiB each (~850 MiB per pod) of
{band_mhz, valid_time} grids — data already on NFS as compact .prop
files. NotifyListener and ScoreCacheReconciler both eagerly materialised
every band into ETS on each Rust completion / 60s sweep, so 5 hot
replicas wasted ~4.2 GiB of redundant cache and OOMed at 6 GiB.
Bound the cache to 32 entries with eviction by oldest valid_time, drop
the eager warm loops, and let LiveView callers lazy-fill via the
existing ScoresFile read path. ScoreCacheReconciler had no remaining
purpose and is deleted; runbook + prom_ex counters updated to match.
Steady-state cache footprint ~60 MiB per pod instead of ~850 MiB.
Wraps the hourly grid_tasks seed with
[:microwaveprop, :propagation, :grid_worker, :perform] so Prometheus
can see duration, success, and failure for the chain-seed cron.
Emits a :seeded event with the queued_tasks count and an :exception
event on GridTaskEnqueuer failure so silent seed errors show up in
the existing PromEx dashboards instead of dying as a lone Logger.info.
Three small wins from a telemetry audit:
1. Propagation.scores_at/3 wrapped every call in Instrument.span, firing
two handler dispatches that dominated the ~10µs ETS lookup on cache
hits. The map's LiveView fires this on every pan + point-click, so
hot-path latency was mostly telemetry. Span now wraps only the miss
branch (where disk IO makes the duration signal meaningful); hit path
still emits the cheap hit/miss counter the cache-ratio panel reads.
2. Oban queue-depth poller: 10s → 30s. Every replica independently
GROUP-BYs oban_jobs, so each extra replica paid ~3 redundant queries
per minute for a gauge that moves on the hour-scale anyway.
3. Dropped summary(phoenix.endpoint.start.system_time): summarizing a
wall-clock timestamp produces no useful aggregate — stop.duration
below it is the meaningful signal.
- Remove the 7-day outlook strip and daily_outlook_at/3. It was
derived from point_forecast, which is capped at HRRR's 18-hour
horizon, so it never had 7 days of data to show.
- factors_for: fall back to the nearest persisted analysis profile
at or before the requested time. Only f00 hours persist profiles,
so clicking any other hour used to yield an empty Analysis and
factor table. After the recent cursor fix lands 'Now' on a
forecast hour more often, this regressed further.
- available_valid_times/1: cap the forward horizon to 18h from now.
Leftover valid_times from prior cycles used to pile on the
timeline without adding information.
Adds spans to 15 previously-unmeasured hot paths so every question we
might ask while tuning has a histogram to answer it:
External I/O:
- iem.fetch_iemre (gridded weather reanalysis)
- mrms.list_latest / mrms.download (precip radar)
- rtma.fetch_observation
- ncei.fetch_metar (historical 5-min METAR backfill)
- solar.fetch_indices (GFZ solar indices)
- swpc.fetch (SWPC Kp/F10.7/X-ray)
- giro.fetch (ionosonde)
- qrz.request, geocoder.geocode (callsign enrichment)
- srtm.download_tile (terrain tile download + gunzip)
- hrrr.download_grib_ranges (parallel byte-range fetch phase)
Subprocess:
- wgrib2.extract_grid / extract_grid_from_file / extract_grid_from_file_mapped
LiveView hot paths:
- propagation.scores_at (map score fetch + cache hit/miss counter)
- propagation.point_forecast (sparkline)
- propagation.point_detail (click-to-inspect)
- propagation.daily_outlook_at (/map outlook strip)
Worker-level end-to-end:
- worker.terrain_profile
- worker.mechanism_classify
- worker.mrms_fetch
Each event is registered in Microwaveprop.PromEx.InstrumentPlugin as
a Prometheus histogram (default / long buckets as appropriate) plus
a counter for the scores_at cache hit/miss ratio. Prometheus at
10.0.15.25 will start seeing the new series on the next scrape after
deploy.
Wires PromEx with its built-in Application / BEAM / Phoenix / Ecto /
Oban plugins plus a custom `Microwaveprop.PromEx.InstrumentPlugin`
that registers Prometheus histograms for every
`Microwaveprop.Instrument` span (HRRR/NEXRAD/IEM/GEFS/NARR/UWYO
/Elevation fetches, DB batch upserts, terrain analysis, radar
aggregation, propagation score band). Queue-depth gauges from the
10-second poller land at `microwaveprop_oban_queue_count{queue,state}`.
Router forwards `/metrics` to MicrowavepropWeb.MetricsPlug which
optionally requires a bearer token (`PROMETHEUS_AUTH_TOKEN` env var in
runtime.exs); leave it unset to expose metrics publicly and restrict
at the ingress / Cloudflare Access layer instead.
Example external Prometheus scrape config:
scrape_configs:
- job_name: microwaveprop
scrape_interval: 30s
metrics_path: /metrics
scheme: https
authorization:
type: Bearer
credentials: <token>
static_configs:
- targets: ['prop.w5isp.com:443']