Two related changes to the contact detail page:
1. The "Flagged invalid" badge was wrapping to two lines on
narrow widths because the daisyUI badge has no explicit
no-wrap class. Add `whitespace-nowrap`, drop the redundant
capitalization on "Invalid", and append the flagger's
callsign so the badge reads "Flagged invalid by W5XYZ".
Hover tooltip carries the timestamp.
2. New `flagged_by_user_id` + `flagged_at` columns on contacts.
`Radio.toggle_flagged_invalid!/2` now takes the admin User
doing the flagging; on unflag it clears both fields so a
subsequent re-flag gets a fresh attribution and the badge
never shows a stale flagger.
3. `Radio.delete_contact!/1` + a Delete button next to Flag.
Admin-only, gated in both the event handler and the template.
Uses LV's `data-confirm` for the destructive-action prompt.
All FK references to `contacts` already declare
`on_delete: :delete_all` (terrain_profiles, contact_edits,
contact_common_volume_radar) so a single Repo.delete!
cascades cleanly.
K8s pods have IPv4-only egress; without :inet in the gen_tcp opts,
BEAM was picking mqtt.pskreporter.info's AAAA record and the
kernel's connect() returned EINVAL, surfacing as :einval and a
30-second retry loop.
emqtt's transitive `quicer` dep needs CMake + msquic submodule
clones to build, which the slim Debian builder image doesn't ship —
the prod CI build was failing inside `mix deps.compile`. Tortoise311
is the obvious pure-Elixir alternative but it bundles a
gen_state_machine + retry/inflight stack we don't use because we're
subscribe-only at QoS 0.
Implementing the wire protocol inline is ~150 lines in
`Microwaveprop.Pskr.Mqtt` (CONNECT/SUBSCRIBE/PINGREQ/DISCONNECT
builders, CONNACK/PUBLISH/SUBACK/PINGRESP parser, varint codec).
Pskr.Client now uses :gen_tcp with `active: :once` framing,
buffers partial reads, schedules its own keepalive, and parses
inbound packets in a tight loop.
Drops `BUILD_WITHOUT_QUIC=1` from the Dockerfile — no longer
relevant since there's no native dep at all in the build chain
now.
Subscribes to mqtt.pskreporter.info:1883 over plain TCP and folds
every reception report into hourly aggregates per
(band, sender_grid_4, receiver_grid_4). Each spot is the empirical
"propagation actually occurred on this path right now" signal we'll
correlate against HRRR atmospheric state to recalibrate the scoring
weights. Aggregating at the path-hour bucket collapses 50-reporter
QSOs to one row instead of 50.
Topic filter anchors on USA (DXCC 291) on either sender or receiver
side so the broker pre-filters out everything outside our HRRR
domain. Bands subscribed: VHF and up (6m, 2m, 70cm, 23cm, +
microwave) — HF is dominated by ionospheric propagation, which this
project doesn't model.
Pskr.Client cluster-elects a singleton via :global.register_name —
all replicas start the GenServer but only one connects to MQTT;
the rest stay in :standby and watch for the leader's nodedown to
re-run election. Off in dev/test, on by default in prod
(PSKR_MQTT_ENABLED=false as kill-switch).
Pskr.Aggregator buffers in memory keyed by Pskr.path_key/1 and
flushes every 60s via Repo.insert_all/3 with an additive
ON CONFLICT (count summed, SNR envelope by GREATEST/LEAST,
modes unioned via unnest+DISTINCT). Idempotent across overlapping
flushes and across leader handover.
Dockerfile sets BUILD_WITHOUT_QUIC=1 — emqtt's transitive quicer
dep wants CMake + OpenSSL headers we'd otherwise have to add to
the builder image just to never use QUIC over plain MQTT. Base
image is unchanged; the new dep compiles cleanly into the existing
prop-base runtime.
phoenix_pubsub_redis 3.1's deprecation message ('Move :url inside the
:redis_opts option') is misleading — wrapping the URL as
`redis_opts: [url: url]` forwards the keyword list to Redix, which
rejects `:url` with `unknown options [:url], valid options are:
[:host, :port, :password, ...]`. New pods on commit b4d6b04 onward
crashed at startup with that NimbleOptions.ValidationError; old pods
running e53a34d (top-level :url) coasted on already-established
connections so the breakage only surfaced when fresh pods rolled.
The library accepts `redis_opts` as either a URL binary OR a
keyword list. Pass the URL as a bare string so it takes the
binary-clause path that hands it directly to Redix.PubSub.start_link/1.
The library deprecated top-level Redis connection keys; passing :url
directly logged a warning on every boot. Functionally identical, just
silences the warning.
Switches the PubSub adapter to Phoenix.PubSub.Redis backed by the
shared Valkey instance (same backend GridCache uses), via the new
phoenix_pubsub_redis dep. Falls back to the default PG2 / Erlang-
distribution adapter when VALKEY_URL is unset, so dev/test boot
unchanged.
Decouples PubSub fan-out from libcluster's BEAM connection. A
flapping Erlang distribution (the OOM-cascade scenario from
2026-05-03 surfaced this — libcluster warnings dominated the logs
while pods CrashLoopBackOff'd) no longer takes broadcasts with it,
so the rest of the cluster keeps publishing while the bad pod
recovers.
node_name uses POD_IP for cluster-wide uniqueness, falling back to
Erlang's node() name in non-K8s environments. Required by the Redis
adapter so loopback filtering doesn't drop a pod's own broadcasts as
if they came from a peer.
Each cached HRRR valid_time is ~32 MiB compressed × cap of 8 = up to
256 MiB per pod. With 4 hot pods + 1 backfill that's ~1.25 GiB of
cluster RAM holding identical content. Replacing the per-pod ETS
replicas with a single shared Valkey copy collapses that to ~256 MiB
total (off-pod) and makes the BEAM heap on each replica meaningfully
smaller — directly addresses the headroom the 2026-05-03 OOM cascade
ate into.
* New `Microwaveprop.Valkey`: single Redix connection (named conn,
sync_connect: false, exit_on_disconnection: false). `start_link/0`
returns `:ignore` when VALKEY_URL is unset so dev/test boot without
Valkey. Helpers wrap GET/MGET/SET-with-TTL/ZADD/ZREVRANGE/SCAN with
`:erlang.term_to_binary` round-tripping under the safe atom flag.
* `Microwaveprop.Weather.GridCache` now branches on
`Valkey.configured?/0`. Public API is identical so callers (Weather,
GridCachePruneWorker, MapLive) don't change. ETS path is preserved
unchanged for dev/test.
* Storage layout when on Valkey:
prop:wg:vts ZSET of valid_time iso8601 strings
prop:wg:<vt>:<lat>:<lon> one chunk per 5°×5° spatial bucket
fetch_bounds MGETs only the chunks intersecting the viewport, so a
regional view is 1-4 round-trips instead of pulling the full grid.
* Per-key TTL = 8 h (matches the ETS valid_time cap). Eviction is
automatic; prune_keep_latest is a no-op on Valkey, prune_older_than
uses ZRANGEBYSCORE + DEL.
* All Valkey errors fall back to disk (ScalarFile.read_bounds), logged
as warnings — page works through Valkey outages, no user-visible
break beyond a one-shot latency bump.
ScoreCache, Microwaveprop.Cache, NexradCache, telemetry/Postgres
protocol/libcluster registry tables remain ETS — they're sub-ms
hot-path lookups where a network hop would be felt.
VALKEY_URL must be present in prop-secrets for production rollout
(applied out-of-band, secret.yaml is git-ignored).
Found while auditing pod memory: :weather_grid_cache is by far the
largest ETS table (~32 MiB compressed per entry, vs ~2 MiB/entry for
ScoreCache) because each cell is a 22-field weather row instead of a
scalar score. Cap of 24 meant the cache could hoard up to ~768 MiB on
a single pod, and hot pods sit at 3-5 GiB against a 6 GiB limit — most
of that headroom belongs to GridCache.
The cap was sized for 'analysis + 18 forecasts + stragglers' but the
real access pattern is current-hour + a short forward scrub. Beyond
that, fall back to ScalarFile.read_bounds (sub-100 ms per file) instead
of permanently parking ~500 MiB on data nobody looks at.
Worst-case GridCache spend drops 768 MiB → 256 MiB. Steady-state win
is smaller (most pods don't fill all 24 slots), but bounds the tail.
The :terrain queue at 3 slots/pod was sized for the lightweight per-QSO
TerrainProfileWorker, but RoverPathProfileWorker now runs the full
PathCompute pipeline (terrain + 9 HRRR profiles + sounding + ionosphere
+ scoring + loss + forecast) — same heap profile as :propagation, which
sits at 1 slot precisely to avoid OOM.
A backfill_paths flood today queued ~200 rover-path jobs. With 4 hot
pods × 3 slots, the cluster stacked 15 concurrent path-computes and
two pods OOM-killed in a loop, taking the libcluster ring with them
(visible as the "unable to connect" warnings on the surviving pods).
New :rover_path queue at 1 slot/pod caps cluster-wide concurrency at
(hot_replicas + 1 backfill), drains the existing 200-job backlog
inside ~30 min, and keeps :terrain free for the cheap workers it was
sized for.
Existing Path rows persisted before the PathCompute extraction shipped
have no "term" key in their result map, so the live /path page hits
the :stale_term branch and falls back to live recompute when the
operator clicks a row. The hourly backfill cron now flips those
:complete rows back to :pending so the path-profile worker rerolls
them with the new full-output term, after which row clicks render
from cache instantly. The flip uses Postgres's jsonb_exists() so it's
a single targeted UPDATE instead of pulling rows into Elixir.
Two stacked wins for the form-submit hot path:
1) create_mission now hands path-matrix population to the same async
reconcile worker update_mission uses, so the request returns
right after the Mission row insert. For a mission with hundreds
of (rover x station x band) tuples this used to block the LiveView
for many seconds.
2) Inside reconcile/enqueue_paths_for, the per-pair Multi.insert
loop is replaced with one Repo.insert_all for the Path rows and
one Oban.insert_all for the worker jobs. N round-trips collapse
to two regardless of pair count.
Three connected changes:
1) Extract PathLive.compute_path/4 + every helper it owned (resolve,
profile grid lookup, sounding/ionosphere readouts, scoring, loss /
power budgets) into Microwaveprop.Propagation.PathCompute. PathLive
now delegates; ~410 lines of dead helpers deleted from PathLive.
2) RoverPathProfileWorker calls PathCompute.compute/4 with the
mission's heights and PathLive's default station params (10 W TX,
30 dBi gains). Stores the full atom-keyed compute output as a
Base64-encoded :erlang.term_to_binary/1 blob alongside the flat
summary fields the rover-planning show table reads. The blob
roundtrips structs and DateTime exactly (Jason.encode would lose
them).
3) PathLive accepts ?rover_path_id=UUID. When set, loads the cached
Path, decodes the term (binary_to_term :safe), assigns it as
@result, and renders normally — no compute_path call. The
rover-planning show table now links rows directly to
/path?rover_path_id=UUID, so a click opens the full Path
Calculator UI from cached data without re-running terrain / HRRR /
sounding lookups.
Bonus prod fixes folded in:
- PathShow's elevation chart attribute (data-* → data-profile JSON)
was crashing the JS hook with 'unexpected character at line 1'.
- Station.changeset now wipes previously-resolved callsign/grid/
lat/lon when the user types into :input — typing 'AA5' early
resolved to a wrong location and locked it; subsequent keystrokes
never re-resolved.
- phx-debounce=600 on the station input so QRZ doesn't get hit on
every keystroke.
Worker now stores the full per-point analysis (beam height,
Fresnel-zone radius, clearance) on Path.result.path_points, and
PathShow renders a real cached page: terrain elevation chart fed by
the same ElevationProfile JS hook /path uses (feet on Y, miles on X,
LOS + Fresnel overlay), distance / bearing / clearance / antenna
heights all in feet, baseline link-budget breakdown, and the cached
midpoint propagation score with computed_at timestamp.
A Recompute live button hands off to /path for the time-varying bits
(current HRRR conditions, scoring breakdown, sounding, ionosphere)
that aren't worth caching.
Two fixes:
1) Form save / rover-site add/remove no longer runs the path-matrix
reconcile inline. update_mission and the show LiveView now
enqueue a single RoverMissionReconcileWorker job and return —
for missions with many rover sites x bands the inline path was
firing hundreds of Multi inserts + Oban.inserts on the request.
2) RoverPathProfileWorker.load_path/1 now strictly matches a
4-tuple (mission_id, rover_location_id, station_id, band_mhz)
with band_mhz as an integer, plus updates the Oban unique key
set to include band_mhz so multi-band missions enqueue per-band
jobs (not just one). Legacy/malformed args are logged and
dropped instead of crashing the queue with MultipleResultsError.
Mission now carries bands_mhz ({:array, :integer}) — operator picks
one or more bands as multi-checkboxes. enqueue_paths_for builds the
cross product (rover x station x band) and persists each tuple as its
own Path row keyed by (mission_id, rover_location_id, station_id,
band_mhz). The path-profile worker reads band_mhz from the path
itself (legacy single-band jobs without band_mhz in args still resolve
to their unique row).
replace_mission_paths/1 is now a thin alias for reconcile_mission_paths/1
which diffs desired vs actual: stale tuples (old band that the user
unchecked, station they removed, rover-site they deleted) get dropped,
new tuples become :pending and enqueue, surviving :complete rows are
left in place — no more wholesale destruction of already-computed
paths on every edit.
The show table gains a Band column, and band_label() in the mission
summary becomes bands_label() (joins the list with commas).
The path-profile worker now also computes the baseline link-loss
budget — free-space path loss, oxygen absorption, baseline H2O loss
(at 7.5 g/m^3 default humidity), plus the already-cached terrain
diffraction — and stores all four numbers in Path.result. The show
page renders the per-row total in a new Loss column. /path still
recomputes against live HRRR weather; this is the offline-readable
cached snapshot.
Adds RoverPlanning.backfill_paths/0 that walks every mission and
re-enqueues RoverPathProfileWorker jobs for any (rover × station) pair
whose Path is missing or not :complete. The path-profile worker
short-circuits on :complete, so this is idempotent.
Wires a new RoverMissionBackfillWorker into both dev (config.exs) and
prod (runtime.exs) crons at HH:30 — heals missions stuck in :pending
or :failed from transient elevation-API errors or interrupted
create_mission enqueues without operator action.
The path-profile worker now also looks up the propagation grid score for
the path midpoint at the mission's band before flipping status to
:complete, so the row never appears 'done' until both terrain and
propagation prediction have run. The show table renders that score as a
red/yellow/green badge in a new Propagation column.
Each rover-location group heading is now a link to /rover-locations/:id
so the user can jump straight to the spot's detail page from the
mission view.
A new "Rover sites" section between Stationary stations and Path
profiles lists the candidate rover locations the mission scores
against, with:
- An add form that takes the same flexible input as station inputs
(callsign, Maidenhead grid, or `lat,lon`); on submit it creates a
global :good rover-location and re-runs the path matrix.
- A per-row trash button visible to the location's owner / admins
that deletes the location and re-runs the matrix.
`RoverPlanning.candidate_rover_locations/1` is now public so the show
page can list exactly what the worker enqueues against. Add/remove
both call `replace_mission_paths/1` so the matrix stays consistent
with the rover-site set.
Anonymous visitors see a sign-in prompt instead of the form.
Renames the rover-location status enum and every label that referenced
it. Existing rows are migrated in place. Also touches the consumers:
- Rover.Location enum + default
- RoverLocationsLive index, status filter, form, show, badges
- Rover.Compute and rover_live (only_ideal_locations → only_good_locations)
- RoverPlanning context candidate filter (status == :good)
- RoverPlanning.Show / Form copy
- All rover-related tests
Adds a new globally-scoped, owner-mutable rover-mission tracker:
- /rover-planning paginated table (LiveTable) with View/Edit/Delete
- /rover-planning/new + /:id/edit form: name, band, antenna heights,
notes, "only check against known good locations" toggle (default on),
and a dynamic list of stationary stations entered as callsigns,
Maidenhead grids, or lat,lon pairs (Station changeset geocodes via
LocationResolver, lat/lon stays editable after add)
- /rover-planning/:id show page renders the station list, scope, and a
matrix of computed path profiles (distance, min clearance, diffraction,
verdict) populated as the worker completes each pairing
After save, RoverPlanning enqueues one RoverPathProfileWorker job per
(rover-location × station) pairing. The worker mirrors PathLive's
synchronous compute (ElevationClient + TerrainAnalysis at the mission's
band + heights) and stores the result on the matching path row.
PubSub broadcast on completion lets the show page live-refresh.
Admins can edit/delete any mission; owners can edit their own.
Rover.update_location/3 and delete_location/2 now resolve any record
when the actor is_admin. The index + show LiveViews surface Edit/Delete
on every row for admin users, matching the existing owner UX.
- Extract PathParser.valid_callsign?/1 and reuse from the calibration
Mix task instead of duplicating the regex.
- Document raise behavior on Aprs.recent_packets_with_paths/1 and
Aprs.station_positions/1 (already noted in moduledoc; @doc reinforces
for callers reading the function head).
- Note in the BandConfig guard comment why it sits before the Oban pause
(so a misconfig exits without leaving queues paused).
- Reword the band-mismatched-negatives note as a v1 limitation (no TODO
tag — codebase convention is zero TODOs).
- decode_packet_row no longer guards Decimal.to_float / DateTime.from_naive
with passthrough clauses — Postgrex returns Decimal for numeric and
NaiveDateTime for timestamp without time zone, so the WHERE clause
filters out the only path that could have hit the alternate branches.
- recent_packets_with_paths shifts :since to UTC before to_naive so a
caller passing a non-UTC %DateTime{} doesn't silently match the wrong
window.
- moduledoc now documents that the module raises on connection/SQL
errors and is unsafe in async contexts without a catch.
Q-construct now matches the explicit APRS-IS set (qAC/qAX/qAU/qAo/qAO/
qAS/qAr/qAR/qAZ/qAI) instead of any qA[A-Za-z]. WIDE/TRACE n-N now allows
n and N up to 7 per the APRS New-N paradigm (was [12]).
Three spec-compliance fixes to PathParser:
1. Q-construct regex now accepts lowercase letters (qAo, qAr, etc.).
The moduledoc lists qAo as a Q-construct to skip but the previous
~r/^qA[A-Z]$/ rejected it, leaving it to be misclassified.
2. self_loop?/2 now compares callsigns only; the spurious
src_pos == dst_pos clause would falsely suppress two distinct
stations that happen to share coarse coordinates (e.g. same building).
3. Strengthen the "lookup returned nil → drop hop, do NOT advance source"
test: extend the path to three tokens so the assertion actually proves
the third hop chains from the previous good source rather than from
the dropped one. The two-token version passed regardless of behavior.
Self-loop fixture updated to trigger via callsign equality (the only
remaining mechanism) instead of position equality.
Pure-Elixir parser that walks an APRS digipeater path left-to-right and
emits verified hops for callsigns carrying the used flag (`*`). Filters
Q-constructs, TCPIP/TCPXX, and routing aliases (WIDE/TRACE/RELAY/etc.);
handles anonymous-digi cases, missing station-position lookups, and
self-loops without advancing source past unattributable hops.
Used by upcoming 144 MHz calibration to attribute aprs.me receives to
real digipeater geometry without persisting any APRS data on our side.
Wire up Microwaveprop.AprsRepo as an optional secondary Ecto repo that
points at the sister aprs.me Postgres for read-only access. This is the
foundation for upcoming 144 MHz calibration work that uses verified APRS
RF hops as ground truth.
* lib/microwaveprop/aprs_repo.ex declares a read_only repo; no schemas,
no migrations, no writes — Microwaveprop never persists APRS data.
* Application supervisor starts the repo only when configured (dev/test
via config blocks, prod via APRS_DATABASE_URL). A missing config logs
a warning and skips startup so the main app still boots.
* config/dev.exs + config/test.exs add local aprsme_dev / aprsme_test
endpoints. Test pool uses Ecto SQL Sandbox; tests that don't touch
APRS data don't need to check the connection out.
* config/runtime.exs guards the prod block on APRS_DATABASE_URL —
unset env logs a warning and leaves the repo unconfigured.
The /weather dual-source timeline picks times from HRRR's hourly
cadence, but HRDPS publishes 4×/day with multi-hour latency, so the
requested time almost never has an exact HRDPS dir on disk. The
controller was returning empty rows for HRDPS, leaving the Canadian
half of the map blank.
Snap to the closest HRDPS valid_time within a 6h window. /weather-ca
keeps exact-time semantics in practice because its timeline is built
from HRDPS-only listings, so the nearest is always the same time.
Recalibrator and FrontalAnalysis use Nx for tensor math, but Nx is
declared `only: [:dev, :test]` in mix.exs because neither module has a
prod entry point — Recalibrator is invoked by hand from IEx and
FrontalAnalysis is research scaffolding. The prod compile would emit
"Nx.to_number/1 is undefined" warnings on every Nx call. Add a
`@compile {:no_warn_undefined, Nx}` directive to both modules so the
prod build is clean without dragging Nx/Axon/EXLA (~hundreds of MB) into
the runtime image.
New LiveView at /weather-ca duplicates the /weather UI but defaults the
viewport to central Canada (55N, -100W, z=4) and renders the map div
with data-source="hrdps". The JS hook reads that attribute and appends
&source=hrdps to every cell fetch; the WeatherTileController routes
those reads through Weather.weather_grid_hrdps_at/2, which only reads
the <vt>.hrdps scalar dir and skips HRRR merging entirely. Timeline is
sourced from ScalarFile.list_valid_times_hrdps/0 so HRDPS-only hours
show up even when no HRRR file exists for the same valid_time.
Also drops "— Unified" from the algo.md H1.
The HRDPS pipeline as written was unusable in production: each chain
step took ~5 hours (not 30-90s as the dev-bench claim) because every
`wgrib2 -lon` batch redundantly JPEG2000-decodes ~150 records. With 57
batches × 18 forecast hours, a single cycle would take days to drain.
The /weather Canadian view only needs the current hour, so the cheapest
fix is to stop seeding the rest of the chain entirely:
* Rust grid: HRDPS_STEP=0.5 (was 0.125) drops point count 16× to ~3.5k.
Coarser cells than HRRR, but /weather gets visible Canadian coverage
inside one chain step (~20 min) instead of never.
* GridTaskEnqueuer.seed_current_hour/2 inserts exactly one forecast row
whose valid_time is closest to `now`. fh is clamped to 1..18 so the
Rust F00Reserved branch never fires.
* HrdpsGridWorker switches to seed_current_hour and drops the
6-hour-cycle cron in favor of HH:35 hourly fires. Reseeds are
idempotent on the 4-column unique key.
Also fixes a pre-existing credo "nested too deep" in
ScalarFile.read_point by extracting lookup_in_chunk + hit_or_false
helpers — happened to be on the read path that surfaces this work, so
folding it into the same commit.
The /weather map now shows HRRR-derived data wherever HRRR has
coverage and falls back to HRDPS-derived data for Canadian cells
outside HRRR's bbox. North of the 49°N HRRR-coverage line the map
stays populated instead of going blank.
Rust:
- weather_scalar_file::dir_for_hrdps + write_atomic_hrdps land HRDPS
scalar chunks at `<vt>.hrdps/` so HRRR's `<vt>/` write isn't
clobbered. The HRRR writer wipes its dir before each write — both
sources need their own dir to coexist.
- pipeline::run_chain_step_hrdps now derives a ScalarRow for each
scored cell (same derive_row that HRRR uses; same wire format) and
writes them via write_atomic_hrdps alongside the .hrdps.prop score
files.
Elixir:
- ScalarFile.dir_for_hrdps + read_bounds + read_point + exists? all
walk both `<vt>/` (HRRR) and `<vt>.hrdps/` (HRDPS). read_bounds
concatenates with HRRR-precedence on overlap (defensive — Rust
pipeline doesn't write HRRR-overlap cells, but cheap to enforce).
- Weather.weather_grid_at unchanged: it routes through ScalarFile,
which now transparently returns the union. GridCache stays
per-valid_time so a single cache entry holds both regions.
Tests cover four scenarios: HRRR-only, HRDPS-only, both-present
union, and HRRR-precedence on the same cell.
Stages 6, 7, 10 of the HRDPS plan, plus the Elixir read-side merge
that makes Rust's `.hrdps.prop` writes show up on /map.
Read side:
- `ScoresFile.path_for_hrdps/2` + `read_hrdps/2` — companions to the
HRRR equivalents. Same decode path; different filename.
- `ScoresFile.read_bounds/3` now reads HRRR + HRDPS files and
concatenates the cells. Files are disjoint by construction
(Grid.hrdps_only_points excludes CONUS) so no de-dup needed.
- `ScoresFile.read_point/4` checks `.prop`, then `.hrdps.prop`, then
legacy `.ntms` — first hit wins. Cells outside their owning region
read as no-data in the other file's grid, so the order doesn't
matter for correctness.
- `Propagation.warm_cache_and_broadcast/2` reads both files and
merges before broadcasting to the cluster ScoreCache. A single
missing file (HRDPS pre-cycle, HRRR briefly absent) is now OK —
the cache warms with whichever side is available.
Activation:
- runtime.exs cron: HRDPS seed worker fires at HH:35 of each cycle
hour (05:35Z, 11:35Z, 17:35Z, 23:35Z) — 5h after cycle, 30 min
staggered from HRRR's */15 schedule.
- HrdpsGridWorker moduledoc updated: no longer dormant; Rust
prop-grid-rs handles source='hrdps' rows.
- map_live North America bbox extended to (24.5, 60.0, -141.0, -52.0)
for the out-of-coverage banner check; visitors anywhere in CONUS
or the Canadian extent now stay banner-free.
Cleanup:
- mix hrdps.probe deleted — its risks are retired and the production
HrdpsClient covers the same surface.
What's still in motion (not blocking the Canadian /map):
- Per-valid_time profile + scalar artifacts for HRDPS (would let
/weather show Canadian temperature/refractivity layers). Rust
pipeline currently skips those so the HRRR-written companion files
aren't clobbered — separate stage when /weather UX is decided.
- FreshnessMonitor HRDPS staleness tracking — HRDPS has its own cron
and a different cadence than HRRR; HRRR-stale logic doesn't apply
cleanly. Defer until prod tells us what alarms we actually want.
Stage 3 (Elixir foundation) of the HRDPS plan. Lands the scaffolding
that lets the Rust prop-grid-rs HRDPS branch be wired up cleanly when
it ships, without further DB or Elixir-side changes.
What ships:
- `grid_tasks.source` column (default 'hrrr', backfill is a no-op).
Replaces the (run_time, forecast_hour, kind) unique index with a
4-column key (..., source) so HRRR and HRDPS analysis/forecast rows
for the same cycle coexist.
- `Grid.hrdps_only_points/0` — Canadian cells inside HRDPS bbox
(49-60°N, -141 to -52°W) but outside HRRR's CONUS bbox. Disjoint
from `conus_points/0` by construction so the two grids never
double-write the same (lat, lon).
- `Grid.point_source/1` — `:hrrr | :hrdps | :neither` classification
for a (lat, lon) pair. Routes per-QSO enrichment to the right NWP
source.
- `GridTaskEnqueuer.seed_with_analysis/2` and `seed/2` accept a
`:source` option; defaults to "hrrr" so existing callers are
unchanged.
- `HrdpsGridWorker` cycle seeder mirroring PropagationGridWorker's
shape but for HRDPS's 4×/day cadence. `:hrdps` queue added to base
Oban config.
What's deferred:
- The Rust `prop-grid-rs` worker doesn't yet branch on
`task.source`. Until that ships, `HrdpsGridWorker` stays out of
runtime.exs cron — it would just produce rows that the Rust worker
mishandles as HRRR. Module doc explicitly notes the dormant state.
- Map overlay extension and FreshnessMonitor HRDPS staleness deferred
to follow-up sessions once real data flows.
Activation path documented in the updated plan file.
Coverage cap recorded as Stage 8 decision: 60°N for v1 (SRTM stops
there). Arctic CDEM deferred to a follow-up plan.
Stage 2 of the HRDPS plan. Mirror of hrrr_profiles for HRDPS-derived
rows so HRDPS can be dropped/reprocessed without disturbing the 4500+
HRRR profiles already in prod. Same column set, same indexes,
PARTITION BY RANGE (valid_time) with quarterly partitions starting at
the deploy quarter.
Schema follows the @primary_key {:id, :binary_id, autogenerate: true}
convention and the same set of derived fields HrrrProfile carries
(surface_refractivity, min_refractivity_gradient, ducting_detected,
duct_characteristics, etc.) so the rest of the pipeline can consume
HRDPS-derived rows interchangeably with HRRR-derived rows.
Unique constraint on (lat, lon, valid_time) is enforced at the DB
level via the partitioned index; schema-level unique_constraint/2
matching against changeset errors doesn't apply because the
partition's index name varies per partition. Production upserts go
through `on_conflict: {:replace_all_except, ...}` per project
convention.
Stage 1 of the HRDPS plan. Mirrors HrrrClient.fetch_grid/3's signature
so PropagationGridWorker can call either polymorphically.
Architectural shape per the probe wedge findings:
- ECCC's date-prefixed datamart URL structure
(dd.weather.gc.ca/{YYYYMMDD}/WXO-DD/model_hrdps/...)
- One-variable-per-file model — fetch ~14 small GRIB2s concurrently,
byte-concatenate into a single multi-record binary, hand to
Wgrib2.extract_points_from_file. wgrib2 reads multi-record files
natively so no -merge step is needed.
- wgrib2's -lon flag handles the rotated-lat/lon reprojection
internally — no manual rotation math.
Variable catalog covers the surface set the propagation scorer reads
plus the 1000-700 mb pressure-level set used by the refractivity
gradient. PWAT is unavailable in HRDPS f000 inventory; the scorer's
PWAT factor will fall back to its default for HRDPS-derived points.
DEPR (T-Td depression) is treated as an alternate primary; when DPT
is absent build_profile_from_extracted derives dewpoint as
temp - depression so downstream consumers see the same shape as HRRR.
Plan task 1.6 (CYYR projection sanity check vs UWYO sounding) is
covered by the live probe at lib/mix/tasks/hrdps_probe.ex which
already validated wgrib2's rotated-grid handling against five
Canadian cities; the production implementation routes through the
same Wgrib2.extract_points_from_file path that the probe used.
Bucket each cache entry into 5°×5° spatial chunks matching the layout
used by Weather.GridCache. fetch_bounds/3 now walks only the chunks
intersecting the viewport instead of the full 92k-cell CONUS map; a
typical zoomed viewport hits ~1500 cells per chunk × the chunks the
viewport overlaps instead of scanning all 92k.
Public API and observable behavior unchanged. Add cross-chunk-boundary
regression tests so future refactors can't silently flatten the layout.
Apply the same overlay refactor we did for /weather to the propagation
map. Score data used to flow as JSON via LiveView push_event
(update_scores + the 18-hour preload_forecast batch); switching bands
or scrubbing time meant a synchronous server roundtrip + ~MB JSON push
per change.
Now:
1. New /scores/cells endpoint returns a "PSCR" binary cell-pack —
columnar f32 lats/lons + u8 scores, ~3-4× smaller than JSON and
parsed via DataView + typed-array views with no JSON.parse.
2. JS hook owns score fetching. mount/moveend/band/time changes
trigger HTTP fetches against /scores/cells. After the initial
paint, every other forecast hour for the current band is fetched
in parallel in the background — timeline scrubs after the prefetch
finishes are pure cache hits with no network at all.
3. LiveView no longer pushes scores. select_band emits update_band_info
with band_mhz so the client knows what to refetch; select_time and
set_selected_time only update server-side state (URL, point detail
bbox). map_bounds, propagation_updated, and advance_now_cursor all
skip the score push — pack_scores, schedule_bounds_update,
schedule_preload_forecast, the :flush_bounds and :preload_forecast
handle_info clauses, and the start_async :initial_scores lifecycle
(incl. retry_initial_scores event) are all deleted.
Net result: instant band toggles after the first paint, no LiveView
WebSocket payloads larger than the timeline metadata, and ~370 fewer
lines in MapLive.
The client moved to the binary cell-pack and canvas overlay last
commit, so the per-tile SVG endpoint, the TileRenderer module, and the
JSON shape of /weather/cells are all dead code. Remove them.
/weather/cells now always returns the binary WCEL pack (the ?format=bin
toggle is gone — there's only one format). The data-weather-tile-url
attribute and the weather_tile_url assign are dropped from the LiveView.
Also fix a latent bug surfaced once tests covered the default-layers
path: nil row values were writing the atom :nan into the binary segment
at runtime. Replaced with the explicit f32 quiet-NaN bit pattern so
Number.isNaN flags them correctly on the JS side.
Each /weather/tiles request was running the full ScalarFile decode path
(File.read → gunzip → Msgpax.unpack → normalize) for any non-analysis
valid_time, because GridCache deliberately skipped forecast hours to keep
memory low. With 21 simultaneous viewport tiles all hitting the same few
chunk files, each tile took 1.5-3.2s of contended IO + CPU.
On a GridCache miss when a ScalarFile exists, hydrate GridCache from the
file once (deduped via the existing claim_fill primitive) and serve all
subsequent reads from ETS. Bound memory with prune_keep_latest/1 — keep
the 24 most-recently-touched valid_times, which covers a full HRRR run
(analysis + 18 forecasts) plus a few stragglers from the previous run.
In the steady state, the first viewport read for a given forecast hour
takes one full ScalarFile decode (~50-100 ms for 92k cells); every
subsequent tile, point lookup, and viewport read for the same hour
serves from ETS in <1 ms.
Move ScalarFile from a dual-writer dual-format setup (ETF on the Elixir
side, materialized via NotifyListener; nothing on the Rust side) to a
single chunked-MessagePack format that both languages produce and
consume.
Elixir side:
- ScalarFile now writes gzipped MessagePack `.mp.gz` per chunk via
Msgpax. Reader decodes with Msgpax + zlib. Atom keys are stripped to
strings on encode and re-atomized via a whitelist on read so callers
see the same shape regardless of which side wrote the bytes.
- DateTime values round-trip through ISO8601 strings.
- New cross-language wire-format test that decodes a hand-rolled
payload mirroring exactly what `rmp_serde::to_vec_named` emits, so a
format regression on either side fails the suite.
Rust side:
- New `prop_grid_rs::weather_scalar_file` module:
- 1:1 port of `Microwaveprop.Weather.WeatherLayers.derive/1` plus
the surface fields `Weather.build_grid_cache_row/4` adds on top
(lapse rate, mid lapse rate, inversion strength + base, 850/700 mb
interpolation, surface RH, surface refractivity, ducting flag).
- `write_atomic` chunks rows into 5°×5° buckets, writes one
`<lat_band>_<lon_band>.mp.gz` per chunk via tmp + rename, drops
pre-existing chunks first to keep writes canonical.
- 5 unit tests covering surface-only rows, upper-air derivation,
write/read round-trip, chunk band boundaries, and the
out-of-physical-range surface-temp filter.
- Pipeline integration: both `run_chain_step` (f01..f18) and the f00
analysis path now build `Vec<ScalarRow>` in the same merged-cell
pass that builds `CellEntry` and `Conditions`, then fire a parallel
`spawn_blocking` write that overlaps with scoring. Awaiting the
scalar future after the score-band writes keeps the failure-surface
symmetric with the existing profile_future.
Result: every new forecast hour lands with the scalar artifact already
on disk, so /weather reads never fall through to ProfilesFile decode +
per-cell SoundingParams + WeatherLayers derivation. The Elixir-side
`NotifyListener.handle_propagation_ready/1` materialization remains as
an idempotent safety net for hours that pre-date this commit.
Mirror ScalarFile's 5°×5° chunk layout inside GridCache. Each ETS entry
is now `%{{lat_band, lon_band} => %{{lat, lon} => row}}` instead of a
flat `%{{lat,lon} => row}`. fetch_bounds/2 walks only the chunks that
intersect the requested viewport, so latest-hour pans become
proportional to visible chunks (~2-4) rather than O(92k cells).
fetch/1, fetch_point/3, put/2 and broadcast_put/2 keep their
pre-existing contracts.
Also closes the open items in docs/weather.md: Phase 3 is now done end
to end. The only remaining future-work item is moving WeatherLayers.derive
itself into the Rust pipeline so the BEAM doesn't pay the one-time
materialization pass per new valid_time — small runtime savings vs.
significant port effort, so left as documented future work.
Move the /weather render path off raw HRRR profile decoding + per-cell
SoundingParams + WeatherLayers derivation. The dominant cost was
re-deriving the same scalar fields on every viewport pan and timeline
scrub.
Server side:
- New Microwaveprop.Weather.ScalarFile persists pre-derived scalar
rows on NFS, bucketed into 5°×5° chunk files under
<base>/weather_scalars/<iso>/<lat_band>_<lon_band>.etf.gz. Viewport
reads only decode the chunks that overlap the requested bounds;
point-detail clicks read exactly one chunk.
- weather_grid_at/2 and weather_point_detail/3 prefer ScalarFile;
ProfilesFile is now the cold-start fallback only.
- warm_grid_cache_and_broadcast/1 and warm_grid_cache_from_latest_profile/0
also persist the scalar artifact, and the cold-derive path kicks
off a per-valid_time-locked async materialization so the next
reader gets the cheap path.
- NotifyListener.handle_propagation_ready/1 fires
Weather.materialize_scalar_file/1 in a detached Task on the Rust
pipeline's NOTIFY propagation_ready, so steady-state forecast hours
arrive with their scalar artifact already on disk.
- available_weather_valid_times/0 unions ScalarFile + ProfilesFile so
the timeline survives an aggressive retention sweep on either side.
Browser side (finding #8):
- Mount the legend Leaflet control once and patch its inner content
on layer changes instead of remove + re-add on every renderLayer.
- renderTimeline now keys off the rendered button set: selection-only
updates take an applyTimelineSelection patch path that restyles
existing buttons in place. Full innerHTML rebuild only fires when
timelineData itself changes.
Also fixes a credo nesting-depth flag in tile_renderer.ex by extracting
interpolate_pair/2 from the inner anonymous function.
Refit on the current contact corpus via mix recalibrate_scorer (val loss
0.140 vs 0.146 train, initial 0.434). Mostly small adjustments — the
notable shifts are time_of_day −0.0116 (now the smallest factor) and
rain +0.0069 (continues to be the largest single weight).
Sums to 1.0 (within rounding); per-band overrides not affected.
- Unify haversine_km / deg_to_rad: Geo is the canonical home (atan2
form for antipodal stability); Radio, common_volume, rain_scatter,
ionosphere, terrain/srtm, terrain/elevation_client now delegate.
- Drop safe_min/safe_max wrappers in favor of Enum.min/max defaults.
- Move parse_float / parse_int to MicrowavepropWeb.LiveHelpers; eme,
path, and rover LiveViews import from there.
- Extract MicrowavepropWeb.LocationResolver from the eme/path
resolve_source / resolve_location duplicate. Standardize the kind
key and "Invalid grid square" error message.
- Convert Scorer cond chains (humidity, td_depression, sky, wind,
rain, pwat, pressure) to multi-clause guarded defs, matching the
existing classify_time_period style.
- extract_profile_fields uses a named map accumulator instead of a
6-tuple; build_path_conditions guards on %{temps: []} / %{dewpoints: []}.
- Collapse scores_file clamp/3 to max |> min.
- humidity_effect_label lives in BandConfig; map_live and path_live
both use it.
Sweep of remaining swallowed-error sites flagged by an audit pass:
- weather/gefs_client.ex: byte-range download task exits log url + reason
before being surfaced as {:error, _} (matches the HRRR fix).
- live/path_live.ex: terrain-profile load failure, native-duct lookup
failure, and ProfilesFile.read failure now log warnings with path
endpoints / midpoint / valid_time. Previously rendered a degraded
result silently.
- live/rover_live.ex: station-mutation {:error, _} now logs id + reason
instead of dropping into {:noreply, socket}.
- application.ex: the boot-time backfill_missing_home_qth Task.start is
now wrapped in try/rescue. A DB issue at boot would have crashed the
task silently (no Logger handler attached pre-Logger crash format).
GefsFetch silently dropped crashed score-computation tasks via
`{:exit, _reason} -> []`, hiding partial-grid outages from k8s logs.
HRRR range downloads mapped exits to `{:error, reason}` without logging
the URL, so the failing range was lost by the time the caller surfaced
the error.
All three sites now Logger.error with enough context (fh, valid_time,
url) to debug from production logs.