Commit graph

1028 commits

Author SHA1 Message Date
8a969e315c
refactor: normalize pos1/pos2 JSONB key to 'lon' everywhere
57,186 prod contacts stored pos1/pos2 with 'lng'; 1,133 used 'lon'.
Every Elixir caller carried a `pos["lon"] || pos["lng"]` fallback
— which just caused a SQL widget to silently miscount 98% of contacts
(count_narr_done used `pos1->>'lon'` directly, no fallback, so every
lng-keyed row returned NULL and failed the coverage check).

- Migration rewrites every pos1/pos2 JSONB in place, renaming 'lng' to
  'lon' and dropping 'lng'.
- Removes all 20+ `|| pos["lng"]` fallbacks across lib/, workers,
  scorer, weather, radio.ex, contact show view, and recalibrator.
- lib_ml/propagation_analyze.ex SQL now reads pos1->>'lon' directly
  (was reading 'lng' only, which would have broken after migration).
- priv/repo/import_contacts.exs one-time seed script now emits 'lon'
  with string keys, matching production shape.
- Test fixtures in 4 test files normalized to 'lon'.
- Two lng-characterization tests deleted — nonsensical post-normalize.
- Updated notebook + old import_weather script to match.
- JS hook contact_map_hook.ts TypeScript type narrowed to 'lon'.
2026-04-17 09:10:32 -05:00
FluxCD
f252d4983b chore: update prop image to git.mcintire.me/graham/prop:main-1776433218-6f452e7 [skip ci] 2026-04-17 13:44:17 +00:00
6f452e7139
feat(map): show CONUS-only banner for visitors outside continental US
Visitor geolocation already flows via the Cloudflare CF-IPLatitude /
CF-IPLongitude headers into the session; MapLive centers the map on
the visitor's location (or DFW as fallback). Add a banner that appears
above the map when those coords are outside the CONUS bbox
(lat 24.5-49.5, lon -125 to -66.5), letting overseas visitors know
propagation data won't be available at their location yet.

Map still centers on the visitor's actual location — this is a note,
not a re-center. Banner is fully absent from the DOM when not needed
(:if conditional).
2026-04-17 08:39:51 -05:00
FluxCD
e57835c955 chore: update prop image to git.mcintire.me/graham/prop:main-1776370515-8f0548b [skip ci] 2026-04-16 20:17:33 +00:00
8f0548b31e
fix(narr): handle cdo 2.5.1 mixed name/code output
parse_cdo_row/1 assumed cdo -outputtab,code,lev,value would always emit
numeric codes in the first column, but Debian's cdo 2.5.1 (the version
in the prod container) emits mixed output:

- 'codeNN' with a literal 'code' prefix for most GRIB1 parameters
- Short GRIB1 names for a few: 2tag2 (PRES sfc), saip (DPT 2m),
  tpag10 (HGT)

Result: every row fell through to :error, yielding a 'no parseable
values' error and 986 discarded NarrFetchWorker jobs.

Fix: normalize the code token by stripping 'code' prefix when present
and mapping known 2.5.1 short names. UGRD/VGRD intentionally omitted
(no discard evidence yet) — comment documents the add-when-they-surface
plan.

Also exposed parse_cdo_outputtab/1 as @doc false for testability.
2026-04-16 15:14:51 -05:00
fbbfe037bc
fix(status): align NARR done-count tolerance with fetcher snap
count_narr_done/0 checked narr_profiles.valid_time within qso_timestamp
± 30 min — but NARR analyses live at 3-hour marks (00Z/03Z/.../21Z) and
NarrClient.snap_to_analysis_hour/1 floors qso_timestamp to the previous
mark before fetching. A QSO at 18:50Z snaps to 18:00Z on the fetcher
side; the profile stored at 18:00Z fell outside the widget's [18:20,
19:20] window, so the status page showed 6/208 done when actual
coverage was 208/208.

Fix: use equality against the Postgres equivalent of the snap
(date_trunc('hour', ts) - make_interval(hours => hour % 3)).
Comment cites NarrClient.snap_to_analysis_hour/1 so future changes keep
them in sync.
2026-04-16 15:14:51 -05:00
FluxCD
432ae83743 chore: update prop image to git.mcintire.me/graham/prop:main-1776369564-63a9417 [skip ci] 2026-04-16 20:01:13 +00:00
63a9417287
fix(admin_task_worker): unknown feature now returns error instead of falling through
Backtest branch's unknown-feature guard was dead code: the if block's
{:error, _} return was discarded and execution fell through, running the
backtest with a just-created atom and writing a report for nonsense
features. Two compounding issues:

- function_exported?/3 returns false for modules not yet loaded in the
  BEAM, so even valid feature names triggered the warning path on a
  cold VM.
- String.to_atom/1 on caller-supplied strings created unbounded atoms.

Fix: Code.ensure_loaded?(Features) before the check, String.to_existing_atom/1
with rescue (Features exports __info__(:functions) so every valid function
atom already exists), and an explicit early return via a with pipeline so
the error tuple actually leaves the worker.
2026-04-16 14:58:55 -05:00
bba860f28d
fix(propagation): dedup PropagationGridWorker enqueues
FreshnessMonitor's comment claimed 'Oban unique constraint on
PropagationGridWorker prevents duplicates' — but the worker declared
no unique: option. During a long outage the monitor's 5-min stale-check
tick would pile up identical jobs (a 2h outage = 24 stacked jobs, each
doing the full f00-f18 HRRR chain).

Put the dedup on the worker (vs. the monitor's insert call) so anything
enqueuing it benefits — FreshnessMonitor, the hourly cron, or a manual
mix propagation_grid run.

unique: [period: 3600, states: [:available, :scheduled, :executing,
:retryable]] aligns with the hourly cron cadence. The seed job args
(%{}) collapse with themselves while chain-step jobs keep distinct
run_time/forecast_hour and remain enqueuable.
2026-04-16 14:58:55 -05:00
221d5516bc
fix(wgrib2): account for 8-byte Fortran record overhead between messages
build_messages_per_message/3 was slicing the wgrib2 -lola bin output as
a packed float32 stream, ignoring the 4-byte record-length header + 4-byte
trailer that Fortran unformatted records prepend/append around each
message's data block. First value of message 0 came back as ~2.24e-44
(little-endian float32 of the int-16 record length), and every subsequent
message was shifted by 8 bytes so values spilled into neighbouring cells.

Fix: propagate the correct offset math already used by parse_lola_binary/3
(record_overhead = 8, stride = bytes_per_message + 8, data_offset =
msg_idx * stride + 4).

Flipped the characterization test from metadata-only to per-cell physical
asserts (200 K < TMP, DPT < 340 K; DPT <= TMP) and a 0.01 K cross-check
against extract_grid/3 output on the same fixture + bbox. Latent bug —
no production callers of extract_grid_messages*.
2026-04-16 14:58:55 -05:00
8c3c097009
test: cover weather/grib2/wgrib2
16 characterization tests against real wgrib2 binary + checked-in HRRR
fixtures: extract_grid/3 (bbox + physical sanity on TMP/DPT + longitude
convention), extract_grid_from_file/3, extract_grid_from_file_mapped/4
reducer plumbing, extract_grid_messages/3 metadata, extract_points_from_file/3
nearest-neighbor snap, undefined-value sentinel filtering, empty-match and
error shapes. Tagged @describetag :slow so they run on 'mix test --include
slow' (project default excludes slow, matches existing HRRR/NARR tests).

Surfaces one real bug (not fixed): extract_grid_messages variants don't
account for the 8-byte Fortran record overhead between messages in wgrib2
-lola output, so per-message values after the first are shifted. The
correct offset math exists in parse_lola_binary/3. Tests characterize
current behavior (metadata-only); worth a follow-up fix + value-range
assertions.
2026-04-16 14:58:55 -05:00
cc480652aa
test: cover repo_listener
10 characterization tests for PostgreSQL NOTIFY → PubSub relay:
contact_status_changed channel → db:contact_status topic, oban_job_changed
→ db:oban_jobs, channel isolation, invalid JSON silent-swallow, catch-all
handling of unknown channels and info messages.
2026-04-16 14:58:55 -05:00
4f956e4606
test: cover radio/edit_notifier
13 characterization tests for deliver_edit_approved + deliver_edit_rejected:
to/from/reply-to, subject format, callsign greeting, APPROVED/REJECTED tags,
band formatting (MHz vs GHz), change rendering, admin-note present/nil/empty,
endpoint URL link, preload tolerance.
2026-04-16 14:58:55 -05:00
58e6fcff0d
test: cover propagation/freshness_monitor
10 characterization tests: fresh vs stale thresholds (>120 min),
boundary behavior, empty scores dir, supervised start, non-dedup of
repeated stale ticks, :ignore when disabled via config.
2026-04-16 14:58:54 -05:00
a6a670aee7
test: cover admin_task_worker dispatcher
18 characterization tests, one per task branch: backtest, backtest_all,
climatology (aggregation math + min_samples + is_grid_point + empty),
native_derive (filters + limit), recalibrate, scorer_diff, unknown-task
handling, and Oban worker metadata. Tests redirect priv/backtest_reports
writes to a tmp cwd to avoid clobbering checked-in baselines.
2026-04-16 14:58:54 -05:00
f73fe2717a
test: cover nexrad_worker
10 characterization tests: happy-path upsert across multiple POIs,
re-run conflict resolution preserving inserted_at, temporal window +
pos1-null filtering, 404/garbage-body error propagation, empty POI
short-circuit, snap-based dedup, legacy 'lng' key support, and the
year/month/day/hour/minute unique constraint.
2026-04-16 14:58:54 -05:00
ea0cbf5205
test: cover mrms_cache + mrms_fetch_worker, add Req.Test seam
Same fix as rtma_client: MrmsClient had a hardcoded req_options/0 —
switched to a Keyword.merge with Application.get_env so tests can stub
HTTP. 13 new characterization tests: cache put/get/clear/broadcast and
worker listing/download/up-to-date paths. Full GRIB2 parse happy path
deferred to the wgrib2 coverage task.
2026-04-16 14:58:54 -05:00
cb82160447
test: cover rtma_client + rtma_fetch_worker, add Req.Test seam
RtmaClient was the only weather client without a Req.Test plug seam —
added the standard req_options() helper matching NarrClient/HrrrClient/
etc. so its HTTP calls can be stubbed in test.

17 characterization tests: URL construction, idx fetch 404/503, hour
truncation, malformed idx, range-GET error propagation, worker's
existing-row short-circuit, client-error propagation. GRIB2 byte
parsing paths deferred to the wgrib2 coverage task.
2026-04-16 14:58:54 -05:00
FluxCD
923575c392 chore: update prop image to git.mcintire.me/graham/prop:main-1776366698-8db3179 [skip ci] 2026-04-16 19:12:03 +00:00
8db31794a1
test: cover qrz/client
13 characterization tests: login success/failure, session caching,
automatic re-login on session expiry with retry, lookup errors
(not found, HTTP 5xx, malformed XML), reset_session/0.
2026-04-16 14:10:50 -05:00
ef0c73cb4d
Fix NARR accounting on status page and stop pre-2014 hrrr_status churn
Status page NARR progress counted `hrrr_status = :unavailable` as the
candidate set, which swept in ~13.6K post-2014 HRRR retry rows and
showed "0 / 13,652". Candidate eligibility is qso_timestamp-based, not
hrrr_status-based — switch both the numerator and denominator to
`qso_timestamp < NarrClient.coverage_end()` so pre-2014 rows are
counted whether their hrrr_status is :queued, :unavailable, or
:complete.

ContactWeatherEnqueueWorker: short-circuit hrrr_points_for_contact/1
for pre-coverage timestamps and mark the contact :unavailable when no
HRRR jobs are built. Was producing a ping-pong where reconcile flipped
pre-2014 :queued → :unavailable and the same cron run flipped it back
to :queued from a rebuilt HRRR job that could never land data.
2026-04-16 14:09:45 -05:00
f693cb5b56
test: override stale distance_km when filling a missing pos
Makes the 'always recompute' intent explicit by removing the dead-code
short-circuit in add_distance_change — the candidate query already
requires at least one nil pos, so any previously stored distance is
tied to a stale grid/pos pair and must be overwritten.
2026-04-16 14:09:45 -05:00
a669e60735
test: expand coverage on ionosphere + space_weather contexts
6 additional characterization tests for empty-list upsert paths, station
filter isolation on latest_observation/1, multi-band xray coexistence,
and newest-across-time/bands latest_* helpers.
2026-04-16 14:09:45 -05:00
c0281e6c55
test: cover accounts user + user_token
80 characterization tests: registration/admin/email/password changesets,
password hashing, admin-flag grant logic, and the full token lifecycle
(session, magic link, confirm, change-email) including expiry boundaries
and cross-context rejection.
2026-04-16 14:09:45 -05:00
391a162bcd
test: cover beacons/beacon schema
46 characterization tests: changeset validation, Maidenhead derivation
(both directions), bearing normalization, keying options, mW/frequency
formatting helpers.
2026-04-16 14:09:45 -05:00
4826cdb2c8
test: cover propagation/rain_scatter
24 characterization tests for bistatic radar geometry: find_scatter_cells
filtering (dBZ threshold, distance floor/ceiling), bearing math, frequency
scaling, and classify/1 buckets.
2026-04-16 14:09:44 -05:00
cb2937f12c
test: cover geo, qrz record, callsign client
34 characterization tests for pure/wrapper modules. Documents existing
behavior — no production code changes.
2026-04-16 14:09:44 -05:00
FluxCD
b148e37c8e chore: update prop image to git.mcintire.me/graham/prop:main-1776361644-ac4e43c [skip ci] 2026-04-16 17:48:38 +00:00
ac4e43c726
Remove 500-limit on BackfillEnqueueWorker cron
10K+ HRRR-pending contacts were starving the 200-ish NARR candidates
out of the 500 per-run slot each 30 min because NARR candidates sort
last by the :hrrr_status priority. Drop the cap so every eligible
contact gets enqueued on each cron run; queue concurrency still paces
the actual fetches. Worker keeps supporting an explicit "limit" arg
for ad-hoc runs and tests.
2026-04-16 12:46:56 -05:00
FluxCD
2871f63cca chore: update prop image to git.mcintire.me/graham/prop:main-1776361263-f93f5c9 [skip ci] 2026-04-16 17:42:38 +00:00
f93f5c9a57
Rename 9cm band 3456 → 3400
The US 9cm amateur allocation was cut from 3300-3500 to 3300-3450 MHz
in 2020, so 3456 (the legacy weak-signal calling frequency) is no
longer in-band. Move the canonical key to 3400, migrate existing
contacts, and keep 3456 input resolving to 3400 via the
nearest-band snap so historical ADIF/CSV logs still import cleanly.
2026-04-16 12:40:38 -05:00
FluxCD
bc7f3f4137 chore: update prop image to git.mcintire.me/graham/prop:main-1776360998-36adf87 [skip ci] 2026-04-16 17:37:37 +00:00
36adf875b7
Add 50/144/222 MHz bands, rename 440→432
Expands submittable-contact bands to include 6m, 2m, 1.25m, and 70cm
(as 432 rather than the old 440 placeholder). Each new band gets an
explicit allocation window in BandResolver.nearest_band so ADIF FREQ
fields near the amateur allocations resolve correctly while 60-900 MHz
frequencies outside those windows are still rejected. Microwave
(>= 900 MHz) snapping is unchanged — nearest-band match across the
full @allowed_bands list.

Also adds BandConfig entries for 50 and 222 (tropo-only config, same
pattern as 144/432). Sporadic-E / F2 / meteor scatter modeling is not
yet in scope — ionosphere data is only used to compute an Es readout
on /path for 50/144/222/432.
2026-04-16 12:36:10 -05:00
FluxCD
90df646c94 chore: update prop image to git.mcintire.me/graham/prop:main-1776360352-11a50bf [skip ci] 2026-04-16 17:27:34 +00:00
11a50bf2c2
Disable PDF export on live_table pages
PDF generation was failing in production. CSV export still works.
2026-04-16 12:25:25 -05:00
FluxCD
5e5663fbb1 chore: update prop image to git.mcintire.me/graham/prop:main-1776360188-01d554f [skip ci] 2026-04-16 17:24:33 +00:00
01d554f2e7
Hourly ContactPositionBackfillWorker fills null pos1/pos2
Safety net for rows that land via direct DB writes (manual fixes, bulk
imports) and bypass Radio.create_contact's grid-resolution requirement.
Runs every hour on the :admin queue with a 500-row per-invocation cap.
Fills whichever side resolves — if one grid is invalid, the other still
gets populated; distance_km is only written when both positions end up
set.
2026-04-16 12:22:34 -05:00
FluxCD
1a8e55dfb0 chore: update prop image to git.mcintire.me/graham/prop:main-1776355088-bbeb95f [skip ci] 2026-04-16 15:59:59 +00:00
bbeb95ff5d
CSV/ADIF import: upsert existing contacts on grid/mode refinement
Widen dedup match to 4-char grid prefix so an upload with a longer grid
(e.g. FN42aa25) finds the existing FN42 contact. Within a match, classify
as a refinement when the upload strictly extends the existing grid or
fills a missing mode; as a contradiction (still skipped) when grids or
modes genuinely disagree. Refinements flow through a new preview bucket,
update the existing row in place, and re-enqueue enrichment when the
position changed.
2026-04-16 10:57:37 -05:00
FluxCD
824143ad64 chore: update prop image to git.mcintire.me/graham/prop:main-1776351221-be6f1a1 [skip ci] 2026-04-16 14:55:32 +00:00
be6f1a1d28
CsvImport: header-driven column mapping
The /contacts LiveTable export produces a 12-column CSV with label-style
headers (Station 1 / Grid 1 / QSO (UTC) / ...) in a different column
order plus extras (id, distance_km, hrrr_status, flagged_invalid,
inserted_at). The old positional import rejected it with "expected 6 or
7 columns, got 12".

Replace the positional parser with header-driven mapping:
- Parse the header row, normalize each cell (downcase + trim), and look
  it up in an alias map that covers both the raw field names and the
  export labels.
- Required columns: station1, station2, grid1, grid2, band, qso_timestamp
  (mode stays optional). If any are missing the header, return
  {:error, {:missing_required_columns, [...]}}.
- Unknown columns are silently ignored so the export's extras don't break
  round-trip.
- Per-row validation stays the same (changeset errors still reported by
  row number).

Existing 6- or 7-column positional CSVs that already have a header still
work — the header parser recognizes `station1`, `station2`, etc.
directly.
2026-04-16 09:53:14 -05:00
c7f9aac3f9
NarrClient: parse cdo numeric codes for cdo-version independence
Prod cdo 2.5.1 (Debian) emits named shortnames (tpag10, 2tag2, saip, code11…)
where cdo 2.6.0 (Mac) emits var11/var33/etc. The parser only recognized
varNN, so every prod NARR fetch failed with "no parseable values in cdo
output" and all 350 in-flight jobs discarded.

Switch cdo args from `-outputtab,name,lev,value` to `-outputtab,code,...`.
cdo then outputs the raw numeric NCEP GRIB1 parameter code regardless of
version, and @cdo_code_to_name maps that integer to the semantic name we
already use downstream (TMP/HGT/SPFH/…).

Verified by reproducing on both a 2010 NARR file (fixture) and a 2007 file
fetched live — both now emit identical numeric code rows.
2026-04-16 09:45:23 -05:00
FluxCD
c1cac4f394 chore: update prop image to git.mcintire.me/graham/prop:main-1776350332-f005ddc [skip ci] 2026-04-16 14:40:29 +00:00
f005ddcaac
NarrClient: include cdo output snippet in parse-failure errors
Prod hit "no parseable values in cdo output" for a 2007 record.
Without the actual cdo output it's guesswork whether it's a different
parameter-table format, an empty result, or something else. Capture
stderr (stderr_to_stdout: true) and include first 500 chars of output
so the error message has enough signal to diagnose.
2026-04-16 09:38:27 -05:00
FluxCD
a2cdee7a31 chore: update prop image to git.mcintire.me/graham/prop:main-1776349903-55d4f28 [skip ci] 2026-04-16 14:33:27 +00:00
55d4f289c2
NarrClient.in_coverage?/1 + callsite guards (1979 → 2014-10-02)
Post-2014 contacts with hrrr_status=:unavailable were being dispatched to
NARR, which 404s because NCEI's archive ends 2014-10-02. 14K of the 14.3K
candidates in prod are post-2014 — these are HRRR's responsibility.

- NarrClient.in_coverage?/1 returns true only inside 1979-01-01 → 2014-10-02
- narr_jobs_for_contact and maybe_enqueue_narr (contact_live) bail out of
  coverage returning []
- BackfillEnqueueWorker.type_filter scopes :narr to qso_timestamp <
  coverage_end so the cron doesn't keep picking post-2014 candidates
2026-04-16 09:31:19 -05:00
FluxCD
f5132a991d chore: update prop image to git.mcintire.me/graham/prop:main-1776349371-3604896 [skip ci] 2026-04-16 14:24:23 +00:00
3604896726
Rename ERA5 → NARR across the codebase
- Schema: Era5Profile → NarrProfile; table renamed via migration
  rename_era5_profiles_to_narr_profiles (ALTER TABLE + rename indexes,
  atomic and instant, no row rewrite)
- Weather helpers: find_nearest_era5/era5_for_contact/era5_profiles_for_path
  → find_nearest_narr/narr_for_contact/narr_profiles_for_path
- BackfillEnqueueWorker: :era5 → :narr type (virtual, no narr_status column);
  reconcile_stale_queued_to_unavailable and status_priority_order now skip
  virtual types via @virtual_types. Fixes the prod crash where the cron tried
  to update a non-existent era5_status column.
- ContactWeatherEnqueueWorker: build_era5_jobs → build_narr_jobs,
  era5_jobs_for_contact → narr_jobs_for_contact
- ContactLive: @era5 → @narr, maybe_enqueue_era5 → maybe_enqueue_narr;
  UI label "ERA5 (0.25°)" → "NARR (32 km)"
- Cron (runtime.exs): args types list now "narr" instead of "era5"
- /status: progress row, status-by-type card, totals.narr_profiles, and
  table-count lookup all target narr_profiles
- Drop obsolete Era5CdsJob schema + era5_cds_jobs table (inspection
  artifact from the retired CDS pipeline; 34 orphan rows in prod)
- Misc docstring/comment cleanups (skew_t, about, wgrib2, propagation_train)

Includes a regression test for the virtual-type crash.
2026-04-16 09:22:23 -05:00
FluxCD
9e94759f7a chore: update prop image to git.mcintire.me/graham/prop:main-1776347847-7857d3b [skip ci] 2026-04-16 13:59:17 +00:00
7857d3bc5a
Status page: relabel ERA5 → NARR, include NARR in backfill cron
- /status progress bar, status-by-type card, and internal map keys now
  read "NARR" (data-progress-key="narr", narr_candidates/narr_done).
  era5_profiles table stays as the DB landing spot (historical artifact).
- BackfillEnqueueWorker cron now passes "era5" in types so pre-2014
  contacts whose hrrr_status is :unavailable actually get NarrFetchWorker
  jobs enqueued every 30 min. The :era5 type key still maps internally
  to NarrFetchWorker via era5_jobs_for_contact/1.
2026-04-16 08:57:01 -05:00