Multi-phase plan to incorporate the meteorologist's feedback on algo.md: hybrid-sigma vertical resolution, boundary-layer turbulence features, ray-traced duct geometry, frontal geometry, radar-based stability, climatological anomalies, regional seasonal tuning, and 5-minute METAR. Every phase has a backtest-gated stop criterion so features that don't show lift against the historical QSO corpus don't ship.
35 KiB
Propagation Modeling Improvements Implementation Plan
For Claude: REQUIRED SUB-SKILL: Use superpowers:executing-plans to implement this plan task-by-task.
Goal: Rework the propagation scorer to incorporate the physical mechanisms surfaced by the April 2026 meteorologist review: fine vertical resolution, boundary-layer turbulence, ray-traced duct geometry, frontal geometry, observed radar-based stability, and regional/climatological corrections. Retire features that don't survive backtesting against the 58k QSO corpus.
Architecture: Add new feature computations on top of the existing HRRR → scoring → propagation_scores pipeline without replacing it. Validate each feature in isolation against historical QSOs. Reweight (and drop) features in a final recalibration pass. Existing scoring must keep working throughout — no feature goes into production until its backtest justifies it.
Tech Stack: Elixir (existing), HRRR native hybrid-sigma GRIB2 (new ingestion), NEXRAD Level II via Iowa Environmental Mesonet (new pipeline), Ecto migrations (additive only), Nx for spatial gradients and ray-tracing math.
Context
The meteorologist's review identified:
- Vertical resolution too coarse — 25 hPa ≈ 250 m; ducts are 50-200 m thick, so the current pressure-level product smooths across the layer we care about.
min_refractivity_gradientconflates two things — the strength of the gradient and the smoothness of the interface. A strong gradient on a turbulent layer does not propagate.- Pressure/wind correlations leak through from unmodeled frontal geometry — findings 2 and 3 are proxies for "south of an advancing cold front, path parallel to it".
- PWAT is a co-variate of frontal position, not a cause — probably drops out entirely when frontal geometry is modeled directly.
- Seasonal/humidity scoring needs regional stratification — Gulf August is dry, Iowa August is the opposite because of corn evapotranspiration.
- NEXRAD clear-air texture is an underused observed BL-stability signal — real-time, free, independent of HRRR's forecast error.
- Elevated ducts above the boundary layer exist — some X-band contacts don't use the BL-top mechanism at all.
- Temperature anomaly vs climatology matters more than absolute temperature for summer afternoon enhancement.
Phase 0 builds the infrastructure to falsify each of these claims against the QSOs we already have. Every downstream phase is a testable hypothesis.
Phase Ordering and Dependencies
Phase 0: Backtest harness ───┬──► all subsequent phases depend on this
│
Phase 1: HRRR hybrid-sigma ──┼──┬── Phase 2: BL turbulence features ──┐
│ │ │
│ └── Phase 4: Ray-traced duct geometry ─┤
│ │
Phase 3: NEXRAD BL stability (parallel, independent) ─────────────────┤
│ │
Phase 5: Frontal geometry ───┴── (requires spatial gradients) │
│
Phase 6: Temperature anomaly ── (requires climatology backfill) ──────┤
│
Phase 7: Regionalized seasons ── (pure config, can start anytime) ────┤
│
Phase 8: 5-min METAR ──── (parallel, independent) ────────────────────┤
│
Phase 9: Feature selection and recalibration ◄────────────────────────┘
- Phase 0 blocks everything.
- Phase 1 blocks Phases 2, 4, and (partially) 5.
- Phase 9 must come last — it's where the scorer actually changes.
- Phases 3, 7, 8 are independent and can run in parallel whenever there's bandwidth.
Each phase has a gate: if backtesting against the QSO corpus doesn't show lift, the feature is parked, not merged. The plan is sized so that a negative result on a phase doesn't block the rest of the work.
Phase 0: Backtest Harness
Why first: We can't evaluate the meteorologist's claims without the ability to score a proposed feature against historical QSO data. Every subsequent phase's gate depends on this.
Goal: Microwaveprop.Backtest — takes a feature function (lat, lon, valid_time) -> float | nil and produces a report comparing feature distribution at QSO times vs. random baseline, lift by distance bin, lift by band.
Key constraint: We don't have non-contact labels. A QSO existing means "propagation was good enough to work at this band between these grids at this time." To create contrast, we use:
- Random-time baseline — same (lat, lon) but a random timestamp in the same month. If a feature has predictive power, its distribution at QSO times should differ from its distribution at random times.
- Distance binning — longer paths on the same band imply better propagation, so a good feature should correlate with
distance_km. - Band stratification — a feature that only works for 10 GHz tells a different story than one that works for 47 GHz.
Task 0.1: Backtest module scaffolding
Files:
- Create:
lib/microwaveprop/backtest.ex - Create:
test/microwaveprop/backtest_test.exs
Step 1: Write failing test
test "evaluate/2 returns counts and distributions for a trivial feature" do
create_qso(band: "10000", qso_timestamp: ~U[2026-01-15 12:00:00Z], pos1: %{...})
feature = fn _lat, _lon, _time -> 1.0 end
report = Backtest.evaluate(feature, sample_size: 10, baseline_size: 10)
assert report.qso_count == 1
assert report.baseline_count == 10
assert %Backtest.Distribution{} = report.qso_distribution
end
Step 2: Implement evaluate/2 pulling QSOs with pos1 != nil, iterating the feature over them, and producing the same stats for random (lat, lon, random-month-hour) triples.
Step 3: Commit. feat(backtest): scaffold backtest harness module
Task 0.2: Distance-binned lift report
Files:
- Modify:
lib/microwaveprop/backtest.ex - Modify:
test/microwaveprop/backtest_test.exs
Add Backtest.lift_by_distance/2:
- Bin QSOs by
distance_km(0-100, 100-250, 250-500, 500-1000, 1000+). - For each bin, compute mean and p50/p90 of the feature value.
- Return a map keyed by distance bin.
A good feature shows monotonically increasing mean with distance bin; a neutral feature shows flat. Write tests with a synthetic feature that has known correlation.
Task 0.3: Band-stratified report
Backtest.lift_by_band/2 — same idea, keyed on band (decimal). Writes a sortable table.
Task 0.4: Random-time baseline generator
Backtest.random_baseline/2 — for a set of (lat, lon) pairs, generate N random timestamps within ±30 days of each QSO's actual timestamp (matched sample). This controls for seasonality.
Task 0.5: CLI runner
Files:
- Create:
lib/mix/tasks/backtest.ex
mix backtest --feature Microwaveprop.Backtest.Features.NaiveGradient — loads a feature module, runs evaluate/2, prints a Markdown report to stdout.
Task 0.6: Baseline feature modules
Create the feature functions for the existing algo factors, so we have known baselines to compare new features against:
Backtest.Features.NaiveGradient— wraps the currentmin_refractivity_gradientscoring.Backtest.Features.TdDepressionBacktest.Features.TimeOfDayBacktest.Features.Pressure
Run mix backtest on each, save reports to priv/backtest_reports/baseline_*.md, commit.
Gate for Phase 0: the baseline reports must plausibly match the published algo findings (e.g. refractivity should show some lift; time-of-day should show stronger lift at night). If they don't, the backtest harness is wrong and needs fixing before moving on.
Estimated effort: ~1 engineer-week.
Phase 1: HRRR Native Hybrid-Sigma Level Ingestion
Why: 25 hPa pressure levels smooth across the ducting layer. Phases 2, 4, 5, and 6 need the native vertical resolution.
Research: HRRR native output is on a hybrid-sigma grid with 50 levels, densest near the surface (the first few sigma levels are <30 m apart). NOAA publishes this as the wrfnatf product alongside the standard wrfsfcf / wrfprsf files on NOMADS and AWS S3. File size is ~4× the pressure-level product.
Goal: Add a parallel ingestion path that pulls and stores hybrid-sigma profiles, without disturbing the existing hrrr_profiles table.
Task 1.1: Spike — download one hybrid-sigma file and inspect
Files:
-
Create:
scripts/hrrr_hybrid_spike.exs -
Pull one
hrrr.tXXz.wrfnatf00.grib2from AWS via the existingHrrrClient. -
Run
wgrib2 -Vto list variables and levels. -
Document: which variables are on which levels, what the level spacing looks like in the lowest 2 km, typical file size.
-
Write findings to
docs/research/hrrr_native_levels.md.
Gate: if hybrid-sigma files aren't reliably available on AWS S3 for historical dates (they rotate off faster than pressure-level files), we need a different data source. Document this first.
Task 1.2: New schema hrrr_native_profiles
Files:
- Create:
priv/repo/migrations/XXXXXXX_add_hrrr_native_profiles.exs - Create:
lib/microwaveprop/weather/hrrr_native_profile.ex
create table(:hrrr_native_profiles, primary_key: false) do
add :id, :binary_id, primary_key: true
add :valid_time, :utc_datetime
add :lat, :float
add :lon, :float
add :level_count, :integer
add :heights_m, {:array, :float} # geopotential height per level
add :temp_k, {:array, :float}
add :dewpoint_k, {:array, :float}
add :pressure_pa, {:array, :float}
add :u_wind_ms, {:array, :float}
add :v_wind_ms, {:array, :float}
add :tke_m2s2, {:array, :float} # if available in HRRR native
timestamps()
end
create unique_index(:hrrr_native_profiles, [:lat, :lon, :valid_time])
Do NOT modify hrrr_profiles. The new table lives alongside it. We flip the scorer only after Phase 9.
Task 1.3: Native-level GRIB2 extraction
Files:
- Modify:
lib/microwaveprop/weather/grib2/wgrib2.ex - Create:
lib/microwaveprop/weather/hrrr_native_client.ex - Create:
test/microwaveprop/weather/hrrr_native_client_test.exs
HrrrNativeClient.fetch_profile(lat, lon, valid_time):
- Uses the byte-range index trick on
wrfnatf00.grib2.idxto pull only the variables and levels we need (keep the transfer small). - Extracts T, Td, P, U, V, Z, TKE on all hybrid levels for the point.
- Returns a normalized struct with all levels sorted by height.
Test with stubbed Req responses and a recorded GRIB2 fixture.
Task 1.4: Ingestion worker
Files:
- Create:
lib/microwaveprop/workers/hrrr_native_fetch_worker.ex
On-demand worker (like the existing HrrrFetchWorker), enqueued per QSO:
use Oban.Worker, queue: :hrrr, max_attempts: 3,
unique: [period: :infinity, keys: [:lat, :lon, :valid_time]]
Fetches, upserts into hrrr_native_profiles, marks the QSO's new hrrr_native_status column. Add that column in a small migration.
Task 1.5: Backfill top-500 most-contacted grid cells
Write a mix task mix hrrr_native_backfill --limit 500 that enqueues the native fetch for the most-used cells so we have enough data to backtest. Run it. Monitor oban_web to see it finish.
Task 1.6: First backtest — does level-0 refractivity improve?
Create Backtest.Features.NativeSurfaceRefractivity using only the lowest native level (vs. the interpolated lowest pressure level the current algo uses). Run backtest. Expected: marginal or no improvement, because the difference is a ~30 m vertical offset.
This is the sanity check that the pipeline works before investing in Phase 2.
Gate: pipeline produces native-level profiles for at least 80% of the attempted backfill cells and the sanity-check backtest doesn't regress vs. the baseline report.
Estimated effort: ~2 engineer-weeks.
Phase 2: Boundary-Layer Turbulence Features
Why: The meteorologist's central claim is that gradient strength and interface smoothness are separate physical factors and the algo only measures one of them. Fix it.
Goal: Compute Richardson number, θₑ gradient, and wind shear at the inversion top from native-level profiles. Store as new columns. Backtest each independently.
Task 2.1: Inversion-top detection
Files:
- Create:
lib/microwaveprop/propagation/inversion.ex - Create:
test/microwaveprop/propagation/inversion_test.exs
def find_inversion_top(profile) do
# Walk the profile from the surface up.
# Find the first height where dT/dz > 0 (temperature inversion)
# or where dN/dz is at its minimum (classic ducting detection).
# Return {:ok, height_m, level_idx} or :none.
end
TDD with three synthetic profiles:
- No inversion (monotonic T cooling with height) →
:none - Surface-based inversion (T increases from surface to 500 m) → top at ~500 m
- Elevated inversion at 2 km → top at ~2 km
Task 2.2: Richardson number computation
Files:
- Modify:
lib/microwaveprop/propagation/inversion.ex
def bulk_richardson(profile, top_level_idx) do
# Ri = (g / theta_ref) * (delta_theta / delta_u^2)
# where delta_theta and delta_u are across the layer from
# surface to top_level_idx.
end
Literature values:
- Ri < 0.25 → dynamically unstable, turbulent
- 0.25 < Ri < 1.0 → marginal
- Ri > 1.0 → stable/laminar (good for ducting)
TDD with synthetic cases: high-shear + weak-inversion → Ri low; calm + strong-inversion → Ri high.
Task 2.3: θₑ computation and gradient
Files:
- Create:
lib/microwaveprop/weather/theta_e.ex - Create:
test/microwaveprop/weather/theta_e_test.exs
Bolton (1980) formulation for equivalent potential temperature from T, Td, P. Property test: θₑ is conserved along a dry adiabat; use stream_data.
Then Inversion.theta_e_jump(profile, top_level_idx) returns ΔθₑK across the inversion.
Task 2.4: Derived fields on hrrr_native_profiles
Files:
- Create:
priv/repo/migrations/XXXXXXX_add_derived_hrrr_native_fields.exs - Modify:
lib/microwaveprop/workers/hrrr_native_fetch_worker.ex
Add columns (all nullable):
inversion_top_m :floatbulk_richardson :floattheta_e_jump_k :floatshear_at_top_ms :floatduct_thickness_m :float(placeholder for Phase 4)
Compute these in the fetch worker after extraction. Backfill the existing rows with a one-shot mix task mix hrrr_native_derive.
Task 2.5: Backtest the turbulence features
Create three feature modules:
Backtest.Features.BulkRichardsonBacktest.Features.ThetaEJumpBacktest.Features.ShearAtTop
Run mix backtest on each. Write the reports to priv/backtest_reports/phase2_*.md.
Gate: at least one of the three features shows meaningful lift (defined as: p90 of the feature at QSO times > p50 at random baseline times, or monotonic lift across distance bins for at least two bands). If none do, either:
- The backtest sample is too small (backfill more cells), or
- The features are correctly computed but the physical claim doesn't hold for our corpus, or
- Our corpus is dominated by contacts that use a different mechanism (elevated ducts — Phase 4 may rescue this).
Document the finding either way.
Estimated effort: ~1.5 engineer-weeks.
Phase 3: NEXRAD Clear-Air Stability (parallel, independent)
Why: Real-time, observed, free. If it works, it's the single highest-ROI new data source and it doesn't depend on HRRR being right.
Goal: Pull NEXRAD Level II composite reflectivity, compute local spatial variance, expose as a feature.
Task 3.1: Spike — download and visualize one frame
Files:
-
Create:
scripts/nexrad_spike.exs -
Use the Iowa Environmental Mesonet composite reflectivity GeoTIFF endpoint.
-
Download one frame for a known dead-band time (e.g. 20Z on a summer afternoon) and one for a known good time (e.g. 10Z at dawn on the same day).
-
Decode with
Imageor shell out togdal_translate. -
Visually inspect: does the dead-band frame look "blotchy" as the meteorologist described?
Gate: if we can't visually see the difference by eye, the feature probably doesn't have signal and Phase 3 stops here.
Task 3.2: NEXRAD ingestion worker
Files:
- Create:
lib/microwaveprop/weather/nexrad_client.ex - Create:
lib/microwaveprop/workers/nexrad_fetch_worker.ex - Create:
priv/repo/migrations/XXXXXXX_create_nexrad_stability.exs
Table:
create table(:nexrad_stability, primary_key: false) do
add :id, :binary_id, primary_key: true
add :valid_time, :utc_datetime
add :lat, :float
add :lon, :float
add :local_variance, :float
add :mean_reflectivity, :float
timestamps()
end
Pulls composite reflectivity at 5-6 minute cadence for a CONUS grid matching the existing propagation grid spacing. Store only the derived variance — we don't need the raw tiles.
Task 3.3: Variance computation
Files:
- Create:
lib/microwaveprop/weather/radar_texture.ex - Create:
test/microwaveprop/weather/radar_texture_test.exs
For each grid point, compute the variance of reflectivity in a 40×40 km window around it. Exclude pixels with precipitation returns (> ~20 dBZ) — we only care about clear-air texture, not actual weather.
Use Nx for the windowed variance to keep it fast.
Task 3.4: Backfill
Mix task mix nexrad_backfill --days 30 — pulls frames for the last 30 days at hourly cadence, computes variance, writes to nexrad_stability.
Task 3.5: Backtest
Backtest.Features.NexradTexture — mean local_variance in a ±30 min window at the QSO midpoint. Run mix backtest.
Gate: NEXRAD texture must show at least as much lift as the best Phase 2 feature to justify carrying a new data pipeline in production. If it's strictly better, it becomes a first-class feature in Phase 9. If it's comparable, ship it alongside as a sanity check.
Estimated effort: ~2 engineer-weeks.
Phase 4: Ray-Traced Duct Geometry
Why: Turns "is it ducty" into "can this band use the duct". Critical for band-specific scoring, and necessary to correctly score elevated ducts.
Goal: Replace min_refractivity_gradient (scalar) with {duct_top_m, duct_base_m, duct_thickness_m, trapped_freq_ghz} — a tuple describing the duct geometry from ray-tracing the native-level profile.
Task 4.1: Refractivity profile from native levels
Files:
- Create:
lib/microwaveprop/propagation/refractivity_profile.ex - Create:
test/microwaveprop/propagation/refractivity_profile_test.exs
def from_native_profile(profile) do
# For each level, compute N = 77.6*P/T + 3.73e5*e/T^2 (ITU-R P.453)
# Return a list of {height_m, N} tuples sorted by height.
end
TDD with the ITU-R sample profiles.
Task 4.2: Duct detection via M-profile
The modified refractivity M = N + 157·h (where h is in km) is the conventional form. A duct exists wherever dM/dh < 0.
Files:
- Modify:
lib/microwaveprop/propagation/refractivity_profile.ex
def detect_ducts(refractivity_profile) do
# Walk the profile computing M at each level.
# Every contiguous region where dM/dh < 0 is a duct.
# Return a list of %Duct{base_m, top_m, thickness_m, m_deficit}.
end
TDD: synthetic profile with a known 100 m surface-based duct and a known 200 m elevated duct at 1.5 km → detector finds both.
Task 4.3: Trapped frequency calculation
The minimum frequency a duct of thickness d can trap is approximately:
f_min_ghz ≈ 3.6 / (d * sqrt(ΔM) / 100) ^ 1.5 (rough engineering approximation)
Find the right formula in Bean & Dutton or ITU-R P.834 and encode it. For each duct, compute f_min_ghz.
Then a band can use the duct iff band_ghz >= f_min_ghz. This gives us per-band scoring derived from physical geometry, not hand-tuned tables.
Task 4.4: Store duct geometry
Files:
- Create:
priv/repo/migrations/XXXXXXX_add_duct_geometry.exs
Add columns to hrrr_native_profiles:
ducts :map(stores the list of %Duct{} structs as JSONB)best_duct_band_ghz :float(lowestf_minacross all detected ducts — cached for fast scoring)
Backfill via mix hrrr_native_derive (reuse task from Phase 2).
Task 4.5: Backtest duct features per band
Create:
Backtest.Features.DuctUsable10GHz— bool: is there a duct with f_min ≤ 10 GHzBacktest.Features.DuctUsable24GHzBacktest.Features.DuctUsable47GHzBacktest.Features.DuctThickness— continuous: max thickness among detected ducts
Run band-stratified backtests. For each band, check:
- At QSO times on 10 GHz, is
DuctUsable10GHztrue more often than at random times? (Expected: yes.) - Is the lift larger for 47 GHz than 10 GHz? (Expected: yes, because 47 GHz is harder, so ducting matters more.)
Gate: at least one of the duct features must outperform the existing NaiveGradient baseline on at least one band. If not, the ray-tracing added no value and we keep the scalar gradient.
Estimated effort: ~2 engineer-weeks, most of it in getting the trapped-frequency formula right and validating against known cases from the literature.
Phase 5: Frontal Geometry (biggest conceptual gain)
Why: The meteorologist's claim is that pressure, calm winds, and PWAT are all proxies for "proximity to and orientation relative to a cold front". Modeling the front directly should replace three weak features with one strong geometric one.
Goal: For every (grid cell, valid_time), compute:
distance_to_front_km— distance to nearest detected cold frontfront_bearing_deg— bearing of the front at the nearest pointpath_front_angle_deg— angle between the QSO path and the front (0 = parallel, 90 = crosses)
Task 5.1: Frontal detection spike
Files:
- Create:
scripts/front_detection_spike.exs
Research first. Candidate approaches:
- Thermal front parameter (TFP) —
-∇|∇θ| · ∇θ / |∇θ|, standard synoptic meteorology tool. Highest negative values mark cold fronts. - Theta-e gradient magnitude directly.
- Maximum of the horizontal wind shift across a grid.
Implement TFP on one HRRR frame with a known strong cold front (pick a date from the NWS archive). Plot the result as a raster on top of the HRRR temperature field. Visual verification: do the detected fronts line up with the NWS surface analysis for that day?
Gate: if the detection doesn't visually match a hand-drawn NWS front, the automated method isn't reliable enough for feature engineering and Phase 5 becomes a research project rather than an implementation task. Document the finding, park the phase.
Task 5.2: Frontal mask pipeline
Files:
- Create:
lib/microwaveprop/weather/frontal_analysis.ex - Create:
test/microwaveprop/weather/frontal_analysis_test.exs
def detect_fronts(hrrr_grid) do
# Input: a 2D grid of HRRR surface fields (T, Td, U, V)
# Output: a 2D boolean mask of front pixels, plus a bearing per pixel
end
Use Nx for the gradient math. Test with a synthetic grid that has a known linear front.
Task 5.3: Per-cell frontal features
Add to propagation_scores or a sidecar table frontal_features:
distance_to_front_kmfront_bearing_deg
Compute these hourly in the propagation pipeline. Cache alongside scores.
Task 5.4: Per-path angle feature
When scoring a QSO (or a potential path) the path_front_angle is a function of both endpoints. This doesn't cache well, so expose it as a helper:
def path_front_angle(path_lat1, path_lon1, path_lat2, path_lon2, valid_time)
It looks up the nearest front at the path midpoint and computes the angle.
Task 5.5: Backtest frontal features
Backtest.Features.DistanceToFront— distance to nearest front at the contact's grid 1 position.Backtest.Features.ParallelToFront— continuous:cos²(path_front_angle)so perpendicular paths score 0 and parallel paths score 1.
Run backtest stratified by distance. Expected: long-distance contacts cluster at paths with low path_front_angle (parallel) and 200-500 km distance_to_front_km (not on the front, but near it).
Gate: ParallelToFront must outperform the existing Pressure baseline, confirming the meteorologist's claim that pressure is a weaker proxy for the same thing. If so, pressure is dropped in Phase 9.
Estimated effort: ~3 engineer-weeks including the research spike.
Phase 6: Temperature Anomaly vs. Climatology
Why: Finding 4 (diurnal) is real but has an important exception: extremely hot summer days produce strong enhancement even in the afternoon. Absolute temperature isn't the right feature; deviation from normal is.
Goal: For every (grid cell, month, hour), compute a climatological baseline temperature, and expose the anomaly as a feature.
Task 6.1: Build climatology from existing hrrr_profiles
Files:
- Create:
lib/mix/tasks/hrrr_climatology.ex - Create:
priv/repo/migrations/XXXXXXX_create_hrrr_climatology.exs
Table:
create table(:hrrr_climatology, primary_key: false) do
add :id, :binary_id, primary_key: true
add :lat, :float
add :lon, :float
add :month, :integer
add :hour, :integer
add :mean_surface_temp_k, :float
add :stddev_surface_temp_k, :float
add :sample_count, :integer
end
create unique_index(:hrrr_climatology, [:lat, :lon, :month, :hour])
We already have 4.5k hrrr_profiles. That's thin for climatology but enough to compute a rough baseline per (lat, lon, month, hour) for the grid cells with the most data. For cells with insufficient samples, fall back to NCEP reanalysis climatology (NARR has 30+ years of data).
Task 6.2: Temperature anomaly feature
def anomaly(lat, lon, valid_time, current_temp_k) do
# Look up climatology.mean_surface_temp_k for (lat, lon, month, hour)
# Return current_temp_k - climatology_mean
end
Task 6.3: Backtest
Backtest.Features.TemperatureAnomaly. Expected: positive anomalies (hotter than normal) should show lift, especially during afternoon hours in summer — which is exactly when the existing TimeOfDay feature has negative lift. The two features should be anti-correlated in this regime, meaning anomaly "rescues" hot-day afternoon scoring.
Gate: anomaly shows positive lift during summer afternoons where the current TimeOfDay suppresses the score.
Estimated effort: ~1 engineer-week.
Phase 7: Regionalized Seasonal Scoring
Why: The current @band_configs seasonal scoring is uniform across CONUS. Gulf August is dry and propagates well; Iowa August is miserable humidity and doesn't. Same month, opposite score.
Goal: Partition CONUS into 6-8 climatological regions and maintain per-region seasonal scoring.
Task 7.1: Regional partition
Files:
- Create:
lib/microwaveprop/propagation/region.ex
Define regions by bounding box (or Köppen climate zone, but bounding boxes are simpler):
@regions [
%{id: :gulf_coast, lat: {25, 32}, lon: {-100, -80}},
%{id: :southern_plains, lat: {32, 38}, lon: {-105, -93}},
%{id: :corn_belt, lat: {38, 45}, lon: {-100, -82}},
%{id: :desert_southwest, lat: {30, 38}, lon: {-120, -105}},
# ... 6-8 total
]
Region.for_point(lat, lon) returns the region id.
Task 7.2: Per-region seasonal tables
Files:
- Modify:
lib/microwaveprop/propagation/band_config.ex
Change @band_configs to index by {band_mhz, region_id} instead of just band_mhz. Start with the existing table replicated across all regions, then tune per region based on backtest.
Task 7.3: Backtest per-region lift
Run the existing Backtest.Features.Season per region, compare against the uniform baseline. Tune the regional seasonal curves until the regional scores fit the regional QSO patterns.
Estimated effort: ~1 engineer-week. Mostly data-driven tuning rather than engineering.
Phase 8: 5-Minute METAR Ingestion
Why: The current surface observation feed is hourly. Propagation windows can open and close on 20-30 minute timescales. Finer-grained surface data (especially temperature and dewpoint) lets the scorer catch transients.
Goal: Ingest NOAA NCEI 5-minute METAR, store alongside the existing surface_observations table.
Task 8.1: NCEI 5-min METAR client
Files:
- Create:
lib/microwaveprop/weather/ncei_metar_client.ex
Use the endpoint from the meteorologist's email (NCEI C00418). Document the format — it's different from the Iowa Mesonet ASOS feed the current pipeline uses.
Task 8.2: Schema extension
Option A: new table metar_5min_observations identical in schema to surface_observations but with 5-minute granularity.
Option B: add a source column to surface_observations and mix them.
A is safer (no risk to existing queries), so do that. Migration adds the new table.
Task 8.3: Ingestion worker
Microwaveprop.Workers.NceiMetar5MinFetchWorker, hourly cron pulling the most recent hour's data.
Task 8.4: Use in scoring
Add a recent_surface_obs/2 helper that prefers the 5-min source when available, falls back to surface_observations otherwise.
Estimated effort: ~1 engineer-week.
Phase 9: Feature Selection and Recalibration
Why: Everything above is additive. The real improvement is cutting the features that don't survive backtesting and reweighting the ones that do.
Goal: Produce a new BandConfig-equivalent weight table that uses only the features with demonstrated lift, and validate end-to-end against the QSO corpus.
Task 9.1: Consolidated backtest report
Run mix backtest --all that evaluates every registered feature and writes a consolidated table: feature name, lift by band, lift by distance, sample count, gate pass/fail.
Task 9.2: Drop dead features
For each existing factor in the scorer (humidity, time_of_day, td_depression, refractivity, sky, season, wind, rain, pressure):
- If the consolidated report shows lift → keep, possibly with new formulation.
- If it shows no lift and there's a new feature measuring the same physics → drop, use the new one.
- If it shows no lift and there's no replacement → drop entirely.
Expected drops (if the meteorologist is right):
pressure— replaced byParallelToFrontwind— same reason- PWAT if it was ever added — replaced by frontal features
Expected keeps:
humidity— core physicstime_of_day— now augmented withTemperatureAnomalytd_depression— core physicsrefractivity— replaced withDuctUsableper bandsky— keep but re-validateseason— keep but regionalized (Phase 7)rain— keep, ITU-R absorption model
Expected adds:
bulk_richardsonortheta_e_jump(from Phase 2)nexrad_texture(from Phase 3)duct_usable_band(from Phase 4)parallel_to_front+distance_to_front(from Phase 5)temperature_anomaly(from Phase 6)
Task 9.3: Recalibrate weights via gradient descent
Files:
- Create:
lib/mix/tasks/recalibrate_scorer.ex
With the new feature set, fit weights by minimizing some loss against the QSO corpus. Simplest loss: negative log-likelihood of QSO distances under a logit link, treating QSOs as positives and random-baseline samples as negatives. Can use Nx.
This is a tiny model (< 15 features), so a few thousand iterations on CPU is fine. Cross-validate by month.
Task 9.4: Side-by-side scoring run
Run the new scorer and the old scorer on the same HRRR frame. Diff the output grids. Look for:
- Any gross regressions (areas where the old scorer said "good" and the new scorer says "dead" with no physical justification)
- Gains in known-good corridors (e.g. Gulf coast → Brazos → DFW on a summer dawn)
Task 9.5: Shadow deployment
Run the new scorer in parallel with the old one for 2 weeks in production. Both publish scores to PubSub but only the old one drives the map. Collect daily diff statistics. Promote if the new one tracks observed QSOs better.
Task 9.6: Cutover
Flip the scorer. Keep the old implementation compilable for one release as a fallback.
Estimated effort: ~2 engineer-weeks.
Total Effort and Sequencing
| Phase | Description | Weeks | Dependencies |
|---|---|---|---|
| 0 | Backtest harness | 1 | — |
| 1 | HRRR hybrid-sigma ingestion | 2 | 0 |
| 2 | BL turbulence features | 1.5 | 1 |
| 3 | NEXRAD clear-air stability | 2 | 0 |
| 4 | Ray-traced duct geometry | 2 | 1 |
| 5 | Frontal geometry | 3 | 0 |
| 6 | Temperature anomaly | 1 | 1 |
| 7 | Regionalized seasonal scoring | 1 | 0 |
| 8 | 5-minute METAR | 1 | 0 |
| 9 | Feature selection + recalibration | 2 | all |
Total: ~17 engineer-weeks if done serially. With some parallelism across independent phases: ~12 calendar weeks.
Suggested execution order (maximizing parallelism and learning rate):
- Weeks 1: Phase 0 (blocking).
- Weeks 2-3: Phase 1 (blocking for most downstream work) + Phase 3 spike in parallel (independent, highest-ROI standalone).
- Weeks 4-5: Phase 3 full + Phase 2 in parallel.
- Weeks 6-7: Phase 4 + Phase 6 in parallel.
- Weeks 8-10: Phase 5 (biggest lift; if the spike doesn't match NWS fronts, park it and move to Phase 7/8 instead).
- Weeks 11-12: Phase 7 + Phase 8 + final Phase 9 recalibration and shadow deployment.
Kill criteria — when to stop mid-plan:
- If Phase 0's baseline reports don't match published algo findings → stop, fix the harness.
- If Phase 1's hybrid-sigma data isn't available historically → switch to ERA5 model levels as a plan B for backtesting; stay on pressure levels for live scoring.
- If Phase 2 shows zero lift for all three turbulence features → the "smoothness" mechanism may not be the dominant one in our corpus; park Phase 2 and let Phase 4 (ducts) do the work.
- If Phase 3's NEXRAD spike shows no visual difference between dead-band and good times → park Phase 3.
- If Phase 5's frontal detection doesn't match hand-drawn NWS fronts → park Phase 5 as a research project.
Each phase's backtest report becomes the justification for shipping or parking its work. Nothing gets merged to production on vibes.
Out of Scope (Explicitly)
- Rewriting the existing scorer in ML — separate track (there's already an Nx skeleton in
lib/microwaveprop/propagation/model.ex). This plan is about better features; ML can consume them later. - Changing the grid resolution — 0.125° stays.
- Real-time radar-based scoring without a model fallback — the existing HRRR-based scorer continues to drive the live map; NEXRAD becomes an additional feature, not a replacement data source.
- Adding new bands — 902 MHz and up, same set.
- Touching terrain analysis — ITU-R P.526-16 is separately accurate and not part of this plan.