prop/CLAUDE.md
Graham McIntire 5e9513bf02
ops(prop): aggressive HPA scale-down + CLAUDE.md guardrails
HPA: prop-grid-rs and hrrr-point-rs were sitting at 4 pods for
~16 min after the hourly chain spike — 10-min scaleDown stabilization
window plus a 1-pod-per-2-min drain policy. Both deployments stayed
visibly over-provisioned even when CPU dropped to 1m/pod.

Tighten scale-down to: 2-min stabilization (long enough that a real
multi-cycle surge doesn't immediately collapse) + Percent: 100 per
60s drain (back to min=1 within 1 minute of stabilization clearing).
Net: spike → scale up to handle queue (unchanged) → back to 1 pod
within ~3 min of spike ending. scaleUp left untouched.

CLAUDE.md updates so this doesn't bite again:

- Document `cargo clippy --all-targets -- -D warnings` as part of
  the local Rust precommit ritual. The HRDPS commits in this branch
  passed `cargo build` locally but failed CI on lints (manual_range_
  contains, iter_kv_map). Adding the recipe under Commands plus a
  Rust subsection in Key Conventions so future Rust touches don't
  ship clippy regressions.
- Added a Testing note about not anchoring relative-time fixtures on
  utc_now-truncated-to-hour — the weather_map_live one-hour-cutoff
  test (fix in 2768fc68) was a textbook example: 50% flake rate
  driven entirely by which minute of the hour the test ran in.
2026-04-29 17:44:17 -05:00

229 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Project Overview
Microwaveprop is a Phoenix 1.8 web application for the North Texas Microwave Society (NTMS) that predicts microwave radio propagation conditions (10-241 GHz). It helps amateur radio operators plan and coordinate microwave contacts by showing real-time and forecast propagation maps.
**Stack:** Elixir ~> 1.15, Phoenix 1.8 + LiveView, PostgreSQL via Ecto, Tailwind CSS v4 + daisyUI, Bandit HTTP server, Nx/Axon/EXLA for ML. esbuild for JS bundling — never npm.
### What It Does
1. **Propagation Map** (`/map`) — Full-screen Leaflet map with colored overlay showing propagation scores (0-100) per grid point across CONUS at 0.125° resolution. Includes 18-hour forecast timeline, band selector (10-241 GHz), Maidenhead grid overlay, point detail with factor breakdown and forecast sparkline, and terrain viewshed analysis.
2. **QSO Submission** (`/submit`) — Users submit microwave contacts. On submission, the system automatically enqueues weather/HRRR/terrain/IEMRE enrichment jobs for that QSO.
3. **Algorithm Documentation** (`/algo`) — Renders `algo.md` as HTML. Documents the scoring algorithm, meteorological foundations, ITU-R models, and calibration data.
4. **Hourly Scoring Pipeline** — Every hour, the PropagationGridWorker fetches HRRR forecast hours f00-f18, computes propagation scores across all bands, and pushes updates to connected clients via PubSub.
### Key Data Sources
- **HRRR** (High-Resolution Rapid Refresh) — 3 km NWP model from NOAA AWS S3. Surface + pressure level profiles. Analysis + 18-hour forecasts.
- **HRDPS** (High Resolution Deterministic Prediction System) — 2.5 km Canadian NWP model from MSC Datamart. 4×/day at 00/06/12/18Z, 48h forecasts. Coverage capped at 60°N for v1 (SRTM stops there). `HrdpsClient` + scaffolding shipped 2026-04-29; full pipeline activates once the Rust `prop-grid-rs` HRDPS branch ships. Plan: `docs/plans/2026-04-29-hrdps-canadian-prop-grid.md`.
- **ASOS** (Automated Surface Observing System) — Surface weather via Iowa Environmental Mesonet.
- **RAOB** (Radiosonde) — Upper-air soundings via IEM.
- **IEMRE** — Gridded hourly reanalysis at 0.125° resolution.
- **SRTM** — 90m terrain elevation for path analysis (local tile cache or API fallback).
- **Commercial Links** — 7 microwave links near DFW polled via SNMP at 5-min intervals.
## Commands
```bash
# Initial setup (deps, DB, assets)
mix setup
# Run dev server (live reloads automatically, no restart needed)
mix phx.server
iex -S mix phx.server # with IEx shell
# Run all tests
mix test
# Run a single test file
mix test test/microwaveprop/terrain/terrain_analysis_test.exs
# Run previously failed tests
mix test --failed
# Format code (runs Styler plugin automatically)
mix format
# Static analysis
mix credo
# Pre-commit check (compile with warnings-as-errors, unlock unused deps, format, test)
mix precommit
# Rust pre-commit (matches the Forgejo CI step). Run from rust/prop_grid_rs/.
# `cargo build` alone does NOT catch clippy errors that gate the CI build.
cd rust/prop_grid_rs && cargo clippy --all-targets -- -D warnings && cargo test --release
# Database
mix ecto.create
mix ecto.migrate
mix ecto.gen.migration migration_name_using_underscores
mix ecto.reset # drop + setup
# Run propagation grid manually
mix propagation_grid
# Assets
mix assets.build
mix assets.deploy # minified + digest
```
## Architecture
### Directory Structure
- `lib/microwaveprop/` — Business logic (contexts, schemas, workers)
- `propagation/` — Scoring engine: `scorer.ex`, `band_config.ex`, `grid.ex`, `model.ex` (ML), `freshness_monitor.ex`
- `weather/` — Data ingestion: `hrrr_client.ex`, `iem_client.ex`, `sounding_params.ex`, GRIB2 decoder
- `terrain/` — Path analysis: `terrain_analysis.ex` (ITU-R P.526-16), `viewshed.ex`, `elevation_client.ex`
- `radio/` — QSO schema, Maidenhead grid conversion
- `commercial/` — SNMP polling for commercial microwave links
- `workers/` — Oban background jobs for all data pipelines
- `lib/microwaveprop_web/` — Web layer
- `live/map_live.ex` — Main propagation map with forecast timeline
- `live/submit_live.ex` — QSO submission form
- `live/algo_live.ex` — Algorithm documentation page
- `live/qso_live/` — QSO list and detail views
- `plugs/remote_ip.ex` — X-Forwarded-For client IP extraction
- `config/` — Environment configs. `runtime.exs` has production Oban config (no backfill queues)
- `assets/js/propagation_map_hook.js` — Leaflet map, canvas heatmap, timeline, grid overlay
- `algo.md` — Full algorithm documentation with meteorological foundations
- `priv/models/` — Saved ML model weights (gitignored)
### Background Job Queues (Oban)
| Queue | Workers | Purpose |
|---|---|---|
| `propagation` | PropagationGridWorker | Hourly HRRR fetch (f00-f18) + score computation |
| `weather` | WeatherFetchWorker | ASOS/RAOB data for QSO enrichment |
| `hrrr` | HrrrFetchWorker | HRRR profiles for individual QSOs |
| `terrain` | TerrainProfileWorker | ITU-R P.526-16 terrain diffraction analysis |
| `iemre` | IemreFetchWorker | IEMRE gridded weather for QSO enrichment |
| `commercial` | PollWorker | SNMP polling of commercial links |
| `solar` | SolarIndexWorker | Daily solar indices |
| `enqueue` | QsoWeatherEnqueueWorker | Batch enqueue enrichment jobs (dev cron only) |
| `hrdps` | HrdpsGridWorker | Canadian propagation chain seed (4×/day). Dormant until Rust `prop-grid-rs` HRDPS branch ships — module exists, cron entry deferred. |
**Production** runs: propagation, commercial, solar, weather, hrrr, terrain, iemre queues. No cron backfill — enrichment is triggered by QSO submission only.
**Dev** runs the same hourly `PropagationGridWorker` + `AsosAdjustmentWorker` (10-minute nudge) cron as production.
### Scoring Algorithm
10-factor weighted composite score (0-100) per grid point per band (recalibrated 2026-04-11 via gradient descent):
| Factor | Weight | Source |
|---|---|---|
| Rain | 13.6% | HRRR precipitation, ITU-R P.838-3 |
| Humidity | 12.4% | HRRR surface temp + dewpoint |
| PWAT | 11.3% | HRRR precipitable water (column-integrated) |
| Season | 11.1% | Month, band-dependent (inverted for 24G+) |
| Refractivity | 10.5% | Native HRRR dM/dh (10-50m), fallback to pressure-level (250m) |
| Pressure | 10.3% | HRRR surface pressure (frontal activity proxy) |
| T-Td depression | 9.8% | HRRR surface |
| Sky cover | 8.0% | HRRR cloud cover |
| Wind | 8.0% | HRRR 10m wind |
| Time of day | 5.0% | Solar-time adjusted diurnal cycle |
Humidity effect reverses by frequency: beneficial at 10 GHz (refractivity), harmful at 24+ GHz (absorption).
### Terrain Analysis (ITU-R P.526-16)
- Knife-edge diffraction loss (Eq. 31)
- Deygout 3-edge method for multiple obstacles
- Dynamic k-factor from HRRR refractivity gradient
- Profiles stored per QSO path (58k+ paths analyzed)
### ML Model (Nx/Axon/EXLA)
Skeleton feed-forward network for future propagation prediction:
- 13 features: 8 atmospheric + 4 cyclical temporal + 1 log-frequency
- Architecture: Dense(64) → Dropout(0.2) → Dense(32) → Dropout(0.1) → Dense(1, sigmoid)
- Weights saved to `priv/models/propagation_v1.nx`
- Not yet trained — scaffolding only
### Data Flow
```
HRRR (NOAA S3) → PropagationGridWorker (hourly, f00-f18)
→ store hrrr_profiles → compute scores → upsert propagation_scores
→ PubSub broadcast → MapLive → JS hook → canvas heatmap + timeline
User submits QSO → enqueue_for_qso() → weather/hrrr/terrain/iemre workers
→ enrich QSO with atmospheric + terrain data
```
## Key Conventions
### Phoenix / LiveView
- LiveView templates must start with `<Layouts.app flash={@flash} ...>` wrapping all content
- Use `<.icon name="hero-x-mark" class="w-5 h-5"/>` for icons, never Heroicons modules
- Use `<.input>` from core_components for form inputs
- Use `to_form/2` for forms, never pass changesets directly to templates
- Use LiveView streams for collections, never `phx-update="append"`/`"prepend"`
- Avoid LiveComponents unless strongly justified
- LiveView names use `Live` suffix: `MicrowavepropWeb.ThingLive`
- Use `<.link navigate={...}>` / `<.link patch={...}>`, never `live_redirect`/`live_patch`
- Router scope aliases prefix automatically; don't add redundant aliases
### HEEx Templates
- Use `{...}` for attribute interpolation, `<%= %>` only for block constructs (if/for/cond) in tag bodies
- Class lists must use `[...]` syntax for conditional classes
- Use `<%!-- comment --%>` for HTML comments
- Use `phx-no-curly-interpolation` on tags containing literal curly braces
- Never use `<% Enum.each %>`, always use `<%= for item <- @collection do %>`
### JS / CSS
- Assets are bundled via esbuild and Tailwind mix tasks — never use npm, npx, or Node.js directly
- Tailwind v4: no `tailwind.config.js`, uses `@import "tailwindcss" source(none)` syntax in `app.css`
- Never use `@apply` in CSS
- No inline `<script>` tags; use colocated JS hooks with `.` prefix names
- Vendor deps must be imported into `app.js`/`app.css`, no external script `src` or link `href`
- LiveView hooks that push `map_bounds` (or any viewport-scoped data request) MUST implement `reconnected()` and a `document.visibilitychange` listener. Server assigns reset on reconnect while the hook keeps its stale client-side grid lookup — Leaflet tiles regenerated outside that extent paint blank. Both callbacks should call `invalidateSize()` + the bounds-push method; clean up the listener in `destroyed()`.
### Elixir
- Write idiomatic Elixir: use pattern matching, pipe operator, and the standard library
- Styler auto-formats on `mix format` (alias sorting, pipe chains, moduledoc enforcement, etc.)
- No index access on lists (`mylist[i]`); use `Enum.at/2`
- Bind results of `if`/`case`/`cond` blocks to variables (immutable rebinding)
- Never nest multiple modules in one file
- Use `struct.field` not `struct[:field]` (structs don't implement Access)
- Predicate functions: `thing?` not `is_thing` (reserve `is_` for guards)
- Use `Req` for HTTP requests, never httpoison/tesla/httpc
- **Always log anything that would normally be swallowed.** Any place that drops `{:exit, _}` from `Task.async_stream` / `Task.Supervisor.async_stream_nolink`, ignores `{:error, _}` clauses, or otherwise discards a failure path MUST `Logger.error`/`Logger.warning` with enough context (input, key, reason) to debug from k8s logs. Silently dropping a tuple in a `flat_map`/`reduce` clause is a bug — production crashes inside an async task otherwise vanish (the path-calculator "0 / 9 HRRR points" outage was exactly this). LiveView `start_async` must have a matching `handle_async(slot, {:exit, reason}, socket)` clause that logs.
### Ecto
- All schemas use `@primary_key {:id, :binary_id, autogenerate: true}`
- Preload associations in queries when accessed in templates
- `field :name, :string` for both string and text columns
- Use `Ecto.Changeset.get_field/2` to access changeset fields
- Programmatic fields (e.g. `user_id`) must not be in `cast` calls
- Generate migrations with `mix ecto.gen.migration`
- Upserts via `on_conflict: {:replace_all_except, [:id, ...]}` with `conflict_target`
### Testing
- Use `start_supervised!/1` for process cleanup
- Use `Process.monitor/1` + `assert_receive {:DOWN, ...}` instead of `Process.sleep`
- Use `LazyHTML` selectors for DOM assertions, never raw HTML matching
- Test against element IDs defined in templates
- Stub HTTP calls with `Req.Test.stub` in test setup
- Oban runs inline in test (`testing: :inline`)
- FreshnessMonitor disabled in test via `config :microwaveprop, start_freshness_monitor: false`
- **Don't anchor relative-time test fixtures on `DateTime.utc_now()` truncated to the hour.** A test that computes `recent_t = (top_of_hour) - 30min` is 30-90 min old of real `utc_now()` depending on what minute the test runs at; any production code that compares against `utc_now()` (e.g. a "drop entries older than 1 h" filter) flakes 50% of the time. Anchor at `top_of_hour` (or use a tiny offset like `now - 1s`) so the test is stable across clock minutes.
### Rust (`prop_grid_rs`)
- **Always run `cargo clippy --all-targets -- -D warnings` before committing Rust changes.** The Forgejo CI step runs clippy with warnings-as-errors and a passing local `cargo build` does NOT catch lints (e.g. `manual_range_contains`, `iter_kv_map`, `useless_format`). A clippy failure blocks the image build and the deploy. The Rust precommit recipe is in the `Commands` section.
- The CI build pipeline runs `cargo clippy --all-targets -- -D warnings && cargo test --release && cargo build --release`. Match this locally — `cargo build` alone is not enough.
### Deployment
- Production runs on Kubernetes (`home-cluster` context, `prop` namespace) — use `kubectl -n prop ...`
- Deployment name `prop` (3 replicas). Inspect with `kubectl -n prop get pods`, `kubectl -n prop logs deploy/prop`, `kubectl -n prop exec deploy/prop -- ...`
- Push to `github` remote for code. The `dokku` remote is obsolete — do not push there.
- Production Oban config in `config/runtime.exs` — no backfill cron, enrichment triggered by QSO submission
- Shared NFS mount `/data` in production container (server: node3 at `10.0.15.103:/data`). SRTM tiles live at `/data/srtm`.