`Pskr.Recalibrator.run/0` reads `pskr_calibration_samples`, bins
each sample per (band × feature), and writes spot-density stats to
`pskr_feature_bins` so an operator can read whether a feature
actually discriminates propagation at the threshold granularity
the scorer uses.
Bins, not regression: `BandConfig` already encodes scoring as
discrete thresholds, so the bin output matches that shape and an
operator can copy adjusted thresholds directly without translating
from regression coefficients.
Self-healing: corpus too thin ⇒ run row written with status
`skipped_insufficient_data` and the analysis is a no-op until next
fire. `min_total_samples = 1000` (≈ 4-5 days of CONUS PSKR
activity); per-band threshold is 100. Both surface in the run row's
`notes`.
Auto-applies nothing. Weight changes still go through human review
of `BandConfig.@band_configs` and a code commit. The recalibrator
is a read-only analyst that stays out of the production scoring
path.
Features binned (matching the scorer's discriminating fields):
* pwat_mm — humidity U-shape candidate
* hpbl_m — boundary layer (mechanism vs scoring re-eval)
* min_refractivity_gradient — refractivity threshold validation
* surface_pressure_mb — pressure-front proxy
* kp_index — aurora boost magnitude tuning
Schema: two tables.
* `pskr_recalibration_runs` — one row per fire with corpus
stats, status, notes
* `pskr_feature_bins` — one row per (run, band, feature, bin)
with sample_count, spot_count_total/avg/p50/p90
Cron: `0 4 * * 0` (Sundays 04:00 UTC, off-peak, post-climatology).
Manual reruns enqueue with no args.
Tests cover the empty-corpus skip path, sub-threshold totals,
per-band threshold gating, the actual bin emission, nil-feature
handling, spot-count averaging, and the always-records-a-run
audit invariant. 8 new tests, 3282 total passing.
Backfill pipeline untouched.
54 lines
1.7 KiB
Elixir
54 lines
1.7 KiB
Elixir
defmodule Microwaveprop.Pskr.RecalibrationRun do
|
|
@moduledoc """
|
|
One execution of the PSKR-driven recalibration analysis.
|
|
|
|
A row is written every time `Pskr.Recalibrator.run/0` fires
|
|
(weekly via `Workers.PskrRecalibrationWorker`). The status field
|
|
separates an actually-completed analysis from runs that bailed
|
|
early because the corpus was too thin to produce useful bins.
|
|
|
|
Status values:
|
|
* `"completed"` — analysis ran and `pskr_feature_bins` rows
|
|
were written for at least one (band, feature)
|
|
* `"skipped_insufficient_data"` — corpus had < threshold
|
|
samples, no bins emitted
|
|
* `"failed"` — exception during analysis (notes carries
|
|
reason)
|
|
"""
|
|
use Ecto.Schema
|
|
|
|
import Ecto.Changeset
|
|
|
|
@primary_key {:id, :binary_id, autogenerate: true}
|
|
@foreign_key_type :binary_id
|
|
|
|
schema "pskr_recalibration_runs" do
|
|
field :run_at, :utc_datetime
|
|
field :status, :string
|
|
field :sample_count, :integer, default: 0
|
|
field :band_count, :integer, default: 0
|
|
field :earliest_sample, :utc_datetime
|
|
field :latest_sample, :utc_datetime
|
|
field :avg_spots_per_sample, :float
|
|
field :notes, :string
|
|
|
|
has_many :feature_bins, Microwaveprop.Pskr.FeatureBin, foreign_key: :run_id
|
|
|
|
timestamps(type: :utc_datetime)
|
|
end
|
|
|
|
@type t :: %__MODULE__{}
|
|
|
|
@cast_fields ~w(run_at status sample_count band_count earliest_sample latest_sample
|
|
avg_spots_per_sample notes)a
|
|
|
|
@required_fields ~w(run_at status)a
|
|
|
|
@spec changeset(t(), map()) :: Ecto.Changeset.t()
|
|
def changeset(record, attrs) do
|
|
record
|
|
|> cast(attrs, @cast_fields)
|
|
|> validate_required(@required_fields)
|
|
|> validate_inclusion(:status, ~w(completed skipped_insufficient_data failed))
|
|
end
|
|
end
|