When the Rust hrrr-point-worker marks a task as `failed` (upstream
`idx HTTP 404` or wgrib2 decode error on old archive formats), the
corresponding contacts sit at `hrrr_status = :queued` until the
generic 3-day stale-queued reconcile flips them. Speed that up: on
every BackfillEnqueue tick, join contacts to `hrrr_fetch_tasks` by
rounded HRRR hour and flip any `:queued` contact whose cycle is
permanently `failed` to `:unavailable` immediately.
Motivation: a handful of 2018-08-04/05 cycles (t20z/t21z/t15z/t17z)
are genuinely missing `wrfsfcf`/`wrfprsf` files on the NOAA S3
archive — other hours on the same dates are fine, so this is an
upstream data gap rather than a structural pre-2019 issue. Users
see an accurate "Partial" marker on those contacts instead of a
three-day `:queued` stall.
Tested: 4 new cases covering rounding, failed-vs-done, and the
`hrrr not in types` skip branch; full suite 2283/0.
Three bugs were letting mechanism/terrain/radar contacts ping-pong
between :complete and :queued on every backfill cron:
1. enqueue_for_contact/2 now filters out types that are already
:complete on the contact before building jobs. Previously
mark_status!/3 demoted every non-empty jobs_by_type entry back to
:queued, overwriting the :complete that workers had just set.
2. reconcile_stale_queued/1 now also flips mechanism_status and
radar_status back to :complete when the companion data
(propagation_mechanism / contact_common_volume_radar) already
exists. Drains the existing ~14k row backlog on the first post-
deploy tick.
3. IemClient.parse_iemre_json/1 accepts "" and nil so a 200-with-
empty-body IEMRE response doesn't FunctionClauseError the worker
through four retry attempts.
The Contact list's 'Partial' badge was sticky: once a worker gave up on
a contact (upstream 404, no nearby station, IEM throttling) or the
stale-queued reconciler flipped it to :unavailable after 3 days, the
backfill cron never looked at it again. Upstream archives heal over
time — IEM backfills gaps, new ASOS stations come online, Rust worker
gains capability — so rows stayed 'Partial' that shouldn't have.
Changes:
* BackfillEnqueueWorker.type_filter now also matches <status>_status =
:unavailable when contacts.updated_at is older than a 24 h cooldown.
enqueue_for_contact is idempotent: if there's still nothing to
fetch it'll flip the status to :complete on the empty-jobs branch;
if new data shows up it builds fresh jobs.
* reconcile_stale_queued_to_unavailable now bumps updated_at when
flipping to :unavailable so the cooldown starts fresh — otherwise
the main loop would pick it right back up on the same cron tick
(stale updated_at < any reasonable cooldown) and immediately
transition :unavailable → :complete, defeating the :unavailable
signal entirely.
Also:
* Drop ScoresFile.point_score/3, dead since read_point moved to the
pread fast path in commit bd3b114.
* Clear Microwaveprop.Cache in WeatherGridTest setup. The new
ProfilesFile read cache is keyed by (base_dir, valid_time); tests
in this file reuse DateTime.utc_now() truncated to seconds and
share the default base_dir, so a prior test's cached read leaks
into the next one's expectations.
10K+ HRRR-pending contacts were starving the 200-ish NARR candidates
out of the 500 per-run slot each 30 min because NARR candidates sort
last by the :hrrr_status priority. Drop the cap so every eligible
contact gets enqueued on each cron run; queue concurrency still paces
the actual fetches. Worker keeps supporting an explicit "limit" arg
for ad-hoc runs and tests.
Contacts whose weather/hrrr/terrain/iemre status has been :queued for
more than 3 days (144 cron cycles) are flipped to :unavailable. The
set_enrichment_status!/3 helper is a no-op when the status is unchanged,
so a stuck row retains a stale updated_at — which is exactly the signal
we need for 'the cron has already tried and the data doesnt exist
upstream' (dead ASOS station, pre-2014 HRRR, out-of-grid IEMRE). Clears
~1,492 pre-existing stuck contacts and stops them from churning the
backfill queue.
The three BackfillEnqueueWorkerTest cases were timing out because the
default types list includes :era5, which cascades through inline Oban
into Era5Client.poll_and_download where Process.sleep blocks for the
full 60s test timeout. Era5Client uses bare Req.get with no plug hook,
so it can't be stubbed via Req.Test the way the other clients are. Pass
explicit non-ERA5 types in the three affected cases — ERA5 has its own
coverage and these tests don't assert anything era5-specific.
Also replace two `length(results) > 0` checks in asos_nudge_test with
`results != []` to silence credo warnings.