The single Era5MonthBatchWorker pinned an Oban slot for the full 30-45
min a CDS month-tile takes to assemble, and lost the work entirely on a
rolling deploy because the in-flight Task died with the pod. This splits
the flow into two tiny workers and persists the CDS job IDs so deploys
survive.
New pipeline:
1. Era5SubmitWorker (:era5_submit queue) — POSTs both CDS requests in
parallel, writes an `era5_cds_jobs` row with the returned job IDs,
and enqueues an Era5PollWorker scheduled +5 min. ~1s of real work.
Short-circuits when the month-tile is already cached or when a row
already exists (in-flight from a previous attempt).
2. Era5PollWorker (:era5_poll queue) — reads the row, calls
Era5Client.check_status/1 for both CDS job IDs, and:
- returns {:snooze, 300} if either job is still running (Oban
re-schedules without counting an attempt and releases the slot
immediately — a pod can keep dozens of tile-months in flight
without pinning workers)
- streams both GRIB files to disk via Req into: File.stream!,
decodes via Wgrib2.extract_grid_messages_from_file, bulk-inserts
via Era5BatchClient.decode_and_insert/6, deletes the row, and
DELETEs both completed jobs from CDS to free server-side quota
- if either leg CDS-reports failed, deletes the row + both CDS
jobs and returns {:error, reason}
Era5Client gains four testable building blocks:
submit_job/2 (bare POST → {:ok, job_id})
check_status/1 (GET → :running | {:done, src} | {:failed, reason})
download_source_to_file/3 (streams {:url, href} or writes {:body, bin})
delete_job/1 (DELETE /jobs/:id, treats 200/202/204/404 as :ok)
All Req calls now route through `era5_req_options` so tests can stub
CDS responses via Req.Test.stub(Era5Client, fn).
Era5MonthBatchWorker is retained as a thin forwarder to Era5SubmitWorker
so any jobs already in the :era5_batch queue on prod pods drain cleanly
on the next rolling deploy. Safe to delete in a follow-up.
Adds era5_cds_jobs table with a unique index on
(year, month, tile_lat, tile_lon) so duplicate submits collapse.
New queue config in runtime.exs:
era5_submit: local_limit 4, rate_limit 30/hour (burst protection)
era5_poll: local_limit 20 (polls are cheap GETs)
era5_batch: kept at 1 for legacy job drain, delete next cycle
28 lines
1.2 KiB
Elixir
28 lines
1.2 KiB
Elixir
defmodule Microwaveprop.Repo.Migrations.CreateEra5CdsJobs do
|
|
use Ecto.Migration
|
|
|
|
# Tracks in-flight CDS requests for the split Era5SubmitWorker / Era5PollWorker
|
|
# design. A row is created when Era5SubmitWorker successfully submits both
|
|
# CDS jobs, and is deleted when Era5PollWorker downloads the results and
|
|
# inserts profiles (or when either leg permanently fails). Persisting the
|
|
# CDS job IDs lets a pod restart resume polling instead of losing the work.
|
|
def change do
|
|
create table(:era5_cds_jobs, primary_key: false) do
|
|
add :id, :binary_id, primary_key: true
|
|
add :year, :integer, null: false
|
|
add :month, :integer, null: false
|
|
add :tile_lat, :integer, null: false
|
|
add :tile_lon, :integer, null: false
|
|
add :single_job_id, :string, null: false
|
|
add :pressure_job_id, :string, null: false
|
|
add :submitted_at, :utc_datetime, null: false
|
|
add :last_polled_at, :utc_datetime
|
|
add :poll_count, :integer, default: 0, null: false
|
|
|
|
timestamps(type: :utc_datetime)
|
|
end
|
|
|
|
# Uniqueness so duplicate submits for the same tile-month collapse.
|
|
create unique_index(:era5_cds_jobs, [:year, :month, :tile_lat, :tile_lon])
|
|
end
|
|
end
|