prop/lib/microwaveprop/weather/nexrad_cache.ex
Graham McIntire f1846c0a53
perf: reduce per-pod RSS and HRRR chain wall time
Telemetry showed the application-master process holding ~830 MiB of
terms from warm_grid_cache_from_latest_profile — the data lives in
the app master's heap and never GCs because the process is idle.
Running it in a Task.start lets the terms die with the task.

Mark GridCache, MrmsCache, NexradCache, and ScoreCache ETS tables
:compressed. The scored-band-map and HRRR grid data are map-heavy;
compression trims hundreds of MiB at a few percent CPU cost.

Memoise HRRR .idx responses in Microwaveprop.Cache. Published idx
files are immutable for a model run, but the hourly chain re-fetches
the same URL dozens of times across forecast hours. Cuts ~10s per
repeat out of hrrr_fetch_idx.

Force a garbage collect at the end of HrrrFetchWorker.perform to
reclaim the refc binary heap held from GRIB2 ranges before the Oban
producer hands the process its next job.
2026-04-19 14:56:48 -05:00

79 lines
2.5 KiB
Elixir
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

defmodule Microwaveprop.Weather.NexradCache do
@moduledoc """
Node-local ETS cache of decoded NEXRAD n0q composite reflectivity frames.
Keyed by a 5-minute rounded timestamp. Stores the raw pixel buffer + image
width so per-point rain-cell extraction can skip the HTTP fetch + PNG decode
(which take 1-5 seconds for a ~5 MB CONUS-wide image).
The cache is populated by `Microwaveprop.Weather.NexradClient.fetch_rain_cells/4`
on its first call per 5-min window, then reused by every concurrent click
until the window rolls over.
"""
use GenServer
@table :nexrad_frame_cache
# Each frame is ~66 MB (12,200 × 5,400 palette-index bytes). 20 frames
# caps at ~1.3 GB — comfortable headroom under a 2 GB pod limit, and
# enough to hit on the common backfill pattern where a worker chews
# through contacts in one timestamp window before moving on. Adjust
# if pod memory limits change or the n0q geometry changes.
@max_entries 20
@type pixels :: binary()
@type width :: pos_integer()
@spec start_link(keyword()) :: GenServer.on_start()
def start_link(opts) do
GenServer.start_link(__MODULE__, opts, name: __MODULE__)
end
@spec fetch(DateTime.t()) :: {:ok, pixels(), width()} | :miss
def fetch(rounded_ts) do
case :ets.lookup(@table, rounded_ts) do
[{_, pixels, width}] -> {:ok, pixels, width}
[] -> :miss
end
end
@spec put(DateTime.t(), pixels(), width()) :: :ok
def put(rounded_ts, pixels, width) do
:ets.insert(@table, {rounded_ts, pixels, width})
enforce_size_cap()
:ok
end
defp enforce_size_cap do
size = :ets.info(@table, :size)
if size > @max_entries do
# Drop the oldest-by-key entries until we're back under the cap.
# O(n log n) but n is ~20, so this runs in microseconds.
excess = size - @max_entries
@table
|> :ets.tab2list()
|> Enum.sort_by(fn {ts, _, _} -> DateTime.to_unix(ts) end)
|> Enum.take(excess)
|> Enum.each(fn {ts, _, _} -> :ets.delete(@table, ts) end)
end
end
@spec prune_older_than(DateTime.t()) :: non_neg_integer()
def prune_older_than(cutoff_ts) do
match_spec = [{{:"$1", :_, :_}, [{:<, :"$1", {:const, cutoff_ts}}], [true]}]
:ets.select_delete(@table, match_spec)
end
@spec clear() :: :ok
def clear do
:ets.delete_all_objects(@table)
:ok
end
@impl true
def init(_opts) do
:ets.new(@table, [:set, :named_table, :public, :compressed, read_concurrency: true])
{:ok, %{}}
end
end