perf(db): add (valid_time, lat, lon) composite index on hrrr[_native]_profiles

Weather.find_nearest_hrrr/3 and Weather.find_nearest_native_profile/3
run a three-range scan gated first by a ±1h valid_time window, then
by a ±0.07° lat/lon box. The existing unique `(lat, lon, valid_time)`
index is only usable on its `lat` prefix for these queries, forcing
Postgres to scan wide lat slices before applying the valid_time
filter — expensive on the Turing Pi 2 node.

Add a `(valid_time, lat, lon)` composite so valid_time leads,
collapsing the search to a narrow btree slice inside a single
partition before the lat/lon box filter kicks in.

`hrrr_profiles` is RANGE-partitioned by valid_time, so
CREATE INDEX CONCURRENTLY is applied per partition and the resulting
indexes are attached to a parent index created ON ONLY the partitioned
table. `hrrr_native_profiles` is a regular table, so the concurrent
build runs directly on it.
This commit is contained in:
Graham McIntire 2026-04-21 13:20:22 -05:00
parent b874020a87
commit c7e7472c57
No known key found for this signature in database
GPG key ID: F4ABF488E6029E59

View file

@ -0,0 +1,78 @@
defmodule Microwaveprop.Repo.Migrations.AddHrrrProfilesValidTimeLatlonIndex do
use Ecto.Migration
# Matches the hot range-scan access pattern in
# `Weather.find_nearest_hrrr/3` and `Weather.find_nearest_native_profile/3`:
# WHERE valid_time BETWEEN ? AND ?
# AND lat BETWEEN ? AND ?
# AND lon BETWEEN ? AND ?
# `valid_time` leads because the ±1h window collapses the search to a
# single partition and a narrow btree slice; lat/lon then filter the
# ~9×9 tile grid within that bucket. The existing unique index
# `(lat, lon, valid_time)` services upsert dedup but its first column
# is lat — it can't serve a three-range scan that is selective only
# on valid_time.
#
# `hrrr_profiles` is partitioned by `valid_time`, so CREATE INDEX
# CONCURRENTLY cannot run directly on the partitioned parent. We use
# the PG 11+ pattern: create an invalid index on ONLY the parent,
# concurrently build a matching index on each partition, then ATTACH
# each partition index to the parent — Postgres validates the parent
# once every child is attached.
#
# `hrrr_native_profiles` is a regular table, so the standard
# `create index(..., concurrently: true)` form applies.
@disable_ddl_transaction true
@disable_migration_lock true
def up do
# Parent index (invalid until partitions attach).
execute """
CREATE INDEX IF NOT EXISTS hrrr_profiles_valid_time_lat_lon_index
ON ONLY hrrr_profiles (valid_time, lat, lon)
"""
for partition <- hrrr_profile_partitions() do
child_idx = "#{partition}_valid_time_lat_lon_index"
execute """
CREATE INDEX CONCURRENTLY IF NOT EXISTS #{child_idx}
ON #{partition} (valid_time, lat, lon)
"""
execute """
ALTER INDEX hrrr_profiles_valid_time_lat_lon_index
ATTACH PARTITION #{child_idx}
"""
end
# Native profiles: not partitioned, so concurrently works directly.
execute """
CREATE INDEX CONCURRENTLY IF NOT EXISTS hrrr_native_profiles_valid_time_lat_lon_index
ON hrrr_native_profiles (valid_time, lat, lon)
"""
end
def down do
execute "DROP INDEX IF EXISTS hrrr_native_profiles_valid_time_lat_lon_index"
execute "DROP INDEX IF EXISTS hrrr_profiles_valid_time_lat_lon_index"
end
defp hrrr_profile_partitions do
{:ok, %{rows: rows}} =
repo().query(
"""
SELECT child.relname
FROM pg_inherits
JOIN pg_class parent ON pg_inherits.inhparent = parent.oid
JOIN pg_class child ON pg_inherits.inhrelid = child.oid
WHERE parent.relname = 'hrrr_profiles'
ORDER BY child.relname
""",
[]
)
Enum.map(rows, fn [name] -> name end)
end
end