Two bugs surfacing 13 h after Stream A cutover:
1. prop-grid-rs pods OOMKilled 16–18× overnight on the analysis step.
f00's native-level GRIB2 (~530 MB + wgrib2 working set) runs on top
of 2 concurrent forecast tasks, briefly reaching ~2 Gi. 2 Gi limit
was too tight — 3 Gi plus parallelism 3→2 gives the analysis step
room without eliminating forecast-lane headroom.
2. hrrr-point-rs pod is CrashLoopBackOff because its container is
running image main-1776640915-65f7963 (pre-Stream-C) which doesn't
contain the hrrr_point_worker binary. The grid-rs CI run for
commit 4fefb81 failed and no newer image got published. Touch the
Dockerfile to re-trigger the workflow so flux picks up a new tag
with both binaries. No functional Dockerfile change, just a
docstring update so the path-filter kicks.
Forecast + analysis both claim per FOR UPDATE SKIP LOCKED, so
reducing per-pod parallelism doesn't break the chain — the two
replicas cover 4 slots cluster-wide, which still drains f01..f18 in
~5 min.
69 lines
2.7 KiB
Docker
69 lines
2.7 KiB
Docker
# syntax=docker/dockerfile:1.7
|
|
#
|
|
# Multi-arch build for `prop-grid-rs`. wgrib2 isn't in Debian apt — the
|
|
# Elixir image builds it from source in its own stage. Rather than
|
|
# rebuild it here (~5 min per CI run), copy the binary + libg2c out of
|
|
# the already-published Elixir image.
|
|
#
|
|
# Runtime uid 65534 matches the fsGroup on the NFS scores mount.
|
|
#
|
|
# This image ships two binaries: `prop-grid-rs` (default ENTRYPOINT,
|
|
# drains grid_tasks) and `hrrr-point-worker` (k8s `command:` override,
|
|
# drains hrrr_fetch_tasks).
|
|
ARG WGRIB2_IMAGE=git.mcintire.me/graham/prop:latest
|
|
|
|
FROM ${WGRIB2_IMAGE} AS wgrib2-src
|
|
|
|
FROM rust:1.94-trixie AS builder
|
|
WORKDIR /src
|
|
|
|
# Cache deps: copy manifests first so dependency changes alone don't
|
|
# invalidate the source-layer cache.
|
|
COPY Cargo.toml Cargo.lock* ./
|
|
COPY src ./src
|
|
COPY tests ./tests
|
|
|
|
# BuildKit cache mounts persist cargo's registry + git + target
|
|
# directories across CI runs, so dependency changes are the only
|
|
# thing that triggers a full rebuild. Lint + test gate the image
|
|
# build — a regressed scorer or a clippy warning never produces a
|
|
# pushable image. `target/` is cached but the final binary is
|
|
# copied out into a non-cached path so the runtime stage can reach
|
|
# it (COPY --from= can't read from a cache mount).
|
|
RUN --mount=type=cache,target=/usr/local/cargo/registry,sharing=locked \
|
|
--mount=type=cache,target=/usr/local/cargo/git,sharing=locked \
|
|
--mount=type=cache,target=/src/target,sharing=locked \
|
|
rustup component add clippy \
|
|
&& cargo clippy --all-targets -- -D warnings \
|
|
&& cargo test --release \
|
|
&& cargo build --release --bin worker --bin hrrr_point_worker \
|
|
&& cp target/release/worker /usr/local/bin/worker \
|
|
&& cp target/release/hrrr_point_worker /usr/local/bin/hrrr_point_worker
|
|
|
|
# Runtime image: slim debian + the shared libs wgrib2 links against.
|
|
# libgfortran5/libaec0/zlib1g/libpng16-16t64/libopenjp2-7 mirror the
|
|
# Elixir runtime stage.
|
|
FROM debian:trixie-slim AS runtime
|
|
RUN apt-get update && apt-get install -y --no-install-recommends \
|
|
ca-certificates \
|
|
tini \
|
|
libgfortran5 \
|
|
libaec0 \
|
|
zlib1g \
|
|
libpng16-16t64 \
|
|
libopenjp2-7 \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
COPY --from=wgrib2-src /usr/local/bin/wgrib2 /usr/local/bin/wgrib2
|
|
COPY --from=wgrib2-src /usr/local/lib/libg2c.so* /usr/local/lib/
|
|
RUN ldconfig
|
|
|
|
COPY --from=builder /usr/local/bin/worker /usr/local/bin/prop-grid-rs
|
|
COPY --from=builder /usr/local/bin/hrrr_point_worker /usr/local/bin/hrrr-point-worker
|
|
|
|
USER 65534:65534
|
|
# Default entrypoint is the grid chain worker. Override with `command:`
|
|
# in the k8s manifest (or `docker run ... hrrr-point-worker`) for the
|
|
# per-QSO point worker.
|
|
ENTRYPOINT ["/usr/bin/tini", "--"]
|
|
CMD ["/usr/local/bin/prop-grid-rs"]
|