Rust workers (prop-grid-rs) that die mid-claim (SIGKILL, OOM, node drain) leave grid_tasks rows stuck in status='running' forever, because claim_next uses FOR UPDATE SKIP LOCKED and nothing resets the orphan. The /status page then shows a permanent spinner with stale f00/f10/f18 badges — currently 21 rows claimed as far back as 2026-04-20. GridTaskEnqueuer.reclaim_stale_running/1 flips rows whose claimed_at is older than 15 minutes back to 'queued'. Rows that have already burned through 5 claim/reclaim cycles become 'failed' with a reclaim-orphan error so the next hourly seed can replace them. Wired into PropagationGridWorker.seed_chain/0 so it runs every :05 cron tick before new rows are seeded. Also rename the status panel "Retrying" column to "Failed" — it was always showing terminal `failed` rows, never retrying ones. |
||
|---|---|---|
| .. | ||
| components | ||
| controllers | ||
| live | ||
| plugs | ||
| endpoint.ex | ||
| gettext.ex | ||
| live_table_footer.ex | ||
| live_table_resource.ex | ||
| metrics_plug.ex | ||
| router.ex | ||
| skew_t.ex | ||
| telemetry.ex | ||
| user_auth.ex | ||