Observed in prod: after the CDS cap guard shipped, jobs started landing but one leg per tile-month would routinely stay in 'accepted' state for 16+ hours while the other completed in ~90 minutes. CDS's per-user queue fairness apparently serializes competing leg requests and the single-level ones got stranded behind higher-priority work. Add an age-based stuck detector in Era5PollWorker: if a row has been submitted for more than 4 hours and neither leg is fully done, treat both legs as terminal, clean them up from CDS, and re-enqueue the submit. Re-uses the resubmit_vanished path. 4h is ~8× the normal completion time so transient queue variance doesn't trip it, but tight enough that stuck jobs self-heal in a working session. |
||
|---|---|---|
| .. | ||
| microwaveprop | ||
| microwaveprop_web | ||
| mix/tasks | ||
| support | ||
| test_helper.exs | ||