PromEx.Plugins.Oban counts job exceptions as a Prometheus counter but
does not tell us which worker/queue failed, whether the job is about to
be retried, or the exception class. This adds a telemetry handler on
[:oban, :job, :exception] that:
- emits a single structured Logger.error with worker, queue, job_id,
attempt, max_attempts, retry_exhausted?, kind, reason class, and
args-keys summary (never values — args can contain PII/tokens);
- dispatches opt-in per-worker recovery callbacks
(on_permanent_failure/1 when retries exhausted,
on_transient_failure/1 otherwise);
- catches callback failures so a bug in one worker's recovery path
cannot detach the handler or break logging for other jobs;
- does not emit a second telemetry event, so PromEx counts are not
double-counted.
Handler is attached from Application.start/2 after the Oban supervisor
starts.