towerops/lib/towerops/error_tracker_ignorer.ex
Graham McIntire a99c327446 fix: apply Oban Pro 1.7 follow-up cleanup + filter transient DB noise
The 1.7 upgrade migration added the new oban_workflows tracking but left
behind legacy state from earlier versions that the upgrade guide
strongly recommends removing
(https://oban.pro/docs/pro/v1-7.html):

  - The uniq_key / partition_key GENERATED ALWAYS columns. The Smart
    engine no longer reads them; uniqueness and partition lookups now
    use expression indexes on meta. Keeping the columns imposes write
    overhead on every oban_jobs insert/update and was contributing to
    intermittent transaction-rollback errors.
  - The oban_jobs_*_old workflow/chain indexes. v1.7 renamed the
    originals and built expression-based replacements; the _old indexes
    are dead weight that still slow down writes.

Also extend ErrorTrackerIgnorer to drop transient
DBConnection.ConnectionError noise ("connection is closed",
"transaction rolling back"). These come from the Repo's intentional
queue-target / disconnect-on-error cycling that recovers from stale SSL
connections — Ecto retries automatically, so logging adds noise without
value.
2026-05-01 16:04:05 -05:00

46 lines
1.6 KiB
Elixir

defmodule Towerops.ErrorTrackerIgnorer do
@moduledoc """
Drops benign errors from ErrorTracker that occur during K8s pod shutdown
or transient DB connection cycling.
Matches the same patterns as `Towerops.HoneybadgerFilter` and
`Towerops.LoggerFilters.drop_shutdown_errors/2`.
"""
@behaviour ErrorTracker.Ignorer
@impl true
def ignore?(%ErrorTracker.Error{reason: reason} = error, _context) do
shutdown_error?(reason) or transient_db_error?(error)
end
defp shutdown_error?(reason) when is_binary(reason) do
String.contains?(reason, "port_died") or
String.contains?(reason, "write_failed") or
String.contains?(reason, "epipe") or
String.contains?(reason, "broken pipe")
end
defp shutdown_error?(_), do: false
# Transient DB connection drops surface as DBConnection.ConnectionError. The
# Repo's aggressive idle/queue cycling (queue_target/queue_interval,
# disconnect_on_error_codes) intentionally tears down stale SSL connections,
# which can interrupt queries mid-flight. Common reasons we don't want to
# log: "connection is closed..." and "transaction rolling back". The query
# is retried by the caller; logging adds noise without value.
@transient_db_phrases [
"connection is closed",
"transaction rolling back"
]
defp transient_db_error?(%ErrorTracker.Error{reason: reason, kind: kind}) do
binary_reason = to_string(reason || "")
binary_kind = to_string(kind || "")
String.contains?(binary_kind, "DBConnection.ConnectionError") and
Enum.any?(@transient_db_phrases, &String.contains?(binary_reason, &1))
end
defp transient_db_error?(_), do: false
end