Adds <meta name='robots'>, twitter:card/title/description meta tags, and WebP favicon reference to the root layout. Fixes title duplication (TowerOps | TowerOps -> Home | TowerOps) by passing page_title from the page controller.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
L2: Remove 2-second Process.sleep from post_startup — the try/catch
already handles transient noproc errors from the TaskSupervisor.
L3: Remove Process.sleep from delete_device — in-flight jobs already
handle the race condition via verify_polling_assignment_unchanged.
The blocking sleep was unnecessary and delayed web requests by 500ms.
Prod paired the Oban Pro Smart engine + unique workers with core
Oban.Plugins.Lifeline, whose rescue has no unique_violation handling. An
orphaned executing CheckExecutorWorker job rescued back to available collides
with its already-scheduled successor on oban_jobs_unique_index (23505),
crashing the plugin every 60s and leaving ~150 orphans stuck. Switch prod to
Oban.Pro.Plugins.DynamicLifeline, which repairs the conflict via
Smart.clear_uniq_violation. Dev keeps core Lifeline (Basic engine).
Add a shallow /health/live endpoint that does not touch db/redis. k8s liveness
and startup probes now target it so a transient dependency outage can't kill or
block boot of a healthy pod; readiness keeps the deep /health to gate the load
balancer. (deployment.yaml probe/replica change pushed separately, after the
image carrying /health/live is live.)
Also drop unused POSTGRES_* keys from the secrets example.
The @session_options module attribute used Application.compile_env, which
baked compile-time placeholders into the endpoint while runtime.exs set
the real values from SESSION_SIGNING_SALT / SESSION_ENCRYPTION_SALT env
vars. Phoenix detected the mismatch and refused to start (failing migrate
Job in k8s).
- Remove hardcoded salts from config/config.exs (no compile-time binding)
- Add stable per-env salts in dev.exs / test.exs so local + CI don't need
the env vars
- Split static cookie opts (@static_session_options) from runtime-resolved
opts in endpoint.ex; expose session_options/0 as an MFA tuple in socket
connect_info so LiveView decodes sessions with the same runtime salts
- New ToweropsWeb.Plugs.RuntimeSession wraps Plug.Session, fetches salts
from app env on first request, and caches the initialized opts in
:persistent_term (zero per-request overhead after warm-up)
- M20: Add check_no_hardlinks and check_no_special_files to MIB upload
archive extraction to prevent hard link, device node, and FIFO attacks
- M21: Change RateLimit ETS table from :public to :protected; route hit/get/reset
through GenServer.call instead of direct ETS access
- M22: Add 8-hour impersonation timeout — store impersonated_at in session and
auto-revoke impersonation when expired
- L14: Add default limit (500) and optional offset to list_checks for pagination
SESSION_SIGNING_SALT, SESSION_ENCRYPTION_SALT, and LIVE_VIEW_SIGNING_SALT
are now loaded from environment variables in production. Dev/test keep the
previous defaults via config.exs. k8s/secrets.yaml has placeholder entries;
the user fills in real values before applying.
Use Repo.get_by(schema, id: id, organization_id: organization_id) instead of
Repo.get + pattern match so that resources from wrong orgs return :not_found
instead of :forbidden, preventing org membership discovery.
- L3: Normalize IPv4-mapped IPv6 addresses (::ffff:x.x.x.x) to IPv4 tuples
when matching against IP whitelist entries
- L7: Verify interface belongs to current org before set/clear capacity
- L12: Guard window.liveSocket behind NODE_ENV check for production safety
- L13: Replace Alert.changeset with Ecto.Changeset.change for simple field
updates (resolved_at, acknowledged_at, gaiia_impact) to prevent
accidental overwrites from the 17-field cast
- M3: Replace blacklist-based valid_return_path? with whitelist of known app
route prefixes, plus URI decoding to prevent %2f encoding bypasses
- M6: Replace full-org device/link load in get_node_detail with targeted
join query filtering by discovered node fields
- M7: Add partial unique index on (organization_id, email) for pending
invitations (WHERE accepted_at IS NULL)
- M12: Add on_mount superuser verification to all 7 admin LiveViews for
defense-in-depth
- M17: default SNMPv3 auth protocol on the four SNMP executors
(interface/storage/processor/sensor) is now SHA-256 instead of MD5.
Devices that explicitly set MD5 still use it; only the fallback for
unconfigured devices changes.
- M23: permissions check now logs a warning when an org's memberships
aren't preloaded so a silent denial is visible at the call site
instead of looking identical to a real authz failure.
- L9: extracted the duplicated should_sync?/1 logic from gaiia, sonar,
splynx, visp, and netbox sync workers into
Towerops.Integrations.due_for_sync?/2. Each worker now passes its
own default interval (10 or 30 minutes) to a single shared helper.
Tests updated to cover the helper directly.
- M2: get_user_by_email/1 uses lower(email) = lower(?) so User@Example.com
and user@example.com resolve to the same account; closes a lookalike-
registration / password-reset confusion risk.
- M8: webhook auth plug no longer echoes "Webhook authentication not
configured" — the misconfiguration is logged server-side and the caller
gets a generic Internal server error so endpoints can't be probed.
- M9: coverages controller logs the underlying KMZ build error and
returns a generic "Failed to build KMZ" string — filesystem paths and
zip internals no longer leak.
- M10: account-data controller no longer crashes with MatchError when an
org has zero or multiple memberships; defaults to "member" and treats
any owner membership as owner.
- M11: ActivityController switches to a hard whitelist
(Atom.to_string/1 comparison) instead of String.to_existing_atom/1;
unknown filter types are dropped silently and the atom table can't be
grown from API input.
- M16: title-tracking MutationObserver is stored on `this` so destroyed()
actually disconnects it — fixes a memory leak per LV navigation.
- H19: /api/v1/mobile/auth/qr/verify no longer returns user_email. Knowing
a QR token now only tells the caller the token is valid; the email is
only revealed by /complete which consumes the token.
- H20: CoverageWorker max_attempts dropped from 3 to 1. The fail/2 path
already returns :ok and writes the failure onto the coverage record,
so the 3-attempt retry policy was never reachable and was misleading.
- H25: ChecksController.create rejects device_ids that don't belong to
the caller's organization (via ScopedResource.fetch). Without this an
API token could attach service checks to devices in another tenant.
Also strip bug ID references from in-code comments (per feedback that
H/M/L numbers don't survive past their bugs.md removal); commit history
keeps the audit trail.
- H12: session and remember-me cookies get http_only + secure (prod-only).
Cookie can no longer be read via document.cookie (XSS exfil defense)
and the Secure flag is set in production via config/prod.exs.
- L2: 404 tracker uses EXPIRE … NX so a sustained probe can't keep
refreshing the 60s window and dodge the threshold ban.
- L5: vault only reads CLOAK_KEY when :env == :prod — a developer with
a prod env var set in their shell won't accidentally encrypt local
data with the production key.
- L6: health endpoint no longer leaks the app version.
- L8: Preseem.dismiss_insight/2 + InsightsLive uses it — dismissing now
requires the insight to belong to the user's org (closes IDOR).
- L10: SidebarCollapse JS hook stores its click handler and removes it
in destroyed(); listeners no longer accumulate across LV navigation.
- L11: WebMCP navigate tool rejects anything that isn't a same-origin
absolute path (blocks javascript:, data:, off-site URLs).
- H23: introduce Towerops.Workers.SyncErrors.transient?/1 classifier;
uisp/preseem/cn_maestro sync workers now retry transient failures (5xx,
timeouts, rate limits) via Oban while still discarding permanent errors
(401/403/404) as :ok. Stale monitoring data is the bigger risk than an
extra retry on a known-bad credential.
- H26: add Monitoring.get_latency_data_for_devices/2 batch counterpart
(currently a stub map but called from SiteLive.Show so any future ping
implementation lands in a single query, not one-per-device).
- H28: Nominatim search popups in the coverage map now bind via a text
node instead of an HTML string, so a compromised/poisoned response
can't execute as JS in the popup.
- H29: yaml_profiles dot-boundary OID prefix match (handles optional
trailing dot from YAML profile keys). Stops 1.3.6.1.4.1.9 from
incorrectly matching 1.3.6.1.4.1.99.
- H6: batch interface stats query in get_site_capacity_summary, replacing
per-interface SELECT with a single grouped query
- H8: explicit expires_at check in MobileAuth and GraphQLAuth plugs as
defense in depth — the query already filters expired rows, but a future
refactor dropping the filter can't quietly re-enable revoked sessions
- H17: Reports LiveView (toggle/delete/run_now) now uses
Reports.get_organization_report/2 to scope by organization_id, fixing
IDOR where a user could manipulate other tenants' reports by ID
- H18: ToweropsWeb.Api.ParamFilter strips identity fields (id,
organization_id, user_id, created_by_id, inserted_at, updated_at) from
user-supplied API params; applied to devices, sites, checks, and
escalation_policies create/update paths to prevent mass-assignment of
tenant ownership fields if a changeset cast list ever drifts
- H1: batch count queries in MobileController (org sites/devices/alerts, site
device counts) — reduces 1+3N+2M queries to 4 + 2
- H2: batch Oban cancel for stop_device_checks — single ANY(?) query instead
of one per check
- H3: resolve_active_alerts_for_device uses update_all + deduped PubSub
broadcasts (was N updates + N broadcasts)
- H4: atomic ON CONFLICT upsert for auto-discovery checks; restore partial
unique index dropped in 20260404000002 (dedup migration included)
- H5: wrap apply_agent_to_all_equipment delete+insert in transaction so a
failed insert_all rolls back the prior delete
- H14: replace String.to_integer/1 on untrusted SNMP OID components with
Integer.parse/1 + validation (neighbor_discovery, discovery storage index,
printer supply index)
- H15: ActivityFeedLive whitelists activity types against @all_types instead
of String.to_existing_atom/1
- H16: defensive page param parsing in user_settings and admin dashboard
LiveViews
- H27: search sanitize/1 delegates to QueryHelpers.sanitize_like/1 which
escapes backslash first
- H30: JobCleanupTask no longer cancels jobs in "executing" state so
in-flight polls complete naturally
- Add organization membership check before API token creation
- Fix admin security allowlist form reading current_user instead of current_scope.user
- Use recursive File.ls instead of Path.wildcard to include hidden files in MIB validation
- Add upload size check before ZIP extraction in MIB controller
- Centralize CSP in SecurityHeaders plug with per-request nonces instead of 'unsafe-inline'
- Enforce upload size limit (100 MB), max extracted files (2000), max
uncompressed size (500 MB), and reject symlinks in MIB archives
- Replace unbounded String.to_atom calls in SNMP tokenizer with
safe_atom/1 that warns when approaching the atom table limit
- Cache MIB vendor listing with persistent_term, invalidated on
upload/delete to avoid blocking synchronous Path.wildcard scans
- Client IP: only trust X-Forwarded-For from RFC 1918 proxy IPs
- Webhook auth: handle nil/blank secret with controlled error, not 500 crash
- Sudo redirect: reuse validated return_path? from login to prevent open redirect
- Map live: remove redundant inline script (ensureLeaflet hook handles loading)
- Bang calls: convert crash-prone exact matches to case in QR live and API controllers
Fix auto_match return tuple order (mac/ip/site were swapped). Add ghost
devices tab for cleaning up stale mappings to deleted devices. Add "Link"
button on untracked devices to manually link to existing Gaiia items.
Add manual refresh button for on-demand reconciliation.
Adds MAC address matching as the primary matching phase:
- Normalizes MACs (strips colons/dots/dashes) from SNMP interface data
- Matches against gaiia_inventory_items.mac_address (synced from Gaiia)
- Phase 1: MAC → Phase 2: IP → Phase 3: Site
Flash now shows breakdown: "Auto-matched 47 devices (12 by MAC, 34 by
IP, 1 by site). 17 remaining."
A bare <select phx-change="..."> sends %{"value" => ...}, not the name
attribute. Wrapping in <form phx-change="..."> fixes this so the handler
receives %{"manufacturer_id" => value}.
- Handle empty manufacturer_id selection (select placeholder)
- Log and flash when model fetch fails or returns empty
- Add require Logger to reconciliation live view
Replaces auto-matching with explicit two-step selection:
1. Select manufacturer from Gaiia's inventory model library (dropdown)
2. Select model for that manufacturer (dropdown, filtered live)
3. Click Create
Manufacturers and models are fetched from Gaiia on demand and cached
in socket assigns. Only one device form is open at a time.
ngettext is a macro that already calls dngettext internally. The
previous code was calling dngettext a second time on the result string,
passing nil as the count — causing a FunctionClauseError.
Adds `Towerops.Gaiia.InventoryCreator` which:
- Matches a device's SNMP-discovered manufacturer/model against Gaiia's
inventory model library to find the right modelId
- Calls `startAddInventoryItemsJob` to create the item, assigned to the
correct Gaiia network site with MAC address from SNMP interfaces
A "Create in Gaiia" button now appears next to each untracked device on
the reconciliation page.
Two improvements to Gaiia reconciliation:
1. Auto-match now uses two-phase matching:
- Phase 1: IP match (high confidence, same as before)
- Phase 2: Site match — if a device is at a Towerops site linked to a
Gaiia network site, match it to unmapped inventory assigned to that site
2. Gaiia sync now automatically maps network sites to Towerops sites by
name during sync. When a Gaiia network site name matches a Towerops
site name, `site_id` is set on the `gaiia_network_sites` row. This
enables the site-based matching in phase 2.
Flash message now shows breakdown: "Auto-matched 47 devices (34 by IP,
13 by site). 17 remaining."
Adds `Gaiia.auto_match_untracked/1` which finds unmapped Gaiia inventory
items whose IP matches an untracked Towerops device and links them.
Button on the reconciliation page's Untracked tab runs the match
inline and shows a flash with matched/remaining counts.
All 4 webhook controllers now verify the signature, insert an Oban job,
and return immediately. Previously, Stripe and PagerDuty did DB queries +
writes inline; Gaiia ran upserts; Agent Release broadcast to all connected
WebSockets — all in the request path.
New workers:
- StripeWebhookWorker — delegates to WebhookProcessor.process/1
- GaiiaWebhookWorker — delegates to Webhooks.process_event/3
- PagerdutyWebhookWorker — resolve/acknowledge alerts by dedup key
- AgentReleaseWebhookWorker — broadcast mass update to agents
Also fix: log error when resolve_alert fails in agent_channel (was silently
discarded, causing "Resolving stuck device_down alert" to repeat in logs).
Adds a parallel LLM path that reviews the whole org snapshot — Preseem
APs, SNMP signals, backhaul, agents, and the insight types the rules
already covered — and emits novel observations as ai_observation
insights (source=ai). Existing rule-based insights and per-insight
enrichment continue unchanged.
- lib/towerops/llm/network_snapshot.ex builds the bounded snapshot
- lib/towerops/llm/network_insight_prompt.ex builds the prompt and
parses the LLM's JSON observations array
- lib/towerops/workers/ai_network_insight_worker.ex runs nightly at
03:30 UTC (after the rule pass + per-insight enrichment) and on
the superuser regenerate button
- preseem/insight.ex extends @valid_types with "ai_observation" and
@valid_sources with "ai"
- insights_live/index.ex includes the new worker in the regen list
and adds an "AI" source badge style
Bumps the SnmpKit Manager timeout-options test threshold from 500ms
to 3000ms (was flaking on loaded dev machines), and scopes
list_unenriched_insights/1 assertion in insights_test:573 to the
test's org so cross-suite leaks don't fail the count.
- ToweropsWeb.AgentChannel now runs sensor readings through
Towerops.Snmp.SensorScale.normalize/2 in both resolve_sensor_value/2
and the GraphQL publisher. Without this, MikroTik mtxrHl* deci-degree
readings (e.g. Culleoka reporting 294 -> 29.4°C) were stored and
alerted on at 294°C. The DevicePollerWorker path already did this;
only the agent path was missing it.
- Demote "Received SNMP result from agent" and "Processing polling
result for X" from info to debug — they fire every poll cycle on
every device and were dominating prod logs.
- /insights now shows a "Regenerate insights" button for superusers
that enqueues all seven insight workers (Preseem baselines,
capacity, device-health, Gaiia, system, wireless, LLM enrichment).
Non-superusers neither see the button nor can trigger the event.
- Dialyzer is clean: added :tools to plt_add_apps, gave
Preseem.Insight a @type t, pinned a few unmatched returns, and
suppressed a known MapSet opaque false-positive in Towerops.Unused.
- e2e: NODE_OPTIONS=--disable-warning=DEP0205 silences the Node 25
deprecation noise from Playwright's internal ESM loader.
UI: the AI-generated footer dropped " · deepseek-v4-pro" — model
identity is irrelevant to operators reviewing the recommendation.
Parser: stop letting raw JSON leak into the summary text when the LLM
returns an almost-but-not-quite-valid response.
- Strip ```json fences (already done) PLUS try a balanced-brace
extraction so a "Here's the JSON: {…}" preamble doesn't break
Jason.decode
- When all decode paths fail, regex-extract `"summary"` and `"action"`
fields from the text rather than dumping the whole blob with braces
intact
- Last-ditch sanitiser strips remaining brace/quote noise so the
rendered summary never shows literal `"summary"` keys
Maintenance: `Towerops.Maintenance.RepairLlmSummaries.run()` clears
`llm_enriched_at` on already-stored rows whose summaries contain JSON
artifacts so the next worker cycle re-enriches them with the new
parser.
Test: insights_live_events_test now asserts the AI-generated label is
present *and* that the model name does not leak into the rendered HTML.
The CSS `capitalize` class only capitalises the first letter, so the
source filter pills and per-card badges showed "Ai" and "Snmp". Switch
to a `source_label/1` helper that returns "AI" / "SNMP" verbatim and
falls back to `String.capitalize/1` for the rest (Preseem, Gaiia,
System).
Adds an "ai" pill to the source-filter strip on /insights and a small
blue "AI" badge on every card whose `llm_enriched_at` is set. Selecting
the pill returns every LLM-enriched insight regardless of underlying
source (preseem / snmp / gaiia / system).
`source: "ai"` is a virtual filter — `list_insights/2` translates it to
`WHERE llm_enriched_at IS NOT NULL` rather than equality on the source
column, so existing real source semantics stay intact.
UI:
- Type-specific evidence cards (red/amber/orange/rose/purple/cyan/
indigo/blue) had `dark:text-{color}-300` labels and
`dark:bg-{color}-800 dark:text-{color}-100` badges that washed out on
the dark background. Bumped 31 label shades to -200 and switched the
three problem badges (Above manufacturer spec, Investigate upstream
first, own fleet, crit) to solid -600 with white text.
- upstream_degradation card got brighter device-link text and a
hover state in dark mode.
DeepSeek:
- Empty-string DEEPSEEK_MODEL / DEEPSEEK_BASE_URL env vars were not
treated as missing — `||` only short-circuits on nil/false — so a
blank secret would override the @default_model and the API would
reject `model: ""`. Both runtime.exs and the client itself now treat
blank values as missing.
Performance:
- HealthController: honor `:timeout` from `:redis` config (was hardcoded
2s) so the 503-when-unreachable test exits immediately
- DevicePollerBranches: drop `collect_events` window 1500ms → 100ms
(PubSub broadcasts are synchronous from `run_perform`)
- Towerops.Unused: cache xref analysis in `:persistent_term` keyed by
BEAM mtimes so repeated `mix unused` invocations skip re-analysis
- DiscoveryWorker: make wait_loop interval configurable via
`:discovery_wait_interval_ms` (25ms in tests vs 500ms in prod)
- ProfileWatcher: configurable `:profile_reloader` so tests can stub
the YAML reload instead of paying the full disk-load cost
- PingExecutor: tag the loopback test `:network` (it shells out to
`ping`); coverage already provided by 4 other `:network`-tagged tests
Flake fix:
- FrequencyChange tests used hardcoded `~U[2026-05-09 12:00:00Z]`,
which now falls outside the rule's 24h lookback window. Switched to
`DateTime.utc_now()` so scans stay inside the window regardless of
clock progression.
Fiber operators care a lot about optical Rx power because connectors
get dirty, splices degrade, and patch panels develop bend-radius
problems. The data is already in snmp_transceiver_readings — this
rule turns it into actionable insights.
- Towerops.Recommendations.Rules.OpticalRxLow — single query selects
the latest reading per transceiver (last hour), filters to those
with rx_power_dbm <= -28, joins back to the device. Critical at
<= -32 dBm. One insight per (device, port) via metadata.dedup_key.
- Insight.@valid_types extended with optical_rx_low.
- LLM prompt hint walks the operator through the standard ladder:
clean SC/LC connector with a one-click cleaner, inspect bend
radius, verify splices, then replace the transceiver.
- UI: cyan card with port, Rx power, Tx power, transceiver type, and
wavelength.
- Wired into RecommendationsRunWorker.
8 new tests.
When 50%+ of the access points at one site simultaneously have
Preseem QoE score <60, root cause is almost always a shared upstream
problem (backhaul, transit, DNS, DHCP) — not per-AP RF. This rule
emits one site-level insight that points the operator at the right
investigation layer instead of letting them chase per-AP RF
symptoms across many APs at the same site.
- Towerops.Recommendations.Rules.UpstreamDegradation — single query
pulls all matched Preseem APs grouped by site_id. Critical when
100% of APs at the site are degraded; warning otherwise. Lists
every degraded AP with its QoE score in metadata.
- Insight.@valid_types extended with upstream_degradation.
- LLM prompt hint tells the model to instruct the operator to
investigate upstream FIRST and to NOT chase per-AP RF symptoms.
- UI: indigo card with degraded/total AP counts, an "Investigate
upstream first" badge, and a clickable list of affected APs.
- Wired into RecommendationsRunWorker FIRST in the @rules list so
the upstream signal surfaces before per-AP RF rules.
Also fixes two flakes:
- device_live/show_test.exs is now async: false. It inserts to
InterfaceStat which is a TimescaleDB hypertable; concurrent inserts
deadlock on chunk-table locks under parallel SQL-sandbox tests
(ERROR 40P01 deadlock_detected in setup).
- Cleared an unused-variable warning in device_overheating_test.exs.
8 new tests for UpstreamDegradation.
Operators frequently misdiagnose thermal throttling as RF problems
and waste truck rolls. This rule reads the hottest temperature
sensor per device and, when the device is linked to a Preseem AP
with degraded QoE, makes the thermal-correlation explicit so the
fix is ventilation/shading rather than antenna alignment.
- Towerops.Recommendations.Rules.DeviceOverheating — single query
pulls the hottest temperature sensor per device above 65 °C; one
insight per device. Critical at 75 °C+. Includes the linked
Preseem qoe_score in metadata when present.
- Insight.@valid_types extended with device_overheating.
- LLM prompt hint tells the model to recommend ventilation/fan/
shading inspection BEFORE any RF action when QoE is also low.
- UI: red card with sensor name, temperature, linked QoE score,
and an "Above manufacturer spec" badge for >=75 °C.
Wired into RecommendationsRunWorker. 8 new tests.