Fixed three compile warnings by removing unreachable error handling:
1. Dynamic.discover_sensors/2 always returns {:ok, sensors}, never errors
- Removed unreachable error clause in discovery.ex
- Updated spec to reflect actual return type
2. parse_detection_rule/2 always returns {:ok, :processed}
- Simplified create_detection_rules/2 to use Enum.each instead of reduce_while
- Removed unreachable error handling in caller
3. import_all_from_directory/2 always returns {:ok, %{...}}
- Removed unreachable error clause in Mix task
- Simplified to direct pattern match
All code now compiles without warnings.
Replace all references with generic terms:
- 'external YAML files' instead of specific project names
- 'device profiles' for database-stored definitions
- 'imported OID definitions' for sensor configurations
This makes the codebase vendor-neutral and focused on the
internal implementation rather than external dependencies.
LibreNMS sensor OIDs contain template variables like '.{{ $index }}'
which need to be removed before performing SNMP walks.
Added clean_oid_template/1 to strip these template variables:
- '.1.3.6.1.4.1.17713.21.2.1.36.{{ $index }}' -> '.1.3.6.1.4.1.17713.21.2.1.36'
- 'sysUptime.{{ $index }}' -> 'sysUptime'
The SNMP walk will now use the base OID and discover all instances
with their actual index values.
LibreNMS OS definitions use MIB symbolic names (e.g.,
CAMBIUM-PMP80211-MIB::cambiumCurrentuImageVersion.0) which require
MIB compilation to resolve to numeric OIDs.
The SNMP client (snmpkit) only supports numeric OIDs, not MIB names.
Disabled OS definition polling for now. The system will use the profile
OS name as firmware_version (e.g., 'epmp') until MIB resolution is
implemented.
Future improvements:
- Add MIB compiler integration (net-snmp, snmptranslate)
- Import numeric OIDs during profile import
- Add MIB -> numeric OID lookup table
Enhance dynamic profile to poll OS definitions from SNMP:
- Add serial_number field to snmp_devices table
- Update Dynamic profile to query profile_os_definitions
- Poll SNMP OIDs for version and serial fields
- Display actual firmware version instead of profile name
For ePMP devices, this will now query:
- CAMBIUM-PMP80211-MIB::cambiumCurrentuImageVersion.0 for firmware
- CAMBIUM-PMP80211-MIB::cambiumEPMPMSN.0 for serial number
The Operating System field will now show actual version like
'Cambium 4.7.1' instead of 'Cambium epmp'.
The Operating System field was showing 'Cambium Verona Networks' because
firmware_version was set to the raw sys_descr SNMP value.
Now uses the profile's OS name (e.g., 'epmp') so it displays as
'Cambium epmp' which is cleaner and more accurate.
Future improvement: Poll actual firmware version from SNMP using the
profile_os_definitions.version OID.
When clicking 'Rediscover Device', navigate to the device details page
where the user can see live updates as discovery progresses.
The device show page already subscribes to PubSub discovery events and
will automatically refresh when discovery completes, showing updated
manufacturer, model, sensors, and interfaces.
LibreNMS uses leading dot format (.1.3.6.1...) but SNMP returns
OIDs without leading dot (1.3.6.1...). Normalize by prepending
dot to sys_object_id before matching, and use starts_with instead
of contains for more accurate prefix matching.
- Configure Phoenix.PubSub to use Redis adapter for pod-resilient messaging
- Add PubSub broadcasts for sensor readings updates in PollerWorker
- Add PubSub broadcasts for interface stats updates in PollerWorker
- Add PubSub broadcasts for neighbor discovery updates in PollerWorker
- Add PubSub broadcasts for monitoring check updates in MonitorWorker
- Broadcast interface and sensor events to device-specific topics
- Add event handlers in DeviceLive.Show for all update types
- Device show page now updates in real-time without polling
- Updates survive pod restarts and work across multiple pods
Added telemetry metrics for monitoring:
Exq Metrics:
- Queue sizes for all queues (default, discovery, polling, monitoring, maintenance)
- Number of busy worker processes
- Total number of worker processes
Redis/Valkey Metrics:
- Connected clients
- Memory usage
- Total commands processed
Metrics are collected every 10 seconds via telemetry_poller and displayed
in LiveDashboard at /dashboard
Script checks:
- Container status in pod
- Valkey logs
- Network connectivity from towerops to Valkey
- Environment variables
- Valkey health
Usage: ./k8s/debug-redis.sh [namespace]
Changes:
- Added startupProbe to towerops container to allow up to 70s for startup
(gives Valkey sidecar time to be ready)
- Changed REDIS_HOST from 'localhost' to '127.0.0.1' for explicit loopback
- Removed initialDelaySeconds from liveness/readiness probes (handled by startup)
- Added comment clarifying Redis connects to Valkey sidecar
This ensures the app waits for Valkey to be ready before accepting traffic.
The Alert schema was updated to use :device_down and :device_up enum values,
but existing database records still had 'equipment_down' and 'equipment_up' values.
Added migration to update all existing alert records to use the new naming
convention.
Updated all references to use the more technically accurate term 'ICMP Latency'
instead of 'Ping Latency' in:
- Graph titles and labels
- HTML comments
- Test assertions
- Added comprehensive tests for GraphLive.Show to ensure:
* Page renders with correct assigns
* Uses @current_organization (not @organization)
* Latency, processor, memory, and traffic graphs work
* Time range selection works correctly
* All required assigns are present
- Reordered device show page so Overall Traffic appears above Ping Latency
Changed device_monitor.ex to always use ICMP ping for monitoring checks,
regardless of whether SNMP is enabled for the device.
Rationale:
- ICMP ping gives actual network latency (round-trip time)
- SNMP checks are for data collection (interfaces, sensors, etc.)
- Both should run independently for SNMP-enabled devices:
* SNMP polling: Collects device metrics (handled by poller_worker)
* ICMP monitoring: Measures latency and availability (handled by device_monitor)
This ensures latency graphs always show actual network response time,
not SNMP query response time.
The monitoring system was creating checks but hardcoding response_time_ms
to nil instead of using the actual ping latency. This caused the latency
graphs to have no data to display.
Changes:
- device_monitor.ex: Capture response time from ping/SNMP check result
- monitor_worker.ex: Use correct field name (response_time_ms) and status enum values (:success/:failure)
After this fix, new monitoring checks will include latency data and the
latency graphs will display properly on both device and site detail pages.
Shows ping latency for all devices in the site over the last 24 hours.
Each device is displayed as a separate line on the graph with its name
as the label.
Graph only appears if there is latency data available for at least
one device in the site.
Created API token system for programmatic access:
- API tokens table with organization scoping
- ApiTokens context for token management
- ApiAuth plug for Bearer token authentication
- Tokens shown once on creation, then hashed for security
Implemented RESTful API v1 endpoints:
- SitesController: CRUD operations for sites
- DevicesController: CRUD operations for devices
- All operations scoped to authenticated organization
- Proper authorization checks via site ownership
Technical details:
- Tokens prefixed with "towerops_" for easy identification
- SHA-256 hashing for token storage
- Last used timestamp tracking (async update)
- Optional token expiration support
- Standard JSON error responses (40x status codes)
Routes:
- /api/v1/sites (GET, POST, PATCH, DELETE)
- /api/v1/devices (GET, POST, PATCH, DELETE)
Authentication:
- Authorization: Bearer towerops_xxxxx header required
- Returns 401 for invalid/expired tokens
- Returns 403 for unauthorized resource access
Created three new Exq workers to handle background jobs:
- PollWorker: SNMP polling operations (queue: polling)
- MonitorWorker: Device monitoring checks (queue: monitoring)
- DiscoveryWorker: SNMP discovery (already existed)
Updated files to use Exq.enqueue instead of Task.start:
- lib/towerops/snmp/poller_worker.ex
- lib/towerops/monitoring/device_monitor.ex
- lib/towerops_web/channels/agent_channel.ex
Each module includes enqueue_* helper functions that:
- Use Task.start in test environment for synchronous execution
- Use Exq.enqueue in dev/prod for proper background job processing
Added "monitoring" queue to Exq configuration in application.ex
All tests passing (824 tests, 0 failures)
- Create DiscoveryWorker for handling SNMP device discovery
- Replace Task.start with Exq.enqueue for discovery operations
- Enqueue jobs to 'discovery' queue with device_id parameter
- Add enqueue_discovery/1 helper that falls back to Task.start in test
- Discovery jobs now persistent and retryable via Exq
- Triggered on: new device creation, device update, manual rediscover button
- Add exq ~> 0.23 dependency with redix and elixir_uuid
- Configure Exq in application supervisor with 4 queues (default, discovery, polling, maintenance)
- Set Exq as included_application to prevent auto-start
- Only start Exq in dev/prod environments, skip in test
- Use :env config to determine runtime environment
- Configure 10 concurrent workers with exq namespace
- Connect to Redis/Valkey sidecar via existing :redis config
- Add Valkey 8.0 alpine container as sidecar in k8s deployment
- Configure with appropriate security context (non-root, no privilege escalation)
- Add resource limits (256Mi memory, 200m CPU) and requests (128Mi, 50m)
- Add health probes using valkey-cli ping command
- Configure Redis connection via REDIS_HOST and REDIS_PORT environment variables
- Add Redis config to runtime.exs, dev.exs, and test.exs
- Valkey runs on localhost:6379 within the same pod as Phoenix app
- Make device name field optional when SNMP is enabled (ICMP Only mode still requires name)
- Add database migration to make name column nullable
- Add conditional validation: name required only when snmp_enabled is false
- Update SNMP discovery to populate device name from sysName when empty
- Add helper text explaining SNMP name population behavior in form
- Add comprehensive tests for device name SNMP population
- Update test mocks to account for NetSnmp profile sensor discovery
- After creating a device, stay on /devices/new instead of navigating to device show page
- Display success flash message with creation confirmation
- Reset form to defaults (preserving selected site) for adding next device
- Allows users to quickly add multiple devices in a row without navigation
- Add tests for staying on page and adding multiple devices sequentially
- All 91 device-related tests passing with no Credo warnings
- Add tab selector in Monitoring Configuration section for choosing monitoring mode
- SNMP & ICMP mode: enables ICMP pings + SNMP discovery (default)
- ICMP Only mode: enables only ICMP ping monitoring, hides SNMP fields
- Monitoring mode controls snmp_enabled field automatically on save
- Test SNMP Connection button shows informational message when agent is configured
(agents handle SNMP testing during their polling cycle)
- Update all tests to use tab selector instead of setting snmp_enabled directly
- All 818 tests passing with no Credo warnings
- Add Monitoring.get_latency_data/2 to query successful checks with response times
- Support 'latency' sensor type in GraphLive.Show with configurable time ranges
- Add latency chart to device show page alongside other metric charts
- Add comprehensive tests for get_latency_data/2 (limit, since, failed checks)
- All 813 tests passing with no Credo warnings
Implemented hierarchical SNMP community string inheritance following TDD:
Schema changes:
- Updated Device.changeset to remove required validation for snmp_community
- Community string is now optional when SNMP is enabled
- Allows inheritance from site or organization level
Inheritance hierarchy:
- Device-level community (highest priority)
- Site-level community
- Organization-level community
- nil (no community set at any level)
Form updates:
- extract_snmp_config now resolves inherited community for SNMP testing
- validate_test_snmp_input provides clearer error message about inheritance
- SNMP test uses effective community from hierarchy
Tests:
- Added tests for creating devices with nil/empty community string
- Added comprehensive SNMP configuration inheritance tests
- Tests verify hierarchy: device > site > org > default
- All 28 device tests passing
All 808 tests passing (2 pre-existing alert notifier failures unrelated to this change)
No Credo warnings
Implemented ability to rename agents through a new edit page following TDD:
Backend changes:
- Added update_agent_token/2 in Agents context for updating agent names
- Added update_changeset/2 in AgentToken schema that only allows name updates
- Security: Token field cannot be changed, only name field is allowed
LiveView implementation:
- Created AgentLive.Edit module with mount, validate, and save handlers
- Created edit.html.heex template with form for renaming agents
- Added route /orgs/:org_slug/agents/:id/edit
- Added Edit button to agent show page header
- Validates organization ownership using Repo.get_by!
Tests:
- Added comprehensive context tests for update_agent_token/2
- Added LiveView tests for edit page functionality
- Tests cover: loading, validation, save/redirect, auth, and org ownership
All 801 tests passing, no Credo warnings.
Replace the revoke functionality with delete to properly clean up agent references.
When an agent is deleted:
- All direct device assignments are removed
- Agent is removed from site defaults
- Agent is removed from organization defaults
- Devices automatically fall back to site/org agents or cloud polling
This ensures no orphaned references and provides better clarity about
what happens to device monitoring when an agent is removed.
- Add delete_agent_token/1 function in Agents context
- Update LiveView to use delete_agent event instead of revoke_agent
- Update UI button from "Revoke" to "Delete" with clearer confirmation message
- Add comprehensive tests for delete functionality and fallback behavior
- All 792 tests passing, no Credo warnings
Change the default filter from 'all' to 'active' on the alerts page.
This shows only active alerts by default, which is more useful for
monitoring current issues. Users can still click 'All Alerts' to see
the complete history.
Prevent WebSocket disconnections during deployments by:
1. Rolling Update Strategy:
- maxSurge: 1 (allow 1 extra pod during rollout)
- maxUnavailable: 0 (keep all pods running)
- minReadySeconds: 10 (wait before continuing rollout)
2. Graceful Shutdown:
- terminationGracePeriodSeconds: 30 (allow Phoenix to drain connections)
3. PodDisruptionBudget:
- minAvailable: 1 (ensure at least 1 pod always available)
- Prevents all pods from being terminated simultaneously
4. Improved Health Checks:
- Faster readiness probe (5s initial, 3s period)
- More aggressive success/failure thresholds
- New pods marked ready faster
This ensures new pods are fully ready and serving traffic before old
pods are terminated, maintaining WebSocket connections and preventing
'something went wrong' errors during deployments.
Change (optional) text color from zinc-500/400 to zinc-400/500 for
better contrast and lighter appearance. Applied to:
- Parent Site field label
- Location field label
- Description field label
- Agent Configuration heading
- SNMP Configuration heading
Add greyed out '(optional)' text to Parent Site, Location, and
Description field labels for better UX. The optional indicator uses
a lighter color (text-zinc-500/400) to distinguish it from the
required field label.