38 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
CRITICAL: Read AGENTS.md First
Before starting any work in this repository, you MUST read AGENTS.md in the project root.
AGENTS.mdcontains comprehensive Phoenix/LiveView/Elixir guidelines that are mandatory for this project- These guidelines take precedence over general development practices when there's a conflict
- Always read
AGENTS.mdat the start of a new conversation or when resuming work - Follow the project-specific patterns, conventions, and constraints documented there
Project Overview
Towerops is a Phoenix 1.8 web application built with Elixir, using Ecto for database operations (PostgreSQL), LiveView for real-time interactions, and Tailwind CSS v4 for styling.
Data Model Relationships
┌─────────────┐
│ User │
│ (Accounts) │
└──────┬──────┘
│ owns/member of (many-to-many)
│
↓
┌──────────────────────────────────────────────────────┐
│ Organization │
│ (Organizations) │
└────┬────────────┬─────────────┬──────────────────────┘
│ │ │
│ has many │ has many │ has many
│ │ │
↓ ↓ ↓
┌─────────┐ ┌──────────┐ ┌────────────────┐
│ Site │ │Equipment │ │ AgentToken │
│ (Sites) │ │(Equipment│ │ (Agents) │
└────┬────┘ └────┬─────┘ └────────┬───────┘
│ │ │
│ has many │ has one │ assigned via
│ │ │
↓ ↓ ↓
┌──────────┐ ┌──────────────┐ ┌──────────────────┐
│Equipment │ │ SNMPDevice │ │ AgentAssignment │◀──┐
│(Equipment│ │(Snmp.Devices)│ │ (Agents) │ │
└────┬─────┘ └──────────────┘ └──────────────────┘ │
│ │ │
│ has many └───────────────┘
├────────────────┬─────────────────┘
│ │ assigns equipment
↓ ↓ to agent
┌──────────────┐ ┌─────────┐
│MonitoringCheck│ │ Alert │
│ (Monitoring) │ │ (Alerts)│
└──────────────┘ └─────────┘
Key Relationships:
- User can own/belong to multiple Organizations
- Organization has default_agent_token_id (optional)
- Equipment belongs to both Site and Organization (denormalized)
- Equipment can be assigned to one AgentToken via AgentAssignment
- Equipment has one SNMPDevice with Sensors and Interfaces
- Equipment has many MonitoringChecks (polling results)
- Equipment has many Alerts (equipment_down, equipment_up)
- AgentToken authenticates remote agents for local SNMP polling
Note to AI assistants: When you make changes to the data model (migrations, schema additions, new contexts), update this diagram to reflect the current state.
Essential Commands
Setup and Development
mix setup- Install dependencies, create/migrate database, setup and build assetsmix phx.server- Start the Phoenix server (accessible at http://localhost:4000)iex -S mix phx.server- Start server with interactive Elixir shell
Testing and Quality
mix test- Run all tests (automatically creates and migrates test database)mix test test/path/to/specific_test.exs- Run a specific test filemix test --failed- Re-run only previously failed testsmix test --cover- Run tests with coverage report (generates HTML report incover/directory)mix precommit- Run before committing: compiles with warnings as errors, unlocks unused deps, formats code, and runs tests
Coverage Target: 90% minimum (currently configured threshold)
- Coverage reports show line-by-line coverage in
cover/directory - View detailed HTML report:
open cover/modules_*.html - Current status: 60.40% overall (as of Jan 12, 2026, 6:35 PM)
- Recent improvements (Session 1):
- Towerops.Snmp: 10.87% → 97.83% (+87%)
- Towerops.Snmp.Profiles.NetSnmp: 56.79% → 92.59% (+35.8%)
- Towerops.Monitoring.Supervisor: 21.21% → 63.64% (+42%)
- Towerops.Snmp.Profiles.Cisco: 20.22% → 57.30% (+37%)
- ToweropsWeb.EquipmentLive.Show: 35.51% → 46.38% (+10.87%)
Coverage Notes:
- Protobuf-generated modules (Towerops.Agent.*) show 0% but don't need tests
- Focus areas for improvement (in priority order):
- Towerops.Monitoring context (33.33%)
- SNMP profiles: NetSnmp (56.79%), Base (69.52%), Mikrotik (77.97%)
- LiveView modules: EquipmentLive.Show (35.51%), EquipmentLive.Form (46.73%)
- Accounts modules: UserCredentialController (33.93%), Accounts context (61.82%)
- 0% coverage modules (Admin, WebAuthn, etc.) - may be intentionally untested
Database
mix ecto.create- Create the databasemix ecto.migrate- Run pending migrationsmix ecto.reset- Drop, create, and migrate databasemix ecto.gen.migration migration_name_using_underscores- Generate a new migration file
Assets
mix assets.build- Build CSS and JS assets (Tailwind + esbuild)mix assets.deploy- Build minified assets for production
Architecture
Application Structure
The application follows standard Phoenix conventions with clear separation between business logic (lib/towerops/) and web interface (lib/towerops_web/):
-
Towerops.Application - OTP application that supervises:
- Telemetry for metrics
- Repo (Ecto) for database
- DNSCluster for service discovery
- PubSub for pub/sub messaging
- Endpoint (web server)
-
ToweropsWeb - Main web module that provides
usemacros for::router- Route definitions:controller- Traditional request/response controllers:live_view- LiveView modules:live_component- LiveView component modules:html- Phoenix.Component modules
Key Configuration Details
- Binary IDs: Generators use binary (UUID) primary keys by default (
binary_id: true) - Timestamps: Use
:utc_datetimefor all timestamps - Web server: Uses Bandit adapter (not Cowboy)
- HTTP client: Uses
:reqlibrary (Req module) - this is the only approved HTTP client - Ecto repos:
[Towerops.Repo]
Web Layer Patterns
All LiveViews, LiveComponents, and HTML modules automatically get these imports/aliases via html_helpers/0:
ToweropsWeb.CoreComponents- Core UI components (<.button>,<.input>,<.form>, etc.)ToweropsWeb.Layouts- Layout components (aliased, no need to re-alias)Phoenix.LiveView.JS- Client-side JS commands- Gettext for translations
- Verified routes with
~psigil
Asset Pipeline
- Tailwind CSS v4: Uses new
@import "tailwindcss"syntax inassets/css/app.css, notailwind.config.jsneeded - esbuild: Bundles JavaScript from
assets/js/app.js - All vendor scripts/styles must be imported into app.js/app.css - cannot reference external src/href in layouts
- No inline
<script>tags in templates - use colocated LiveView hooks instead
Development Environment
- Dev routes enabled for:
- LiveDashboard at
/dev/dashboard - Swoosh mailbox preview at
/dev/mailbox
- LiveDashboard at
- Phoenix LiveReload watches for file changes
- Code reloader enabled via
listeners: [Phoenix.CodeReloader]
Background Job Architecture
The application uses Oban for all background job processing with PostgreSQL-backed queuing for cluster-wide coordination and resilience.
Job Types:
-
Self-Scheduling Workers (per-device recurring jobs):
DeviceMonitorWorker- Health check polling (60s default interval)DevicePollerWorker- SNMP data collection (60s default interval)- Schedule next run before completing via
schedule_next_poll/2 - One job per enabled device, automatically created/cancelled when device settings change
-
Oban Cron Workers (cluster-wide periodic maintenance):
NeighborCleanupWorker- Hourly cleanup of stale neighbors/ARP/MAC entries (runs every hour via cron)StaleAgentWorker- Detects agents that haven't checked in for 10+ minutes (runs every minute)AgentLatencyEvaluator- Latency-based agent reassignment (runs every 5 minutes)JobHealthCheckWorker- Safety net to recover missing monitor/poller jobs (runs every 10 minutes)- Scheduled via
Oban.Plugins.Croninconfig/dev.exsandconfig/runtime.exs - Run cluster-wide (only one instance executes at a time across all pods)
Queues:
default(10 workers) - General background tasksdiscovery(10 workers) - SNMP discovery operationspollers(50 workers) - SNMP polling jobs (one per device)monitors(50 workers) - Health check jobs (one per device)maintenance(5 workers) - Periodic cleanup and health check workers
Resilience:
- Oban Cron jobs run cluster-wide via PostgreSQL-based locking
- If a pod dies, another pod immediately picks up scheduled cron jobs
- Self-scheduling workers (monitor/poller) are recovered by
JobHealthCheckWorkerevery 10 minutes - All jobs visible in Oban dashboard at
/dev/dashboard→ Oban tab - Failed jobs automatically retried with exponential backoff
Configuration:
- Dev:
config/dev.exs- Oban config with cron plugin - Prod:
config/runtime.exs- Same cron schedule as dev - Application:
lib/towerops/application.ex- Oban supervision, no GenServer workers for periodic tasks
LiveDashboard and Telemetry
LiveDashboard is available at /dashboard (requires authentication in production, /dev/dashboard in development).
Available Metrics:
- Phoenix Metrics: Request duration, route timings, channel/socket metrics
- Database Metrics: Query time, decode time, queue time, idle time
- VM Metrics: Memory usage, run queue lengths
- Oban/Background Jobs:
- Queue sizes (default, discovery, pollers, monitors, maintenance)
- Currently executing jobs
- Available jobs waiting to be processed
- Redis/Valkey:
- Connected clients
- Memory usage (MB)
- Total commands processed
Metrics are collected every 10 seconds via telemetry_poller and published using :telemetry.execute/3.
To view metrics:
- Navigate to
/dashboard(or/dev/dashboardin development) - Click the "Metrics" tab
- Scroll to see Oban and Redis sections
Oban Web Dashboard
Oban Web provides a dedicated dashboard for job monitoring and management at /oban (superuser access required).
Access Control: Only superusers can access the Oban Web dashboard. The route is protected by :require_superuser plug.
Features:
- Real-time job queue monitoring
- Job filtering and search
- Job details and execution history
- Queue statistics and insights
- Job cancellation and retry
- Worker and plugin configuration
Router Configuration:
# lib/towerops_web/router.ex
scope "/" do
pipe_through [:browser, :require_authenticated_user, :require_superuser]
oban_dashboard "/oban"
end
Vendored Dependencies:
All Oban packages are vendored in vendor/ to avoid requiring authentication during CI/CD:
oban_pro(1.6.11) - Commercial Oban Pro featuresoban_web(2.11.7) - Web dashboardoban_met(1.0.5) - Metrics support
See vendor/README.md for update instructions.
Admin Features
User Impersonation
Location: /admin/users (superuser only)
Files:
- LiveView:
lib/towerops_web/live/admin/user_live/index.ex - Template:
lib/towerops_web/live/admin/user_live/index.html.heex - Controller:
lib/towerops_web/controllers/admin_controller.ex - Auth:
lib/towerops_web/user_auth.ex(start_impersonation/2, stop_impersonation/1)
Implementation Notes:
- Impersonate button is disabled only when
user.id == @current_scope.user.id(prevents impersonating yourself) - Superusers CAN impersonate other superusers (security feature, not bug)
- Backend checks in
UserAuth.start_impersonation/2prevent self-impersonation - All impersonation events are logged to audit_logs table via
Towerops.Admin.create_audit_log/1 - Stop impersonation: "Stop Impersonating" link appears in user menu when impersonating
GeoIP Database Management
Purpose: Import MaxMind GeoLite2-City database for IP-based country/city detection (used for GDPR cookie consent).
Files:
- Controller:
lib/towerops_web/controllers/api/v1/geoip_controller.ex - Mix Task:
lib/mix/tasks/geoip.import.ex - Schemas:
lib/towerops/geoip/location.ex,lib/towerops/geoip/block.ex - Lookup:
lib/towerops/geoip.ex - Migration:
priv/repo/migrations/*_create_geoip_tables.exs
Database Tables:
geoip_locations(~130,000 rows) - Countries, cities, regions with coordinatesgeoip_blocks(~3.5M rows) - IP ranges mapped to locations
Import Methods:
- Local Import (Development):
make geoip-import DIR=~/Downloads/GeoLite2-City-CSV_20260127/
- Production Import via API (requires
TOWEROPS_KEYenv var):
# Using Makefile (recommended)
make geoip-import-prod DIR=~/Downloads/GeoLite2-City-CSV_20260127/
# Or directly with mix
mix geoip.import ~/Downloads/GeoLite2-City-CSV_20260127/ --production
How It Works:
- Local mode (
make geoip-import): Direct database import via Ecto - Production mode (
make geoip-import-prodor--productionflag): Processes CSVs locally, sends data to API in batches of 5,000 records - API endpoint:
POST /admin/api/geoip/import(superuser only) - No need to upload large CSV files to production - only processed data is sent over HTTPS
- Production mode requires
TOWEROPS_KEYenvironment variable (superuser API token) - Default API URL:
https://towerops.net(override withTOWEROPS_URLenv var)
Implementation Notes:
- Uses batch inserts of 5,000 rows to stay within PostgreSQL's 65,535 parameter limit
- Automatically filters orphaned IP blocks (those without matching location references)
- Mix task processes CSVs locally and sends batches via API when
TOWEROPS_KEYis set - CSV files are NOT committed to git - imported separately in each environment
- Download GeoLite2-City CSV from MaxMind: https://dev.maxmind.com/geoip/geolite2-free-geolocation-data
- Required files:
GeoLite2-City-Blocks-IPv4.csv,GeoLite2-City-Locations-en.csv
Project-Specific Constraints
Key constraints from AGENTS.md (see that file for complete details):
- Use
mix precommitbefore committing changes - Use
:req(Req) for all HTTP requests - never use:httpoison,:tesla, or:httpc - Never use
daisyUI- write custom Tailwind components for world-class design - LiveView templates must start with
<Layouts.app flash={@flash}> - Use
<.icon name="hero-x-mark">component for icons, never Heroicons modules - Use LiveView streams for collections to avoid memory issues
- Forms must use
to_form/2in LiveView and<.form for={@form}>in templates - Never access changesets directly in templates - always use the form assign
- Tailwind v4 uses new import syntax - never use
@applyin CSS - when adding/updating hex modules, only use https://hex.pm
- assets are rebuilt on save and don't need to be built with mix assets.build
- be sure to always run mix format after you modify an elixir file
- when you run mix format, don't specify a file, let it format everything
Database Schema Critical Notes
Binary UUID Primary Keys: All tables MUST use :binary_id for primary keys, not bigint.
When creating new migrations with tables that have UUID primary keys:
create table(:table_name, primary_key: false) do
add :id, :binary_id, primary_key: true
# ...
end
Issue History: The snmp_neighbors table was initially created with a bigint primary key instead of :binary_id, causing DBConnection.EncodeError when trying to insert records. This happened because the migration file was correct but the table was created before the fix. Always verify the actual database schema matches the migration:
psql towerops_dev -c "\d table_name" # Check actual schema
If a table has the wrong primary key type, create a fix migration that drops and recreates the table.
SNMP Binary Data Handling: All SNMP values (chassis IDs, port IDs, system names) must be converted to printable strings before saving to the database. Non-printable binaries (like MAC addresses in binary format) should be converted to colon-separated hex format using String.printable?/1 checks. See lib/towerops/snmp/neighbor_discovery.ex:sanitize_string_field/1 for the pattern.
API Documentation
The API documentation is available at /docs/api and documents all public API endpoints.
Location:
- Controller:
lib/towerops_web/controllers/api_docs_controller.ex - HTML Module:
lib/towerops_web/controllers/api_docs_html.ex - Template:
lib/towerops_web/controllers/api_docs_html/index.html.heex - Route:
get "/docs/api", ApiDocsController, :indexinrouter.ex
Template Structure:
- Adapted from Tailwind UI Protocol template
- Sidebar navigation with section links
- Two-column layout: descriptions on left, code examples on right
- Sections: Introduction, Authentication, Errors, Sites API, Devices API
IMPORTANT: When updating API endpoints:
- Update the corresponding controller documentation (
@doccomments in controller) - Update
/docs/apitemplate with new endpoint documentation - Follow the existing pattern:
- Add endpoint to navigation sidebar if it's a new resource
- Document the endpoint with HTTP method badge (GET/POST/PATCH/DELETE)
- Include request example with
curl - Include JSON response example
- Document all parameters (required and optional)
- Wrap all code examples in
<%= raw(~S"""...""") %>to prevent HEEx from parsing curly braces as Elixir code
Code Example Pattern:
<pre class="p-4 text-sm text-zinc-100 overflow-x-auto"><code><%= raw(~S"""
curl https://towerops.net/api/v1/sites \
-H "Authorization: Bearer {token}"
""") %></code></pre>
JSON Example Pattern:
<pre class="p-4 text-sm text-zinc-100 overflow-x-auto"><code><%= raw(~S"""
{
"id": "uuid",
"name": "Site Name"
}
""") %></code></pre>
Current Documented Endpoints:
- Sites: GET /api/v1/sites, POST /api/v1/sites, GET /api/v1/sites/:id, PATCH /api/v1/sites/:id, DELETE /api/v1/sites/:id
- Devices: GET /api/v1/devices, POST /api/v1/devices, GET /api/v1/devices/:id, PATCH /api/v1/devices/:id, DELETE /api/v1/devices/:id
Background Job Architecture (Oban)
The application uses Oban (PostgreSQL-backed job queue) for all background processing. Jobs are distributed across the cluster and automatically retried on failure.
Oban Queues:
default(10 workers) - General background tasksdiscovery(10 workers) - SNMP discovery jobspollers(50 workers) - SNMP polling (one job per device)monitors(50 workers) - Device health monitoring (one job per device)maintenance(5 workers) - Cleanup and housekeeping tasks
Key Oban Workers:
-
DevicePollerWorker- SNMP polling for time-series data- Runs every 60 seconds per device (configurable via
check_interval_seconds) - Collects: sensor readings, interface stats, neighbor topology, ARP/MAC tables
- Uses Oban
uniqueconstraint to prevent duplicate jobs - Self-schedules next poll after completion
- Broadcasts PubSub events for real-time UI updates
- Runs every 60 seconds per device (configurable via
-
DeviceMonitorWorker- Health monitoring via ping- Runs every 30 seconds per device
- Performs ICMP ping checks to detect device up/down status
- Creates alerts on status changes
- Self-schedules next check after completion
-
DiscoveryWorker- One-time SNMP discovery- Triggered manually or on device creation
- Collects: system info, interfaces, sensors, neighbors
- Creates/updates SNMPDevice record with full topology
-
NeighborCleanupWorker- Maintenance task- Runs periodically to clean up stale neighbor data
- Deletes neighbors older than 24 hours
Automatic Job Management:
- Jobs are automatically started when
monitoring_enabledorsnmp_enabledis set totrue - Jobs are automatically stopped when disabled or device is deleted
- All job lifecycle managed in
Devicescontext (create_device, update_device, delete_device)
SNMP Polling Architecture
The application has two separate SNMP collection mechanisms:
1. Discovery (Towerops.Snmp.Discovery) - One-time or manual collection
- Triggered manually by user or via
DiscoveryWorkerOban job - Collects: system info, device identification, interfaces, sensors, neighbors
- Creates/updates the device record and all associated data
- Typically runs when equipment is first added or when user clicks "Rediscover"
2. Polling (DevicePollerWorker) - Continuous time-series collection
- Runs automatically every 60 seconds (configurable per equipment)
- Collects: sensor readings, interface statistics, neighbor topology
- Updates time-series data without recreating device records
- Keeps neighbor data fresh by polling LLDP/CDP every minute
- Architecture: Direct Oban worker (no GenServer or coordinator layer)
Neighbor Data Lifecycle:
- Discovery saves initial neighbors with 5-minute stale threshold
- Poller continuously updates neighbors every 60 seconds with 5-minute stale threshold
NeighborCleanupWorkerdeletes neighbors older than 24 hours (safety net)- This ensures neighbors stay current and topology changes are detected quickly
Polling Scheduling Details
How Polling Works (lib/towerops/workers/device_poller_worker.ex):
-
Job Creation:
DevicePollerWorker.start_polling/1(lines 66-70) - Creates initial Oban job for a device- Called from
Devices.create_device/1(line 275-278) whensnmp_enabled=true - Called from
Devices.update_device/2(line 379-396) when SNMP is enabled
-
Self-Scheduling Pattern:
- After each poll completes,
schedule_next_poll/2(lines 1066-1070) creates the next job - Uses
new(schedule_in: interval_seconds)to schedule the next run - Interval comes from
device.check_interval_secondsfield (default 300 seconds / 5 minutes) - Minimum interval enforced by
Application.get_env(:towerops, :snmp_min_poll_interval, 300)
- After each poll completes,
-
Unique Job Constraint:
- Uses Oban unique constraint with 60-second period (lines 53-63)
- Ensures only one polling job per device exists across the cluster
- Prevents duplicate jobs if multiple pods try to schedule simultaneously
-
Current Problem - Synchronized Polling:
- All devices start polling at the same time when first created
- Each device reschedules exactly
check_interval_secondsafter previous poll completes - Results in "thundering herd" - all devices poll at clock intervals (1:00, 1:05, 1:10, etc.)
- This creates load spikes on the server and Oban queue
-
Job Recovery (
lib/towerops/workers/job_health_check_worker.ex):- Runs every 10 minutes via Oban Cron
- Finds devices with
snmp_enabled=truebut no active polling job - Re-creates missing jobs (lines 63-80)
- Provides resilience after pod failures/crashes
Key Configuration:
- Default interval: 300 seconds (5 minutes) - defined in
lib/towerops/devices/device.ex:31 - Validation: 1-3600 seconds allowed (line 95)
- Oban queue:
:pollerswith 50 concurrent workers (config/dev.exs:39) - Queue concurrency can handle ~50 devices polling simultaneously
- Load Distribution - Staggered Polling:
- Offset Calculation: Uses deterministic hash-based offsets via
PollingOffset.calculate_offset/2 - Formula:
offset = :erlang.phash2(device_id) |> rem(interval_seconds) - Distribution: Devices spread evenly across full interval (0 to interval-1 seconds)
- Example: With 300s interval and 100 devices:
- Device A → offset 94s → polls at :01:34, :06:34, :11:34
- Device B → offset 171s → polls at :02:51, :07:51, :12:51
- Device C → offset 247s → polls at :04:07, :09:07, :14:07
- Deterministic: Same device always gets same offset (stable across restarts)
- Applied to: Both
DevicePollerWorker(SNMP polling) andDeviceMonitorWorker(health checks) - Implementation:
lib/towerops/workers/polling_offset.ex
- Offset Calculation: Uses deterministic hash-based offsets via
Related Files:
- Offset calculator:
lib/towerops/workers/polling_offset.ex - Poller worker:
lib/towerops/workers/device_poller_worker.ex - Monitor worker:
lib/towerops/workers/device_monitor_worker.ex - Device context:
lib/towerops/devices.ex(lines 264-285, 379-416) - Device schema:
lib/towerops/devices/device.ex - Config:
config/dev.exs(lines 31-61),config/runtime.exs(lines 103-131)
MIB Name Resolution (Rust NIF)
The application uses a Rust NIF (Native Implemented Function) via Rustler for fast, reliable SNMP MIB name resolution. This replaces unreliable Erlang SNMP modules for converting MIB names (like IEEE802dot11-MIB::dot11manufacturerProductName) to numeric OIDs (like 1.2.840.10036.3.1.2.1.3).
Architecture:
- NIF Module:
lib/towerops_native.ex- Elixir wrapper for Rust NIF functions - Rust Implementation:
native/towerops_native/src/mib.rs- MIB resolution logic usingsnmptranslatecommand - Crate Config:
native/towerops_native/Cargo.toml- Rustler dependency only (no snmptools) - Build Integration:
mix.exs- Rustler compiler configured viarustler_crates
How It Works:
-
Initialization (
lib/towerops/application.ex:31-40):- Application startup calls
ToweropsNative.load_mib_directory/1withpriv/mibspath - Sets environment variables for net-snmp MIB directories
- Logs success or warning if MIB loading fails
- Application startup calls
-
MIB Resolution (
lib/towerops/snmp/client.ex:406-427):- Client calls
resolve_mib_name/1before SNMP operations - Tries
ToweropsNative.resolve_oid/1first (Rust NIF) - Falls back to
SnmpKit.resolve/1if NIF fails - Returns numeric OID or original name if resolution fails
- Client calls
-
NIF Implementation (
native/towerops_native/src/mib.rs):- Uses
snmptranslate -Oncommand-line tool for MIB resolution - Auto-detects system MIB directories via
net-snmp-config --default-mibdirs - Falls back to common paths (
/usr/share/snmp/mibs, Homebrew paths on Mac) - Handles
MODULE::nameformat by prependingSNMPv2-MIB::if no module specified
- Uses
Performance:
- Resolution: ~2-10ms per OID (subprocess overhead from
snmptranslate) - Memory safe: Rustler catches Rust panics, preventing BEAM crashes
- Fallback: SnmpKit provides backward compatibility if NIF fails
System Dependencies:
- Development (macOS):
brew install net-snmp(keg-only, requires PKG_CONFIG_PATH) - Production (Docker):
apt-get install libsnmp-dev snmp-mibs-downloader - Runtime:
snmptranslatecommand must be in PATH
MIB Files:
- Located in
priv/mibs/directory (570+ vendor and standard MIB files) - Loaded from custom directory + system paths (combined via
MIBDIRSenv var) - Standard MIBs: SNMPv2-MIB, IF-MIB, etc. (from net-snmp)
- Vendor MIBs: Cisco, Ubiquiti, IEEE, etc. (in
priv/mibs/)
Testing:
- Unit tests:
test/towerops_native_test.exs - Tests marked with
@tag :skiprequire MIB files to be properly configured - Manual testing: Use IEx to call
ToweropsNative.resolve_oid("sysDescr")
Important Notes:
- SnmpKit is still used for actual SNMP protocol operations (get, walk, get_next, get_bulk)
- The NIF only handles MIB name resolution, not SNMP communication
- Rust compilation happens automatically via Mix when running
mix compile - Clean Rust builds with
cd native/towerops_native && cargo clean
Kubernetes Deployment
Prerequisites for Talos Kubernetes Cluster
The application requires the following infrastructure components to be installed in the cluster before deployment:
1. Core Infrastructure
cert-manager - Manages SSL/TLS certificates
- Required for automatic Let's Encrypt certificate provisioning
- Must have ClusterIssuers configured:
letsencrypt-prodandletsencrypt-staging
Traefik - Ingress controller
- Handles HTTP/HTTPS traffic routing
- Must be configured with
web(HTTP) andwebsecure(HTTPS) entrypoints - The application uses IngressRoute CRDs (not standard Ingress resources)
MetalLB - Load balancer for bare metal clusters
- Provides LoadBalancer service type support in on-premises environments
- Required for Traefik to get an external IP
FluxCD - GitOps continuous delivery
- Automatically syncs Kubernetes resources from Git repository
- Required components: source-controller, kustomize-controller, helm-controller, notification-controller
- Must have GitRepository resource pointing to this repository
- Must have Kustomization resource for the
k8s/directory
2. Storage
NFS Provisioner - Dynamic persistent volume provisioning
- Provides StorageClass for persistent volume claims
- Required for stateful workloads (if needed in the future)
3. CI/CD Integration
GitLab Agent - Connects cluster to GitLab for CI/CD
- Enables GitLab CI/CD pipelines to deploy to the cluster
- Required for automated deployments from GitLab CI
- Agent configuration managed in
.gitlab/agents/directory - Agent must be registered in GitLab project settings
4. Network Services
Newt - Pangolin tunnel forwarding
- Enables external access to internal services
- Required for Pangolin protocol forwarding
- Configuration specific to network topology
CoreDNS - Cluster DNS
- Standard Kubernetes DNS service (included in Talos by default)
Tailscale Operator - VPN mesh networking (optional)
- Provides secure access to cluster services
- Useful for development and administrative access
Secrets Management
All secrets are stored in 1Password and must be created manually in the cluster before deployment.
Required Secrets in towerops Namespace
-
gitlab-registry- Docker registry credentials- Type:
kubernetes.io/dockerconfigjson - Used to pull images from GitLab Container Registry
- Retrieve from 1Password or GitLab project settings
- Type:
-
towerops-secrets- Application secrets- Type:
Opaque - Fields:
RELEASE_COOKIE- Erlang distributed cookie for node clusteringSECRET_KEY_BASE- Phoenix secret key base for session encryption
- Stored in 1Password: "Towerops K8s Secrets"
- Type:
-
towerops-db- Database connection- Type:
Opaque - Fields:
POSTGRES_HOST- Database hostnamePOSTGRES_PORT- Database portPOSTGRES_DB- Database namePOSTGRES_USER- Database usernamePOSTGRES_PASSWORD- Database passwordDATABASE_URL- Full connection URL
- Stored in 1Password: "Towerops K8s Database"
- Type:
-
towerops-aws- AWS credentials- Type:
Opaque - Fields:
AWS_ACCESS_KEY_ID- AWS access keyAWS_SECRET_ACCESS_KEY- AWS secret keyAWS_REGION- AWS region (us-east-1)
- Stored in 1Password: "Towerops K8s AWS"
- Type:
Creating Secrets from 1Password
# Set 1Password account
export OP_ACCOUNT=YOOATCZZSVGH7AD6VABUVPORLI
# Create towerops-secrets
kubectl create secret generic towerops-secrets -n towerops \
--from-literal=RELEASE_COOKIE="$(op item get '74ske37wsm5tb4mjd4kiphwtji' --vault Private --account=$OP_ACCOUNT --fields label=RELEASE_COOKIE)" \
--from-literal=SECRET_KEY_BASE="$(op item get '74ske37wsm5tb4mjd4kiphwtji' --vault Private --account=$OP_ACCOUNT --fields label=SECRET_KEY_BASE)"
# Create towerops-db
kubectl create secret generic towerops-db -n towerops \
--from-literal=POSTGRES_HOST="$(op item get 'w2rg6bbstm5bxlskcrmn3g3on4' --vault Private --account=$OP_ACCOUNT --fields label=server)" \
--from-literal=POSTGRES_PORT="$(op item get 'w2rg6bbstm5bxlskcrmn3g3on4' --vault Private --account=$OP_ACCOUNT --fields label=port)" \
--from-literal=POSTGRES_DB="$(op item get 'w2rg6bbstm5bxlskcrmn3g3on4' --vault Private --account=$OP_ACCOUNT --fields label=database)" \
--from-literal=POSTGRES_USER="$(op item get 'w2rg6bbstm5bxlskcrmn3g3on4' --vault Private --account=$OP_ACCOUNT --fields label=username)" \
--from-literal=POSTGRES_PASSWORD="$(op item get 'w2rg6bbstm5bxlskcrmn3g3on4' --vault Private --account=$OP_ACCOUNT --fields label=password)" \
--from-literal=DATABASE_URL="$(op item get 'w2rg6bbstm5bxlskcrmn3g3on4' --vault Private --account=$OP_ACCOUNT --fields label=DATABASE_URL)"
# Create towerops-aws
kubectl create secret generic towerops-aws -n towerops \
--from-literal=AWS_ACCESS_KEY_ID="$(op item get 's4xwql6bp4huqxvjs6jkx5by6q' --vault Private --account=$OP_ACCOUNT --fields label=AWS_ACCESS_KEY_ID)" \
--from-literal=AWS_SECRET_ACCESS_KEY="$(op item get 's4xwql6bp4huqxvjs6jkx5by6q' --vault Private --account=$OP_ACCOUNT --fields label=AWS_SECRET_ACCESS_KEY)" \
--from-literal=AWS_REGION="$(op item get 's4xwql6bp4huqxvjs6jkx5by6q' --vault Private --account=$OP_ACCOUNT --fields label=AWS_REGION)"
Deployment Process
FluxCD Automatic Deployment
The application is automatically deployed via FluxCD GitOps:
- FluxCD monitors the Git repository for changes
- When changes are detected, it applies the manifests in
k8s/directory - Kustomize builds the manifests and applies them to the cluster
GitRepository Configuration:
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: towerops
namespace: flux-system
spec:
interval: 1m0s
url: ssh://git@gitlab.com/graham/towerops.git
ref:
branch: main
secretRef:
name: towerops-git
Kustomization Configuration:
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: towerops-app
namespace: flux-system
spec:
interval: 1m0s
path: ./k8s
prune: true
sourceRef:
kind: GitRepository
name: towerops
targetNamespace: towerops
Manual Deployment
If FluxCD is not available or for testing:
# Apply all resources using kustomize
kubectl apply -k k8s/
# Or apply individually
kubectl apply -f k8s/namespace.yaml
kubectl apply -f k8s/deployment.yaml
kubectl apply -f k8s/service.yaml
kubectl apply -f k8s/service-headless.yaml
kubectl apply -f k8s/certificate.yaml
kubectl apply -f k8s/ingressroute.yaml
Cluster Setup Checklist
When recreating a Talos cluster from scratch:
- Install Talos on nodes and bootstrap cluster
- Install MetalLB and configure IP address pool
- Install Traefik ingress controller
- Install cert-manager and configure ClusterIssuers
- Install FluxCD and configure GitRepository source
- Install GitLab Agent and connect to GitLab project
- Install Newt for Pangolin forwarding
- Install NFS provisioner (if using persistent storage)
- Create
toweropsnamespace - Create all required secrets from 1Password
- Create FluxCD Kustomization resource for
k8s/directory - Verify deployment with
kubectl get pods -n towerops - Check certificate provisioning with
kubectl get certificate -n towerops - Test ingress with
curl https://towerops.net/health
Troubleshooting
FluxCD reconciliation errors:
- Check kustomization status:
kubectl get kustomization -n flux-system - View detailed status:
kubectl describe kustomization towerops-app -n flux-system - Force reconciliation:
kubectl annotate kustomization towerops-app -n flux-system reconcile.fluxcd.io/requestedAt="$(date +%Y-%m-%dT%H:%M:%S%z)" --overwrite
Certificate issues:
- Check certificate status:
kubectl describe certificate towerops-net-cert -n towerops - View cert-manager logs:
kubectl logs -n cert-manager deployment/cert-manager - Verify ClusterIssuer:
kubectl get clusterissuer letsencrypt-prod -o yaml
Pod not starting:
- Check pod status:
kubectl describe pod -n towerops -l app=towerops - View logs:
kubectl logs -n towerops deployment/towerops - Verify secrets exist:
kubectl get secrets -n towerops
Testing Patterns
SNMP Mocking with Mox
The application uses Mox for SNMP client mocking in tests. Key patterns:
SNMP Adapter Mock Format:
snmp_adapter().get/3returns{:ok, value}or{:error, reason}snmp_adapter().walk/3returns{:ok, [%{oid: "...", value: ...}]}(list of maps, NOT a map)- Values from
get/3are already extracted (no type wrapper needed in mocks)
Mock Expectations:
# Mock get_multiple - it calls get/3 once per OID
expect(SnmpMock, :get, 6, fn _target, oid, _opts ->
case oid do
"1.3.6.1.2.1.1.1.0" -> {:ok, "Cisco IOS Software"} # sysDescr
"1.3.6.1.2.1.1.2.0" -> {:ok, [1, 3, 6, 1, 4, 1, 9]} # sysObjectID
"1.3.6.1.2.1.1.3.0" -> {:ok, 12_345} # sysUpTime (integer)
"1.3.6.1.2.1.1.4.0" -> {:ok, "admin@example.com"} # sysContact
"1.3.6.1.2.1.1.5.0" -> {:ok, "test-device"} # sysName
"1.3.6.1.2.1.1.6.0" -> {:ok, "Test Location"} # sysLocation
end
end)
# Mock walk - returns list of OID/value maps
expect(SnmpMock, :walk, fn _target, _oid, _opts ->
{:ok, [
%{oid: "1.3.6.1.4.1.9.9.91.1.1.1.1.1.1000", value: 8}
]}
end)
# Empty walk result (no data found)
expect(SnmpMock, :walk, fn _, _, _ ->
{:ok, []} # NOT {:ok, %{}}
end)
Common Pitfalls:
- ❌ Returning
{:ok, %{}}from walk - should be{:ok, []} - ❌ Returning type wrappers like
{:integer, 123}- values are already extracted - ❌ Not matching OIDs in get expectations - use pattern matching on OID string
- ❌ Expecting wrong number of calls -
get_multiple/2callsget/3once per OID
TimescaleDB Tests:
- Tests that query continuous aggregates should be tagged with
@tag :skip - Reason: Test database doesn't have TimescaleDB continuous aggregates configured
- Example:
get_hourly_stats/3,get_daily_stats/3,get_uptime_percentage/1
Test Organization
- Use
DataCasefor tests that need database access - Use
ConnCasefor controller/LiveView tests - Use
async: truefor tests that can run in parallel (most unit tests) - Use
async: falsefor tests with shared state (supervisor tests, integration tests) - never open the test coverage html files
- remember when working in rust to always run cargo fmt before committing
- never try to use npm to install js things, use esbuild built in to phoenix