602 lines
21 KiB
Markdown
602 lines
21 KiB
Markdown
# TowerOps - Network Monitoring & Alerting Platform
|
|
## Implementation Plan
|
|
|
|
## Overview
|
|
Multi-tenant network monitoring and alerting application for tracking network equipment health via ping monitoring.
|
|
|
|
## Core Requirements
|
|
1. User accounts (email as username)
|
|
2. Organizations (multi-tenant)
|
|
3. Organization membership with permission levels
|
|
4. Site hierarchy
|
|
5. Equipment management (IP-based)
|
|
6. Automated ping monitoring (5-minute intervals)
|
|
7. Alerting system
|
|
|
|
---
|
|
|
|
## Data Model Design
|
|
|
|
### Users
|
|
- `id` (binary_id, PK)
|
|
- `email` (string, unique, username)
|
|
- `hashed_password` (string)
|
|
- `confirmed_at` (utc_datetime, nullable)
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
|
|
### Organizations
|
|
- `id` (binary_id, PK)
|
|
- `name` (string)
|
|
- `slug` (string, unique) - for URLs
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
|
|
### OrganizationMemberships
|
|
- `id` (binary_id, PK)
|
|
- `organization_id` (binary_id, FK -> organizations)
|
|
- `user_id` (binary_id, FK -> users)
|
|
- `role` (enum: owner, admin, member, viewer)
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
- Unique constraint on (organization_id, user_id)
|
|
|
|
**Permission Levels:**
|
|
- `owner` - Full control, can delete org, manage all settings
|
|
- `admin` - Can manage users, sites, equipment, view all
|
|
- `member` - Can add/edit equipment, view sites
|
|
- `viewer` - Read-only access
|
|
|
|
### Sites
|
|
- `id` (binary_id, PK)
|
|
- `organization_id` (binary_id, FK -> organizations)
|
|
- `parent_site_id` (binary_id, FK -> sites, nullable) - for hierarchy
|
|
- `name` (string)
|
|
- `description` (text, nullable)
|
|
- `location` (string, nullable)
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
|
|
### Equipment
|
|
- `id` (binary_id, PK)
|
|
- `site_id` (binary_id, FK -> sites)
|
|
- `name` (string)
|
|
- `ip_address` (string)
|
|
- `description` (text, nullable)
|
|
- `status` (enum: up, down, unknown)
|
|
- `last_checked_at` (utc_datetime, nullable)
|
|
- `last_status_change_at` (utc_datetime, nullable)
|
|
- `monitoring_enabled` (boolean, default: true)
|
|
- `check_interval_seconds` (integer, default: 300) - 5 minutes
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
|
|
### MonitoringChecks (historical log)
|
|
- `id` (binary_id, PK)
|
|
- `equipment_id` (binary_id, FK -> equipment)
|
|
- `status` (enum: success, failure)
|
|
- `response_time_ms` (integer, nullable)
|
|
- `checked_at` (utc_datetime)
|
|
- Index on (equipment_id, checked_at)
|
|
|
|
### Alerts
|
|
- `id` (binary_id, PK)
|
|
- `equipment_id` (binary_id, FK -> equipment)
|
|
- `alert_type` (enum: equipment_down, equipment_up)
|
|
- `triggered_at` (utc_datetime)
|
|
- `acknowledged_at` (utc_datetime, nullable)
|
|
- `acknowledged_by_id` (binary_id, FK -> users, nullable)
|
|
- `resolved_at` (utc_datetime, nullable)
|
|
- `email_sent_at` (utc_datetime, nullable)
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
|
|
### OrganizationInvitations
|
|
- `id` (binary_id, PK)
|
|
- `organization_id` (binary_id, FK -> organizations)
|
|
- `email` (string)
|
|
- `role` (enum: admin, member, viewer) - cannot invite as owner
|
|
- `token` (string, unique) - secure random token for invite link
|
|
- `invited_by_id` (binary_id, FK -> users)
|
|
- `accepted_at` (utc_datetime, nullable)
|
|
- `accepted_by_id` (binary_id, FK -> users, nullable)
|
|
- `expires_at` (utc_datetime) - invites expire after 7 days
|
|
- `inserted_at` / `updated_at` (utc_datetime)
|
|
|
|
---
|
|
|
|
## Architecture Components
|
|
|
|
### 1. Authentication & Authorization
|
|
|
|
**Use Phoenix.LiveView built-in patterns:**
|
|
- Email/password authentication
|
|
- Session-based auth with LiveView
|
|
- Password reset via email
|
|
|
|
**Authorization Strategy:**
|
|
- Context-based permissions (organization-scoped)
|
|
- Plugs for organization membership verification
|
|
- LiveView mount hooks for permission checks
|
|
- Helper functions: `can?(user, :action, resource)`
|
|
|
|
### 2. Multi-Tenancy
|
|
|
|
**Organization Scoping:**
|
|
- All queries scoped by current_organization
|
|
- LiveView assigns: `@current_user`, `@current_organization`, `@current_membership`
|
|
- Router organization switcher for users in multiple orgs
|
|
- URL structure: `/orgs/:org_slug/sites`, `/orgs/:org_slug/equipment`
|
|
|
|
### 3. Monitoring System
|
|
|
|
**Background Job Architecture:**
|
|
- Use GenServer or DynamicSupervisor for monitoring workers
|
|
- One worker per equipment (or batched by site)
|
|
- Quantum or similar for scheduling (or custom OTP solution)
|
|
- Use `:gen_icmp` or System.cmd("ping") for ping checks
|
|
|
|
**Monitoring Flow:**
|
|
1. Worker wakes up every N seconds (configurable per equipment)
|
|
2. Pings equipment IP address
|
|
3. Records result in monitoring_checks table
|
|
4. Updates equipment status if changed
|
|
5. Creates alert if status transitions (up->down or down->up)
|
|
6. Broadcasts status change via PubSub for real-time UI updates
|
|
|
|
### 4. Real-Time Updates
|
|
|
|
**Phoenix PubSub Topics:**
|
|
- `organization:#{org_id}:equipment:#{equipment_id}` - Equipment status
|
|
- `organization:#{org_id}:alerts` - New alerts
|
|
- `organization:#{org_id}:sites` - Site changes
|
|
|
|
**LiveView Integration:**
|
|
- Subscribe to relevant topics on mount
|
|
- Handle PubSub messages to update assigns
|
|
- Stream-based updates for lists
|
|
|
|
### 5. UI Structure
|
|
|
|
**LiveView Pages:**
|
|
- `/login`, `/register` - Authentication
|
|
- `/orgs` - Organization list/switcher
|
|
- `/orgs/:slug/dashboard` - Overview, active alerts, recent changes
|
|
- `/orgs/:slug/sites` - Site hierarchy tree view
|
|
- `/orgs/:slug/sites/:id` - Site detail with equipment list
|
|
- `/orgs/:slug/equipment` - All equipment list
|
|
- `/orgs/:slug/equipment/:id` - Equipment detail with check history
|
|
- `/orgs/:slug/alerts` - Alert history
|
|
- `/orgs/:slug/settings` - Org settings, members
|
|
|
|
**Components:**
|
|
- Organization switcher (header)
|
|
- Site tree navigator
|
|
- Equipment status badge
|
|
- Alert list/feed
|
|
- Permission-based action buttons
|
|
|
|
---
|
|
|
|
## Implementation Stages
|
|
|
|
### Stage 1: Foundation & Authentication
|
|
**Goal**: User authentication and basic org structure
|
|
**Success Criteria**:
|
|
- Users can register/login with email
|
|
- Users can create organizations
|
|
- Users can switch between organizations
|
|
- Basic navigation structure
|
|
- Tests passing
|
|
|
|
**Detailed Tasks**:
|
|
|
|
#### 1.1 User Authentication
|
|
- [ ] Run `mix phx.gen.auth Accounts User users` to scaffold auth system
|
|
- [ ] Review and customize generated code (email as username)
|
|
- [ ] Update user registration to create first organization
|
|
- [ ] Add tests for auth flows
|
|
|
|
**Files Created**:
|
|
- `lib/towerops/accounts.ex` - User context
|
|
- `lib/towerops/accounts/user.ex` - User schema
|
|
- `lib/towerops/accounts/user_token.ex` - Session tokens
|
|
- `lib/towerops_web/user_auth.ex` - Auth plugs
|
|
- `lib/towerops_web/controllers/user_session_controller.ex`
|
|
- `lib/towerops_web/controllers/user_registration_controller.ex`
|
|
- `lib/towerops_web/controllers/user_reset_password_controller.ex`
|
|
- `lib/towerops_web/controllers/user_settings_controller.ex`
|
|
- Migrations for users and user_tokens tables
|
|
|
|
#### 1.2 Organizations & Memberships
|
|
- [ ] Create migration: `mix ecto.gen.migration create_organizations`
|
|
- [ ] Create migration: `mix ecto.gen.migration create_organization_memberships`
|
|
- [ ] Create migration: `mix ecto.gen.migration create_organization_invitations`
|
|
- [ ] Create `lib/towerops/organizations.ex` context
|
|
- [ ] Create `lib/towerops/organizations/organization.ex` schema
|
|
- [ ] Create `lib/towerops/organizations/membership.ex` schema
|
|
- [ ] Create `lib/towerops/organizations/invitation.ex` schema
|
|
- [ ] Add slug generation for organizations (use Ecto changeset)
|
|
- [ ] Add tests for organizations context
|
|
|
|
**Migration Details**:
|
|
```elixir
|
|
# organizations table
|
|
create table(:organizations, primary_key: false) do
|
|
add :id, :binary_id, primary_key: true
|
|
add :name, :string, null: false
|
|
add :slug, :string, null: false
|
|
timestamps(type: :utc_datetime)
|
|
end
|
|
create unique_index(:organizations, [:slug])
|
|
|
|
# organization_memberships table
|
|
create table(:organization_memberships, primary_key: false) do
|
|
add :id, :binary_id, primary_key: true
|
|
add :organization_id, references(:organizations, type: :binary_id, on_delete: :delete_all)
|
|
add :user_id, references(:users, type: :binary_id, on_delete: :delete_all)
|
|
add :role, :string, null: false
|
|
timestamps(type: :utc_datetime)
|
|
end
|
|
create unique_index(:organization_memberships, [:organization_id, :user_id])
|
|
create index(:organization_memberships, [:user_id])
|
|
```
|
|
|
|
#### 1.3 Multi-Org Navigation
|
|
- [ ] Create organization switcher LiveView component
|
|
- [ ] Add current_organization plug to router
|
|
- [ ] Create `/orgs` LiveView (list user's organizations)
|
|
- [ ] Create `/orgs/new` LiveView (create organization)
|
|
- [ ] Update router with organization-scoped routes
|
|
- [ ] Add breadcrumb navigation component
|
|
- [ ] Add tests for organization switching
|
|
|
|
**Router Structure**:
|
|
```elixir
|
|
scope "/", ToweropsWeb do
|
|
pipe_through [:browser, :require_authenticated_user]
|
|
|
|
live "/orgs", OrgLive.Index
|
|
live "/orgs/new", OrgLive.New
|
|
end
|
|
|
|
scope "/orgs/:org_slug", ToweropsWeb do
|
|
pipe_through [:browser, :require_authenticated_user, :load_current_organization]
|
|
|
|
live "/", DashboardLive
|
|
# Future: sites, equipment, etc.
|
|
end
|
|
```
|
|
|
|
#### 1.4 Authorization System
|
|
- [ ] Create `lib/towerops/organizations/policy.ex` for permission checks
|
|
- [ ] Add `can?/3` helper function
|
|
- [ ] Create `:load_current_organization` plug
|
|
- [ ] Add permission checks to LiveView mount callbacks
|
|
- [ ] Add tests for authorization
|
|
|
|
**Permission Matrix**:
|
|
| Action | Owner | Admin | Member | Viewer |
|
|
|--------|-------|-------|--------|--------|
|
|
| View org | ✓ | ✓ | ✓ | ✓ |
|
|
| Edit org settings | ✓ | ✓ | ✗ | ✗ |
|
|
| Delete org | ✓ | ✗ | ✗ | ✗ |
|
|
| Manage members | ✓ | ✓ | ✗ | ✗ |
|
|
| Add/edit equipment | ✓ | ✓ | ✓ | ✗ |
|
|
| View equipment | ✓ | ✓ | ✓ | ✓ |
|
|
|
|
#### 1.5 Basic UI & Layouts
|
|
- [ ] Update `layouts.ex` to include org switcher in header
|
|
- [ ] Create organization badge component
|
|
- [ ] Style navigation with Tailwind
|
|
- [ ] Add flash message styling
|
|
- [ ] Ensure responsive design
|
|
|
|
### Stage 2: Sites & Equipment Management
|
|
**Goal**: CRUD for sites and equipment
|
|
**Success Criteria**:
|
|
- Users can create/edit/delete sites
|
|
- Sites can have parent sites (hierarchy)
|
|
- Users can add equipment to sites
|
|
- Equipment has IP address and basic info
|
|
|
|
**Tasks**:
|
|
1. Generate sites schema and LiveViews
|
|
2. Build site hierarchy tree component
|
|
3. Generate equipment schema and LiveViews
|
|
4. Add IP address validation
|
|
5. Permission checks on all actions
|
|
|
|
### Stage 3: Monitoring System ✓ COMPLETE
|
|
**Goal**: Automated ping monitoring with TimescaleDB
|
|
**Success Criteria**:
|
|
- ✓ Equipment is pinged at configurable intervals
|
|
- ✓ Status updates in real-time via PubSub
|
|
- ✓ Monitoring checks are logged
|
|
- ✓ TimescaleDB hypertable for efficient time-series storage
|
|
- ✓ Automatic data retention and compression
|
|
- ✓ Continuous aggregates for dashboard metrics
|
|
|
|
**Tasks**:
|
|
1. ✓ Design monitoring worker architecture (GenServer + DynamicSupervisor)
|
|
2. ✓ Implement ping functionality (System.cmd with OS-specific args)
|
|
3. ✓ Create monitoring_checks schema with TimescaleDB hypertable
|
|
4. ✓ Build GenServer workers for monitoring
|
|
5. ✓ Add PubSub broadcasting
|
|
6. ✓ Update LiveView to receive real-time updates
|
|
7. ✓ Configure TimescaleDB retention policies
|
|
8. ✓ Create continuous aggregates for metrics
|
|
|
|
**Completed**: 2025-12-21
|
|
**Files Created**:
|
|
- `lib/towerops/monitoring/check.ex` - MonitoringCheck schema
|
|
- `lib/towerops/monitoring.ex` - Monitoring context
|
|
- `lib/towerops/monitoring/ping.ex` - Ping functionality
|
|
- `lib/towerops/monitoring/equipment_monitor.ex` - GenServer for monitoring individual equipment
|
|
- `lib/towerops/monitoring/supervisor.ex` - Supervisor for managing monitor workers
|
|
- `priv/repo/migrations/*_create_monitoring_checks.exs` - Database migration with TimescaleDB hypertable
|
|
|
|
**TimescaleDB Integration**:
|
|
- `monitoring_checks` converted to hypertable partitioned by `checked_at`
|
|
- Retention policy: Keep raw data for 90 days, then automatically delete
|
|
- Compression policy: Compress chunks older than 7 days
|
|
- Continuous aggregates: Hourly and daily rollups for dashboard performance
|
|
|
|
### Stage 4: Alerting ✓ COMPLETE
|
|
**Goal**: Alert generation and management
|
|
**Success Criteria**:
|
|
- ✓ Alerts created on status changes
|
|
- ✓ Users can view alert history
|
|
- ✓ Users can acknowledge alerts
|
|
- ✓ Real-time alert notifications via PubSub
|
|
|
|
**Tasks**:
|
|
1. ✓ Create alerts schema
|
|
2. ✓ Build alert creation logic in monitoring workers
|
|
3. ✓ Create alert LiveView pages
|
|
4. ✓ Add alert acknowledgment
|
|
5. ✓ Alert notifications in UI
|
|
6. ✓ Dashboard integration with active alerts
|
|
|
|
**Completed**: 2025-12-21
|
|
**Files Created**:
|
|
- `priv/repo/migrations/*_create_alerts.exs` - Alerts table migration
|
|
- `lib/towerops/alerts/alert.ex` - Alert schema
|
|
- `lib/towerops/alerts.ex` - Alerts context
|
|
- `lib/towerops_web/live/alert_live/index.ex` - Alerts listing page
|
|
- `lib/towerops_web/live/alert_live/index.html.heex` - Alerts UI
|
|
|
|
**Files Modified**:
|
|
- `lib/towerops/monitoring/equipment_monitor.ex` - Alert creation on status changes
|
|
- `lib/towerops_web/live/dashboard_live.ex` - Active alerts display
|
|
- `lib/towerops_web/live/dashboard_live.html.heex` - Dashboard UI updates
|
|
- `lib/towerops_web/router.ex` - Alert routes
|
|
|
|
### Stage 5: Polish & Production ✓ COMPLETE
|
|
**Goal**: Production-ready application
|
|
**Success Criteria**:
|
|
- ✓ All tests passing (154 tests, 60.83% coverage)
|
|
- ✓ Polished UI with Tailwind (removed daisyUI, custom components)
|
|
- ✓ Email notifications configured
|
|
- ✓ Production deployment ready
|
|
|
|
**Tasks**:
|
|
1. ✓ Comprehensive test coverage (100% on core business logic)
|
|
2. ✓ UI/UX improvements (Dashboard redesigned with custom Tailwind)
|
|
3. ✓ Email alert notifications (SMTP-based, sent to owners/admins)
|
|
4. ✓ Performance optimization (Added monitoring_enabled index)
|
|
5. ✓ Documentation (Complete DEPLOYMENT.md guide)
|
|
|
|
**Completed**: 2025-12-24
|
|
**Files Created**:
|
|
- `lib/towerops/alerts/alert_notifier.ex` - Email notification service
|
|
- `DEPLOYMENT.md` - Comprehensive production deployment guide
|
|
- `priv/repo/migrations/*_add_email_sent_at_to_alerts.exs` - Email tracking
|
|
- `priv/repo/migrations/*_add_monitoring_enabled_index.exs` - Performance optimization
|
|
|
|
**Files Modified**:
|
|
- `lib/towerops/alerts/alert.ex` - Added email_sent_at field
|
|
- `lib/towerops/alerts.ex` - Added send_alert_notification/1 function
|
|
- `lib/towerops/organizations.ex` - Added list_organization_notification_recipients/1
|
|
- `lib/towerops/monitoring/equipment_monitor.ex` - Integrated email sending
|
|
- `lib/towerops_web/live/dashboard_live.html.heex` - Modern Tailwind UI
|
|
|
|
**Email Notification System**:
|
|
- Sends emails to organization owners and admins
|
|
- Equipment down alerts with details
|
|
- Equipment recovery notifications
|
|
- Disabled in test environment to avoid connection issues
|
|
- Tracks email_sent_at timestamp
|
|
|
|
**UI Improvements**:
|
|
- Removed all daisyUI classes
|
|
- Custom Tailwind components with dark mode support
|
|
- Modern card-based dashboard layout
|
|
- Proper semantic color scheme (zinc, blue, red, green, amber)
|
|
- Improved accessibility and responsiveness
|
|
|
|
**Performance Optimizations**:
|
|
- Added index on equipment.monitoring_enabled for fast filtering
|
|
- All existing queries already optimized with proper indexes
|
|
- TimescaleDB compression and retention in production
|
|
|
|
**Production Readiness**:
|
|
- Complete deployment documentation
|
|
- Environment-specific configurations
|
|
- SSL/TLS setup guide
|
|
- Docker and systemd examples
|
|
- Database backup procedures
|
|
- Troubleshooting guide
|
|
|
|
### Stage 6: Distributed Monitoring Agents (Future)
|
|
**Goal**: Deploy-able Rust-based monitoring agents for internal network monitoring
|
|
**Status**: PLANNED - Not yet started
|
|
|
|
**Overview**:
|
|
Customer-deployable Rust binary that runs inside customer networks to monitor equipment from within their own infrastructure. Agents authenticate using deployment keys and report monitoring results back to the TowerOps platform.
|
|
|
|
**Success Criteria**:
|
|
- Rust agent binary can be deployed on customer networks (Linux, Windows, macOS)
|
|
- Agent authenticates with deployment key tied to organization
|
|
- Agent pings equipment reachable on local network
|
|
- Results reported back to TowerOps API
|
|
- Equipment can be configured to use agent-based or platform-based monitoring
|
|
- Agent status visible in TowerOps dashboard
|
|
|
|
**Architecture**:
|
|
- **Rust Agent**: Lightweight binary for customer deployment
|
|
- Ping functionality using OS-native ICMP
|
|
- Configurable check intervals
|
|
- Local caching/queue for offline resilience
|
|
- Secure API communication over HTTPS
|
|
- Auto-update capability
|
|
|
|
- **Deployment Keys**: Scoped API credentials
|
|
- Per-organization deployment keys
|
|
- Limited scope (can only submit monitoring results)
|
|
- Revocable from TowerOps dashboard
|
|
- Multiple keys per organization for different sites/networks
|
|
|
|
- **API Endpoints**: Backend support for agent communication
|
|
- POST /api/v1/monitoring/checks - Submit monitoring results
|
|
- GET /api/v1/monitoring/config - Fetch equipment list for agent
|
|
- Authentication via deployment key header
|
|
|
|
**Tasks**:
|
|
1. Design deployment key schema and API
|
|
2. Create agent registration and key management in TowerOps
|
|
3. Build Rust monitoring agent
|
|
- ICMP ping implementation
|
|
- API client for result submission
|
|
- Configuration management
|
|
- Error handling and retry logic
|
|
4. Add equipment assignment to agents
|
|
5. Agent status monitoring in dashboard
|
|
6. Documentation for agent deployment
|
|
|
|
**Database Changes Needed**:
|
|
```elixir
|
|
# deployment_keys table
|
|
- id (binary_id)
|
|
- organization_id (FK)
|
|
- name (string) - Human-readable name for the key
|
|
- key_hash (string) - Hashed version of the key
|
|
- last_used_at (utc_datetime)
|
|
- created_by_id (FK -> users)
|
|
- revoked_at (utc_datetime, nullable)
|
|
|
|
# monitoring_agents table
|
|
- id (binary_id)
|
|
- organization_id (FK)
|
|
- deployment_key_id (FK)
|
|
- name (string)
|
|
- version (string) - Agent version
|
|
- last_seen_at (utc_datetime)
|
|
- status (enum: active, inactive, error)
|
|
|
|
# equipment table additions
|
|
- monitoring_agent_id (FK -> monitoring_agents, nullable)
|
|
- monitoring_source (enum: platform, agent) - Where monitoring happens
|
|
```
|
|
|
|
**Notes**:
|
|
- This is a significant feature and will be implemented much later
|
|
- Current platform-based monitoring (Stage 3) remains as fallback
|
|
- Allows monitoring of internal/private networks not accessible from internet
|
|
- Agent runs continuously, not triggered from platform
|
|
|
|
---
|
|
|
|
## Technical Decisions
|
|
|
|
### Database
|
|
- PostgreSQL with Ecto
|
|
- Use binary_id (UUID) for all primary keys (already configured)
|
|
- Indexes on foreign keys and frequently queried fields
|
|
|
|
**TimescaleDB for Time-Series Data**:
|
|
- TimescaleDB extension enabled on PostgreSQL
|
|
- `monitoring_checks` table converted to hypertable
|
|
- Automatic partitioning by time (checked_at column)
|
|
- Retention policy: 90 days for raw data
|
|
- Compression policy: Compress chunks older than 7 days
|
|
- Continuous aggregates for dashboard performance:
|
|
- `monitoring_checks_hourly` - Hourly rollups (avg response time, success rate)
|
|
- `monitoring_checks_daily` - Daily rollups for long-term trends
|
|
|
|
**Why TimescaleDB**:
|
|
- Built as PostgreSQL extension (not separate database)
|
|
- Works seamlessly with Ecto
|
|
- Optimized for time-series queries
|
|
- Automatic data lifecycle management
|
|
- Fast aggregations for dashboards
|
|
- Scales to millions of monitoring checks
|
|
|
|
### Background Jobs
|
|
**Options:**
|
|
1. Custom OTP solution with GenServer + Process.send_after
|
|
2. Quantum scheduler
|
|
3. Oban (more heavyweight but robust)
|
|
|
|
**Recommendation**: Start with custom OTP, migrate to Oban if needed
|
|
|
|
### Real-time Communication
|
|
- Phoenix PubSub for server-side messaging
|
|
- LiveView for UI updates
|
|
- No WebSocket client code needed
|
|
|
|
### Ping Implementation
|
|
**Options:**
|
|
1. `:gen_icmp` library (requires raw sockets, might need permissions)
|
|
2. System.cmd("ping") - simpler, cross-platform
|
|
3. HTTP health checks (future enhancement)
|
|
|
|
**Recommendation**: System.cmd("ping") for MVP, abstract for future protocols
|
|
|
|
---
|
|
|
|
## Decisions Made
|
|
|
|
1. **Email notifications**: ✓ Yes - Users receive email alerts on status changes
|
|
2. **Multi-org users**: ✓ Yes - Users can belong to multiple organizations
|
|
3. **Invite system**: ✓ Email invitation system (secure token-based invites)
|
|
4. **Check intervals**: ✓ Customizable per equipment (stored in equipment.check_interval_seconds)
|
|
|
|
## Additional Schema Needed
|
|
|
|
Based on decisions:
|
|
- **OrganizationInvitations** table for email invite workflow
|
|
- **email_sent_at** field in Alerts for tracking email delivery
|
|
|
|
## Open Questions
|
|
|
|
1. **Alert escalation**: Any escalation policies (e.g., page admin if down > X minutes)?
|
|
2. **Retention**: How long to keep monitoring_checks history?
|
|
3. **Equipment types**: Just ping for now, or plan for SNMP, HTTP, etc.?
|
|
4. **Site hierarchy depth**: Any limit on site nesting levels?
|
|
|
|
---
|
|
|
|
## Status: Stage 5 Complete - Production Ready! 🚀
|
|
|
|
**Completed Stages**:
|
|
- ✓ Stage 1: Foundation & Authentication (2025-12-21)
|
|
- ✓ Stage 2: Sites & Equipment Management (2025-12-21)
|
|
- ✓ Stage 3: Monitoring System with TimescaleDB (2025-12-21)
|
|
- ✓ Stage 4: Alerting (2025-12-21)
|
|
- ✓ Stage 5: Polish & Production (2025-12-24)
|
|
|
|
**Current Status**: All tests passing (154/154), production-ready
|
|
|
|
**Features Implemented**:
|
|
- Multi-tenant organizations with role-based access
|
|
- Hierarchical site management
|
|
- Equipment monitoring with ping checks
|
|
- TimescaleDB time-series optimization (production-ready)
|
|
- Automatic alerting on status changes
|
|
- Email notifications to owners/admins
|
|
- Real-time dashboard with LiveView
|
|
- Alert acknowledgment system
|
|
- Modern Tailwind UI (no daisyUI)
|
|
- Comprehensive deployment documentation
|
|
- Performance optimizations (database indexes)
|
|
|
|
**Next Steps**:
|
|
- Stage 6: Distributed Monitoring Agents (Future - see above for details)
|
|
- Deploy to production following DEPLOYMENT.md guide
|
|
|
|
Last updated: 2025-12-24
|