network/CLAUDE.md
2026-07-19 14:42:02 -05:00

19 KiB
Raw Blame History

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

A small fleet of CLIs for documenting, auditing, and operating the vntx WISP — a multi-tower MikroTik backbone with cnMaestro and UISP managing the radios/CPE behind it. Source of truth for network documentation is NetBox at https://netbox.vntx.net/ (API token in env var NETBOX_KEY); source of truth for the routers is whatever the live .rsc exports say.

Repo Layout

The repo is a collection of independent CLIs that each operate on the same physical fleet but talk to different control planes. Each subproject has its own README/CLAUDE.md with details; only repo-wide context lives here.

Subdir Lang Talks to Purpose
mikrotik-tool/ Go MikroTik API-SSL + SSH Fan-out to every router in routers.yaml: list, export (saves <name>.rsc per router), inventory (rewrites inventory.yaml from live data + SNMP), api/schema (ad-hoc API introspection on one router), script (push a .rsc file via SSH and stream output). Authoritative inventory.yaml, subnets.yaml, radios.yaml, ipv6.md, and mpls.md live here.
mikrotik-tool-rs/ Rust MikroTik API-SSL + system ssh Newer rewrite of a subset of mikrotik-tool (list, export, run <cmd…>) with -r name1,name2 filter to scope a fan-out command. Has its own routers.yaml. See mikrotik-tool-rs/CLAUDE.md.
uisp/ Go (main.go) + Python (uisp.py) UISP API at uisp.vntx.net List/approve/upgrade UISP-managed devices. Auth via UISP_KEY env var. UNKNOWN-model devices are filtered by default (use --include-unknown).
cnmaestro/ Python (uv) cnMaestro Cloud OAuth2 Cleanup offline devices and bulk-upgrade firmware. uv sync + .env with client id/secret. See cnmaestro/README.md.

routers.yaml at the repo root is the config consumed by mikrotik-tool/ (Go); mikrotik-tool-rs/ has its own copy. Keep them aligned when adding/removing routers.

poll_snmp_site.sh is a one-off helper for SNMP-polling a list of IPs at a site (./poll_snmp_site.sh <site> <ip>...). The *.rsc files at the repo root are stale exports — fresh ones come out of mikrotik-tool or mikrotik-tool-rs.

Common Commands

# mikrotik-tool (Go)
cd mikrotik-tool && go build .
./mikrotik-tool list
./mikrotik-tool export                       # writes <name>.rsc per router into cwd
./mikrotik-tool api verona /ip/route/print
./mikrotik-tool schema verona /routing/bgp/connection
./mikrotik-tool script 494 494-v6-apply.rsc
./mikrotik-tool inventory                    # re-renders inventory.yaml in place
go test ./...                                # inventory_test.go

# mikrotik-tool-rs (Rust)
cd mikrotik-tool-rs && cargo build --release
cargo run -- list
cargo run -- -r verona,982 export            # filter to a subset before the verb
cargo run -- run /system identity print
cargo test                                   # full suite
cargo test encode_length                     # single test by name substring
cargo clippy --all-targets && cargo fmt

# uisp (Go binary, also a Python script)
cd uisp && go build .
./uisp list
./uisp approve <device>
./uisp upgrade
# python equivalent:
python3 uisp.py list

# cnmaestro (Python, uv)
cd cnmaestro && uv sync
uv run cnmaestro cleanup --dry-run
uv run cnmaestro upgrade --product ePMP,PMP --dry-run
uv run pytest
uv run pytest tests/test_devices.py::test_specific_thing   # single test

UISP_KEY and NETBOX_KEY need to be exported in the shell. cnmaestro reads cnmaestro/.env.

NetBox

  • URL: https://netbox.vntx.net/, API base https://netbox.vntx.net/api/
  • Auth: Authorization: Token $NETBOX_KEY
  • Convention when creating sites/devices: site slugs are lowercase-hyphenated, status active; routers are Manufacturer=MikroTik / Role=Router / primary IP = loopback /32 (10.254.254.x/32).
  • Prefix roles in use: Infrastructure, Customer, Management, Loopback.

MikroTik Router Access

Read-only fan-out (used by all the above tools): API-SSL on port 8729, user grahamro / password in CLAUDE.md history. The committed routers.yaml actually uses graham (not grahamro) because /export requires the ftp policy in addition to read,api, which the read-only group lacks. New routers: ensure the user's group has read,api,ftp (or extend the read group: /user group set read add-policy=ftp).

Every router's loopback is 10.254.254.x/32; see the Fleet Topology table below for the full mapping.

Network Topology Patterns

Access Point Placement

  • Access points are always placed in the top /24 of the management subnet for each tower
  • Example: For management subnet 10.10.16.0/20, APs are in 10.10.31.0/24 (the last /24 in that range)
  • Formula: For subnet X.Y.Z.0/20, APs are in X.Y.(Z+15).0/24

Ubiquiti MAC Prefixes

Common MAC address prefixes for Ubiquiti devices:

  • 00:04:56 (legacy)
  • 00:27:22 (legacy)
  • 04:18:D6
  • 24:A4:3C
  • 68:72:51
  • 80:2A:A8
  • F0:9F:C2
  • FC:EC:DA

MPLS / LDP

FastTrack is incompatible with MPLS on RouterOS 7

FastTrack bypasses the IP forwarding path that MPLS push/pop runs on, so any flow that gets fasttracked on a router whose path uses an MPLS-enabled interface can break — packets either hit the wrong interface or never get labeled, which presents as black-holing for specific source subnets that weren't fasttracked before. Symptoms: pings/SSH/TCP from one source IP work but the same destination is unreachable from another source on the same router; loopback-sourced traffic works but vlan-interface-sourced doesn't.

Fix: before each action=fasttrack-connection rule in chain=forward, add accept rules that match the MPLS-bound interface(s) so those flows never enter the fasttrack path:

/ip firewall filter
add chain=forward action=accept in-interface=<mpls-iface>  comment="bypass fasttrack for MPLS spine (in)"  place-before=<fasttrack-id>
add chain=forward action=accept out-interface=<mpls-iface> comment="bypass fasttrack for MPLS spine (out)" place-before=<fasttrack-id>

Customer→internet flows continue to fasttrack normally; only flows traversing the MPLS spine bypass it.

LDP doesn't label OSPF Type-5 externals by default

Prefixes redistributed via redistribute=connected (e.g., a /27 customer WAN handoff like 204.110.191.0/27) appear as Type-5 external LSAs and don't get LDP label bindings. Forward path to a labeled destination still works, but the return path is plain IP. If you need labeled bidirectional reach for a redistributed prefix, configure an LDP advertise-filter that explicitly includes it.

MPLS-MTU is the labeled-frame cap, not the IP-payload cap

mpls-mtu=1500 caps the labeled frame at 1500 bytes, which means an inner IP payload is limited to 1496 bytes — so 1500-byte DF customer traffic gets icmp-frag-needed. Use mpls-mtu=1508 for a 1500-byte IP payload + 4-byte label, with 4 bytes of headroom for one more stacked label. The AF11/AF24 radio l2mtu is 2024, so 1508 fits comfortably.

Fleet-wide MPLS topology

LDP runs IPv4-only across every backbone link in the network. Every backbone port has mpls-mtu=1508 set explicitly and a fasttrack-bypass pair (in/out) above the fasttrack-connection rule on both endpoints. Documented in mikrotik-tool/mpls.md.

verona ──AF11── climax ──AF24── core ──AF11── culleoka
                  │              │              │
                  │ 5GHz         │ AF11         │ AF11 (DOWN: power injector unplugged)
                  │              │              │
                 494           newhope ──AF24── lowry
                                 │
                                 │ 60 GHz
                                 │
                                982

Wait — that diagram's links are: climax↔494 (5 GHz airMAX, not AF11 — see below), core↔newhope (AF11), core↔982 (60 GHz), newhope↔lowry (AF24). The climax↔culleoka direct AF11 is currently down at the radio (physical issue), so culleoka traffic transits via core.

Fleet Topology

Routers and loopbacks

All ROS7 routers run RouterOS 7.21.4 long-term (post-2026-05-08 fleet upgrade). Edge runs ROS 6.49.18 (legacy, no MPLS, ignore for the spine).

Router Loopback (10.254.254.x) Hardware Site name
verona .101 CCR2004-16G-2S+ (arm64) verona
climax .102 CCR2004-16G-2S+ (arm64) climax
culleoka .104 CCR1009-7G-1C-1S+ (tile) culleoka
newhope .108 CCR1009-7G-1C-1S+ (tile) newhope
lowry .109 (tile) lowrycrossing
982 .110 (CCR, tile) 982
494 .111 (CCR, tile) 494
core .253 (CCR, arm64) at 380 core/380
edge .254 (legacy, ROS 6.49.18) edge

Tile-arch boxes can run MPLS but not ZeroTier (no .npk for tile).

Every link below has IPv4 LDP enabled at both ends, mpls-mtu=1508, and fasttrack-bypass rules in both directions on both routers.

Link Type A-side iface B-side iface /29 subnet l2mtu
verona↔climax AF11 verona ether3-climax-11ghz climax ether6-verona-11ghz 10.250.1.24/29 2024
climax↔core AF24 climax ether4-380-airfiber24 core ether5-climax 10.250.1.88/29 2024
climax↔494 5 GHz airMAX climax ether5-494 494 ether2-climax 10.250.1.64/29 1580
climax↔culleoka AF11 climax ether3-culleoka-11ghz culleoka ether1-climax-11ghz 10.250.1.8/29 2024 (link DOWN)
core↔culleoka AF11 core ether6-culleoka-11ghz culleoka ether6-380-11ghz 10.250.1.48/29 2024
core↔newhope AF11 core ether4-newhope newhope ether2-380 10.250.1.56/29 9000
core↔982 60 GHz core ether1-982-60ghz 982 ether7-380 10.250.1.32/29 9000
newhope↔lowry AF24 newhope ether6-lowrycrossing lowry ether1-newhope 10.250.1.104/29 9000
core↔edge wired core sfp-sfpplus1-edge-preseem + ether3-edge-direct edge ports 204.110.191.x n/a

l2mtu mismatches across the fleet are intentional per platform: AF11 base ports default to 2024 on CCR2004 / 1580 on smaller CCRs; jumbo-capable links (60 GHz, AF24-with-jumbo, fiber) go to 9000. Always raise both sides symmetrically when changing l2mtu — single-side raises usually work because Ethernet receivers accept anything ≤ their cap, but symmetric is the rule.

climax↔494 backhaul is 5 GHz airMAX, not AF11

Old docs called this hop AF11; it is actually a Ubiquiti airMAX AC PtP pair (SSID vntx_pr_494, WPA2, 40 MHz wide): AP "Climax to 494" (PBE-5AC-500) at 10.250.1.69 on the climax side, station "494 to Climax" (PBE-5AC-400) at 10.250.1.66 on the 494 side. Both answer SNMP v1 (kdyyJrT0Mm, airMAX MIB 1.3.6.1.4.1.41112.1.4); added to radios.yaml 2026-07-17. The station's mgmt IP (and the whole 494 tower) is only reachable across this RF link — when changing the channel, first make sure the station's frequency scan list covers the target, then change the AP (climax side), which stays reachable for rollback either way.

2026-07-17 flap incident: link ran on 5260 MHz (DFS); associations dropped for 4090 s every 45130 min (OSPF 40 s dead-interval timeouts on both routers; ethernet ports never dropped, radios never rebooted — association-uptime SNMP counters matched each OSPF flap). Local 5 GHz survey: climax APs on 5335/5545/5575 (airMAX) + 5750/5775 (ePMP), 494 ePMP omni on 5800 → U-NII-1 (51705250) was empty at both towers and is non-DFS, so the fix was moving the PtP there.

IGP / routing

  • OSPFv2 area backbone-v2 (id 0.0.0.0) on all spine links, SHA-512 auth with auth-id=1 and a shared key. PTP type. BFD is off on every wireless backbone interface (AF11/AF24/60GHz) — the global timers (/routing/bfd/configuration = 200ms×5 = 1s detection) are too aggressive for backhaul RF: a single >1s burst tore down OSPF on climax↔494 every ~1540 min until use-bfd=false was applied to both ends 2026-05-09. Wired fiber links may keep BFD if desired. If sub-second failover on a wireless link is genuinely needed, also loosen the BFD timers (e.g. min-rx/min-tx=500ms, multiplier=3) — do not flip use-bfd=true alone.
  • OSPFv3 area backbone-v3 for IPv6 (some interfaces only).
  • All instances redistribute=connected with passthrough filters (/routing filter rule chain=ospf-out rule="accept;").
  • Verona has a static default to 10.250.1.30 (climax) backing up the OSPF default — keep this; bouncing OSPF on verona doesn't blackhole it.
  • Distance-1 static routes also exist on climax for 204.110.191.0/27 so the home /27 has guaranteed return path even if OSPF redistribution hiccups.

Management subnets per tower

10.10.x.0/20 per site, top /24 reserved for APs (see Access Point Placement section). Authoritative mapping is in mikrotik-tool/inventory.yaml. Quick reference:

  • verona: 10.10.0.0/20
  • altoga (behind verona, no router): 10.10.16.0/20
  • climax: 10.10.48.0/20
  • core/380: 10.10.64.0/20
  • culleoka: 10.10.96.0/20
  • 982: 10.10.128.0/20
  • newhope: 10.10.144.0/20
  • 494: 10.10.160.0/20
  • lowry: 10.10.80.0/20

CGNAT pools: 100.64.x.x/22 per tower (see inventory.yaml / subnets.yaml).

graham's home network gotcha

graham's home connects to verona via vlan9_sfpplus1 carrying 204.110.191.0/27 (home router at .1, verona at .30). This /27 is a subnet of the verona hotspot's covered range (204.110.188.0/22). After any verona reboot, ensure /ip hotspot ip-binding has an entry: address=204.110.191.0/27 type=bypassed comment="graham home /27" — without it, hotspot drops all 204.110.191.x traffic in hs-unauth-to chain with icmp-host-prohibited. Symptom is "I can reach verona but nothing past it" from the home network.

graham's home network — multi-WAN failover (recursive check-gateway)

Home (RB5009, 10.0.19.254, not in routers.yaml — use a temp config with the same default creds to reach it via mikrotik-tool) fails over between TMO (distance 1), VNTX/verona (distance 2), and Starlink (distance 3, disabled) using the standard ROS recursive pattern, applied 2026-07-17 via mikrotik-tool/home-recursive-failover.rsc:

  • Probe pins: 4.2.2.1/32→192.168.12.1%ether6-tmobile, 4.2.2.2/32→204.110.191.30%ether5-vntx-static, 4.2.2.3/32→192.168.1.1%ether7-starlink, all scope=10, no check-gateway on the pins — they stay active whenever the interface has link, gluing probes to their WAN regardless of route state.
  • Defaults: gateway=4.2.2.x target-scope=11 check-gateway=ping — the check pings the probe IP through the pin; ~20s to go inactive, first success to return. tmo table default (eweka policy routing) has the same check so it falls back to main when TMO upstream dies.

Do not reintroduce netwatch enable/disable failover scripts. The old design (netwatch 2s/1s single-packet → /ip route disable) deadlocked: disabled=yes persists across reboots, and the probe was only tied to the VNTX path while its pin was active — so the verona hotspot-binding loss (verona pingable, transit dead) left the route disabled forever, and probes leaked out other WANs giving false "up" (TMO netwatch read up with the TMO interface physically down). Same lesson as the fleet BFD incident: 1-2s detection on AF RF with >1s bursts causes false failover; 1020s check-gateway damping is intentional.

IPv6 plan

Per-tower /44s + central server LAN at 2606:1c80::/64 on edge. Full allocation plan in mikrotik-tool/ipv6.md. NetBox has these as IPAM prefixes.

IPv6 forward firewall — must accept the whole /32

Every tower router's /ipv6 firewall filter chain=forward needs action=accept src-address=2606:1c80::/32 before the default deny. A per-tower /44 rule is wrong — customer traffic transits other routers to reach edge, so each transit hop blackholes other towers' /44s. Classic symptom: ICMPv6 ping works (there's an accept protocol=icmpv6 rule) but TCP/HTTP times out at connect (forwarded TCP hits default deny). The original ipv6.md template used in-interface-list=customer, but that list was never created on any router so the rule matched nothing — fixed fleet-wide 2026-05-14, and the now-redundant per-tower /44 and internal-to-internal forward rules were removed at the same time. edge has a permissive forward chain (no default-deny) so it didn't need the rule. Resolver/DNS roll-out (2606:1c80::240/::250 in /ip dns + the v6-dns DHCPv6 option) is also fleet-wide as of 2026-05-14.

graham's home network — IPv6

Home (RB5009, identity graham, 10.0.19.254) takes a /56 from verona via DHCPv6-PD on ether5-vntx-static (verona's verona-wired-pd dhcp-server on vlan9_sfpplus1, drawing from verona-cust-pd-1). Currently 2606:1c80:1001:e00::/56. Key config on home:

  • /ipv6 dhcp-client on ether5-vntx-static: request=prefix, pool-name=home-pd-vntx, pool-prefix-length=64 (must be 64, not 56, so the pool yields /64s — changing this only takes effect on a client disable/enable, not live), default-route-tables=v6-verona.
  • /ipv6 address with from-pool=home-pd-vntx address=::1 advertise=yes on bridge (→ :e00::1/64) and ether3-servers (→ :e01::1/64). Remaining 254 /64s of the /56 stay in the pool for internal use.
  • Egress policy: two /routing rules, order matters — (1) src-address=2606:1c80::/32 dst-address=2606:1c80:1001:e00::/56 action=lookup-only-in-table table=main so vntx-sourced traffic destined to the home LANs resolves via main's connected routes, then (2) src-address=2606:1c80::/32 action=lookup-only-in-table table=v6-verona for everything else. The verona /56 is single-homed, so traffic sourced from it must egress via verona only (T-Mobile/Starlink would BCP38-drop it). No v6 WAN failover for the internal /56, by design. The PD client's default lands in the v6-verona table, not main. Without rule (1) the broad rule also catches transit packets (echo replies, ICMPv6 time-exceeded) sourced from internal routers and destined back to the home LAN, and hairpins them to verona in a loop — symptom: home reaches the v6 internet fine but cannot ping any internal router and internal hops show * in traceroute6. Rule (1)'s dst-address is tied to the current delegated /56 — update it if the PD lease ever changes.
  • Stateless DHCPv6 servers (home-bridge-stateless, home-servers-stateless) carry the v6-dns option (Cloudflare).
  • Stale tunnerbroker /ipv6 pool (HE.net 2001:470:ba50::/48) is unused — leave or remove.
  • The mikrotik-tool script SSH path can't handle bare ROS verbs like release/disable/enable or :delay (flattenRSC only knows add/set/remove/print/get) — use set [find] disabled=yes/no instead, and sequence delays from the shell side.

Claude Assistant Guidelines

  • Any time Claude learns something new, automatically add it to CLAUDE.md

Development Best Practices

  • When making scripts, keep them as generic and reusable as possible