354 lines
No EOL
19 KiB
Markdown
354 lines
No EOL
19 KiB
Markdown
# CLAUDE.md
|
||
|
||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||
|
||
## Project Overview
|
||
|
||
A small fleet of CLIs for documenting, auditing, and operating the vntx WISP — a multi-tower MikroTik backbone with cnMaestro and UISP managing the radios/CPE behind it. Source of truth for network documentation is NetBox at `https://netbox.vntx.net/` (API token in env var `NETBOX_KEY`); source of truth for the routers is whatever the live `.rsc` exports say.
|
||
|
||
## Repo Layout
|
||
|
||
The repo is a collection of independent CLIs that each operate on the same physical fleet but talk to different control planes. Each subproject has its own README/CLAUDE.md with details; only repo-wide context lives here.
|
||
|
||
| Subdir | Lang | Talks to | Purpose |
|
||
|--|--|--|--|
|
||
| `mikrotik-tool/` | Go | MikroTik API-SSL + SSH | Fan-out to every router in `routers.yaml`: `list`, `export` (saves `<name>.rsc` per router), `inventory` (rewrites `inventory.yaml` from live data + SNMP), `api`/`schema` (ad-hoc API introspection on one router), `script` (push a `.rsc` file via SSH and stream output). Authoritative `inventory.yaml`, `subnets.yaml`, `radios.yaml`, `ipv6.md`, and `mpls.md` live here. |
|
||
| `mikrotik-tool-rs/` | Rust | MikroTik API-SSL + system `ssh` | Newer rewrite of a subset of `mikrotik-tool` (`list`, `export`, `run <cmd…>`) with `-r name1,name2` filter to scope a fan-out command. Has its own `routers.yaml`. See `mikrotik-tool-rs/CLAUDE.md`. |
|
||
| `uisp/` | Go (`main.go`) + Python (`uisp.py`) | UISP API at `uisp.vntx.net` | List/approve/upgrade UISP-managed devices. Auth via `UISP_KEY` env var. UNKNOWN-model devices are filtered by default (use `--include-unknown`). |
|
||
| `cnmaestro/` | Python (uv) | cnMaestro Cloud OAuth2 | Cleanup offline devices and bulk-upgrade firmware. `uv sync` + `.env` with client id/secret. See `cnmaestro/README.md`. |
|
||
|
||
`routers.yaml` at the repo root is the config consumed by `mikrotik-tool/` (Go); `mikrotik-tool-rs/` has its own copy. Keep them aligned when adding/removing routers.
|
||
|
||
`poll_snmp_site.sh` is a one-off helper for SNMP-polling a list of IPs at a site (`./poll_snmp_site.sh <site> <ip>...`). The `*.rsc` files at the repo root are stale exports — fresh ones come out of `mikrotik-tool` or `mikrotik-tool-rs`.
|
||
|
||
## Common Commands
|
||
|
||
```bash
|
||
# mikrotik-tool (Go)
|
||
cd mikrotik-tool && go build .
|
||
./mikrotik-tool list
|
||
./mikrotik-tool export # writes <name>.rsc per router into cwd
|
||
./mikrotik-tool api verona /ip/route/print
|
||
./mikrotik-tool schema verona /routing/bgp/connection
|
||
./mikrotik-tool script 494 494-v6-apply.rsc
|
||
./mikrotik-tool inventory # re-renders inventory.yaml in place
|
||
go test ./... # inventory_test.go
|
||
|
||
# mikrotik-tool-rs (Rust)
|
||
cd mikrotik-tool-rs && cargo build --release
|
||
cargo run -- list
|
||
cargo run -- -r verona,982 export # filter to a subset before the verb
|
||
cargo run -- run /system identity print
|
||
cargo test # full suite
|
||
cargo test encode_length # single test by name substring
|
||
cargo clippy --all-targets && cargo fmt
|
||
|
||
# uisp (Go binary, also a Python script)
|
||
cd uisp && go build .
|
||
./uisp list
|
||
./uisp approve <device>
|
||
./uisp upgrade
|
||
# python equivalent:
|
||
python3 uisp.py list
|
||
|
||
# cnmaestro (Python, uv)
|
||
cd cnmaestro && uv sync
|
||
uv run cnmaestro cleanup --dry-run
|
||
uv run cnmaestro upgrade --product ePMP,PMP --dry-run
|
||
uv run pytest
|
||
uv run pytest tests/test_devices.py::test_specific_thing # single test
|
||
```
|
||
|
||
`UISP_KEY` and `NETBOX_KEY` need to be exported in the shell. `cnmaestro` reads `cnmaestro/.env`.
|
||
|
||
## NetBox
|
||
|
||
- URL: `https://netbox.vntx.net/`, API base `https://netbox.vntx.net/api/`
|
||
- Auth: `Authorization: Token $NETBOX_KEY`
|
||
- Convention when creating sites/devices: site slugs are lowercase-hyphenated, status `active`; routers are Manufacturer=MikroTik / Role=Router / primary IP = loopback `/32` (`10.254.254.x/32`).
|
||
- Prefix roles in use: Infrastructure, Customer, Management, Loopback.
|
||
|
||
## MikroTik Router Access
|
||
|
||
Read-only fan-out (used by all the above tools): API-SSL on port 8729, user `grahamro` / password in `CLAUDE.md` history. The committed `routers.yaml` actually uses `graham` (not `grahamro`) because `/export` requires the `ftp` policy in addition to `read,api`, which the read-only group lacks. New routers: ensure the user's group has `read,api,ftp` (or extend the read group: `/user group set read add-policy=ftp`).
|
||
|
||
Every router's loopback is `10.254.254.x/32`; see the **Fleet Topology** table below for the full mapping.
|
||
|
||
## Network Topology Patterns
|
||
|
||
### Access Point Placement
|
||
- Access points are always placed in the top /24 of the management subnet for each tower
|
||
- Example: For management subnet 10.10.16.0/20, APs are in 10.10.31.0/24 (the last /24 in that range)
|
||
- Formula: For subnet X.Y.Z.0/20, APs are in X.Y.(Z+15).0/24
|
||
|
||
### Ubiquiti MAC Prefixes
|
||
Common MAC address prefixes for Ubiquiti devices:
|
||
- 00:04:56 (legacy)
|
||
- 00:27:22 (legacy)
|
||
- 04:18:D6
|
||
- 24:A4:3C
|
||
- 68:72:51
|
||
- 80:2A:A8
|
||
- F0:9F:C2
|
||
- FC:EC:DA
|
||
|
||
## MPLS / LDP
|
||
|
||
### FastTrack is incompatible with MPLS on RouterOS 7
|
||
FastTrack bypasses the IP forwarding path that MPLS push/pop runs on, so any
|
||
flow that gets fasttracked on a router whose path uses an MPLS-enabled
|
||
interface can break — packets either hit the wrong interface or never get
|
||
labeled, which presents as black-holing for specific source subnets that
|
||
weren't fasttracked before. Symptoms: pings/SSH/TCP from one source IP work
|
||
but the same destination is unreachable from another source on the same
|
||
router; loopback-sourced traffic works but vlan-interface-sourced doesn't.
|
||
|
||
**Fix:** before each `action=fasttrack-connection` rule in `chain=forward`,
|
||
add `accept` rules that match the MPLS-bound interface(s) so those flows
|
||
never enter the fasttrack path:
|
||
```
|
||
/ip firewall filter
|
||
add chain=forward action=accept in-interface=<mpls-iface> comment="bypass fasttrack for MPLS spine (in)" place-before=<fasttrack-id>
|
||
add chain=forward action=accept out-interface=<mpls-iface> comment="bypass fasttrack for MPLS spine (out)" place-before=<fasttrack-id>
|
||
```
|
||
Customer→internet flows continue to fasttrack normally; only flows traversing
|
||
the MPLS spine bypass it.
|
||
|
||
### LDP doesn't label OSPF Type-5 externals by default
|
||
Prefixes redistributed via `redistribute=connected` (e.g., a /27 customer
|
||
WAN handoff like 204.110.191.0/27) appear as Type-5 external LSAs and don't
|
||
get LDP label bindings. Forward path to a labeled destination still works,
|
||
but the return path is plain IP. If you need labeled bidirectional reach
|
||
for a redistributed prefix, configure an LDP advertise-filter that
|
||
explicitly includes it.
|
||
|
||
### MPLS-MTU is the labeled-frame cap, not the IP-payload cap
|
||
`mpls-mtu=1500` caps the *labeled* frame at 1500 bytes, which means an inner
|
||
IP payload is limited to 1496 bytes — so 1500-byte DF customer traffic gets
|
||
icmp-frag-needed. Use `mpls-mtu=1508` for a 1500-byte IP payload + 4-byte
|
||
label, with 4 bytes of headroom for one more stacked label. The AF11/AF24
|
||
radio l2mtu is 2024, so 1508 fits comfortably.
|
||
|
||
### Fleet-wide MPLS topology
|
||
LDP runs IPv4-only across every backbone link in the network. Every backbone
|
||
port has `mpls-mtu=1508` set explicitly and a fasttrack-bypass pair (in/out)
|
||
above the `fasttrack-connection` rule on both endpoints. Documented in
|
||
`mikrotik-tool/mpls.md`.
|
||
|
||
```
|
||
verona ──AF11── climax ──AF24── core ──AF11── culleoka
|
||
│ │ │
|
||
│ 5GHz │ AF11 │ AF11 (DOWN: power injector unplugged)
|
||
│ │ │
|
||
494 newhope ──AF24── lowry
|
||
│
|
||
│ 60 GHz
|
||
│
|
||
982
|
||
```
|
||
Wait — that diagram's links are: climax↔494 (5 GHz airMAX, **not** AF11 — see below), core↔newhope (AF11),
|
||
core↔982 (60 GHz), newhope↔lowry (AF24). The climax↔culleoka direct AF11
|
||
is currently down at the radio (physical issue), so culleoka traffic
|
||
transits via core.
|
||
|
||
## Fleet Topology
|
||
|
||
### Routers and loopbacks
|
||
All ROS7 routers run RouterOS 7.21.4 long-term (post-2026-05-08 fleet
|
||
upgrade). Edge runs ROS 6.49.18 (legacy, no MPLS, ignore for the spine).
|
||
|
||
| Router | Loopback (10.254.254.x) | Hardware | Site name |
|
||
|--|--|--|--|
|
||
| verona | .101 | CCR2004-16G-2S+ (arm64) | verona |
|
||
| climax | .102 | CCR2004-16G-2S+ (arm64) | climax |
|
||
| culleoka | .104 | CCR1009-7G-1C-1S+ (tile) | culleoka |
|
||
| newhope | .108 | CCR1009-7G-1C-1S+ (tile) | newhope |
|
||
| lowry | .109 | (tile) | lowrycrossing |
|
||
| 982 | .110 | (CCR, tile) | 982 |
|
||
| 494 | .111 | (CCR, tile) | 494 |
|
||
| core | .253 | (CCR, arm64) at 380 | core/380 |
|
||
| edge | .254 | (legacy, ROS 6.49.18) | edge |
|
||
|
||
Tile-arch boxes can run MPLS but **not** ZeroTier (no .npk for tile).
|
||
|
||
### Backbone links
|
||
Every link below has IPv4 LDP enabled at both ends, `mpls-mtu=1508`, and
|
||
fasttrack-bypass rules in both directions on both routers.
|
||
|
||
| Link | Type | A-side iface | B-side iface | /29 subnet | l2mtu |
|
||
|--|--|--|--|--|--|
|
||
| verona↔climax | AF11 | verona `ether3-climax-11ghz` | climax `ether6-verona-11ghz` | 10.250.1.24/29 | 2024 |
|
||
| climax↔core | AF24 | climax `ether4-380-airfiber24` | core `ether5-climax` | 10.250.1.88/29 | 2024 |
|
||
| climax↔494 | 5 GHz airMAX | climax `ether5-494` | 494 `ether2-climax` | 10.250.1.64/29 | 1580 |
|
||
| climax↔culleoka | AF11 | climax `ether3-culleoka-11ghz` | culleoka `ether1-climax-11ghz` | 10.250.1.8/29 | 2024 (link DOWN) |
|
||
| core↔culleoka | AF11 | core `ether6-culleoka-11ghz` | culleoka `ether6-380-11ghz` | 10.250.1.48/29 | 2024 |
|
||
| core↔newhope | AF11 | core `ether4-newhope` | newhope `ether2-380` | 10.250.1.56/29 | 9000 |
|
||
| core↔982 | 60 GHz | core `ether1-982-60ghz` | 982 `ether7-380` | 10.250.1.32/29 | 9000 |
|
||
| newhope↔lowry | AF24 | newhope `ether6-lowrycrossing` | lowry `ether1-newhope` | 10.250.1.104/29 | 9000 |
|
||
| core↔edge | wired | core `sfp-sfpplus1-edge-preseem` + `ether3-edge-direct` | edge ports | 204.110.191.x | n/a |
|
||
|
||
l2mtu mismatches across the fleet are intentional per platform: AF11 base
|
||
ports default to 2024 on CCR2004 / 1580 on smaller CCRs; jumbo-capable
|
||
links (60 GHz, AF24-with-jumbo, fiber) go to 9000. **Always raise both
|
||
sides symmetrically when changing l2mtu** — single-side raises usually work
|
||
because Ethernet receivers accept anything ≤ their cap, but symmetric is
|
||
the rule.
|
||
|
||
### climax↔494 backhaul is 5 GHz airMAX, not AF11
|
||
Old docs called this hop AF11; it is actually a Ubiquiti airMAX AC PtP
|
||
pair (SSID `vntx_pr_494`, WPA2, 40 MHz wide): AP **"Climax to 494"**
|
||
(PBE-5AC-500) at `10.250.1.69` on the climax side, station **"494 to
|
||
Climax"** (PBE-5AC-400) at `10.250.1.66` on the 494 side. Both answer
|
||
SNMP v1 (`kdyyJrT0Mm`, airMAX MIB `1.3.6.1.4.1.41112.1.4`); added to
|
||
`radios.yaml` 2026-07-17. The station's mgmt IP (and the whole 494
|
||
tower) is only reachable **across this RF link** — when changing the
|
||
channel, first make sure the station's frequency scan list covers the
|
||
target, then change the AP (climax side), which stays reachable for
|
||
rollback either way.
|
||
|
||
2026-07-17 flap incident: link ran on **5260 MHz (DFS)**; associations
|
||
dropped for 40–90 s every 45–130 min (OSPF 40 s dead-interval timeouts
|
||
on both routers; ethernet ports never dropped, radios never rebooted —
|
||
association-uptime SNMP counters matched each OSPF flap). Local 5 GHz
|
||
survey: climax APs on 5335/5545/5575 (airMAX) + 5750/5775 (ePMP), 494
|
||
ePMP omni on 5800 → U-NII-1 (5170–5250) was empty at both towers and
|
||
is non-DFS, so the fix was moving the PtP there.
|
||
|
||
### IGP / routing
|
||
- OSPFv2 area `backbone-v2` (id 0.0.0.0) on all spine links, SHA-512 auth
|
||
with `auth-id=1` and a shared key. PTP type. **BFD is off on every
|
||
wireless backbone interface (AF11/AF24/60GHz)** — the global timers
|
||
(`/routing/bfd/configuration` = 200ms×5 = 1s detection) are too
|
||
aggressive for backhaul RF: a single >1s burst tore down OSPF on
|
||
climax↔494 every ~15–40 min until `use-bfd=false` was applied to both
|
||
ends 2026-05-09. Wired fiber links may keep BFD if desired. If
|
||
sub-second failover on a wireless link is genuinely needed, also
|
||
loosen the BFD timers (e.g. min-rx/min-tx=500ms, multiplier=3) — do
|
||
not flip `use-bfd=true` alone.
|
||
- OSPFv3 area `backbone-v3` for IPv6 (some interfaces only).
|
||
- All instances `redistribute=connected` with passthrough filters
|
||
(`/routing filter rule chain=ospf-out rule="accept;"`).
|
||
- Verona has a static default to `10.250.1.30` (climax) backing up the OSPF
|
||
default — keep this; bouncing OSPF on verona doesn't blackhole it.
|
||
- Distance-1 static routes also exist on climax for `204.110.191.0/27` so
|
||
the home /27 has guaranteed return path even if OSPF redistribution
|
||
hiccups.
|
||
|
||
### Management subnets per tower
|
||
`10.10.x.0/20` per site, top /24 reserved for APs (see Access Point
|
||
Placement section). Authoritative mapping is in
|
||
`mikrotik-tool/inventory.yaml`. Quick reference:
|
||
- verona: 10.10.0.0/20
|
||
- altoga (behind verona, no router): 10.10.16.0/20
|
||
- climax: 10.10.48.0/20
|
||
- core/380: 10.10.64.0/20
|
||
- culleoka: 10.10.96.0/20
|
||
- 982: 10.10.128.0/20
|
||
- newhope: 10.10.144.0/20
|
||
- 494: 10.10.160.0/20
|
||
- lowry: 10.10.80.0/20
|
||
|
||
CGNAT pools: `100.64.x.x/22` per tower (see `inventory.yaml` /
|
||
`subnets.yaml`).
|
||
|
||
### graham's home network gotcha
|
||
graham's home connects to verona via `vlan9_sfpplus1` carrying
|
||
`204.110.191.0/27` (home router at `.1`, verona at `.30`). This /27 is a
|
||
subnet of the verona hotspot's covered range (`204.110.188.0/22`). After
|
||
any verona reboot, ensure `/ip hotspot ip-binding` has an entry:
|
||
`address=204.110.191.0/27 type=bypassed comment="graham home /27"` —
|
||
without it, hotspot drops all `204.110.191.x` traffic in `hs-unauth-to`
|
||
chain with `icmp-host-prohibited`. Symptom is "I can reach verona but
|
||
nothing past it" from the home network.
|
||
|
||
### graham's home network — multi-WAN failover (recursive check-gateway)
|
||
Home (RB5009, `10.0.19.254`, **not** in `routers.yaml` — use a temp
|
||
config with the same default creds to reach it via mikrotik-tool) fails
|
||
over between TMO (distance 1), VNTX/verona (distance 2), and Starlink
|
||
(distance 3, disabled) using the standard ROS recursive pattern, applied
|
||
2026-07-17 via `mikrotik-tool/home-recursive-failover.rsc`:
|
||
- Probe pins: `4.2.2.1/32→192.168.12.1%ether6-tmobile`,
|
||
`4.2.2.2/32→204.110.191.30%ether5-vntx-static`,
|
||
`4.2.2.3/32→192.168.1.1%ether7-starlink`, all `scope=10`, **no
|
||
check-gateway on the pins** — they stay active whenever the interface
|
||
has link, gluing probes to their WAN regardless of route state.
|
||
- Defaults: `gateway=4.2.2.x target-scope=11 check-gateway=ping` — the
|
||
check pings the probe IP through the pin; ~20s to go inactive, first
|
||
success to return. `tmo table default` (eweka policy routing) has the
|
||
same check so it falls back to main when TMO upstream dies.
|
||
|
||
**Do not reintroduce netwatch enable/disable failover scripts.** The old
|
||
design (netwatch 2s/1s single-packet → `/ip route disable`) deadlocked:
|
||
`disabled=yes` persists across reboots, and the probe was only tied to
|
||
the VNTX path while its pin was active — so the verona hotspot-binding
|
||
loss (verona pingable, transit dead) left the route disabled forever,
|
||
and probes leaked out other WANs giving false "up" (TMO netwatch read up
|
||
with the TMO interface physically down). Same lesson as the fleet BFD
|
||
incident: 1-2s detection on AF RF with >1s bursts causes false failover;
|
||
10–20s check-gateway damping is intentional.
|
||
|
||
|
||
### IPv6 plan
|
||
Per-tower /44s + central server LAN at `2606:1c80::/64` on edge. Full
|
||
allocation plan in `mikrotik-tool/ipv6.md`. NetBox has these as IPAM
|
||
prefixes.
|
||
|
||
### IPv6 forward firewall — must accept the whole /32
|
||
Every tower router's `/ipv6 firewall filter chain=forward` needs
|
||
`action=accept src-address=2606:1c80::/32` before the `default deny`.
|
||
A per-tower `/44` rule is **wrong** — customer traffic transits other
|
||
routers to reach `edge`, so each transit hop blackholes other towers'
|
||
`/44`s. Classic symptom: ICMPv6 ping works (there's an `accept
|
||
protocol=icmpv6` rule) but TCP/HTTP times out at `connect` (forwarded
|
||
TCP hits `default deny`). The original `ipv6.md` template used
|
||
`in-interface-list=customer`, but that list was never created on any
|
||
router so the rule matched nothing — fixed fleet-wide 2026-05-14, and
|
||
the now-redundant per-tower `/44` and `internal-to-internal` forward
|
||
rules were removed at the same time. `edge` has a permissive forward
|
||
chain (no default-deny) so it didn't need the rule. Resolver/DNS
|
||
roll-out (`2606:1c80::240`/`::250` in `/ip dns` + the `v6-dns` DHCPv6
|
||
option) is also fleet-wide as of 2026-05-14.
|
||
|
||
### graham's home network — IPv6
|
||
Home (`RB5009`, identity `graham`, `10.0.19.254`) takes a **`/56`** from
|
||
verona via DHCPv6-PD on `ether5-vntx-static` (verona's `verona-wired-pd`
|
||
dhcp-server on `vlan9_sfpplus1`, drawing from `verona-cust-pd-1`).
|
||
Currently `2606:1c80:1001:e00::/56`. Key config on home:
|
||
- `/ipv6 dhcp-client` on `ether5-vntx-static`: `request=prefix`,
|
||
`pool-name=home-pd-vntx`, **`pool-prefix-length=64`** (must be 64, not
|
||
56, so the pool yields /64s — changing this only takes effect on a
|
||
client disable/enable, not live), `default-route-tables=v6-verona`.
|
||
- `/ipv6 address` with `from-pool=home-pd-vntx address=::1 advertise=yes`
|
||
on `bridge` (→ `:e00::1/64`) and `ether3-servers` (→ `:e01::1/64`).
|
||
Remaining 254 /64s of the /56 stay in the pool for internal use.
|
||
- Egress policy: **two** `/routing rule`s, order matters —
|
||
(1) `src-address=2606:1c80::/32 dst-address=2606:1c80:1001:e00::/56
|
||
action=lookup-only-in-table table=main` so vntx-sourced traffic
|
||
*destined to the home LANs* resolves via main's connected routes,
|
||
then (2) `src-address=2606:1c80::/32 action=lookup-only-in-table
|
||
table=v6-verona` for everything else. The verona /56 is single-homed,
|
||
so traffic sourced from it must egress via verona only
|
||
(T-Mobile/Starlink would BCP38-drop it). No v6 WAN failover for the
|
||
internal /56, by design. The PD client's default lands in the
|
||
`v6-verona` table, not `main`. **Without rule (1)** the broad rule
|
||
also catches transit packets (echo replies, ICMPv6 time-exceeded)
|
||
sourced from internal routers and destined back to the home LAN, and
|
||
hairpins them to verona in a loop — symptom: home reaches the v6
|
||
internet fine but cannot ping any internal router and internal hops
|
||
show `*` in `traceroute6`. Rule (1)'s `dst-address` is tied to the
|
||
current delegated /56 — update it if the PD lease ever changes.
|
||
- Stateless DHCPv6 servers (`home-bridge-stateless`,
|
||
`home-servers-stateless`) carry the `v6-dns` option (Cloudflare).
|
||
- Stale `tunnerbroker` `/ipv6 pool` (HE.net `2001:470:ba50::/48`) is
|
||
unused — leave or remove.
|
||
- The `mikrotik-tool script` SSH path can't handle bare ROS verbs like
|
||
`release`/`disable`/`enable` or `:delay` (flattenRSC only knows
|
||
add/set/remove/print/get) — use `set [find] disabled=yes/no` instead,
|
||
and sequence delays from the shell side.
|
||
|
||
## Claude Assistant Guidelines
|
||
|
||
- Any time Claude learns something new, automatically add it to CLAUDE.md
|
||
|
||
## Development Best Practices
|
||
- When making scripts, keep them as generic and reusable as possible |