Files
meshcore-analyzer/AGENTS.md
T
efitenandClaude Opus 5.5 d232072f85 feat: optional user accounts (part A: foundation) (#2129)
Part A of #2128: optional, off-by-default user accounts. With the
feature off nothing changes; with it on, visitors can register and log
in, and admins manage users and use the operator actions without the API
key.

PR #2130 (settings sync) builds on this one. The two are meant to be
merged together.

## The situation

- Operator actions (geofilter save and prune, backup, perf reset) need
the shared `apiKey`. There is no per-person right.
- Nothing in CoreScope knows who a visitor is, so the requests in #2128
that need that (#1835, #2092, #1508, #730) have nothing to build on.

## What this PR adds

**Two new Go modules**
- `internal/users`: a separate `users.db` (SQLite through
`modernc.org/sqlite`) with users, sessions, single-use tokens, an audit
log and a mail log. Passwords use argon2id.
- `internal/mailer`: a `Mailer` interface with a Brevo client (send,
delivery events, webhook parsing) and an in-memory fake for tests.

**Server (`cmd/server`)**, active only with `userManagement.enabled`
- 24 routes, all documented in OpenAPI under the `users` tag
([`auth_routes.go`](https://github.com/efiten/CoreScope/blob/feat/user-management/cmd/server/auth_routes.go)):
  - auth: register, activate, login, logout, me, forgot, reset;
- account: profile, password, email change with confirmation, sessions,
self-delete;
- admin: list, detail, disable, enable, delete, role, resend activation,
manual activation, mail status refresh;
  - a Brevo webhook, registered only when `mail.webhookSecret` is set.
- `requireAdmin` replaces `requireAPIKey` at the 7 operator call sites:
the API key **or** an admin session. With the feature off it is the old
API-key gate (`TestRequireAdminWithoutUserManagementIsAPIKeyGate`).
- `/api/config/client` gets `userManagement: {enabled: true}` only when
the service started; with the feature off the response is
byte-identical.

**Frontend**
- `auth.js` (header account control, request helper that adds the CSRF
header), `account.js` (login, register, activate, forgot, reset, confirm
email, my account), `admin-users.js` (`#/admin/users`, deep-linked
filters), `account.css` (theme tokens only).
- On phones the top-bar control is hidden, so a conditional entry goes
into the bottom-nav "More" sheet and the nav drawer.
- The customizer geofilter tab and the Perf "Reset stats" button use the
admin session when there is one.

**Config.** A `userManagement` block (`config.example.json`,
[`docs/user-guide/accounts.md`](https://github.com/efiten/CoreScope/blob/feat/user-management/docs/user-guide/accounts.md)).
The Brevo key can come from `CORESCOPE_BREVO_API_KEY`. The server
refuses to start when the block is enabled but incomplete.

## Security choices

- Session cookie `cs_session`: HttpOnly, SameSite=Lax, Secure when
`publicBaseUrl` is https. Every cookie-authenticated state change needs
the `X-CS-CSRF` header and a matching Origin.
- Activation needs the token **and** the account password. Without the
password, an attacker who keeps re-registering a known address could get
the owner to activate an account that carries the attacker's password.
- Register, forgot and email change answer identically for known and
unknown addresses. A password reset ends all sessions, a password change
ends all other sessions, and both end outstanding email-change links.
- Rate limits: login 10 per 15 minutes, register and forgot 5 per hour,
per IP and per address. The bucket count is capped. `trustedProxies`
makes the per-IP limits see real client IPs behind a proxy.
- Server logs carry `#<user id>`, never addresses, tokens or passwords;
mail-provider error texts are redacted before logging.

## Performance

No change to an existing hot path with the feature off. With it on:
- One `users.db` lookup per authenticated request (session by token
hash).
- The admin user table rebuilds its `tbody` on each filter change.
`users.List` caps the result at 1000 rows (`internal/users/users.go`),
which bounds the rebuild.
- `map[string]interface{}` in `openapi.go`: 79 before, 78 after.

## Verification

- `internal/users`, `internal/mailer` and `cmd/server`: `go vet` and `go
test -race` pass locally. 121 new Go tests.
- `cmd/server` with `-tags e2etest`: vet and the e2e hook tests pass.
- `sh test-all.sh` exits 0. `tests/unit/test-user-management-ui.js`: 67
passing (vm, real modules).
- `tests/e2e/test-user-management-e2e.js` (6 steps) passed locally
against an `e2etest` build with the fake mailer and against a
feature-off build. CI builds the `e2etest` binary and runs the suite on
a second server (`deploy.yml`).
- On a staging instance with a real Brevo key: register, activation mail
delivered, activate, admin table, "Refresh status" showing sent,
deferred, delivered, opened and clicked.

## Not in this PR

- Settings sync (#2130), the admin dashboard, approval flows and
notifications (parts B to E of #2128).
- A `requireReadAuth` mode (#1835). Sessions from this PR are what such
a mode would accept.
- Binary size and build time with `modernc.org/sqlite` linked next to
`mattn/go-sqlite3` were not measured. Their driver names do not collide.
#1992 discusses the driver choice.
- No Brevo webhook was configured on staging; delivery status there came
from "Refresh status".

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 09:25:29 +02:00

26 KiB
Raw Blame History

AGENTS.md — CoreScope

Guide for AI agents working on this codebase. Read this before writing any code.

Architecture

Go backend + static frontend. No build step. No framework. No bundler.

⚠️ The Node.js server (server.js) is DEPRECATED and has been removed. All backend code is in Go. ⚠️ DO NOT create or modify any Node.js server files. All backend changes go in cmd/server/ or cmd/ingestor/.

cmd/server/        — Go API server (REST + WebSocket broadcast + static file serving)
  main.go          — Entry point, flags, SPA handler
  routes.go        — All /api/* endpoints
  store.go         — In-memory packet store + analytics + SQLite queries
  config.go        — Configuration loading
  decoder.go       — MeshCore packet decoder
cmd/ingestor/      — Go MQTT ingestor (separate binary, writes to shared SQLite DB)
public/            — Frontend (vanilla JS, one file per page) — ACTIVE, NOT DEPRECATED
  app.js           — SPA router, shared globals, theme loading
  roles.js         — ROLE_COLORS, TYPE_COLORS, health thresholds, shared helpers
  nodes.js         — Nodes list + side pane + full detail page
  map.js           — Leaflet map with markers, legend, filters
  packets.js       — Packets table + detail pane + hex breakdown
  packet-filter.js — Wireshark-style filter engine (standalone, testable)
  customize.js     — Theme customizer panel (self-contained IIFE)
  analytics.js     — Analytics tabs (RF, topology, hash issues, etc.)
  channels.js      — Channel message viewer
  live.js          — Live packet feed + VCR mode
  home.js          — Home/onboarding page
  hop-resolver.js  — Client-side hop prefix → node name resolution
  style.css        — Main styles, CSS variables for theming
  live.css         — Live page styles
  home.css         — Home page styles
  index.html       — SPA shell, script/style tags with __BUST__ placeholder (auto-replaced at server startup)
test-fixtures/     — Real data SQLite fixture from staging (used for E2E tests)
scripts/           — Tooling (coverage collector, fixture capture, frontend instrumentation)

Data Flow

  1. MQTT brokers → Go ingestor (cmd/ingestor/) ingests packets → decodes → writes to SQLite
  2. Go server (cmd/server/) polls SQLite for new packets, broadcasts via WebSocket
  3. Frontend fetches via REST API (/api/*), filters/sorts client-side

Read/Write Separation Invariant (#1283)

  • All DB writes live in cmd/ingestor/. INSERT / UPDATE / DELETE / VACUUM / schema migrations / retention all run in the ingestor process.
  • cmd/server/ never writes measurement data. It opens the analyzer DB with mode=ro and must not acquire a write lock on it. Adding a write-side helper (e.g. a cachedRW-style RW connection) regresses this invariant and races the ingestor → SQLITE_BUSY.
  • Single exception: users.db (optional user management). When userManagement.enabled, the server owns a separate SQLite file through internal/users only. users.Open refuses the analyzer DB path, and TestUsersOpenIsTheOnlyServerWritePath pins the one call site. Account data never goes into the analyzer DB, and measurement writes never go through internal/users.
  • Enforcement: cmd/server/readonly_invariant_test.go reflect-asserts that PruneOldPackets, PruneOldMetrics, and RemoveStaleObservers are NOT methods on the server's *DB. If you need a new write, add it to cmd/ingestor/.

What's Deprecated (DO NOT TOUCH)

The following were part of the old Node.js backend and have been removed:

  • server.js, db.js, decoder.js, server-helpers.js, packet-store.js, iata-coords.js
  • All test-server-*.js, test-decoder*.js, test-db*.js, test-regional*.js files
  • If you see references to these in comments or docs, they're stale — ignore them

Rules — Read These First

0. Performance is a feature — not an afterthought

Every change must consider performance impact BEFORE implementation. This codebase handles 30K+ packets, 2K+ nodes, and real-time WebSocket updates. A single O(n²) loop or per-item API call can freeze the UI or stall the server.

Before writing code, ask:

  • What's the worst-case data size this code will process?
  • Am I adding work inside a hot loop (render, ingest, WS broadcast)?
  • Am I fetching from the server what I could compute client-side?
  • Am I recomputing something that could be cached/incremental?
  • Does my change invalidate caches more broadly than necessary?

Hard rules:

  • No per-item API calls. Fetch bulk, filter client-side.
  • No O(n²) in hot paths. Use Maps/Sets for lookups, not nested array scans.
  • No full DOM rebuilds. Diff or virtualize — never innerHTML entire tables.
  • No unbounded data structures. Every map/slice/array must have eviction or size limits.
  • No expensive work under locks. Copy data under lock, process outside.
  • Cache expensive computations. Invalidate surgically, not globally.
  • Debounce/coalesce rapid events. WebSocket messages, scroll, resize — never fire raw.

If your change touches a hot path (packet rendering, ingest, analytics), include a perf justification in the PR description: what the complexity is, what the expected scale is, and why it won't degrade.

Perf claims require proof. "This is faster" without data is not acceptable. Every PR claiming to fix or improve performance MUST include one of:

  • A benchmark test (before/after timings with realistic data sizes)
  • Profile output or timing measurements (e.g. "renderTableRows: 450ms → 12ms on 30K packets")
  • A test assertion that enforces the perf characteristic (e.g. "filters 30K packets in <50ms") No proof = no merge.

1. No commit without tests

Every change that touches logic MUST have tests. For Go backend: cd cmd/server && go test ./... and cd cmd/ingestor && go test ./..., or make test for all 14 modules. For frontend: node tests/unit/test-packet-filter.js && node tests/unit/test-aging.js && node tests/unit/test-frontend-helpers.js. If you add new logic, add tests. No exceptions.

The SQLite driver (github.com/mattn/go-sqlite3) is cgo. CGO_ENABLED=0 still builds — that is the trap. It links a stub, and the binary dies on the first query with go-sqlite3 requires cgo to work. This is a stub, so a green build proves nothing. A plain GOOS=linux go build likewise cannot cross-compile it. Use make build / make crossbuild — the latter needs zig as the cross C compiler and produces static musl binaries. Never re-add CGO_ENABLED=0.

2. No commit without browser validation

After pushing, verify the change works in an actual browser. Use browser profile=openclaw against the running instance. Take a screenshot if the change is visual. If you can't validate it, say so — don't claim it works.

3. Cache busters are automatic — do NOT manually edit them

Cache busters are injected automatically by the Go server at startup. The __BUST__ placeholder in index.html is replaced with a Unix timestamp when the server reads the file. No manual bumping needed — every server restart picks up new asset versions. Do NOT replace __BUST__ with hardcoded timestamps.

4. Verify API response shape before building UI

Before writing client code that consumes an API endpoint, check what the endpoint ACTUALLY returns. Use curl or check the server code. Don't assume fields exist — grouped packets (groupByHash=true) have different fields than raw packets. This has caused multiple breakages.

5. Plan before implementing

Present a plan with milestones to the human. Wait for sign-off before starting. The plan must include:

  • What changes in each milestone
  • What tests will be written
  • What browser validation will be done
  • What config/customizer implications exist (see rule 8)

Do NOT start coding until the human says "go" or "start" or equivalent.

6. One commit per logical change

Don't push half-finished work. Don't push "let me try this" experiments. Get it right locally, test it, THEN push ONE commit. The QR overlay took 6 commits because each one was pushed without looking at the result. That's 6x the review burden for one visual change.

7. Understand before fixing

When something doesn't work as expected, INVESTIGATE before "fixing." Read the firmware source. Check the actual data. Understand WHY before changing code. The hash_size saga (21 commits) happened because we guessed at behavior instead of reading the MeshCore source.

8. Config values belong in the customizer eventually

If a feature introduces configurable values (thresholds, timeouts, display limits), note in the plan that these should be exposed in the customizer in a later milestone. It's OK to hardcode initially, but don't forget — track it in the plan.

9. Explicit git add only

Never use git add -A or git add .. Always list files explicitly: git add file1.js file2.js. Review with git diff --cached --stat before committing.

10. Don't regress performance

The packets page loads 30K+ packets. Don't add per-packet API calls. Don't add O(n²) loops. Client-side filtering is preferred over server-side. If you need data from the server, fetch it once and cache it.

11. PR descriptions must be clean markdown

When opening a pull request, the description must be valid, readable markdown. Use real newlines (not \n literals), proper code fences, and correct heading syntax. Write it using --body-file - (piped from a heredoc or file), never inline --body with escaped characters. If the description renders as garbage, fix it before requesting review. This is the first thing reviewers see.

12. Post a follow-up comment when review feedback is addressed

When you push fixes for review comments, post a comment on the PR listing what was changed and the commit hash. Reviewers should not have to dig through commits to find what was fixed. Format: "Review feedback addressed (commit abc1234)" followed by a numbered list of what was done.

13. Use git worktrees for parallel work — never pollute the main checkout

Multiple agents work in parallel. The main clone (C:\Projects\meshcore-analyzer\) must stay on master and never be modified directly.

Implementation agents must create a dedicated worktree before making any changes:

git worktree add _wt-<branch-name> -b <branch-name> origin/master
cd _wt-<branch-name>
# ... do all work here ...

After PR is merged, clean up: git worktree remove _wt-<branch-name>

Review agents must NEVER read files from the working tree. Use git commands to read the remote branch directly:

git fetch origin <branch>
git show origin/<branch>:<path/to/file>      # read a specific file
git diff origin/master..origin/<branch>       # see the full diff

The working tree may have a different branch checked out. Reading it will give you wrong code.

Periodic cleanup: run git worktree prune to remove stale worktree references.

MeshCore Firmware — Source of Truth

The MeshCore firmware source is cloned at firmware/ (gitignored — not part of this repo). This is THE authoritative reference for anything related to the protocol, packet format, device behavior, advert structure, flags, hash sizes, route types, or how repeaters/companions/rooms/sensors behave.

Before implementing any feature that touches protocol behavior:

  1. Check the firmware source in firmware/src/ and firmware/docs/
  2. Key files: Mesh.h (constants, packet structure), Packet.cpp (encoding/decoding), helpers/AdvertDataHelpers.h (advert flags/types), helpers/CommonCLI.cpp (CLI commands), docs/packet_format.md, docs/payloads.md
  3. If firmware/ doesn't exist, clone it: git clone --depth 1 https://github.com/meshcore-dev/MeshCore.git firmware
  4. To update: cd firmware && git pull

Do NOT guess at protocol behavior. The hash_size saga (21 commits) and the advert flags bug (room servers misclassified as repeaters) both happened because we assumed instead of reading the firmware source. The firmware is C++ — read it.

MeshCore Protocol

Do not memorize or hardcode protocol details from this file. Read the firmware source.

  • Packet format: firmware/docs/packet_format.md
  • Payload types & structures: firmware/docs/payloads.md
  • Advert flags & types: firmware/src/helpers/AdvertDataHelpers.h
  • Route types & constants: firmware/src/Mesh.h
  • CLI commands & behavior: firmware/docs/cli_commands.md
  • FAQ (advert intervals, etc.): firmware/docs/faq.md

If you need to know how something works — a flag, a field, a timing, a behavior — open the file and read it. Don't rely on comments in our code, don't rely on what someone told you, don't guess. The firmware C++ source is the only thing that matters.

Frontend Conventions

Theming

All colors MUST use CSS variables. Never hardcode #hex values outside of :root definitions. The customizer controls colors via THEME_CSS_MAP in customize.js. If you add a new color, add it as a CSS variable and map it in the customizer.

Shared Helpers (roles.js)

  • getNodeStatus(role, lastSeenMs) → 'active' | 'stale'
  • getHealthThresholds(role) → { staleMs, degradedMs, silentMs }
  • ROLE_COLORS, ROLE_STYLE, TYPE_COLORS — global color maps

Shared Helpers (nodes.js)

  • getStatusInfo(n) → { status, statusLabel, explanation, roleColor, ... }
  • renderNodeBadges(n, roleColor) → HTML string
  • renderStatusExplanation(n) → HTML string

last_heard vs last_seen

  • last_seen = DB timestamp, only updates on adverts/direct upserts
  • last_heard = from in-memory packet store, updates on ALL traffic
  • Always prefer n.last_heard || n.last_seen for display and status calculation

Packet Filter (packet-filter.js)

Standalone module. No dependencies on app globals (copies what it needs). Testable in Node.js:

node tests/unit/test-packet-filter.js

Uses firmware-standard type names (GRP_TXT, TXT_MSG, REQ) with aliases for convenience.

Testing

Test Pipeline

npm test                    # all backend tests + coverage summary
npm run test:unit           # fast: unit tests only (no server needed)
npm run test:coverage       # all tests + HTML coverage report
npm run test:full-coverage  # backend + instrumented frontend coverage via Playwright

Test Files

# Backend (deterministic, run before every push)
sh test-all.sh                               # every suite in tests/unit
node tests/unit/test-packet-filter.js        # filter engine
node tests/unit/test-aging.js                # node aging system
node tests/unit/test-frontend-helpers.js     # frontend logic (via vm.createContext)

# Frontend E2E (requires running server or Playwright)
node tests/e2e/test-e2e-playwright.js       # Playwright browser tests (default: localhost:3000)

Rules

ALL existing tests must pass before pushing. No exceptions. No "known failures."

Every new feature must add tests. Unit tests for logic, Playwright tests for UI changes. Test count only goes up.

Coverage targets: Backend 85%+, Frontend 42%+ (both should only go up). CI reports both and updates badges automatically.

When writing a new feature

  1. Write the feature code
  2. Write unit tests for the logic
  3. Write/update Playwright tests if it's a UI change
  4. Run npm test — all tests must pass
  5. Run node tests/e2e/test-e2e-playwright.js against a local server — E2E must pass
  6. THEN push to master

Testing infrastructure

  • Backend coverage: c8 tracks server-side code in-process
  • Frontend coverage: Istanbul instruments public/*.js → Playwright exercises them → window.__coverage__ extracted → nyc reports. Instrumented files are generated fresh each CI run, never checked in.
  • CI pipeline: backend tests + coverage → instrument frontend → start local server → Playwright E2E + coverage collection → badges update → deploy (only if all pass)
  • Playwright tests default to localhost:3000 — NEVER run against prod. CI sets BASE_URL=http://localhost:13581. Running locally: start your server, then node tests/e2e/test-e2e-playwright.js
  • ARM machines: Basic Playwright tests work with system chromium (CHROMIUM_PATH=/usr/bin/chromium-browser). Heavy coverage collection scripts may crash — use CI for those.

Tests that need live mesh data can use https://analyzer.00id.net — all API endpoints are public, no auth required.

What Needs Tests

  • Parsers and decoders (packet-filter, decoder)
  • Threshold/status calculations (aging, health)
  • Data transformations (hash size computation, field resolvers)
  • Anything with edge cases (null handling, boundary values)
  • UI interactions that exercise frontend code branches

Engineering Principles

These aren't optional. Every change must follow these principles.

DRY — Don't Repeat Yourself

If the same logic exists in two places, it MUST be extracted into a shared function. We had 5 separate implementations of hash prefix disambiguation across the codebase — that's a maintenance nightmare and a bug factory. One implementation, imported everywhere.

Before writing new code, search the codebase for existing implementations. grep -rn 'functionName\|pattern' public/ server.js takes 2 seconds and prevents duplication.

SOLID Principles

  • Single Responsibility: Each function does ONE thing. A 200-line function that fetches, transforms, renders, and caches is wrong. Split it.
  • Open/Closed: Add behavior by extending, not modifying. Use callbacks, options objects, or configuration — not if (caller === 'live') branches inside shared code.
  • Dependency Injection: Functions should accept their dependencies as parameters, not reach into globals. resolveHops(hops, nodeList) — not resolveHops(hops) where it secretly reads window.allNodes. This makes functions testable in isolation.
  • Interface Segregation: Don't force callers to depend on things they don't need. If a function returns 20 fields but the caller uses 3, consider a simpler return shape or let the caller pick.

Code Reuse

  • Shared helpers go in shared files. Frontend: roles.js, hop-resolver.js. Backend: server-helpers.js, decoder.js.
  • Don't copy-paste between files. If live.js needs the same algorithm as packets.js, import it from a shared module. If the shared module doesn't exist yet, create one.
  • Parameterize, don't duplicate. If two callers need slightly different behavior, add a parameter — don't fork the function.

Testability

  • Write functions that are easy to test. Pure functions (input → output, no side effects) are ideal. If a function reads from the DOM, the DB, and localStorage, it's untestable without mocking everything.
  • Dependency injection enables testing. Pass the node list, the map reference, the API function as parameters. Tests can substitute fakes.
  • Test the real code, not copies. Don't paste a function into a test file and test the copy. Import/require the actual module. If the module isn't importable (IIFE, browser-only), refactor it so it is — or use vm.createContext like tests/unit/test-frontend-helpers.js does.
  • Every bug fix gets a regression test. If it broke once, it'll break again. The test proves it stays fixed.

Type Safety (without TypeScript)

  • Cast at the boundary. Data from the DB, API, or localStorage may be strings when you expect numbers. Cast early: Number(val), parseInt(val), String(val). Don't let type mismatches propagate deep into logic where they cause cryptic .toFixed is not a function errors.
  • Null-check before method calls. val != null ? Number(val).toFixed(1) : '—' — not val.toFixed(1).

Performance Awareness

  • No per-item API calls. Fetch bulk data once, filter/transform client-side.
  • No O(n²) in hot paths. The packets page has 30K+ rows. A nested loop over all packets × all nodes = 20 billion operations. Use Maps/Sets for lookups.
  • Cache expensive computations. If you compute the same thing on every render, cache it and invalidate on data change.

XP (Extreme Programming) Practices

Test-First Development

Write the test BEFORE the code. Not after. Not "I'll add tests later." The test defines the expected behavior, then you write the minimum code to make it pass.

Flow: Red (write failing test) → Green (make it pass) → Refactor (clean up).

This prevents shipping bugs like .toFixed on a string — if the test existed first with string inputs, the bug could never have been introduced. Every bug fix starts by writing a test that reproduces the bug, THEN fixing it.

YAGNI — You Aren't Gonna Need It

Don't build for hypothetical future requirements. Build the simplest thing that solves the current problem. The 5 separate disambiguation implementations happened because each page rolled its own "just in case" version instead of importing the one that already existed.

If you're writing code that handles a case nobody asked for: stop. Delete it. Add it when there's a real need.

Refactor Mercilessly

When you touch a file and see duplication, dead code, unclear names, or structural mess — clean it up in the same commit. Don't leave it for "later." Later never comes. Tech debt compounds.

The Boy Scout Rule: Leave every file cleaner than you found it.

Simple Design

The simplest solution that works is the correct one. Complexity is a bug. Before building something, ask:

  1. Does this already exist somewhere in the codebase?
  2. Can I solve this with an existing function + a parameter?
  3. Am I over-engineering for a case that doesn't exist yet?

If the answer to any of these is yes, simplify.

Pair Programming (Human + AI Model)

For this project, pair programming means: subagent writes the code → parent agent reviews and tests locally → THEN pushes to master. The subagent is the "driver," the parent is the "navigator."

What this means in practice:

  • Subagent output is NEVER pushed directly without review
  • Parent agent runs the tests, checks the diff, verifies the behavior
  • If the subagent's work is wrong, parent fixes it before pushing — not after
  • "The subagent said it works" is not verification. Running the tests is.

Continuous Integration as a Gate

CI must pass before code is considered shipped. But CI is the LAST line of defense, not the first. The process is:

  1. Test locally (unit + E2E)
  2. Review the diff
  3. Push
  4. CI confirms

If CI catches something you missed locally, that's a process failure — figure out why your local testing didn't catch it and fix the gap.

10-Minute Build

Everything must be testable locally in under 10 minutes. If local tests are broken, flaky, or crashing — that's a P0 blocker. Fix the test infrastructure before shipping features. Broken tests = no tests = shipping blind.

Collective Code Ownership

No file is "someone else's problem." Every file follows the same patterns, uses the same shared modules, meets the same quality bar. live.js doesn't get to be a special snowflake with its own reimplementation of everything. If it drifts from the shared patterns, bring it back in line.

Small Releases

One logical change per commit. Each commit is deployable. Each commit has its tests. Don't bundle "fix A + feature B + cleanup C" into one push — if B breaks, you can't revert without losing A and C.

Common Pitfalls

Pitfall Times it happened Prevention
Forgot cache busters 7 Now automatic — __BUST__ replaced at server startup
Grouped packets missing fields 3 curl the actual API first
last_seen vs last_heard mismatch 4 Always use last_heard || last_seen
CSS selectors don't match SVG 2 Manipulate SVG in JS after generation
Feature built on wrong assumption 5+ Read source/data before coding
Pushed without testing 5+ Run tests + browser check every time
Tests defaulting to prod 2 Always default to localhost, never prod
Gave up testing locally 2 Basic tests work on ARM — only heavy coverage scripts crash
Copy-pasted functions for "coverage" 1 Test the real code, not copies in a helper file
Subagent timed out mid-work 4 Give clear scope, don't try to run slow pipelines locally

File Naming

  • Tests: test-{feature}.js in tests/unit/ (standalone, listed in test-all.sh) or tests/e2e/ (browser, listed in scripts/non-unit-tests.json) — never in repo root
  • No build step, no transpilation — write ES2020 for server, ES5/6 for frontend (broad browser support)

Deep Linking

All new UI states that a user might want to share or bookmark MUST be reflected in the URL hash. This includes: tabs, filters, selected items, view modes. Use query parameters on the hash (e.g., #/packets?observer=ABC&timeRange=24h) for filter state. Existing patterns: #/nodes/{pubkey}?section=node-neighbors, #/analytics?tab=collisions, #/packets/{hash}.

What NOT to Do

  • Don't check in private information — no names, API keys, tokens, passwords, IP addresses, personal data, or any identifying information. This is a PUBLIC repo.
  • Don't introduce new map[string]interface{} in API response builders, handler returns, or internal data structures that cross domain boundaries. Use a named Go struct with explicit JSON tags. CoreScope already carries 694 occurrences (see #1383); the count must monotonically decrease. If your change adds even one new occurrence in a touched file, the PR is wrong-shaped — fix the design, don't paper over with interface{}. Exempt: third-party library boundaries that genuinely return interface{}, and ad-hoc test fixture assertions.
  • Don't add npm dependencies without asking
  • Don't create a build step
  • Don't add framework abstractions (React, Vue, etc.)
  • Don't hardcode colors — use CSS variables
  • Don't make per-packet server API calls from the frontend
  • Don't push without running tests
  • Don't start implementing without plan approval