31 Commits
Author SHA1 Message Date
n30nex 29095658e9 fix(nodes): explain packet and observation counts (#2132)
Red commit: `d8a6d42` ([CI
assertion](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37695889461)).
Review-regression red: `f73643b` ([CI
assertion](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37703995494)).

Fixes #2131. Partial fix for #1154: U2 only; other audit items remain
open.

Both node views explain distinct transmissions versus recorded
observations. The shared help is reachable by keyboard and touch, stays
readable at either scroll edge, and remains open when the pointer enters
its text. Escape dismisses help while preserving node context and focus;
a second Escape keeps the existing navigation behavior.

Count expressions and API requests are preserved. One controller per
nodes page handles one active help element and one queued animation
frame. Scroll/resize work is coalesced; listeners are removed on page
teardown. Work is constant relative to packet/node counts, with no
dataset scans or new settings/dependencies.

E2E assertion added: `tests/e2e/test-e2e-playwright.js:114`. Thirteen
cases cover exact counts, accessible description, keyboard/touch, hover
transfer, Escape, and viewport/scroll bounds. Existing metric assertions
are unchanged.

Browser verified: http://localhost:63508 using Chromium, controlled
fixtures and actual API data. Desktop/mobile screenshots retained
locally; additional checks exercised resizing and repeated page cleanup.

Local validation: 189 standalone suites; core browser 148 passed, 3
fixture-dependent skips (existing manual skips unchanged);
location/metric suite 6 passed; syntax, theme variables and lint passed
(0 errors, 92 existing warnings).

## Preflight overrides
- OpenClaw preflight/profile unavailable; repository checks and local
Chromium used.
- Windows ingestor symlink fixture lacks permission. Backend files are
unchanged; full Linux Go CI is required.
2026-10-08 22:37:15 +02:00
efitenandClaude Opus 5.5 2456762179 feat: account data export and users.db backup (#2141)
A small follow-up to parts A to E of #2128: users can download all data
the instance holds about them, and the server backs up `users.db`.

Stacked on #2140. Review the commits after that branch.

## The situation

- An instance with accounts holds personal data, but a user has no way
to get a copy (GDPR articles 15 and 20). Self-delete already exists.
- `users.db` is the only file the server writes, and nothing backs it
up: `GET /api/backup` snapshots the analyzer database only.

## What this PR adds

**Data export.** `GET /api/account/export` and a "Download my data"
button on the account page give one JSON file with:
- the profile, including activation and a pending email change;
- sessions (times and user agent);
- synced settings, own proposals, notification prefs and watches;
- audit rows where the user is actor or target, with other accounts as
an id only;
- the mail log.

Password hash, session and link tokens and the unsubscribe token are
left out: they are credentials. Each export is audited as `user.export`.

**users.db backup**
- Daily `VACUUM INTO` snapshots through `internal/users`, on by default
when user management is on. They go to `backups/` next to `users.db`,
the 7 newest are kept, the directory is 0700 and the files 0600.
Rotation only touches `users-<timestamp>.db` files and never the
snapshot it just wrote. Config: `userManagement.backup {enabled, dir,
keep}`.
- `GET /api/admin/users-backup`: an admin downloads a fresh snapshot to
keep off the server; audited as `user.backup`.
- `docs/user-guide/accounts.md` gains a restore procedure. It covers
what a restore brings back (accounts deleted after the snapshot, old
passwords, revoked sessions) and how to clean that up afterwards.

Spec:
[`docs/specs/2026-10-08-account-export-and-users-backup-design.md`](https://github.com/efiten/CoreScope/blob/feat/account-export-backup/docs/specs/2026-10-08-account-export-and-users-backup-design.md).

## Verification

- `internal/users` with `-race` and the `cmd/server` suite pass locally,
30 new Go tests. The file-mode tests skip on Windows.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 27 of 27 steps locally.
- On our production instance since 8 October 2026; the first snapshot
there was 143 kB.

## Not in this PR

- Importing an export into another instance.
- Off-site upload of the snapshots. The admin download covers that by
hand.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 22:16:49 +02:00
efitenandClaude Opus 5.5 de364a8bb6 feat: mail notifications for watched nodes (user management part E) (#2140)
Part E of #2128: a logged-in user watches nodes and gets one mail when a
watched node goes offline, comes back, or reports a low battery. Admins
can add instance events: a new foreign node, and an observer going
offline or back. This covers the per-user mail part of #775.

Stacked on #2139 (part D, which builds on #2138). Review the commits
after that branch.

## The situation

A sysop learns that a repeater went silent, or that its battery is
running down, only by opening CoreScope. #775 asks for notifications.
Accounts (part A) now give a verified address to mail and a mailer with
delivery status.

## What this PR adds

**Events.** They use the instance thresholds, so a mail does not
disagree with the node page:
- `node.offline`: silent for the role's `healthThresholds` window. Last
heard comes from the packet store, and repeaters and rooms count relayed
traffic (#1598).
- `node.battery`: advert telemetry below `batteryThresholds.lowMv`,
recovered at `lowMv + 100`.
- Admins only: `foreign.new` (once per node) and `observer.offline`.

**Store.** Schema v5 with `notification_prefs`, `notification_watches`
and `notification_state`.

**Server** (only with `userManagement.notifications.enabled`)
- A notifier next to the janitor, every `intervalMinutes` (5). The
evaluator is a pure function of watches, prefs, stored states and node
snapshots.
- No burst after a restart or a new watch: a subject's first evaluation
stores its state without mailing. The loop waits for the startup load.
While ingest is stale (newest packet older than 30 minutes) the offline
checks pause, and after recovery they wait one silent window.
- One mail per user per check. Limits: 20 per user and 100 per instance
per rolling 24 hours, 50 watches per user. A change over a limit is
recorded and never mailed later.
- Every mail has a one-click unsubscribe (`List-Unsubscribe` and
`List-Unsubscribe-Post`). The GET only redirects to a confirm page, so
link scanners cannot unsubscribe anyone. Node names are cleaned of
control and bidi characters before they go into a mail.
- Routes: `GET`/`PUT /api/account/notifications`, `PUT`/`DELETE
/api/account/notifications/watches/{pubkey}`, `POST
/api/account/notifications/watch-my-nodes` (copies the synced "my nodes"
list), `GET`/`POST /api/notifications/unsubscribe`. `mailer.Message`
gains `Headers`.

**Frontend.** A "Notify me" toggle on the node pages, a Notifications
section on the account page (`#/account?section=notifications`), an
unsubscribe page, and the notification mail count on the admin overview.

Spec:
[`docs/specs/2026-10-07-node-notifications-design.md`](https://github.com/efiten/CoreScope/blob/feat/notifications/docs/specs/2026-10-07-node-notifications-design.md).

## Performance

Each check reads watches, prefs and states from `users.db` (one query
each), the watched nodes from the analyzer DB in chunks of 500, and
last-heard times from the packet store.

| Measurement | Result |
|---|---|
| One check on our staging instance, 130,000 packets in memory, 1
watcher | 6.9 ms, of which 0.93 ms holding the store read lock |
| Evaluator benchmark, 100 users with 50 watches each, 2,000 nodes |
1.75 ms |

No work is added to ingest, broadcast or any request path.

## Verification

- `internal/users` and `internal/mailer` with `-race`; the `cmd/server`
suite plus `-race` on the notifier tests; 74 new Go tests. The same
Windows-only failure as noted in #2138 applies.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 26 of 26 steps locally. With notifications off
the new steps fail, so they do test the flag.
- On our staging and production instance since 7 October 2026.

## Not in this PR

- Other channels (Discord, Telegram, webhooks). Detection is separate
from delivery, so one can be added.
- The topology, RF and anomaly alerts of #775.
- If the whole server starts after a feed outage that already ended, the
grace window is unknown and a watcher can get one wrong offline mail.
The user guide says so.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 22:02:39 +02:00
efitenandClaude Opus 5.5 abbe6846ff feat: propose and approve hashtag channels (user management part D) (#2139)
Part D of #2128: logged-in users propose a hashtag channel, an admin
approves, rejects or later revokes it, and the ingestor decrypts
approved channels without a restart. It is the first consumer of a
generic proposal table, and answers the request in #2092.

Stacked on #2138 (part C). Review the commits after that branch; GitHub
shows C's commits here until it is merged.

## The situation

- An instance decrypts only the hashtag channels listed in
`hashChannels`. Users who want a city or club channel shown have to ask
the operator, who edits `config.json` and restarts the ingestor.
- #2092 asked for a suggest-and-approve flow; a downstream fork
(dborup/CoreScope#99) runs one with anonymous suggestions and a file
queue.

## What this PR adds

**Name rules (`internal/channel`).** `ValidateHashtagName`, mirrored in
the browser:
- at most 31 UTF-8 bytes including `#` (firmware
`ChannelDetails.name[32]`), case preserved, `Public` refused in any
case;
- control, bidi, separator and format characters refused except ZWJ,
plus invisible fillers (`Other_Default_Ignorable_Code_Point`, U+2800,
non-ASCII spaces).

Go and JS give the same verdict for every code point.

**Store.** Schema v4, one generic `proposals` table (kind, subject,
status pending/approved/rejected/revoked, proposer, reviewer, note). One
row per (kind, subject); limits are checked in the same transaction as
the change.

**Server** (only with `userManagement.channelProposals.enabled`)
- `POST /api/proposals`, `GET /api/account/proposals`, `GET
/api/admin/proposals`, `POST
/api/admin/proposals/{id}/approve|reject|revoke`.
- 400 for an invalid name. 409 for a duplicate, a rejected name, or a
name already in `hashChannels`/`channelKeys` ("this channel is already
decrypted on this instance"). 429 over a limit (`maxPending` 100,
`maxApproved` 128, `perUserPerDay` 5).
- `GET /api/channels` gains `approvedChannels`, so approved channels are
listed before they have traffic. With the feature off the response is
byte-identical.

**Ingestor.** Opens `users.db` with `mode=ro` every 60 seconds,
re-validates the names, and swaps its key map through an
`atomic.Pointer`. Configured channels always win, and a failed read
keeps the last good set. The ingestor never writes `users.db`; AGENTS.md
now says so.

**Frontend.** "Propose for everyone" in the add-channel dialog, an
"approved" marker in the channel list, "My proposals" on the account
page, and an admin "Proposals" tab with a warning on approve. The packet
detail pane now escapes the channel name: until now only the operator
chose those names.

Spec:
[`docs/specs/2026-10-07-channel-proposals-design.md`](https://github.com/efiten/CoreScope/blob/feat/channel-proposals/docs/specs/2026-10-07-channel-proposals-design.md).

## Performance

Every key is tried on each GRP_TXT that no key opens (`decodeGrpTxt`).
Benchmark of that worst case:

| Keys | Time per packet |
|---|---|
| 320 configured | 204 to 216 µs |
| 320 configured + 128 approved | 272 to 320 µs |

That is about +33%, linear in the key count and capped by `maxApproved`.
With the feature off the decoder gets the identical map and no ticker
runs.

## Verification

- `internal/channel`, `internal/users` (with `-race`), `cmd/server` and
`cmd/ingestor` pass locally, 53 new Go tests. The ingestor's
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced` needs symlink rights on
Windows and fails on `master` too.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 21 of 21 steps locally (propose, approve, the
channel appears, revoke removes it).
- On our staging and production instance since 7 October 2026.

## Not in this PR

- Private (PSK) channels: their keys stay in the browser (#725).
- Anonymous proposals.
- Names that need ZWNJ (some Persian and Urdu spellings) and subdivision
flags are refused, because ZWNJ and tag characters are format
characters. Allowing them later only widens the rule, so no approved
name gets stranded.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 21:51:42 +02:00
efitenandClaude Opus 5.5 18a634c115 feat: admin dashboard with overview and audit log (user management part C) (#2138)
Part C of #2128: an admin area with an overview, the user table and a
global audit log. Off by default like the rest of user management: with
`userManagement` off nothing changes, and non-admins get no new UI.

This is the first of 4 stacked PRs (C, D, E, then a small export/backup
follow-up). Each later one contains this branch; review them in order.

## The situation

- An admin could manage users one by one in `#/admin/users`, but nothing
showed whether the instance needs attention: accounts stuck in
activation, bouncing mail, someone guessing a password, an MQTT source
that dropped.
- The audit log existed (`internal/users/audit.go`) but could only be
read per user, and logins were not recorded.

## What this PR adds

**Store (`internal/users`)**
- Schema v3: an index on `audit_log(at)`.
- `AuditList` with filters (action or group prefix like `user.login.*`,
user as actor or target, period) and keyset pagination; `PruneAudit`;
`Stats` for the user figures.

**Server**
- `GET /api/admin/stats` (typed struct): accounts by status, admins,
registrations and active users over 7 and 30 days, logins and failed
logins in 24 hours, mail by final status, and the attention items
computed server-side.
- `GET /api/admin/audit`: filtered, newest first, `next` cursor.
- `GET /api/admin/users` gains `bouncing=1`.
- Logins are audited as `user.login` and `user.login.failed` (reason
`wrong_password`, `pending` or `disabled`). An unknown address writes no
row. The writes are asynchronous, so the login response does not wait on
`users.db`. Login rows are pruned after 90 days.

**Frontend**
- `#/admin?tab=overview|users|audit`: `admin.js` (tab shell),
`admin-overview.js` ("Needs attention", Users card, System card from the
existing health, MQTT and observer endpoints), `admin-audit.js` (filters
in the URL, "Load more"). `admin-users.js` becomes the Users tab; the
old `#/admin/users` link rewrites to it.
- `/api/healthz` is read on open and on Refresh only, not on the
60-second timer, because it walks every packet under a read lock.

Spec:
[`docs/specs/2026-10-07-admin-dashboard-design.md`](https://github.com/efiten/CoreScope/blob/feat/admin-dashboard/docs/specs/2026-10-07-admin-dashboard-design.md).

## Verification

- `internal/users` and `cmd/server`: `go vet` and `go test` pass
locally, 26 new Go tests. One upstream test,
`TestSaveGeoFilterPreservesFileMode`, also fails on Windows on plain
`master` (file modes) and is unrelated.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E with the `e2etest` build: 13 of 13 steps locally.
- Running on our staging and production instance since 7 October 2026.

## Not in this PR

- Restricting existing pages (perf, MQTT status) to logged-in users or
admins. The admin area adds no access mechanism of its own, so that
stays a later router rule plus endpoint check.
- Server-enforced customizer tab restrictions (#1508).
- The 24-hour, 5-attempt and 10-minute attention thresholds are
constants for now (AGENTS.md rule 8: customizer later).

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 21:38:30 +02:00
efitenandClaude Opus 5.5 a4445b4985 ci: run E2E, Go tests and the image build side by side (#2135)
Closes #2134.

## What changes

The test jobs and the image build now run side by side. Only the GHCR
push waits for all of them.

```
changes ─┬─ go-test ─────────────────────────┐
         ├─ race-test (ingestor changes only)  │
         ├─ e2e-shard ×3 ── e2e-test (gate) ───┼─ build-and-publish → deploy, publish
         └─ image-check (two-arch build) ──────┘
```

- **`e2e-shard` (new, matrix of 3).** It needs only `changes`. The suite
list stays in the workflow: each line now reads `suite <shard> <file>
[VAR=value ...]`, with the same environment variables as before.
- Shards are balanced on the measured suite times, about 6 minutes each.
- A shard number outside 1..3 fails the step, so a typo cannot drop a
suite silently.
- `test-issue-1648-m4-icons-e2e.js` stays in the same shard as, and
after, the a11y suite. Its distance-tab check only counts once the lazy
distance index (#1011) has been built, and the a11y suite is what first
requests it.
- **`e2e-test`** is now a gate under the old check name. It runs only if
all three shards passed, checks that three results arrived, and builds
the same `e2e-badges` artifact as before.
- **`image-check` (new).** It runs the two-arch build (no push) and the
arm64 QEMU smoke on every event, beside the tests. It also computes the
build metadata once, which the push reuses.
- **`build-and-publish`** now needs `go-test`, `e2e-test` and
`image-check`. On a PR every step is skipped. On push and tag refs it
only pushes, with the same build args as `image-check`, so every layer
should be a cache hit.
- **The unused `docker compose build` of the staging image is gone.**
`docker compose config --quiet` still validates the file.
- `AXE_SCREENSHOT_DIR` now actually reaches the a11y suite.
- `tests/unit/test-issue-1956-release-routing.js` follows the new jobs,
and checks more than before:
- a failure in `go-test`, `e2e-shard`, `e2e-test` or `image-check` skips
`build-and-publish`;
- the only non-publishing build lives in `image-check` and runs on every
event;
  - the push takes its build args from `image-check`.

## Measurements

The old numbers are from upstream; the new ones are from PR runs on my
fork (a branch on top of this commit, plus one temporary commit that
also triggers the workflow for PRs into a test base branch).

| | Before (PR run 37766553643) | After (fork runs 37812660655 /
37814554422) |
|---|---|---|
| Go Build & Test | 4.9 min | 7.2 / 7.5 min, in parallel |
| E2E | 18.4 min | 3 shards of 7.1–9.0 min, in parallel |
| Image build | 6.9 min, after E2E | 5.9 / 5.4 min, in parallel |
| **Total** | **30.5 min** | **9.6 / 10.6 min** |

**Same tests:** I compared the second fork run with the E2E log of the
upstream run, suite by suite. All 118 suites ran, and each one passed
the same number of checks (1161 in total).

**Cost:** more runner minutes, because each shard repeats about 2.5
minutes of setup. Standard runners are free for a public repository.

## Not verified here

- **A master push.** On a fork the GHCR push cannot run. The expectation
is that `Build and push to GHCR` gets cache hits from `image-check` and
finishes in a minute or two instead of 4–4.5. The first master run after
merge will show whether it does.
- **Branch protection.** If it requires other job names than `🎭
Playwright E2E Tests` and `🏗️ Build & Publish Docker Image`, those may
need adding. I cannot read the protection settings, so I could not
check.

## Tests run locally

- `sh test-all.sh` passes (on Windows with `PYTHONUTF8=1`).
- `test-issue-1956-release-routing.js` passes, and fails as expected
when `go-test` is removed from the `needs` of `build-and-publish`.
- actionlint 1.7.7 reports only the existing self-hosted runner label.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-08 20:13:09 +02:00
efitenandClaude Opus 5.5 f3ae8ac25a feat: sync a logged-in user's settings across devices (part B) (#2130)
Part B of #2128: a logged-in user's settings follow them across devices.
Log in on a phone and your own nodes, favorites, customizer and filters
are there; a change on one device reaches the others within a minute or
when you return to the tab.

**This PR builds on #2129.** Until that one is merged, the diff here
includes it. The commits for this part start at `docs(specs): settings
sync for optional user management (sub-project B)`.

## The situation

Everything a visitor sets up lives in one browser's `localStorage`
(about 100 keys in `public/`). A second device or a cleared cache starts
from zero (#895).

## What this PR adds

**Storage.** `users.db` schema v2: one JSON document per user in
`user_settings`, with a revision number and a generation id. A write
succeeds only when the client's revision and generation match the stored
ones, so two devices cannot overwrite each other silently.

**Server.** `GET`, `PUT` and `DELETE /api/account/settings`, behind the
same session and CSRF checks as the account routes.
- The server owns the list of synced keys (61 keys,
[`settings_allowlist.go`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/cmd/server/settings_allowlist.go))
and sends it to the client, so the two cannot drift.
- A hard denylist, checked first, refuses `meshcore-api-key`, every
`corescope_channel_*` key and `live-channel-colors` (#725). The colour
map is keyed by channel hash, and for a user-added channel that hash is
`user:<name>`, which would expose hashtag channel names.
- Documents are capped at 256 KiB, measured like `JSON.stringify`. PUT
is limited to 60 requests per hour per user. A stale revision gets 409
with the current document.

**Client**
([`settings-sync.js`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/public/settings-sync.js)).
Inert unless the feature is on and someone is logged in.
- It wraps `localStorage.setItem` and `removeItem` for allowlisted keys
only and pushes 2 seconds after the last change.
- It pulls on login, page load, tab focus and every 60 seconds while the
tab is visible.
- **Merge:** three-way, against a per-device baseline that belongs to
one user and one document generation. Lists (own nodes, favorites, saved
filters) merge per item, so an item added anywhere is kept and an item
removed on one device does not come back from another. Single values:
the profile wins unless only this device changed it.
- Remote changes are written without a push, theme and colour-blind
preset are re-applied, and the current page re-renders (skipped on
account pages and while the geofilter editor is open).

**UI.**
- Logout asks: keep my settings on this device (default), remove them
from this device, or cancel. Channel keys are never removed: no copy
exists anywhere else.
- The account page gets a "Settings sync" section: last synced time,
"Sync now", what is and is not synced, and "Delete synced settings from
my account".

## Not synced

Layout and device state (panel and column widths, collapsed panels, map
positions, geofilter drafts), channel data (#725), the API key, and all
`sessionStorage`. The full list is in the
[spec](https://github.com/efiten/CoreScope/blob/feat/settings-sync/docs/specs/2026-10-06-user-settings-sync-design.md).

## Performance

- One GET per page load, tab focus and minute while visible; one
debounced PUT per burst of changes.
- The `setItem` wrapper costs one Set lookup per write for non-synced
keys. A synced write reads one small revision key, not the stored
document.
- The server reads or writes one row per request.

## Verification

- `internal/users` and `cmd/server`: `go vet` and `go test` pass locally
(22 new Go tests), including a test that every allowlisted key still
occurs in `public/`, and denylist tests.
- `tests/unit/test-settings-sync.js`: 79 passing (vm, real module). The
cases cover the merge table, two tabs sharing one storage, stale answers
after a push, delete while a push is in flight, and logout while the
final push fails.
- `sh test-all.sh` exits 0.
- `tests/e2e/test-user-management-e2e.js` (10 steps, 4 of them new)
passed locally with two browser contexts as two devices: a favorite and
the packet time window travel from device 1 to device 2, a removal does
not come back, and "remove from this device" clears the synced keys
while a channel key stays.
- Checked by hand on a staging instance with a desktop and a phone on
one account.

## Not in this PR

- On a shared browser where the previous user chose "keep", the next
user's first login merges those settings into their own account. The
user guide says to choose "remove" on shared computers.
- Saved filter expressions are synced as typed, including any channel
names written in them. The guide says so.
- Realtime push between devices; the minute pull is the sync interval.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 11:20:44 +02:00
efitenandClaude Opus 5.5 d232072f85 feat: optional user accounts (part A: foundation) (#2129)
Part A of #2128: optional, off-by-default user accounts. With the
feature off nothing changes; with it on, visitors can register and log
in, and admins manage users and use the operator actions without the API
key.

PR #2130 (settings sync) builds on this one. The two are meant to be
merged together.

## The situation

- Operator actions (geofilter save and prune, backup, perf reset) need
the shared `apiKey`. There is no per-person right.
- Nothing in CoreScope knows who a visitor is, so the requests in #2128
that need that (#1835, #2092, #1508, #730) have nothing to build on.

## What this PR adds

**Two new Go modules**
- `internal/users`: a separate `users.db` (SQLite through
`modernc.org/sqlite`) with users, sessions, single-use tokens, an audit
log and a mail log. Passwords use argon2id.
- `internal/mailer`: a `Mailer` interface with a Brevo client (send,
delivery events, webhook parsing) and an in-memory fake for tests.

**Server (`cmd/server`)**, active only with `userManagement.enabled`
- 24 routes, all documented in OpenAPI under the `users` tag
([`auth_routes.go`](https://github.com/efiten/CoreScope/blob/feat/user-management/cmd/server/auth_routes.go)):
  - auth: register, activate, login, logout, me, forgot, reset;
- account: profile, password, email change with confirmation, sessions,
self-delete;
- admin: list, detail, disable, enable, delete, role, resend activation,
manual activation, mail status refresh;
  - a Brevo webhook, registered only when `mail.webhookSecret` is set.
- `requireAdmin` replaces `requireAPIKey` at the 7 operator call sites:
the API key **or** an admin session. With the feature off it is the old
API-key gate (`TestRequireAdminWithoutUserManagementIsAPIKeyGate`).
- `/api/config/client` gets `userManagement: {enabled: true}` only when
the service started; with the feature off the response is
byte-identical.

**Frontend**
- `auth.js` (header account control, request helper that adds the CSRF
header), `account.js` (login, register, activate, forgot, reset, confirm
email, my account), `admin-users.js` (`#/admin/users`, deep-linked
filters), `account.css` (theme tokens only).
- On phones the top-bar control is hidden, so a conditional entry goes
into the bottom-nav "More" sheet and the nav drawer.
- The customizer geofilter tab and the Perf "Reset stats" button use the
admin session when there is one.

**Config.** A `userManagement` block (`config.example.json`,
[`docs/user-guide/accounts.md`](https://github.com/efiten/CoreScope/blob/feat/user-management/docs/user-guide/accounts.md)).
The Brevo key can come from `CORESCOPE_BREVO_API_KEY`. The server
refuses to start when the block is enabled but incomplete.

## Security choices

- Session cookie `cs_session`: HttpOnly, SameSite=Lax, Secure when
`publicBaseUrl` is https. Every cookie-authenticated state change needs
the `X-CS-CSRF` header and a matching Origin.
- Activation needs the token **and** the account password. Without the
password, an attacker who keeps re-registering a known address could get
the owner to activate an account that carries the attacker's password.
- Register, forgot and email change answer identically for known and
unknown addresses. A password reset ends all sessions, a password change
ends all other sessions, and both end outstanding email-change links.
- Rate limits: login 10 per 15 minutes, register and forgot 5 per hour,
per IP and per address. The bucket count is capped. `trustedProxies`
makes the per-IP limits see real client IPs behind a proxy.
- Server logs carry `#<user id>`, never addresses, tokens or passwords;
mail-provider error texts are redacted before logging.

## Performance

No change to an existing hot path with the feature off. With it on:
- One `users.db` lookup per authenticated request (session by token
hash).
- The admin user table rebuilds its `tbody` on each filter change.
`users.List` caps the result at 1000 rows (`internal/users/users.go`),
which bounds the rebuild.
- `map[string]interface{}` in `openapi.go`: 79 before, 78 after.

## Verification

- `internal/users`, `internal/mailer` and `cmd/server`: `go vet` and `go
test -race` pass locally. 121 new Go tests.
- `cmd/server` with `-tags e2etest`: vet and the e2e hook tests pass.
- `sh test-all.sh` exits 0. `tests/unit/test-user-management-ui.js`: 67
passing (vm, real modules).
- `tests/e2e/test-user-management-e2e.js` (6 steps) passed locally
against an `e2etest` build with the fake mailer and against a
feature-off build. CI builds the `e2etest` binary and runs the suite on
a second server (`deploy.yml`).
- On a staging instance with a real Brevo key: register, activation mail
delivered, activate, admin table, "Refresh status" showing sent,
deferred, delivered, opened and clicked.

## Not in this PR

- Settings sync (#2130), the admin dashboard, approval flows and
notifications (parts B to E of #2128).
- A `requireReadAuth` mode (#1835). Sessions from this PR are what such
a mode would accept.
- Binary size and build time with `modernc.org/sqlite` linked next to
`mattn/go-sqlite3` were not measured. Their driver names do not collide.
#1992 discusses the driver choice.
- No Brevo webhook was configured on staging; delivery status there came
from "Refresh status".

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 09:25:29 +02:00
efitenandClaude Opus 5.5 000d9ab030 feat(coverage): RF noise-floor layer on the Mobile RX coverage page (#2113)
The ingestor already stores the noise floor that CoreDrive RX companions
report with each GPS fix (`client_rf_samples`, opt-in through
`clientRfSamples`, `cmd/ingestor/client_rf_sample.go:62`), but nothing
reads it back. An operator who enables it collects the data and cannot
see it. This adds the read side: `GET /api/rf-noise` and a Signal/Noise
toggle on the Mobile RX coverage page.

## What it does

- **`GET /api/rf-noise?bbox=&z=&days=`** returns a GeoJSON hex grid with
the median, quietest and noisiest noise floor per cell, in the same
shape as `/api/rx-coverage`. It is registered always and 404s unless
`clientRfSamples.enabled` is true, the same pattern as the coverage
routes.
- **Stationary samples are excluded.** A parked companion logs hundreds
of readings at one point, which would otherwise define its cell.
- **Coverage page:** a Signal/Noise toggle, rendered only when
`/api/config/client` reports `clientRfSamples: true`. The noise layer
reuses the coverage colour tokens with the axis inverted, because a
lower dBm is quieter. Tiers are at -115 and -108 dBm.
- **Deep link:** `#/rx-coverage?layer=noise` opens on the noise layer.
- **Empty and failed answers** are labelled on the map ("No RF samples
in this view yet", or a retry hint), so a blank map never reads as
"feature off".

No new configuration key: it reads the existing `clientRfSamples`
section that the ingestor already uses. Default off, so nothing changes
for an instance that has not opted in.

## Where it comes from

Ported from the efiten/CoreScope fork, where it has run on
analyzer.on8ar.eu since September (fork commits `42d09f0e`, `80581c43`,
`cde95078`). The cherry-picks conflicted with upstream's newer
`routes.go`, `types.go` and `rx-coverage.js`, so the final state was
ported by hand. The fork-only `/scopes` route that sat next to it in the
same hunk is deliberately left out.

## Performance

The query is bounded by `sampled_at` (indexed, `idx_crf_prune`) and the
bbox, aggregation is one pass plus a per-cell sort, and the response is
capped at 5000 cells. Measured on analyzer.on8ar.eu, all of Belgium
(`bbox=49.4,2.4,51.6,6.5&z=9`), from a client in Belgium, so network
time included:

| window | samples | cells | response time |
|---|---|---|---|
| 7 days | 10,835 | 241 | 0.37 s |
| 30 days | 43,303 | 429 | 0.39 s |

That table holds 45,270 rows in total. It is only read when someone
opens the noise layer, never on ingest or WebSocket paths.

## Tests

- `cmd/server/rf_noise_test.go`: 8 tests for the aggregation (median and
extremes, stationary exclusion, the cell cap, empty input) and the gate
(404 when off, even with data present). All pass, and the full
`cmd/server` suite passes locally.
- `tests/unit/test-rx-coverage-noise.js` (new, in `test-all.sh`): the
colour axis runs the right way, including both tier bounds and a string
median from the API. Slices the real function out of `rx-coverage.js`
and fails on master's copy, which has no thresholds.
- `tests/e2e/test-rx-coverage-noise-e2e.js` (new, wired in `deploy.yml`
and `scripts/non-unit-tests.json`): no toggle when the flag is off;
Noise fetches `/api/rf-noise`, draws the cells, swaps legend and
subtitle, puts `layer=noise` in the hash; a deep link opens on the noise
layer and an empty answer shows the message. 3/3 locally; against
master's `rx-coverage.js` and `roles.js`, 1/3 (only the "flag off" case
passes, as it should).
- `test-rx-coverage-viewport-e2e.js` still passes. ESLint 8 reports 0
errors on the changed files.

## Not done

- The thresholds (-115 / -108 dBm) are fitted to the fork's own data
(1,241 moving samples at the time) and are constants in
`rx-coverage.js`. Per AGENTS.md rule 8 they belong in the customizer
eventually; not in this PR.
- The E2E suite mocks the API. The real endpoint is covered by the Go
tests and by the measurements above, not by a fixture with seeded
samples.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 22:11:34 +02:00
efitenandClaude Opus 5.5 ef12535364 test(map): wait for the new route sidebar before measuring it (#2109)
Step (9) of `tests/e2e/test-path-inspector-e2e.js` fails intermittently
on master with `(3-9) fixture route coverage: 0 !== 700`: 3 of the 18
runs that executed it since 2026-09-30, including master runs
[37120155356](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37120155356)
and
[37156362836](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37156362836).
A red master run skips the image build, so `:edge` is not rebuilt.

## Cause

The step clicks a candidate and then waits for any `.mc-rt-sidebar`. The
previous route's sidebar survives the hash-only `goto` to `#/map`, and
`drawPacketRoute` (`public/map.js:1080`) replaces it only after
`/api/resolve-hops` answers. So the wait matched the old sidebar, and
the checks that follow measured either the old sidebar or, if the
replacement landed between the locator resolving and the evaluation, a
detached element, whose `getBoundingClientRect()` width is 0.

This is a test defect. The app keeps the previous route on screen while
the next one loads, which is intended.

## Fix

Remember the sidebar before the click and wait until a different one is
in the DOM. Test-only, one file.

## Evidence

Locally against the fixture, with `/api/resolve-hops` delayed through
`page.route` in a throwaway copy of the test, measuring which sidebar
the step-9 assertion sees:

| Delay | Before | After |
|---|---|---|
| 0 ms | new sidebar, pass | new sidebar, pass |
| 80 to 100 ms | old sidebar, `isConnected: false`, width 0 (3 of 5
runs) | new sidebar, pass (5 of 5) |
| 1500 ms | old sidebar, pass: the step checked the wrong page | new
sidebar, pass |

The unmodified fixed test then passed 3 of 3 runs.

## Not done

- I did not look for the same wait pattern in other suites.
- The delay probe is not committed. It needs a timing window to be
useful, and that window would be a fixed sleep in CI.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:27:02 +02:00
n30nex 9e49db9725 feat(nodes): group adverts by observed route evidence (#2085)
Red commit: `ac60248` ([two intended label
failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37137429333)).
Original grouping red: `767da9b`
([CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36353796690)).

Fixes #2073.

#2088 is merged. Rebased onto master `415362c`; all eight patches are
unchanged, including the test-first history.

Recent Adverts groups the bounded sample by authoritative `advert_kind`:
Flood, Mixed flood / direct (empty path), Direct (empty path), and Other
/ unknown. Missing evidence stays unknown. Each advert appears once,
preserving reception metadata, order, packet links, counts and
origin/history explanations. Wording describes observed remaining paths
without inferring original send mode or RF distance.

Validation at `6317baf`:
- All 185 standalone suites passed; syntax, XSS and lint passed (zero
errors, 92 existing warnings).
- Browser verified: `http://127.0.0.1:51827` (temporary Go fixture
server, now stopped), desktop pane/full view and mobile. Mixed evidence,
exact links and caveats checked again after pushing.
- The expanded fixture exposes the existing #2104 failure in the
unmodified core harness. Supplemental integration with #2105's exact
test correction passes 135 checks, with three existing skips. Both PRs
remain separate.
- [Fresh green
CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37138366726)
passed; three independent automated reviews found no must-fix issues.

E2E assertion added: `tests/e2e/test-e2e-playwright.js:142`.

Grouping is O(n) over the existing 20-advert sample. No new requests,
settings, colors or dependencies.

## Preflight overrides
External OpenClaw runner/profile unavailable; repository checks and
actual local Chromium were used.
2026-10-03 23:49:08 +02:00
n30nex 99a326a6b3 test(packets): verify observation hex after rerenders (#2105)
Red commit: `ca553e4` ([expected assertion
failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37137426691)).

Fixes #2104.

The observation hex test now opens the existing grouped fixture
directly, selects A → B → A with stable-ID locators, and asserts the
selected ID and exact raw bytes after each render. Missing fixture data
fails explicitly. Distinct observation frames are seeded after schema
migration; this test's fixed sleeps and silent skips are removed.

Rebased onto master `415362c` after #2090 merged. Both patches are
unchanged. The previous #1122 layout blocker now passes against the
actual current server and assets.

Validation at `5db8017`:
- All 185 standalone frontend suites passed.
- Core browser: 133 passed, three existing skips; #1122 layout 6/6,
filter 11/11, grouped-collapse 4/4.
- Syntax, YAML, fixture SQL, XSS and lint passed (zero errors, 92
existing warnings).
- Fresh red CI reached the intended detached-observation assertion.
[Green
CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37138187496)
passed; three independent automated reviews found no must-fix issues.

Only test code and fixture setup change; no configuration/customizer
implications.

## Preflight overrides
External OpenClaw runner/profile unavailable; repository checks and
actual local Chromium were used.
2026-10-03 23:48:32 +02:00
efitenandClaude Opus 5.5 415362c5af fix(packets): resolve path hops with the observer that heard them (#2097) (#2099)
Fixes #2097.

Every 1-byte path prefix on this network is shared. Measured on the live
deployment: 254 prefixes cover all 2043 nodes, nine of them on a single
prefix, and not one node has a prefix to itself. The packets page
printed one of those nine as a certainty.

It could already have said otherwise. `hop-resolver.js` marks such a hop
`ambiguous` and returns its full candidate list; `hop-display.js`
renders a warning badge with a count and a candidate popover;
`renderHop(h, observerId)` reads a per-observer cache key. None of it
fired, because `HopResolver.resolve()` takes six parameters and
`packets.js` passed one.

Without `observerId`, `packetIata` is null, `nodeInRegion()` never runs,
no candidate is flagged `regional`, `globalFallback` stays false, and
`hop-display.js` computes a badge count of 0. The whole chain stayed
silent while the data said it was a guess.

## What this changes

**Resolution carries the observer.** `resolveHops()` takes it and writes
the per-observer cache key that `renderHop()` has always read and
nothing ever wrote. `resolveHopsForPackets()` groups a multi-packet call
by observer so each group is filtered on its own. One site was the bug
in miniature: it wrote its result under `hopNameCache[k + ':' +
pkt.observer_id]` after resolving without that observer, so the key
promised something the value was not.

**The observer's position is the anchor.** `resolve()` has always
accepted `observerLat`/`observerLon` as the anchor at the receiving end:
when nothing later in the path is resolved, the observer is the next
known position, which is what `pickByAffinity` needs.
`observerPosition()` is the single lookup, used by both resolve paths.

This matters more than the IATA route, which turns out to be dead on
this network. `nodeInRegion()` looks the observer's code up in
`/api/iata-coords`; **0 of 42 observers have their code in that table**
(`ANR BRU GNE HEP KJK LGG MST NRW OBL OST` are all missing, while its 54
entries are `AMS`, `APC`, `ATL` and the like). So the 300 km filter has
never excluded anything here. 27 of 42 observers report lat/lon
directly, and that works. Filed separately as #2098, since it is a
network-wide lookup that resolves nothing and hop resolution may not be
its only consumer.

**Each badge belongs to its pill.** The badge is a sibling after the
pill, so a path read `A [8] → B [6] → C` with nothing to say which name
the 8 qualified. Pill and badge now render inside one `white-space:
nowrap` `.hop-group`, leaving the separator outside the pair. Only a hop
that carries a badge is wrapped; wrapping every hop would change the
layout of every path for nothing. That markup predates this work, but
these commits are what made it visible — before them the badge
essentially never appeared on this page.

**The list summarises, the detail pane does not.** Every 1-byte hop
getting its own badge turned a table row into a line of warning
triangles. `renderPath()` takes `{ summary: true }` for the three list
call sites: names with no per-hop badge, and one indicator reading "N of
M hops have more than one candidate". The two detail-pane sites are
unchanged and keep a badge per hop, because that is where the candidates
actually get read. `HopDisplay` gains `opts.badge === false` for it, and
the hop keeps its `hop-ambiguous` class either way so CSS can still mark
it.

## Scope

Deliberately not included: chaining the next resolved hop to narrow a
candidate set to one. `pickByAffinity` already scores on graph edges and
would do it; it needs neighbouring context this call does not yet
supply, and that belongs in its own change to `hop-resolver.js`. What is
here makes the uncertainty visible and picks a defensible candidate
rather than the first in index order.

## Verification

30 unit tests in `tests/unit/test-issue-2097-hop-ambiguity-badge.js`,
registered in `test-all.sh`. Full frontend suite exits 0.

Four of them are structural guards over `packets.js`, and each was
mutation-checked by reverting the change it protects:

- no `HopResolver.resolve()` call passes the hops alone
- no call passes an observer id and drops its position
- one helper does the observer lookup
- the detail pane calls `renderPath` without `summary`

The behavioural test for the anchor runs with `iataCoords` deliberately
empty, which is the live state, and lists the distant candidate first:
without an anchor the resolver keeps candidate order and picks it, with
one it does not, and the hop stays reported as ambiguous with all
candidates listed.

**Browser validated** on a staging deployment of this branch, against
live data:

| | list | detail pane |
|---|---|---|
| per-hop badges | 0 | 24 |
| path indicators | 1 | 0 |
| badges outside a `.hop-group` | — | 0 |

On the packet that prompted the issue, the first hop now resolves to a
repeater 10.6 km from the transmitter instead of one 126 km away, and
the second hop resolves to the only candidate that has a recorded
neighbour edge matching the next hop in the path.

## Note for reviewers

The last commit exists because staging disagreed with the tests. After
the anchor landed, the detail pane still picked the distant candidate:
it resolves through its own `HopResolver.resolve` call, which had the
observer id but a null position. The guard that now fails on exactly
that shape is what would have caught it before the deploy.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:40 +02:00
n30nex 3b415fb007 fix(live): enable node filtering after its handlers are ready (#2103)
Red commit: `8974616` ([CI assertion
failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37064461392)).

Fixes #2094.

The Live node-filter input now starts natively disabled and becomes
editable after its input, keyboard and blur handlers attach. Previously
it accepted text while initial node loading was pending, then never
processed that text without another keystroke.

The change adds one HTML attribute and one initialization assignment.
Failed initial loading still reaches the existing error handler and
installs the controls. Saved filters, URL precedence, autocomplete and
keyboard selection keep their existing behavior. No new settings,
styling, dependencies, requests or packet-processing work.

## Validation

- The existing CI-wired Live browser suite delays or aborts the real
initial node request. It verifies rejected premature typing, later
editability, actual suggestions, keyboard selection and URL state. Both
new cases failed on the red commit; all six cases pass locally on
`11fab46`.
- All 183 standalone frontend suites passed, plus lint and XSS gates.
Separate browser checks passed saved-filter restoration and URL
override; desktop/mobile loading/ready screenshots were inspected.
- A broader local core run exposed the unchanged stale
observation-handle test tracked in #2104. It is outside this Live fix.
Full
[CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37067337557)
passed at `11fab46`, including Go, browsers and both container
architectures; three independent reviews are clear.

E2E assertion added: `tests/e2e/test-1110-live-filter.js:69`.

Preflight overrides: external script/OpenClaw profile unavailable;
repository checks and local Chromium used.
2026-10-03 12:18:33 +02:00
n30nex b9fb3d244c feat(analytics): split adverts using recorded route evidence (#2088)
Red commits: `dc7dcbe`, `76e3217` ([assertion
failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37059977385)).

Fixes #2041; supplies #2085's backend contract.

Records at most two route bits per retained advert. Node flood counts
now use the same evidence, including mixed and transport-flood adverts.
Analytics failures log/count errors while core observation, path and
liveness writes continue. The unrelated recent-advert limit clamp is
removed.

Direct adverts (empty path) describes the observed remaining path across
valid hash widths. Firmware can forward an advert down to an empty path;
this cannot establish original send mode or RF distance. The `zero_hop`
API spelling remains compatible.

Successful post-upgrade evidence writes survive observation replacement
and restart. Older overwritten frames and failed writes can leave gaps.
The UI explains automatic background backfill.

## Measured validation

Synthetic Windows workloads; timings are not production guarantees:

- 300K retained adverts, one-hour window: median 64.35→39.84 ms; 13.98
MB→145 KB allocated. Seven-day allocations unchanged; five paired
samples.
- 1M transmissions/11M observations: backfill 5m43s. Simple 10 Hz
writer: all 3,423 observations persisted; p99 28 ms, maximum 323 ms.
- Simulated pre-backfill cursor: retained evidence ready 4m21s; full
catch-up 6m41s, 122 cache invalidations. Normal completed-migration
startup bulk-loads current evidence.

Portable opt-in scale tests and window benchmarks are included. All 183
frontend suites, full server tests, targeted races, lint and real-API
desktop/mobile checks passed. Full [Linux
CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37063210229)
passed at `36def9c`; three independent reviews are clear. Local Windows
testing hit the unchanged symlink-privilege test. No configuration
changes.

Preflight overrides: external script/OpenClaw profile unavailable;
repository checks and Chromium used.
2026-10-03 12:16:55 +02:00
Sylvain Rabot a4a447acea fix(packets): pack short columns, one expand arrow, Full Names toggle (#2090)
The packets table wastes most of its width on short columns, draws two
arrows on grouped rows, and cuts every observer and hop name short. This
fixes all three and adds a way to read the full path when it does not
fit.

## Why the columns were wide

`makeColumnsResizable` measured each column once on load and stored the
widths as **percentages** of the table. Percentages grow with the
screen: on a 2000px display "17s ago" sat in a 180px column. The
measurement also read the virtual-scroll spacer row — a single `colspan`
cell as wide as the table — as column 0, so the expand column came out
160px wide locally and about 950px on a live instance with grouped rows.

Measured on the fixture at 2000px:

| Column | Before | After |
|---|---|---|
| Expand | 162–230px | 32px |
| Time | 111px | 67px |
| Size | 64px | 46px |
| HB | 44px | 29px |
| Scope | 80px | 48px |
| Path / Details | ~180px each | ~650px each |

## The change

**Column packing.** A new `fitColumnsToContent` (`app.js`) sizes every
column except Path and Details to its widest cell, in px; those two
split what is left. Every visible column gets an explicit width and the
table gets their sum: a spacer row keeps a slot for each hidden column,
and leaving any column `auto` hands those slots a share of the slack as
a blank strip at the right edge.

When space is short (detail panel open, narrow window) the widest fixed
columns give way first, then Path and Details shrink to 60px, and only
then does the wrapper scroll. The split itself is a pure function,
`distributeColumnWidths`, with unit tests.

It re-lays out when the container resizes (detail panel), when
TableResponsive hides or reveals columns (a new `table-columns-changed`
event), and after sorting or a column toggle. After each render `grow()`
widens columns for the rows just inserted and never narrows them, so
nothing jitters while scrolling. Drag handles store px; double-clicking
one resets it. Path and Details have no handle.

`makeColumnsResizable`, still used by the nodes, observers and analytics
tables, now skips `colspan` rows when measuring; other rows are measured
as before.

**One expand arrow.** A CSS `::before` "▶" sat next to the SVG caret,
which pointed up when collapsed. The glyph is gone and the caret points
right, then down.

**Full Names.** A toolbar toggle next to Hex Paths shows observer and
hop names untruncated. It is saved per browser; a link can carry
`fullNames=1`, which applies to that page without overwriting the
visitor's saved preference. Paths stay on one line: the virtual scroller
positions rows at one fixed height, so wrapping would make scrolling
jump.

**The "+N" path pill, and a full-path popover.** Hops past the column
edge were meant to go behind a "+N" pill (#1124), but the pill was
appended after the overflowing chips and so sat past the clipped edge:
it never showed. It is now sticky at the right edge, opaque on hover,
and counts hop chips only (it used to count warning icons too). Hovering
or focusing it shows the full path, one hop per line and numbered;
click, Enter or Space pins it until an outside click or Escape, and it
closes when the page changes or its row scrolls away. Clicking the pill
no longer selects the row underneath.

**Two related fixes.**
- Rows drawn before `/api/observers` resolves show raw 64-char pubkeys.
`loadObservers` now re-renders and re-measures, instead of leaving
Observer at the pubkey's width (~600px under Full Names).
- Group child rows indented every cell by 20px, which misaligned them
with the headers and widened every column on expand. Only the Time cell
is indented now.

Anyone who had dragged packet columns gets the new defaults once: the
old percentage key is dropped.

## Performance

`grow()` runs after every `renderVisibleRows`, so it is on the hot path.
Measured in Chromium on the fixture, 300 iterations, three runs:

| Case | Cost |
|---|---|
| Full render, 63 rows measured | +0.88–0.93ms over the forced reflow
alone |
| Scroll step (incremental render) | ~73 `Range` rect reads, only the
inserted rows |
| Large jump / re-render | ~569 rect reads |

Before narrowing `grow()` to inserted rows, one scroll step with 92
rendered rows made 5,472 rect reads. The DOM row count is bounded by the
virtual-scroll window, so the cost does not grow with the packet count.
`distributeColumnWidths` is a binary search over at most ~10 columns.

## Tests

- `tests/e2e/test-packets-compact-columns-e2e.js`, 16 cases: short
columns stay under fixed px bounds at 1920px; Path and Details take the
freed width; hiding both still fills the wrapper; no blank strip and no
horizontal scroll, with and without the detail panel, at 1920, 1100,
1000 and 375px; responsive reveal and re-hide refit; stale percentage
widths dropped; drag follows the pointer and double-click resets;
expanding a group leaves widths alone; Full Names untruncates, keeps
paths on one line, persists and round-trips through the URL; the
Observer column shrinks back when observer names arrive late; the pill
popover lists every hop vertically, closes when the pointer leaves, pins
on click without selecting the row, closes on Escape and on navigation,
and opens on keyboard focus. Registered in `scripts/non-unit-tests.json`
and `deploy.yml`.
- On `master` the suite fails where it should, and restoring each fix's
old behaviour (child-row indent, the late-observer refit) turns its case
red.
- `tests/unit/test-packets.js`: `distributeColumnWidths` (room to spare,
rounding, widest-first capping, floors, dragged columns, the 60px
minimum, overflow), the single chevron and no `::before` glyph, Full
Names on flat and grouped observer cells, and the Full Names CSS.
- `tests/unit/test-xss-escape-sinks.js` now builds the real
`obsCellName` from source: escaping moved into it. Removing its
`escapeHtml` turns both observer-cell tests red.
- `test-issue-1189-composed-cell.js` and
`test-issue-2012-clear-filters-selection.js` evaluate packets.js
fragments with injected helpers and needed `obsCellName` /
`showFullNames` added.
- Existing packets E2E suites pass on this branch and on master against
a CI-prepared fixture: scope column, #1128 (layout and multi-viewport),
#1122, #1188, #1648, table sort, #1692, #1486.
- Unit: 182 of 183 suites pass locally;
`test-issue-1956-release-routing.js` could not create a temp directory
in my sandbox and is unrelated.

## Browser validation

Checked in Chromium against the fixture at 2000, 1920, 1280 (panel open
and closed), 1100 and 375px, and with the seeded grouped row expanded
and collapsed. Screenshots taken at each step; the table fills its
container exactly, grouped rows show one arrow, and the pill popover
lists full names.

## Not done, deliberately

- **No width cap on Observer under Full Names.** A cap brings back the
truncation the toggle removes, and Observer is already the first column
to give way when space is short.
- **Customizer.** The floors (60px flex minimum, 64px shrink floor, 32px
expand, 58px Rpt) are hardcoded, per AGENTS rule 8; exposing them in the
customizer can follow.
- **Wrapping paths.** Rejected because of the fixed row height above;
the pill popover is the way to read a long path.
2026-10-03 12:15:52 +02:00
n30nex 093e320c2b fix(map): keep Path Inspector candidates accessible and cover route replacement (#2082)
## Fix

Fixes #2060. Fixes #2081.

Path Inspector route drawing and candidate replacement now have
deterministic browser coverage. An idempotent SQL seed adds four
synthetic repeaters and four edges after fixture migration, without
changing packet ordering. The test uses the real API, normal clicks,
visible Leaflet geometry, and removal of every previous route object.

Restoring coverage exposed a desktop layout bug: drawing the first route
moved the map over the next candidate button. The route sidebar now
participates in the existing flex layout, including resizing and
collapse, while preserving the mobile bottom sheet.

The historical zero-candidate result did not reproduce on the current
baseline. No search thresholds or API behavior changed. No new
dependencies or customizer settings.

## Validation

- Local assertion-red history: `042db2b` requires seeded candidates;
`ca1d22a` exposes the blocked second click. Subsequent commits repair
layout and narrow-window behavior.
- All 183 standalone suites passed. Unchanged-base Go failures: #2083
readiness race (fixed in #2084) and Windows symlink privilege.
- Real Chromium: 11 Path Inspector/layout checks, 33 related map checks,
and core E2E (131 passed, 3 existing fixture skips).
- CSS-variable checks, XSS diff check, inventory, whitespace and PII
checks passed. Seed remains valid after periodic graph refresh.

Local browser validation used Chromium with a 60-second navigation
budget; the OpenClaw profile and external preflight script were
unavailable. Repository checks ran directly. [CI run
36343343569](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36343343569)
records the initial assertion-red commit. Final CI must pass before
merge.
2026-09-30 11:40:49 +02:00
n30nex ded27a8312 fix(packets): stabilize fixture tests and preserve mobile empty states (#2080)
Fixes #2054 and #2079.

Packet-row tests could outlive the fixture's default 15-minute window.
An audit of 115 wired suites identified four needing explicit fixture
windows: 180 minutes on mobile, 1440 on desktop. Mobile rejects larger
saved windows, so a desktop-only preference was insufficient.
Packet-icon checks now require a real path-bearing packet and both
detail controls.

Default-window coverage exposed #2079: responsive column hiding also hid
the spanning empty-message cell. Skip structural single-cell spanning
rows when hiding individual columns, preserving status messages and
virtual-scroll spacer height. Ordinary and partial-span data cells
retain their behavior; application defaults stay unchanged.

- Red evidence: `bae5e78` fails the aged-fixture row assertion;
`9b75cf6` fails spanning-row visibility. Guards/setup make them pass.
- Validation: 183 standalone suites; 82 targeted browser cases on a
fixture aged over 30 minutes (one pre-existing node-gesture skip).
Broader browser run: 131 passed, 3 fixture-dependent skips.
- Browser verified: fixture-backed Go server and Chromium;
`coverage/2079-mobile-empty.png`. Real mobile scrolling retains spacer
height and has no horizontal overflow.
- E2E assertions: `test-observer-iata-1188-e2e.js` default-window cases
and `test-table-fluid-e2e.js:58` shared-helper regression, both under
`tests/e2e/`.
- Three independent reviews completed; packet-icon coverage feedback
addressed in `760057d`.
- Populated mobile axe coverage remains part of #1995.
- One constant-time guard per already-visited row; no new requests or
collections.

## Preflight overrides

- External `run-all.sh` unavailable; repository checks run directly.
Local runs use UTF-8, the release harness's `GITHUB_REF_NAME`, and a 60s
navigation budget; repository timeouts are unchanged.
2026-09-30 11:40:36 +02:00
n30nex 410c82c02c fix(ui): use consistent 48px navigation and channel buttons (#2078)
Fixes #2052.

Navigation (`.nav-btn`) and channel icon (`.ch-icon-btn`) buttons now
use the shared 48px house minimum. Remove competing 44px
component/coarse-pointer rules and the legacy 32px mobile rule, and
correct comments that confused the house preference with accessibility
requirements.

At 768px, user-added channel rows already clipped Remove; the larger
Share target exposed the same constraint. Allow only those rows'
controls to wrap so both actions remain within the sidebar. Navigation
height and normal network-channel rows are unchanged.

- Red `7787d9b`: rendered target assertions fail at 375/390/768/1280px.
Green `e6ed9c8` changes CSS only.
- Validation: all 183 standalone suites; 33 computed touch-target
checks; 15 actual-page layout/click checks. Broader browser suite: 131
passed, 3 fixture-dependent skips. Three independent reviews found no
required changes.
- Browser verified: local Chromium and fixture-backed Go server;
screenshots `coverage/2052-targets-375.png` and
`coverage/2052-targets-768.png`.
- E2E assertion added: `tests/e2e/test-channel-fluid-e2e.js:112`,
covering dimensions, containment, overflow and normal Filter/Share
clicks. Existing CI selection is retained; Chromium-required runs cannot
silently skip the touch suite.
- Test-only follow-up `5144503` waits for packet readiness before
measuring navbar controls. Five fresh mobile contexts passed; size and
click assertions remain intact.
- No new settings, requests, dependencies or runtime JavaScript.

## Preflight overrides

- External `run-all.sh` unavailable; repository checks run directly.
- Local unit runner uses UTF-8 and `GITHUB_REF_NAME=local-validation`
for the existing release harness. Browser navigation uses a 60s local
budget for slow assets; repository assertions/timeouts are unchanged.
2026-09-30 11:40:22 +02:00
n30nex f1edbbef3f fix(packets): preserve observation selection in detail URLs (#2093)
Red commit: `c165087` ([CI assertion
failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36644030012)).

Fixes #2091.

Packet detail links now retain the selected observation when page
initialization or filter changes rebuild the URL. Initialization also
restores `obs` from the complete hash after the router strips its query.
Explicit ID links render the requested observation and retain it in Copy
Link state.

The shared updater reads selection from the current route, so returning
to the list or selecting another packet cannot resurrect an old
observation. Existing filter serialization and Clear Filters behavior
remain intact. No new requests, configuration, dependencies or layout
changes.

## Validation

- Real fixture browser checks use nondefault observation 502: hash/ID
load, type/observer/time-window changes, refresh, Clear, and another
refresh. Both URL and selected row are asserted.
- All 183 standalone frontend suites passed; focused filter browser
suite 11/11.
- Broader local core run reached the unrelated Live input readiness bug
#2094, reproduced on unchanged master. The complete CI browser suite
passed.
- ESLint 8: zero errors; 91 existing warnings. XSS diff, syntax and
whitespace checks passed.
- Three independent reviews found no production defects; assertions
additionally prove filter controls changed before checking selection.

E2E assertion added: `tests/e2e/test-filter-ux-e2e.js:179`.

OpenClaw profile/external preflight were unavailable; Chromium and
repository checks ran directly. Local navigation used a 60-second
budget. Final CI passed at `280ced5`, including Go, browsers, coverage
and both container architectures
([run](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36646175352)).
2026-09-30 11:40:00 +02:00
n30nex dc4db17c48 fix(nodes): remove unsupported region fields from Heard By (#2077)
Fixes #2062.

The node health API does not emit observer region data, but Heard By
rendered a Region column containing only dashes and a Regions summary
that could never appear. Remove that column, its sort control, and the
unsupported region displays from the full node page and side panel.

Keep Observer, Packets, Avg SNR and Avg RSSI, along with existing
escaping, signal placeholders, relay counts and badges. No API,
dependency or configuration changes.

- Red commit `4396492` fails on the unwanted Region header and summary;
`a360711` removes the unsupported fields.
- Validation: 183 standalone suites and 7 focused browser checks passed.
Broader browser run: 131 passed, 3 fixture-dependent skips. Independent
reviews completed; the sorting-test finding was addressed in `1452981`.
- Unit coverage evaluates the real templates. The existing CI-selected
browser suite checks four-column alignment and sorting on desktop/mobile
using the actual API field names.
- Browser verified: local Chromium against a fixture-backed Go server;
screenshots recorded in `coverage/issue-2062-heard-by-1400.png` and
`coverage/issue-2062-heard-by-390.png`.
- E2E assertion added:
`tests/e2e/test-issue-1151-orphan-separators-e2e.js:125`, the two #2062
desktop/mobile cases.
- No added requests, loops or data structures.

## Preflight overrides

- External `run-all.sh` unavailable; repository syntax, whitespace,
CSS-variable, XSS and PII checks run directly.
- Local unit execution supplies UTF-8 settings and
`GITHUB_REF_NAME=local-validation`, required by the existing
release-routing test harness.


- Local browser navigation budget raised to 60s for slow local asset
responses; repository assertions and timeout settings unchanged.
2026-09-30 11:39:51 +02:00
n30nex 31744c6da2 test(channels): isolate WS assertions from initial loading (#2087)
Fixes #2086.

The #1468 WebSocket browser checks could fail when initial channel
loading completed between their separate before/after evaluations. An
orphan message added no channel, yet unrelated loading changed the count
from 0 to 5; the positive control could also lose its sentinel when
loading replaced the list.

Each check now captures before state, processes its packet, and captures
after state in one synchronous browser evaluation. All four original
assertions and both message payloads are unchanged. This updates one
test file only, with no production, dependency or configuration changes.

## Validation

- Delayed real API loading reproduces both failures on unchanged master.
- Both corrected checks pass with loading held and with loading
completed; each uses one evaluation.
- Parent independently ran the delayed-loading harness and full browser
suite: 131 passed, three existing fixture skips.
- All 183 standalone frontend suites passed. Syntax, inventory,
whitespace and privacy checks passed.

TDD justification: test synchronization repair only. Existing assertions
demonstrate the baseline failure; no production behavior or manufactured
failing test was added.

Browser checks used Chromium against unchanged production source.
OpenClaw profile/external preflight were unavailable; repository checks
ran directly, with a 60-second local navigation budget. Final CI must
pass before merge.
2026-09-30 11:39:34 +02:00
efitenandClaude Opus 5 3caa847323 fix(nodes): the advert section is called "Recent Adverts", and says why (#2071)
Closes #2042, reported by @damn-simple-scripts.

## Which of the two fixes

The report offered widening the section to all packets the node
originated, or renaming it, and proposed the rename. Widening sounds
like the better fix, so I measured before agreeing. On a production
database:

| | |
|---|---|
| transmissions total | 1,251,967 |
| with `from_pubkey` populated | 191,237 |
| of those, `payload_type = 4` (ADVERT) | **191,237** |
| with `from_pubkey` and **not** an advert | **0** |

The ingestor fills that column for adverts only. Attributing a relayed
CHAN or TXT packet back to its sender is the path-resolution problem,
not a filter this section could apply — so "show all packets from this
node" is not a small change, it is a different feature resting on
attribution the data does not carry.

So the rename is correct, and the numbers say so rather than my
preference.

## The change

Both copies renamed — the full node page and the side pane. Both already
read `nodeData.recentAdverts`, so the field feeding them said "adverts"
while the heading said "packets".

The heading also gained a `title` naming `from_pubkey` as the reason it
is adverts only. Renaming without explaining invites the same report
from the next reader; the tooltip is where that explanation costs
nothing.

## Tests

`tests/unit/test-issue-2042-recent-adverts-label.js`, four cases: both
copies present and headed "Recent Adverts", no copy back to "Recent
Packets", the advert field still feeding them, and the explanation still
in place. The third matters most — it ties the label to its data source,
so pointing this section at a different field in future fails here
rather than silently making the label wrong again.

## Not done, deliberately

The reporter's follow-up idea, splitting flood adverts from zero-hop
adverts into two lists. They wrote that it can be dropped if an issue
should tackle one thing, so it is not here. It is feasible now: a
zero-hop advert is `ROUTE_TYPE_DIRECT` with an empty path, established
while fixing #2064. That deserves its own issue rather than a paragraph
in this one — say the word and I will open it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:08:28 +02:00
liquidraverandClaude Opus 5 b695b979a2 fix(nodes): keep paginating past a page that post-LIMIT filtering shortened (#2061)
## Problem

`handleNodes` runs the geo-filter, `nodeBlacklist`, `hiddenNamePrefixes`
and area
passes **after** the SQL `LIMIT/OFFSET`, and rewrites `total` to the
filtered
length. A page that loses a row is therefore short **without being the
last
page**, and neither the page length nor `total` can tell a client
whether to ask
for another page.

#1606 added the pagination loop and chose the page length as the
canonical stop.
That is correct only where nothing is ever filtered. Everywhere else the
list
truncates at the first filtered page boundary and strands every node
behind it —
the #1598 symptom reached by a different route: a node that is relaying
right
now simply stops being in the list.

The comment at `app.js:240` rejects `total` for exactly the right
reason, then
picks the signal the same code path also breaks.

## Measured on a live 2346-node deployment

Page sizes for the query the map issues:

```
offset=0     returned=500     ← full, loop continues
offset=500   returned=499     ← one row filtered AFTER the LIMIT → loop STOPS
offset=1000  returned=500     ← never requested
offset=1500  returned=500     ← never requested
offset=2000  returned=345     ← never requested
```

| stop rule | requests | nodes reached |
|---|---:|---:|
| short page (master) | 2 | **999** |
| `has_more`, else empty page | 6 | **2344** |

**1341 nodes, 57%, unreachable through the UI.**

### One hidden node truncates the whole list

The deployment this came from has no `geoFilter`
(`/api/config/geo-filter`
returns `polygon: null`) and no `nodeBlacklist`. It has a single
`hiddenNamePrefixes` entry — a deliberate operator choice — and exactly
one node
whose name starts with it:

```
public_key   d4a46ea2…1054      (64 clean hex chars)
name         🚫🔥☀️
role         repeater
last_seen    2026-09-22T10:09:38Z
```

`handleNodes` drops that row in the `IsNameHidden` pass, which runs
after the SQL
`LIMIT`. The row is counted by the `LIMIT` and by `COUNT(*)`, so the
page it lands
in comes back exactly one short — and stops every client that treats a
short page
as the end.

Isolated against SQL on the same database, seconds apart:

```
SELECT lower(public_key) FROM nodes ORDER BY last_seen DESC LIMIT 500 OFFSET 500
  -> 500 rows
GET /api/nodes?limit=500&offset=500
  -> 499 rows
comm -23 sql.txt api.txt
  -> d4a46ea2e99cab132a3286ef3d9cce9099318790af7f25671fe83de453721054
```

Deterministic — `offset=500` returned 499 on three consecutive requests.
Not a
CDN artifact either: `cf-cache-status: DYNAMIC`, origin `cache-control:
no-store`, no `age` header, and four requests with deliberately unique
cache keys
all returned 499.

So **one deliberately hidden node makes 1341 of 2344 nodes
unreachable.** The
hiding feature does exactly what it was asked to do for that one node,
and takes
57% of the network with it, silently. A single `hiddenNamePrefixes`
entry is
enough; no geo-filter, blacklist or area filter is needed to reach this
state.

### The cutoff moves, which is why this reads as intermittent

The visible set is the sum of the pages up to and including the first
short one,
so the boundary sits wherever the unreturnable row currently sorts by
`last_seen`, and jumps a whole page as ingest reorders the list. Same
deployment, same code, same config, ~2h apart:

| dropped row's rank | first short page | nodes visible |
|---|---|---:|
| inside 0–499 | page 1 | 499 |
| inside 500–999 | page 2 | 999 |

A node is visible or invisible purely by where it lands relative to that
moving
line, so affected nodes appear to vanish and return on their own. Two
operators
on this deployment reported exactly that, independently, while I was
measuring.

### A named reproduction

`HU-ZA-Lentihegy` (`5287a33f…`), reported missing from the map by an
operator
whose companion had logged its advert at 04:20 local the same morning.

Ingest was fine. The row is in `nodes` with `last_seen`
`2026-09-22T02:20:35Z` — the same advert, to the second — valid GPS,
role
`repeater`, 1033 adverts, and `/api/nodes/search?q=lentihegy` returns
it.

```
rank by last_seen : 1081
cutoff at the time:  999
```

It missed by 82 positions. Walking the same live endpoint, same moment:

| stop rule | requests | nodes reached | Lentihegy |
|---|---:|---:|---|
| short page (master) | 2 | 999 | **not reached** |
| `has_more`, else empty page | 6 | 2340 | reached |

The practical shape of this on a busy mesh: 1081 nodes had been heard
more
recently than 9.4 hours, so on that deployment **anything last heard
more than
~9 hours ago was invisible**, alive or not.

`#/nodes` compounds it — its search box filters client-side over the
truncated
set, so the server-side `?search=` never runs and an operator cannot
find the
node by searching for it either, even though the endpoint would return
it.

## Change

**Server** — `NodeListResponse` gains `has_more`, computed from the raw
SQL page
against the real `COUNT(*)` before the filter passes run, so it survives
them:

```go
hasMore := offset+len(nodes) < total
```

Always emitted (no `omitempty`) so a client can tell `false` from an old
server.
No extra request in the fixed path: `has_more` ends the loop exactly,
where the
old rule needed a probe page.

**Clients** — `app.js` `fetchAllNodes`, `nodes.js` `loadNodes` and
`area-map.html`'s inline helper stop on `has_more`, falling back to a
zero-length
page against a server that predates it. An empty page always ends the
loop, so a
`has_more` against a concurrently-shrinking table cannot spin to
`safetyCap`.

Left alone: the three loops are still three copies. Collapsing them onto
`fetchAllNodes` is a bigger change than this fix needs, and `nodes.js`
has its
own inter-page progress UI. Happy to do it separately if you want it.

## Testing

- **Unit** (`tests/unit/test-fetch-all-nodes-pagination.js`): the
fixture now
models the real handler — a row counted by the LIMIT and by `COUNT(*)`,
then
removed from the page. Three new cases. Fails on the old rule at 499 of
1199.
- **E2E** (`tests/e2e/test-map-nodes-pagination-e2e.js`, already wired
into
  `deploy.yml`): the mock drops a page-1 row and emits `has_more`.
  Mutation-checked — restoring master's stop rule fails 3 of its steps.
- **Go** (`cmd/server/nodes_pagination_has_more_test.go`): asserts
`has_more`
stays true on a page filtering shortened. Mutation-checked — recomputing
it
  after the filter block fails the test.
- Full server suite `go test -race`: ok, 41.4s. `gofmt` clean, `go vet`
passes.
- **Against a real binary**, not just mocks: fixture DB migrated with
`corescope-migrate`, `hiddenNamePrefixes: ["SKCE"]`, `limit=3`. Page 1
returns
2 of 3 with `total` rewritten to 2 and `has_more=true`. Walking the real
server
with master's rule reaches 2 nodes; with `has_more`, all 199 visible of
200,
the hidden one still hidden. The real frontend against that server loads
199
  with no JS errors.

Two existing expectations changed, both deliberate:

1. `surfaces ALL nodes past the 500 server cap` — 3 → 4 requests. That
mock emits
no `has_more`, so the 200-row final page can no longer end the loop (a
short
page is exactly what a filtered page looks like) and a zero-length probe
   follows. Against a current server `has_more` still ends it at 3.
2. `rows missing public_key are NOT collapsed into one` — its stub
returned a
constant body, which would now be paged to `safetyCap`. It serves one
page
   then empties.

Local `test-all.sh` exits 1 on two XSS-gate self-tests
(`good-2-tested.js`, `good-4-tested.js`) via a `UnicodeEncodeError`
printing an
emoji under Windows cp1252. Identical on clean `origin/master` in a
scratch
worktree, so it is pre-existing and platform-local, not this branch.

There is a second identical filter block further down `routes.go` on
another list
endpoint. Likely the same class; not touched here.

If you would rather land your own version of this, say so and I will
close mine.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 09:45:18 +02:00
efitenandClaude Opus 5 25f8426d32 test: make every suite in tests/e2e run, and tell the truth when it does (#2053)
Closes #2037.

Step 1 (#2045) wired in the nine suites that already passed. This is
steps 2 and 3: the four that ran and failed, and the five that could not
run at all. After this, the count of suites in `tests/e2e` invoked by
nothing goes from 18 to 0.

## Step 2 — the four that ran and failed

Triaged against a fixture server in CI with full output kept, not the
four-line tail the first probe saved.

| suite | verdict |
|---|---|
| `test-packets-scope-column.js` | passes on master today. It failed on
2026-09-18, so something between the two fixed it. Wired in unchanged
rather than investigated. |
| `test-node-reach-e2e.js` | test wrong. It waited for the reach map
whenever any *link* had GPS; `public/node-reach.js` builds the map only
when the *node* has coordinates. Guaranteed 10s timeout on a node with
positioned neighbours and no position of its own. |
| `test-channel-modal-e2e.js` | both failures test-side. The Add
button's visible label was shortened to "+ Add" with the accessible name
moved to `aria-label` (`public/channels.js:746`); and
`.ch-section-mychannels` is conditional on the visitor having added a
channel, which has not happened at that point in the suite. |
| `test-touch-targets.js` | three test-side, two a product finding. |

The three test-side touch-target failures were the harness measuring
controls that are not shown: `.compare-btn` (the CTA was removed in
#1646, and `style.css` says so), `.ch-back-btn` (`display:none` outside
the mobile channels layout), and `.filter-toggle-btn` (`display:none` on
mobile since #1461; the control shown is the navbar mirror, which
`mobile-page-actions.js:70` builds as a `.nav-btn`, so it was already
measured). All three are dropped from the table with the reason recorded
in the file.

The remaining two are **not** a test problem: `.nav-btn` and
`.ch-icon-btn` are each declared twice in `public/style.css`, 48px in
the touch-target block and 44px in their own component rule, and the
later one wins. Rather than lower the blanket or hide the failures, the
suite now has `DEFAULT_MIN = 48` plus a `MIN_OVERRIDES` table holding
those two at their effective 44, so a third selector dropping to 44
still fails the build. The contradiction is **#2052**, with both ways
out costed; the override entries should go when it is settled.

## Step 3 — the five that could not run

None is deleted. I checked each selector and seam against the product
before deciding, and every one still targets a surface that exists and
that nothing else covers.

Four were written against `@playwright/test`, a runner the project
neither installs nor uses anywhere else. Adopting a second runner for
twelve tests costs more than porting them, and the precedent is already
set: `test-path-inspector-coverage-e2e.js` exists, as its own header
says, because `test-path-inspector-e2e.js` could not run. So they are
ported to the plain-node Chromium pattern the other 109 suites use.

- **`test-issue-1522-trace-url-sync-e2e.js`** — the trace hash in the
URL, both directions. `test-e2e-playwright.js` covers that the page
loads and searches; it never looks at the URL, which is the whole of
#1522.
- **`test-marker-outline-weight.js`** — the canvas pulse ring never
thins below 2px. There is no CSS rule to read and axe cannot see inside
a canvas, so sampling the seam is the only way. Added a guard that the
ring was actually visible, so the weight check cannot pass vacuously on
a pulse that never rendered.
- **`test-pr-1490-live-map-gpu-animations-e2e.js`** — the queue drains,
the engine sleeps again, the fading trails stay under the cap of 5, and
the canvas sits on `animationsPane` rather than under the markers.
- **`test-path-inspector-e2e.js`** — reduced to what nothing else
covers: the map side pane, the `/#/traces/<hash>` redirect, the tools
landing. Its standalone-page test duplicated the wired coverage suite
and is dropped. Its "switching candidate clears prior polyline" case
ended after the click with a comment and no assertion, which is the same
green-but-empty problem this issue is about; it now compares path
counts, and skips loudly when the fixture yields too few candidates.

The fifth, **`test-table-sort.js`**, needed `jsdom`, which was declared
nowhere. It is a unit test of `public/table-sort.js` filed under
`tests/e2e`, so: `jsdom` is a devDependency (lockfile updated, `npm ci`
stays consistent), the file moved to `tests/unit/`, and the
`domIntegration` group in `scripts/non-unit-tests.json` is gone with its
only member. It runs 22 tests. 20 passed immediately; 2 had rotted,
because #1648 M2 replaced the up/down glyphs with Phosphor sprites and
the direction moved out of `textContent` into the `<use href>`. Those
two now read the sprite ref and the `aria-sort` value, so they also
guard the accessible announcement.

## Verification

`tests/unit/test-table-sort.js` 22/22 and `test-test-inventory.js` pass
locally; the E2E suites need a fixture server, which I cannot build here
(no cgo toolchain since #1992), so CI is their first run as committed.
The triage above was measured in CI, not assumed.

## Not done

The per-assertion skips named in the second comment on #2037 are
untouched: the two flaky packet-detail cases, the fixture-data ones, and
the two `clientRxCoverage` suites that skip wholesale while reporting
success. Those need a fixture deployment with coverage enabled, which is
its own change. I have not opened it.

Option 2 from the issue, making `test-test-inventory.js` require a
`deploy.yml` line for every `tests/e2e` file, is also not here. It is
the right guard and it is now enforceable, since the list is finally at
zero, but it belongs in its own change where a red build means what it
says.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 11:45:11 +02:00
efitenandClaude Opus 5 5c016a210a fix(analytics): show a building state on the distance index's 202, and stop caching it (#2051)
Closes #1997.

## What was wrong

`/api/analytics/distance` answers `202 {status:"building",
retry_after_seconds:5}` with no `summary` until the lazy index (#1011)
has been built.

1. `renderDistanceTab` read `data.summary.totalHops` straight away. The
TypeError is caught by the tab's own try/catch, so it never reaches
`window.onerror`: it is painted into the tab as `Failed to load distance
analytics: Cannot read properties of undefined (reading 'totalHops')`.
2. `api()` caches any `res.ok` body, and `res.ok` is true for 202, so
the placeholder was stored for the `analyticsRF` TTL. Even a correct
retry read the cached "building" body back. That is the half that made
the broken state outlast the index build.

## What changed

- `public/app.js:173` skips the cache write when `res.status === 202`. A
200 still caches, unchanged.
- `public/analytics.js` renders a building notice and retries itself,
honouring `retry_after_seconds` clamped to [1s, 30s]. The timer is
cleared on tab switch and in `destroy()`, and a new render supersedes a
pending retry, so two renders cannot write into the same tab.

The notice uses `.text-center`/`.text-muted` rather than the `.spinner`
class used at `analytics.js:1683`, because `.spinner` has no CSS
anywhere in the repo and renders nothing.

## Tests

`tests/unit/test-issue-1997-distance-building.js` (7 assertions, wired
into `test-all.sh`) pins the two pure decisions the renderer makes and
`api()`'s refusal to cache a 202 while still caching a 200. Red-run on
the unfixed sources: 6 of 7 fail, and the "a 200 is still cached"
control stays green.

`tests/e2e/test-issue-1997-distance-building-e2e.js` (classified in
`scripts/non-unit-tests.json`, invoked from `deploy.yml` with
`CHROMIUM_REQUIRE=1`) serves both responses by route interception, so it
does not depend on whether the server under test has an index built. It
asserts: the building state appears, the tab does not paint the error
text, a retry arrives with no interaction, the retry replaces the
placeholder once the server answers 200, and no retry fires after
leaving the tab.

Check (2) deliberately asserts on the rendered text and not on
`pageerror`: the TypeError is caught, so a `pageerror` assertion would
pass on the broken build too.

## Verification

No local cgo toolchain here since #1992, so I could not build a server
to run the E2E against. Instead I ran its five steps in Playwright
against a live instance with this branch's `public/app.js` and
`public/analytics.js` injected in place of the deployed ones (both files
are byte-identical between that instance and upstream master, so the
injection is faithful):

| check | deployed build | this branch |
|---|---|---|
| (1) building state shown | fail | pass |
| (2) not rendered as data | fail | pass |
| (3) retried on its own | fail (1 request) | pass (2 requests) |
| (4) real payload after retry | fail | pass |
| (5) no retry after leaving the tab | pass | pass |

(5) passes on the broken build too: it schedules no retry at all, so it
is a control and only means anything together with (3).

The committed E2E suite has not been run as committed. CI is its first
real run.

## Not done

- The server still recomputes on every 202 poll rather than signalling
readiness.
- No other analytics tab was audited for the same assume-a-summary
pattern.
- One pre-existing unit suite (`test-preflight-xss-gate.js`) fails on
this Windows machine with a cp1252 `UnicodeEncodeError` from its Python
helper, on a clean tree as well as with this change. Unrelated, and
green on Linux CI.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 17:04:37 +02:00
liquidraver 84eba41592 fix(live): anchor the legend toggle to .live-page so the VCR bar stops eating its clicks (#2049)
The PACKET TYPES legend could not be dismissed on Live: clicking the palette toggle did nothing, while document.querySelector('#legendToggleBtn').click() from the console worked. That asymmetry was the diagnosis, since a synthesised click skips hit testing.

Residual half of #1833. That fix corrected the button's offset to calc(var(--vcr-bar-height) + 10px) but left the anchor: .legend-toggle-btn was position: fixed, so the offset resolved against the viewport, while the .vcr-bar it clears is absolute inside .live-page, whose height subtracts --bottom-nav-reserve (56px + safe-area at <=768). The two anchors disagreed by exactly the reserve, dropping the button into the bar's band, and .vcr-bar at z-index 1000 against the button's 500 took every real click. The window is 641-768px: below 641 both buttons are display:none, above 768 the reserve is 0, which is why neither the desktop nor the phone layout test saw it.

Both buttons are now position: absolute, resolving against the same containing block as the bar. .feed-show-btn carried the identical defect but missed the bar by 16px because --legend-toggle-stack parks it a row higher, so that change prevents a break rather than repairs one.

Three review rounds, each verified rather than argued:
- The new E2E was classified in non-unit-tests.json but had no deploy.yml line, so it would never have run (#2037). Now wired in, and it covers both buttons.
- The '.feed-show-btn is display:none at <=768' claim was wrong (it is <=640), which the author measured and corrected.
- Its first real execution failed on its own setup: getComputedStyle().getPropertyValue() on an unregistered custom property returns the unevaluated declaration ('calc(56px + 0px)'), so the non-zero check could not work. It now asserts the measured gap instead, and deliberately does not pin 56 because the gap measures 58 on CI.

Final run: 14 passed, 0 failed. The discriminating assertion is (f), each button's offset net of the bar being identical at 720 and 1440 (11px and 123px slack); under the old anchor those differ by the reserve. (g) pins the desktop layout unmoved.

Merged by the interim maintainer without a second human reviewer: CI and the review above are the independent checks.

Fixes #1833
2026-09-20 14:00:45 +02:00
n30nex 7f4b357d89 fix(map): keep inspector controls clickable and ignore stale loads (#2048)
Fixes #1998 and #2030.

Map Controls overlapped the Path Inspector toggle on desktop and intercepted its clicks; the controls are now offset by the collapsed and expanded inspector widths, with the mobile layout unchanged.

Leaving the Map page while its node data was loading could log an invalidateSize error or let an old response overwrite a replacement map. Node-loading continuations and the deferred initialization callbacks are now tied to the Leaflet instance that started them, and obsolete results return before touching state or rendering.

Reviewed against the code: the identity check is sound because destroy() clears the reference (map.remove() then map = null, public/map.js:2428-2433), so every guard is false after teardown and after replacement. Each await in loadNodes is followed by a check, as are the two deferred callbacks, the error branch and the finally. Moving 'nodes = data.nodes' after the observers await is deliberate: shared state is only mutated once the load is known to still own the map.

The tests are the strong half. Overlap is asserted through elementFromPoint on the toggle's centre across five viewport widths and both pane states, rather than a forced click that would pass over an overlapping control, and 641/640 pin both sides of the breakpoint. Both extended suites are invoked by deploy.yml, one line each, which was checked rather than assumed. The commit chain is test-then-fix twice: e675100/7d736d1 and fd56ea2/823668b.

One nit left to the author: the CSS repeats the pane widths from style.css:4532-4533 four times, so changing the pane width would silently misalign the controls, caught only as a waitForFunction timeout.

Merged by the interim maintainer without a second human reviewer: CI and the review above are the independent checks.
2026-09-19 21:11:53 +02:00
efiten 6c4fa041de test(rx-coverage): budget the viewport assertions in pixels, not a flat 0.001° (#2040)
assertViewport allowed 0.001 degrees between the centre it asked for and the centre Leaflet reports back. Leaflet keeps the centre as a pixel coordinate, so a read-back centre is only accurate to the pixel it landed on, and one pixel spans 360/(256*2^zoom) degrees: 0.02197 at zoom 6, 0.00137 at zoom 10. The flat budget demanded 1/20 of a pixel at zoom 6 while its own comment said 'allow sub-pixel coordinate rounding'.

It passes on master because the container rounds favourably. It fails as soon as the layout changes: on the ON8AR fork, whose coverage page carries one extra toolbar row, the first assertion reported {lat:11.996338401936226, lng:33.99169921875001, zoom:6} against an expected 12/34 — 0.167 px of latitude and 0.378 px of longitude, so the view was right and the assertion was wrong. Any container-height change can trip it, fork or not.

The budget is now one pixel at the asserted zoom via a degPerPixel(zoom) helper, applied to both centre assertions and to the waitForFunction on the hash (zoom 10, where 0.001 was 0.73 px and merely lucky). For latitude the Mercator scale is the longitude figure times cos(lat), so the longitude figure is the looser bound there, deliberately: this asserts the view is where we asked, not a projection identity. A one-tile miss at zoom 6 is 5.6 degrees and still fails.

Verified by arithmetic rather than locally (the suite needs a served instance and a browser): the observed drift falls inside the new budget and a one-tile miss does not. CI run 35273657579 is green, with 'RX coverage viewport browser regressions OK' in the Playwright job.

Merged by the interim maintainer without a second human reviewer: CI is the independent check.
2026-09-18 10:51:34 +02:00
liquidraver bbf54cfe1b fix(rx-coverage): open at the configured map default, with its own saved viewport (#2033)
The coverage page opened at a hardcoded [51.0, 4.8] zoom 8 regardless of deployment, ignoring /api/config/map (#2032). It now follows the same precedence as the main map (URL hash, then saved position, then /api/config/map, then [37.6, -122.1] zoom 9) and persists its own position across visits, syncing lat/lon/zoom into the hash so a view is shareable.

The saved position lives under its own key, rx-coverage-view, and the page never reads or writes the main map's map-view: sharing the configured default was the bug, sharing the session position was not. Both suites assert map-view stays untouched after a pan, so reintroducing a shared write fails instead of passing quietly.

Also fixed here: selectedRx is now percent-encoded into the hash, and a generation counter stops a late /api/config/map response or a stale 150ms layout timer from building a map for a page that was already left.

Reviewed twice. Verified by mutation rather than by reading: writing map-view too, ignoring the saved coverage position, and dropping the /api/config/map fetch each make the unit suite exit 1, so it covers the feature, the fix and the rejected alternative. The deploy.yml invocation was confirmed to land inside Run Playwright E2E tests (fail-fast) by parsing the workflow, and CI run 35261689694 is the E2E suite's first real execution: Go prints 'RX coverage viewport regressions OK' and Playwright prints 'RX coverage viewport browser regressions OK'.

Worth recording for the next reviewer: that E2E asserted localStorage.getItem('map-view'), so wiring it into deploy.yml without updating the assertion would have turned the job red on its first ever run. It was registered in scripts/non-unit-tests.json but invoked by nothing, the gap tracked as #2037.

Merged by the interim maintainer without a second human reviewer: CI and the mutation checks above are the independent checks.

Fixes #2032
2026-09-17 22:02:36 +02:00
Alex B 893773338e chore(tests): move root test-*.js into tests/unit and tests/e2e (#2036)
Moves 290 root test-*.js into tests/unit (177, listed in test-all.sh) and tests/e2e (113, classified in scripts/non-unit-tests.json), per #1981 and PR-D of #1385. Root goes from 348 entries to 48. test-all.sh and test-fixtures/ stay put. The inventory guard now fails if a test reappears in the root or sits in the wrong folder.

Verified independently of the diff: the invoked sets are unchanged (test-all.sh 177 before and after, deploy.yml 96 before and after, both identical as sets), and a full local run of test-all.sh on master and on the branch produced 4702 output lines each whose only differences are absolute paths, stack-trace line numbers shifted by the REPO_ROOT line, the inventory wording and two perf ratios. The guard was mutation-checked: a test back in the root, a unit suite in tests/e2e, and a suite dropped from test-all.sh each make it exit 1. CI run 35246304316 ran 97 suites from tests/e2e and is green.

Follow-up 9335c51d finished the instruction files: no bare root test command is left in AGENTS.md, the squad charters, .github or docs, and every tests/ path they name resolves.

Merged by the interim maintainer without a second human reviewer: CI and the local runs above are the independent checks.

Known and deliberately out of scope: 18 of the 113 files in tests/e2e are invoked by no runner at all, and one of them cannot run anywhere because it requires jsdom, which is not a declared dependency. Tracked separately.
2026-09-17 19:06:07 +02:00