Commit Graph
2967 Commits
Author SHA1 Message Date
n30nex f1edbbef3f fix(packets): preserve observation selection in detail URLs (#2093)
Red commit: `c165087` ([CI assertion
failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36644030012)).

Fixes #2091.

Packet detail links now retain the selected observation when page
initialization or filter changes rebuild the URL. Initialization also
restores `obs` from the complete hash after the router strips its query.
Explicit ID links render the requested observation and retain it in Copy
Link state.

The shared updater reads selection from the current route, so returning
to the list or selecting another packet cannot resurrect an old
observation. Existing filter serialization and Clear Filters behavior
remain intact. No new requests, configuration, dependencies or layout
changes.

## Validation

- Real fixture browser checks use nondefault observation 502: hash/ID
load, type/observer/time-window changes, refresh, Clear, and another
refresh. Both URL and selected row are asserted.
- All 183 standalone frontend suites passed; focused filter browser
suite 11/11.
- Broader local core run reached the unrelated Live input readiness bug
#2094, reproduced on unchanged master. The complete CI browser suite
passed.
- ESLint 8: zero errors; 91 existing warnings. XSS diff, syntax and
whitespace checks passed.
- Three independent reviews found no production defects; assertions
additionally prove filter controls changed before checking selection.

E2E assertion added: `tests/e2e/test-filter-ux-e2e.js:179`.

OpenClaw profile/external preflight were unavailable; Chromium and
repository checks ran directly. Local navigation used a 60-second
budget. Final CI passed at `280ced5`, including Go, browsers, coverage
and both container architectures
([run](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36646175352)).
2026-09-30 11:40:00 +02:00
n30nex dc4db17c48 fix(nodes): remove unsupported region fields from Heard By (#2077)
Fixes #2062.

The node health API does not emit observer region data, but Heard By
rendered a Region column containing only dashes and a Regions summary
that could never appear. Remove that column, its sort control, and the
unsupported region displays from the full node page and side panel.

Keep Observer, Packets, Avg SNR and Avg RSSI, along with existing
escaping, signal placeholders, relay counts and badges. No API,
dependency or configuration changes.

- Red commit `4396492` fails on the unwanted Region header and summary;
`a360711` removes the unsupported fields.
- Validation: 183 standalone suites and 7 focused browser checks passed.
Broader browser run: 131 passed, 3 fixture-dependent skips. Independent
reviews completed; the sorting-test finding was addressed in `1452981`.
- Unit coverage evaluates the real templates. The existing CI-selected
browser suite checks four-column alignment and sorting on desktop/mobile
using the actual API field names.
- Browser verified: local Chromium against a fixture-backed Go server;
screenshots recorded in `coverage/issue-2062-heard-by-1400.png` and
`coverage/issue-2062-heard-by-390.png`.
- E2E assertion added:
`tests/e2e/test-issue-1151-orphan-separators-e2e.js:125`, the two #2062
desktop/mobile cases.
- No added requests, loops or data structures.

## Preflight overrides

- External `run-all.sh` unavailable; repository syntax, whitespace,
CSS-variable, XSS and PII checks run directly.
- Local unit execution supplies UTF-8 settings and
`GITHUB_REF_NAME=local-validation`, required by the existing
release-routing test harness.


- Local browser navigation budget raised to 60s for slow local asset
responses; repository assertions and timeout settings unchanged.
2026-09-30 11:39:51 +02:00
n30nex 31744c6da2 test(channels): isolate WS assertions from initial loading (#2087)
Fixes #2086.

The #1468 WebSocket browser checks could fail when initial channel
loading completed between their separate before/after evaluations. An
orphan message added no channel, yet unrelated loading changed the count
from 0 to 5; the positive control could also lose its sentinel when
loading replaced the list.

Each check now captures before state, processes its packet, and captures
after state in one synchronous browser evaluation. All four original
assertions and both message payloads are unchanged. This updates one
test file only, with no production, dependency or configuration changes.

## Validation

- Delayed real API loading reproduces both failures on unchanged master.
- Both corrected checks pass with loading held and with loading
completed; each uses one evaluation.
- Parent independently ran the delayed-loading harness and full browser
suite: 131 passed, three existing fixture skips.
- All 183 standalone frontend suites passed. Syntax, inventory,
whitespace and privacy checks passed.

TDD justification: test synchronization repair only. Existing assertions
demonstrate the baseline failure; no production behavior or manufactured
failing test was added.

Browser checks used Chromium against unchanged production source.
OpenClaw profile/external preflight were unavailable; repository checks
ran directly, with a 60-second local navigation budget. Final CI must
pass before merge.
2026-09-30 11:39:34 +02:00
n30nex 248d2045fd test(server): await indexes before node-path regression requests (#2084)
Fixes #2083.

Node-path regression tests could request `/paths` while background
indexes were still loading, intermittently receiving HTTP 503 instead of
exercising hop resolution or sorting. Seven fixture loads now await the
existing bounded `WaitIndexesReady` signal. The anchor-bias test uses
the same signal instead of polling.

This changes five test files only. Production readiness behavior and
every HTTP/content assertion are preserved. No dependencies,
configuration or customizer changes.

## Validation

- Unchanged baseline: 94 passed / 6 failed across 100 targeted
executions; failures were actual HTTP 503 assertions.
- Fixed setup: 200/200 targeted executions passed.
- Parent independently ran the entire server suite: exit 0, 83.4
seconds.
- All 183 standalone frontend suites and 20 real Chromium route-map
checks passed.
- Formatting, whitespace and PII checks passed.

TDD justification: test-fixture synchronization repair, with no
production logic change. The existing unchanged behavioral assertions
supplied the before/after failure proof; no fabricated failing test was
added.

Local browser checks used Chromium against the unchanged production
source; OpenClaw profile/external preflight were unavailable. The
unrelated ingestor symlink test requires a Windows privilege absent
locally; Linux CI validates that suite. Final CI must pass before merge.
2026-09-30 11:39:25 +02:00
efitenandClaude Opus 5 9eb3098867 fix(release): publish the release notes as the release body (#2076)
Three releases went out with an empty release body. Checked with `gh
release view --json body`:

| release | body | assets |
|---|---|---|
| v3.9.1 | 1206 bytes | 2 |
| v3.9.2 | 2768 bytes | 0 |
| v3.10.1 | **0** | 2 |
| v3.11.0 | **0** | 2 |
| v3.12.0 | **0** | 2 |

v3.9.1 and v3.9.2 were written by hand, which this repository then
stopped allowing because published releases are immutable. Nothing
replaced them, so the `docs/release-notes/` convention was never wired
to the release page and three releases shipped with a description of
nothing.

Reported by the fork operator, who went looking on the v3.12.0 page for
the list of fixed issues that older releases carried.

## The cause

`action-gh-release` in this workflow was given `files` and
`fail_on_unmatched_files` and nothing else. No `body`, no `body_path`,
no `generate_release_notes`.

## The change

A step resolves `docs/release-notes/${GITHUB_REF_NAME}.md` and passes it
as `body_path`, with `generate_release_notes: true` so GitHub's pull
request list lands underneath the hand-written notes. A missing notes
file is a warning rather than a failure: the release then still gets the
generated list, which is more than an empty body.

## Also in this PR

The v3.12.0 notes file gains the two sections it should have had:

- **Issues closed**, 13 of them. Collected from each pull request's
`closingIssuesReferences`, not from commit text, which is why 30 pull
requests yield 13 issues.
- **Contributors**, separating pull request authors (@efiten 22,
@liquidraver 3, @A13xB0 2, @n30nex 1, @dborup 1, @anieto 1) from commit
co-authors (@nullrouten, @anieto, SaarMesh-Bot, Openclaw) from issue
reporters (@efiten 5, @n30nex 4, @anieto 2, @liquidraver 1,
@damn-simple-scripts 1). It says outright that 22 of the 30 pull
requests are the interim maintainer's own, which is the shape of a
release cut while the owner is unreachable, not a healthy ratio.

v3.12.0's published body has already been set to exactly this content by
hand, so the release page and the file agree.

## Worth recording

`gh release edit <tag> --notes-file <f>` updates a published release
body even though `action-gh-release` cannot. The immutability that
permanently burns a tag name does not extend to the description, so a
thin release body is recoverable. That was not obvious from the existing
comments, which describe releases as immutable without qualifying which
parts.

## Verification, and its limit

`yaml.safe_load` parses the file, but that proves little: it silently
accepts duplicate mapping keys that GitHub rejects outright, which is
how a previous workflow edit passed a local check and produced no runs
at all. The real validator is this PR's own pipeline, and the change
cannot be exercised end to end until the next tag is pushed.

## Not done

No backfill for v3.10.1 or v3.11.0. Neither has a notes file, and
reconstructing them is the same retrospective work deliberately skipped
for the 3.11.0 changelog entry.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-26 18:45:50 +02:00
efitenandClaude Opus 5 6d5686dbb4 docs: release notes for v3.12.0, and the 3.11.0 entry that was never written (#2075)
Prepares the v3.12.0 release. No tag is pushed by this PR: the release
procedure requires the tagged commit to carry an `:edge` image whose
revision label matches it, so that check belongs after this merges,
against the merge commit.

## Why minor, not patch

29 commits since v3.11.0: 15 fix, 5 test, 3 perf, **3 feat**, 2 ci, 1
chore. The feats are #2047, #2067 and #2068.

## What the notes lead with

Two changes alter what a running instance does without anyone asking it
to, so they are at the top rather than in a list:

- **#2058**: the first start against a database that has never been
`ANALYZE`d builds planner statistics and stalls ingest while it does.
Measured at 3m43.9s on 9.4 GB, once per database, buffered with nothing
dropped. `db.analysisLimit` set negative skips it.
- **#2035**: `maxMemoryMB` eviction now fires where it previously did
not, because the footprint compared against the limit was undercounted.
An instance that set the limit and never saw eviction will start seeing
it.

## The 3.11.0 gap

`CHANGELOG.md` had no entry for 3.11.0 and `[Unreleased]` was empty, so
52 shipped commits were undocumented. Added as a short entry that says
outright it was written after the fact and has no notes file. The
alternative of reconstructing 52 commits for a superseded version is
error-prone work with little value, and leaving the gap silent is worse
than naming it.

## One fix beyond documentation

`deploy.yml` carried a comment that would mislead the next person
cutting a release. It still said documentation-only commits skip the
workflow "(see the `paths-ignore` above)", which is exactly how v3.10.0
lost its `:edge` image and then its tag name permanently. That filter no
longer exists: the `changes` job forces `code=true` for anything that is
not a pull request (lines 61-73), so every master commit gets an image
and a documentation commit is safe to tag. The history stays in the
comment; the false present tense does not.

## Verified rather than asserted

Both went into the notes as upgrade advice, so both were checked in the
tree:

- `node_declared_regions` is created at boot with `CREATE TABLE IF NOT
EXISTS` (`cmd/ingestor/db.go:435`), added by #2047, which found the
table was read by `region_keys.go` and `config.go` and created by
nothing. So "no manual migration step" is accurate.
- The cgo dependency attributed to #1992 in the 3.11.0 entry is the one
that stops this repo building with `CGO_ENABLED=0` today.

## Not done

- No tag, no release. Next steps, after this merges: confirm the merge
commit's pipeline is green, confirm `:edge`'s revision label is that
commit, then tag `v3.12.0` annotated, push the tag only, dispatch `CI/CD
Pipeline` on the tag ref, and verify the published digests in the
registry rather than in a green job.
- No `docs/release-notes/v3.11.0.md`, deliberately, as above.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
v3.12.0
2026-09-26 12:03:31 +02:00
efitenandClaude Opus 5 6e121ff8e6 perf(db): build planner statistics at startup when there are none (#2058) (#2074)
Follow-up to #2072, which closed #2058 but left one gap named in its own
description: the refresh ticker waits 2 minutes before its first run,
and a query arriving in that window against a database with no
statistics gets the bad plan.

Deployed to staging to measure it rather than reason about it, with
`sqlite_stat1` dropped first so the build path actually ran. That
changed two of the numbers in #2072, both in the expensive direction.

## The gap is once per database, not once per restart

`sqlite_stat1` is an ordinary table, so once `ANALYZE` has written it
the statistics stay in the file. Checked four ways:

- they survive closing the connection that wrote them
- a `mode=ro` handle reads them back, which is how `cmd/server` opens
the database
- reopening the same path through a second `OpenStore` finds them and
skips the rebuild (`TestPlannerStatsSurviveReopen_Issue2058`)
- on staging they survived a full redeploy to a different build that has
no refresh ticker at all, and that build still gets the good plan

So the window opens once, on the first start after this lands, and never
again for that database.

## The cost, corrected

#2072 said 2.0s. Observed on staging, 9.4 GB, commit `4500cfa6`:

```
13:51:26  [analyze] planner statistics refresh scheduled every 24h (analysis_limit=10000)
13:55:10  [analyze] planner statistics built in 3m43.874s (analysis_limit=10000, first run against this database)
```

**3m43.9s.** Every `ANALYZE` duration in #2072's ladder was timed warm,
run after run; cold, on a freshly started container, the same statement
takes nearly four minutes. That is the same warm/cold split #2058 work
already established for the query itself, 56.7s against 0.80s, and I
then repeated it for the `ANALYZE`. Every `2.0s` in the tree is now
marked warm and points at the cold figure.

It holds the single write connection throughout, so ingest stalls and
buffers. Per minute in `observations`:

| minute | rows |
|---|---|
| 13:49 | 220 |
| 13:50 | 106 |
| 13:51 | 0 |
| 13:52 | 0 |
| 13:53 | 0 |
| 13:54 | 0 |
| 13:55 | **1027** |
| 13:56 | 154 |

Nothing was dropped. The burst is about four minutes of traffic at the
surrounding rate, and the only ingest-buffer line in the log is the
startup one reporting `0 dropped`. The cost is a four-minute write
stall, once, not data loss.

## This cost is not introduced here

The ticker merged in #2072 pays the identical 3m43.9s two minutes later
on any database with no statistics. **Live has none, so #2072 as merged
will stall live ingest for about four minutes on its first run, with or
without this branch.** This only moves it earlier, into the startup
burst the ingest buffer is already sized for. Flagging it on #2072 as
well.

## The change

`Store.EnsurePlannerStats(analysisLimit)` checks before it builds:

- database has statistics: one `sqlite_master` query. This is every
restart after the first.
- database has none: one `ANALYZE`, and a warning first.

The warning is the part that earns its place operationally. Four minutes
of stalled ingest with no explanation in the log looks exactly like a
hang, so `EnsurePlannerStats` now says why the write path is about to
pause, what it measured on 9.4 GB, and that it happens once per
database. It stays silent on a restart, because a warning on every boot
would be worse than none.

It runs on the refresh goroutine, not the startup path, so no boot step
waits for it.

`hasPlannerStats` now gates a decision instead of only wording a log
line, so its comment says what the swallowed error costs: a query
failure reads as "no stats", which spends one unnecessary `ANALYZE`
rather than skipping a necessary one.

## Verification on staging

- before: plan drove from `idx_transmissions_payload_type`, no
`sqlite_stat1`
- after: 50 rows in `sqlite_stat1`, plan drives from
`idx_tx_channel_hash`
- dropping the table first flipped the plan back, so the causality holds
in both directions

## Tests

13 in the file. New here: builds when absent, skips when present,
disabled on a negative limit, survives close-and-reopen, warns before
building, stays quiet when statistics exist.

The reopen test is the guard on the whole design: if statistics ever
stopped living in the file, `EnsurePlannerStats` would quietly run a
four-minute `ANALYZE` on every restart and nothing else would notice.

Run locally: 13/13 on the `Issue2058` tests, `go vet` clean, `gofmt`
clean, and the rest of `cmd/ingestor` green apart from
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on
`os.Symlink` with "A required privilege is not held by the client" on
Windows without elevation, in a file this branch does not touch.

## Not done

- No query rewrite, same as #2058 and #2072.
- **No live deploy.** Live still has no statistics, so the four-minute
stall is ahead of it whenever #2072 ships there. Worth picking the
moment.
- Staging has been returned to its own fork build; the statistics it
built remain, so its next start exercises the skip path rather than the
build path.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-26 10:08:43 +02:00
efitenandClaude Opus 5 a5aa3cdc41 perf(db): refresh SQLite planner statistics with a bounded ANALYZE (#2058) (#2072)
Closes #2058.

`ANALYZE` has never run against these databases, so `sqlite_stat1` does
not exist and the planner works from built-in guesses. On the channel
queries it guesses wrong: it drives from the plain
`idx_transmissions_payload_type` instead of `idx_tx_channel_hash`, the
partial index (`WHERE payload_type = 5`) the schema already carries for
that exact filter.

@anieto's report did the diagnosis and the arithmetic. This adds the
maintenance operation that was missing, at a value measured rather than
assumed.

## The diagnosis transfers, the remedy needed measuring

Measured on our 9.4 GB staging database: 1,250,489 transmissions,
14,169,329 observations, 2.7x and 7x the reported database.
Region-filtered `GetChannels` produces the identical plan reported in
#2058, down to both temp b-trees, so the problem is the same one.

Wall time is the wrong metric here. The same query and the same plan
measure **56.7s cold and 0.80s warm** on that file, so the OS page cache
dominates. Counting page-cache misses instead:

| analysis_limit | ANALYZE | driving index | page misses |
|---|---|---|---|
| none (no statistics) | - | `idx_transmissions_payload_type` | 143,442
|
| 400 | 171 ms | `idx_transmissions_payload_type` | 143,449 |
| 1000 | 171 ms | `idx_transmissions_payload_type` | 143,450 |
| **10000** | **2.0 s** | **`idx_tx_channel_hash`** | **107,429** |
| 0 (unbounded) | 242.9 s | `idx_tx_channel_hash` | 107,429 |

400, the value SQLite's documentation offers for the bounded form,
changes nothing on this data: it samples too few rows to separate the
126,336-row partial index from the 920,700-row plain one. 10000 buys the
entire plan change for 2.0 s, and the four-minute unbounded `ANALYZE`
buys nothing beyond it.

## What it is worth, as measured

25% fewer pages read per query, 143,442 to 107,429, about 147 MB less at
a 4 KB page. Warm wall time does not move: 0.80s either way. The gain
lands on the cold path, the one that measured 56.7s, so the claim here
is fewer pages read, not a warm speedup.

This is smaller and differently shaped than the 3-4x in #2058. I cannot
reproduce that ratio on a database of this size and am not claiming it.

## The change

- `Store.RefreshPlannerStats(analysisLimit)` in `cmd/ingestor/db.go`:
`PRAGMA analysis_limit=N` then `ANALYZE`, logging the duration and
whether this was the first run.
- Wired in `cmd/ingestor/main.go` next to the existing WAL checkpoint
ticker: 24h, staggered 2 minutes past startup because it takes the write
lock.
- `db.analysisLimit` in `internal/dbconfig`, default 10000, negative
disables it.

It runs in the ingestor, not the server: `cmd/server/db.go:145` opens
`mode=ro`, and `ANALYZE` writes. This respects the read/write separation
invariant in AGENTS.md.

`analysis_limit=0` means *no* limit to SQLite rather than "use a
default", so an unset config maps to 10000 and a test covers that
specific case.

## Two faults the measurement caught in my own first commit

Both are in the history rather than hidden, because the second commit is
the one that measured:

1. **`PRAGMA optimize` was the wrong statement.** It analyzes only
tables the calling connection has itself queried during the session, and
a maintenance call has queried none. Run against staging it wrote
nothing and left `sqlite_stat1` absent; `PRAGMA optimize(0x03)` returned
no statements at all. Verified on an empty database too (SQLite 3.45.1):
`ANALYZE` creates `sqlite_stat1`, `PRAGMA optimize` does not. That
difference is what makes the behavioural test a guard instead of a
no-op.
2. **`analysis_limit=400` was the wrong value**, per the table above.

## Tests

Six cases in `cmd/ingestor/refresh_planner_stats_test.go`:

- statistics are actually written (the guard against returning to
`PRAGMA optimize`)
- the pragma reaches the connection, read back through `PRAGMA
analysis_limit`
- a negative limit leaves `sqlite_stat1` absent
- two consecutive refreshes, since a ticker calls this repeatedly
- the config default and the JSON round trip
- the default is above the range measured ineffective, so lowering it
back to 400 fails

## Not done, and one caveat

- **Correction to an earlier version of this description**, which said
the Go tests could not run locally because this box has no C compiler.
That was wrong: `CGO_ENABLED=0` and `gcc` being absent from `PATH` is
not the same as no compiler, and a mingw-w64 toolchain is installed
here. Run properly, all six tests pass locally, and so does the rest of
`cmd/ingestor` apart from
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on
`os.Symlink` with "A required privilege is not held by the client" on
Windows without elevation and lives in `stats_file_test.go`, a file this
branch does not touch. CI agrees: Go Build & Test and the ingestor race
detector are both green.
- The SQLite behaviours above were measured against 3.45.1 on the
server, not against the amalgamation `mattn/go-sqlite3` bundles.
- **No query rewrite.** #2058 explicitly left that out and so does this;
the correctness caveats it lists (per-channel most-recent-message
semantics, v2/v3 branches, `enc_` exclusion) are untouched here.
- Staging carries limit-10000 statistics, matching what this code
produces. Reversible: `DROP TABLE sqlite_stat1` was verified on a
scratch database before any of it ran.
- Whether a cold-start `ANALYZE` should also run before the 2 minute
stagger is not addressed. The first query after a restart is the
expensive one, and it can arrive first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:08:54 +02:00
efitenandClaude Opus 5 3caa847323 fix(nodes): the advert section is called "Recent Adverts", and says why (#2071)
Closes #2042, reported by @damn-simple-scripts.

## Which of the two fixes

The report offered widening the section to all packets the node
originated, or renaming it, and proposed the rename. Widening sounds
like the better fix, so I measured before agreeing. On a production
database:

| | |
|---|---|
| transmissions total | 1,251,967 |
| with `from_pubkey` populated | 191,237 |
| of those, `payload_type = 4` (ADVERT) | **191,237** |
| with `from_pubkey` and **not** an advert | **0** |

The ingestor fills that column for adverts only. Attributing a relayed
CHAN or TXT packet back to its sender is the path-resolution problem,
not a filter this section could apply — so "show all packets from this
node" is not a small change, it is a different feature resting on
attribution the data does not carry.

So the rename is correct, and the numbers say so rather than my
preference.

## The change

Both copies renamed — the full node page and the side pane. Both already
read `nodeData.recentAdverts`, so the field feeding them said "adverts"
while the heading said "packets".

The heading also gained a `title` naming `from_pubkey` as the reason it
is adverts only. Renaming without explaining invites the same report
from the next reader; the tooltip is where that explanation costs
nothing.

## Tests

`tests/unit/test-issue-2042-recent-adverts-label.js`, four cases: both
copies present and headed "Recent Adverts", no copy back to "Recent
Packets", the advert field still feeding them, and the explanation still
in place. The third matters most — it ties the label to its data source,
so pointing this section at a different field in future fails here
rather than silently making the label wrong again.

## Not done, deliberately

The reporter's follow-up idea, splitting flood adverts from zero-hop
adverts into two lists. They wrote that it can be dropped if an issue
should tackle one thing, so it is not here. It is feasible now: a
zero-hop advert is `ROUTE_TYPE_DIRECT` with an empty path, established
while fixing #2064. That deserves its own issue rather than a paragraph
in this one — say the word and I will open it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:08:28 +02:00
efitenandClaude Opus 5 2cfe9cbbc5 fix(a11y): raise the Scope Audit observed-chip contrast above 4.5:1 (#2070)
Closes #1996.

## Reproduced first

The chips tint their own background — `color-mix(in srgb,
var(--status-green) 16%, transparent)` — which darkens whatever surface
sits behind them. With `--status-green-text` (green-700, `#15803d`) on
top:

| surface | composited chip | ratio |
|---|---|---|
| `--surface-0` `#f4f5f7` | `#d2eddf` | **4.04:1** |
| `--card-bg` `#ffffff` | `#dcf6e5` | **4.38:1** |

Both under the 4.5:1 the #1719 gate requires for normal text. The first
row is the same ratio **and the same hex** @n30nex measured in the
browser, which is how I know the model in the test agrees with what a
visitor actually sees.

## Fixed on the text, not by thinning the tint

Dropping the tint to 12% reaches 4.52:1. That is two hundredths above
the line, and a margin that thin fails again the next time a surface
value moves — the same trap #2039 hit with a threshold sitting inside
the healthy band. One palette step darker gives **6.23:1** on white and
**5.74:1** on `--surface-0`, and keeps the tint that makes a chip read
as a chip rather than as plain text.

- `--palette-green-800: #166534` added. The greens ran 300–700 while the
blues already reach 900, so this fills the scale rather than inventing a
colour.
- `--sa-chip-observed-fg` defined per theme: green-800 in light, and the
bright `#22c55e` **kept** in dark, where the chip already passed at
6.48:1 / 5.73:1. Dark is deliberately untouched.
- The chip reads the variable, so the customizer still governs it and no
literal enters a component.

`.sa-chip-verified` needed no change: `scope-audit.js:106` only ever
adds it alongside `sa-chip-observed` or `sa-chip-unobserved`, and it
contributes an underline. So the fix covers both classes the issue names
— worth stating, since the title mentions both.

## Tests

`tests/unit/test-a11y-1996-scope-audit-chips.js`, 7 cases: both themes ×
both surfaces, that the chip still tints (so the suite cannot pass by
testing nothing), that its colour comes from a variable rather than a
literal, and a guard asserting green-700 **would** still fail — so a
quiet revert to `--status-green-text` turns this red instead of passing.

**Red-run confirmed:** with the old colour restored it fails at exactly
4.04:1 and 4.38:1, naming the composited `rgb(210,237,223)`.

Two deliberate choices in the test:

- **A separate suite, not a case in `test-a11y-1719`.** That suite's
`parseColor` handles hex and `rgb()` only; teaching it `color-mix()` is
a larger change than this fix. These chips are the only
contrast-critical user of the function today. If a second appears, the
two should merge, and the file says so.
- **Derived from the stylesheets, not a rendered page.** The default
Scope Audit fixture renders no chips at all, which is precisely why the
existing browser coverage missed this. A stylesheet-derived check cannot
be defeated by a fixture that shows nothing.

`check-css-vars` passes (180 definitions, 0 undefined) and
`test-test-inventory` passes with the new file classified.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:07:57 +02:00
d14317b008 fix(qa): query blacklist retention by from_pubkey (#2069)
## Summary

- query `transmissions.from_pubkey`, the column present in CoreScope's
real schema, instead of the nonexistent `from_node`
- preserve the existing SQLite parameter binding and stdin transport
- derive the unit-test table from the ingestor's committed `CREATE
TABLE` definition rather than a hand-written schema
- add positive, negative, case-normalization, injection-shaped input and
missing-column coverage
- verify the query against the committed staging-captured E2E fixture

## Why

The parameter-binding change correctly removed SQL interpolation, but
its count query and test fixture both used `from_node`. The production
schema defines `transmissions.from_pubkey`; therefore the live QA probe
could only fail with `no such column: from_node`, while the synthetic
unit fixture remained green.

The ingestor stores attributed ADVERT pubkeys as lowercase hex. The
query now uses:

```sql
SELECT COUNT(*)
FROM transmissions
WHERE from_pubkey = lower(:pubkey);
```

The value remains a bound parameter. No production schema or runtime
code changes.

## Verification

- `bash -n qa/scripts/blacklist-test.sh`
- `bash -n qa/scripts/test-blacklist-sql.sh`
- `bash qa/scripts/test-blacklist-sql.sh` — 72 passed, 0 failed
- schema fixture extracted directly from `cmd/ingestor/db.go`
- committed E2E fixture returns the same attributed-row count through
the helper query and a direct control query
- legacy fixture containing only `from_node` fails non-zero with `no
such column: from_pubkey`
- injection-shaped, empty, whitespace, multibyte and long values remain
literal bound values
- `git diff --check`

This PR intentionally contains only the schema correction and its
regression coverage.

Co-authored-by: Openclaw <openclaw@Openclaws-Mac-mini.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:32:54 +02:00
435ac25dd6 feat(nav): show the running version in the nav-drawer footer (#2068)
Takes over #1985 by @SaarMesh-Bot, as offered there on 2026-09-16 and
2026-09-17. **Their commit is the first here, unchanged and under their
authorship**; the rest clears the two review points. Closing #1985 in
favour of this so the rebase and the fixes travel together, not to
reassign the work.

## The feature, unchanged

The frontend never surfaced which build was running, though
`/api/health` has reported `{version, commit, buildTime}` all along. A
footer on the nav drawer now renders `CoreScope <version>`, linking to
the releases page, with commit and build time in the tooltip. Colours
come from existing CSS variables, the label is set with `textContent`,
and a failed health call leaves a neutral label rather than an empty
footer — all as the author wrote it.

## Review point 1: fetch on open, not on page load

The `/api/health` call sat in `buildDom()`, which runs on page load. The
drawer may never be opened, and **cannot** be opened at ≤768px, where
the module is disabled by design. So every visitor's browser was
requesting the endpoint to fill a footer most of them would never see.

Moved into `open()`, after the width gate. `fetchVersion` still caches
its promise for the page lifetime, so re-opening costs nothing.

## Review point 2: the fallback is pinned

`tests/unit/test-nav-drawer-version-footer.js`, six cases:

- a rejected fetch, a non-ok response, and a 200 without `version` each
leave the neutral `CoreScope` label — not a blank footer and not
`CoreScope undefined`, which is what an instance shows exactly when
someone is trying to read its version
- a normal response renders the version, with commit and build time in
the tooltip
- a version carrying markup lands verbatim in `textContent`, and
`innerHTML` is never touched
- the endpoint is requested once however often the footer is filled,
pinning the cache from the other side

It **slices `fetchVersion` and `fillVersion` out of the shipped
`public/nav-drawer.js`** and evaluates them rather than copying them
into the test, so it exercises what ships — same approach as
`tests/unit/test-direct-rf-heard-by.js`. Both slice markers are
asserted, so a rename fails loudly instead of quietly testing nothing.
Each case re-evaluates the slice, because `versionPromise` caches for
the page lifetime and a shared sandbox would hand the second case the
first case's answer.

Listed in `test-all.sh`, which per #2036 is the only frontend runner.

## A note on the third commit

The first version of the test used `setImmediate` to settle the promise
queue. It ran fine under node and **failed eslint**, which treats these
as browser code. I ran eslint and committed in the same command and
pushed without reading its output. Fixed in the commit after, using
`setTimeout(r, 0)`. Recording it because the PR would otherwise show a
lint failure in its history with no explanation.

## Verification

6 of 6 in the new suite, eslint clean on both changed files,
`test-test-inventory.js` passes with the new file classified.
Cherry-picked cleanly onto current master.

---------

Co-authored-by: SaarMesh-Bot <bot@saarmesh.de>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 10:58:36 +02:00
291393dcc0 feat(ingestor): log a throttled warning when the IATA whitelist drops a region (#2067)
Takes over #2008 by @nullrouten0, as offered there on 2026-09-16 and
2026-09-17. **Their commit is the first of the two here, unchanged and
under their authorship**; the second is only the fix for the one
blocker. Closing #2008 in favour of this so the rebase and the fix
travel together, not to reassign the work.

## The feature, unchanged

`observerIATAWhitelist` dropped non-whitelisted regions silently. An
allow-list fails in the dangerous direction: a legitimate but unlisted
region vanishes with nothing to show for it. One line per dropped region
now, re-logged at most every `iataWarnIntervalSec` (new optional key,
default 6h) for as long as that region keeps arriving.

The periodic re-log rather than a strict log-once is the author's call
and it is the right one: a single edge event rolls out of any scrape
window, leaving an actively-dropping region indistinguishable from a
healthy one.

## The blocker, now fixed

`ShouldWarnIATADrop` keyed its throttle map on a topic segment the
**publisher** controls and never evicted it — a remote memory sink.
Measured on the original branch: 200,000 distinct codes retained 200,000
entries and 15.1 MB of heap.

My review offered two shapes. This takes the cap rather than
shape-validation, and the reason matters: **nothing in this codebase
constrains an IATA code's shape.** It is uppercased and trimmed in
`config.go` and `db.go` and never validated. Rejecting by shape would
invent a rule operators have not agreed to, and would silently drop the
warning for anyone whose code does not fit it — the same failure mode,
one level down.

So `iataWarnMaxTracked = 512`: far above any real deployment (the
reference instance runs 43 observers across a handful of regions) and
small enough that a hostile feed gains nothing.

**Past the cap the drop is still logged**, throttled on one shared
timestamp instead of a per-code one. Swallowing it there would
reintroduce exactly the silent failure this feature exists to fix.

## Tests

The author's `iata_drop_warn_test.go` plus three:

- the map stops growing when fed 2048 distinct codes
- a new code past the cap still warns once, is then throttled, and
speaks again after the interval elapses
- an already-tracked code's throttling is unchanged, so the cap does not
alter the normal path

## Verification

`gofmt` clean, cherry-picked cleanly onto current master
(`cmd/ingestor/main.go` auto-merged). Go tests not run locally: no cgo
toolchain here since #1992, and per AGENTS.md `CGO_ENABLED=0` links a
stub that proves nothing. CI is their first run.

## Not done

The `iataWarnIntervalSec` key is undocumented outside the struct
comment. If there is a config reference that should list it, say where
and I will add it.

---------

Co-authored-by: nullrouten <nullrouten@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:58:27 +02:00
efitenandClaude Opus 5 50c4d9615b test(ingestor): wait for the boot migrations before handing over a test store (#2066)
Closes #2065. Master's `🏁 Race detector (ingestor)` job has been red
since the pushes at 2026-09-22 21:55 and 21:56.

## Correcting my own diagnosis

The issue says the fault is a goroutine outliving its test and racing a
later one, and proposes making it joinable. That was wrong, and it
matters because it changes the fix: `Close()` **already** waits on
`backfillWg` (`cmd/ingestor/db.go`), and `newTestStore` registers it as
`t.Cleanup`. The goroutines are joined before the next test starts.

The race is inside a single test.

`OpenStore` schedules two async migrations — `obs_observer_ts_idx_v1`
and `tx_last_seen_backfill_v1` — whose goroutines log while they run.
`TestHandleMessageDecodeErrorLog_PII_Issue1211` then points the standard
logger at a `bytes.Buffer` and reads it, so its **own** store's
migrations write into the buffer it reads:

```
Write by goroutine 760:  RunAsyncMigration.func1   async_migration.go:124  (log.Printf)
Read  by goroutine 757:  ...PII_Issue1211          decode_error_log_test.go:37 (buf.String)
```

## Why the helper rather than the one test

These tests capture the standard logger in **21 places across 7 files**.
Any of them that also builds a store is exposed to the same thing; the
decode-error test is just the one whose timing lost. So `newTestStore`
now waits after `OpenStore` instead of only at cleanup, and no test body
can run while a migration is in flight.

Checked before touching a shared helper: no test references either boot
migration by name, and the `pending_async` assertions in
`async_migration_test.go` use their own names with a blocking `fn`, so
they are unaffected. Cost is a few milliseconds against an empty temp
database.

## Tests

The race detector only catches this when the scheduler cooperates — it
sat latent from 2026-09-03, when those files were last touched, until it
surfaced three weeks later, and a re-run would have made it look like a
flake. So both new tests are deterministic:

- **`TestNewTestStoreWaitsForBootMigrations`** —
`tx_last_seen_backfill_v1` is scheduled unconditionally by `OpenStore`,
so on a fresh temp database it is pending at that instant and can only
read `done` if something waited. Remove the wait and this fails every
run.
- **`TestCapturedLogIsFreeOfMigrationOutput`** — asserts a captured
buffer holds no `[migration/async]` or `[async-migration]` output, which
is the failing test's own situation stated as an assertion.

## Verification

`gofmt` clean. Go tests were not run locally: no cgo toolchain on this
machine since #1992, and per AGENTS.md `CGO_ENABLED=0` links a stub that
proves nothing. CI is their first run, and the race-detector job is the
one that matters here.

## Not done

The decode-error path still logs through the standard logger, so a
future test capturing it while any other goroutine logs will race again.
An injectable logger would close that class properly. This closes the
store-boot case, which is the one that exists today.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:23:59 +02:00
6d3da77b67 perf(channels): coalesce concurrent GetChannels/GetEncryptedChannels cache misses (#2059)
Fixes #2029.

## What was wrong

`GetChannels`/`GetEncryptedChannels` (`cmd/server/db.go`) cache their
region-scoped result for 60s but had no request coalescing on a cache
miss, so every request that arrived while the cache was cold or expired
ran the region-scoped `GROUP BY` scan itself. Measured on production: 5
concurrent requests for the same never-cached region each took ~8s, no
cheaper than 5 independent runs.

`statsSF`/`regionMembershipSF` already fix the identical bug class
elsewhere in this file (#1910), so this wraps both functions'
query-build/execute/cache-populate block in a `singleflight.Group` the
same way, keyed per region, double-checking the cache inside the flight
in case a previous winner already refreshed it.

## Tests

The review on #2029 pointed out that a timing-based check ("finish
within ~2ms of each other") doesn't actually prove coalescing happened —
it would pass on a fast machine even without singleflight.
`db_channels_singleflight_test.go` uses a call counter instead, same
pattern as `TestEnsureNeighborGraph_Singleflight` (#1203 Pair A):

- `TestGetChannels_SingleflightCoalescesQueries` /
`TestGetEncryptedChannels_SingleflightCoalescesQueries`: 10 concurrent
callers against a cold cache, asserting the real query runs exactly
once. A test-only hook (`channelsQueryHook`/`encChannelsQueryHook`, nil
in production, same contract as `bgLoaderEntryHook`) increments the
counter right where the query executes, since these functions hit
`db.conn.Query` directly rather than going through an injectable builder
function.
- `TestGetChannels_SingleflightPerRegion`: two regions queried
concurrently (5 callers each) assert 2 queries, not 1 — pins that the
flight is keyed per-region and a caller for one region can't receive
another region's coalesced result.

Anti-tautology: reverting `channelsSF.Do`/`encChannelsSF.Do` back to a
bare call makes the coalescing tests observe N instead of 1.

`go build ./...`, `go vet ./...`, `gofmt -l .` clean. Full `cmd/server`
suite (race-enabled for the new concurrency tests) run in a
`golang:1.22-alpine` container, mounted repo, workdir `cmd/server` so
the sibling `internal/*` replace directives resolve:

```
=== RUN   TestGetChannels_SingleflightCoalescesQueries
--- PASS: TestGetChannels_SingleflightCoalescesQueries (0.06s)
=== RUN   TestGetChannels_SingleflightPerRegion
--- PASS: TestGetChannels_SingleflightPerRegion (0.05s)
=== RUN   TestGetEncryptedChannels_SingleflightCoalescesQueries
--- PASS: TestGetEncryptedChannels_SingleflightCoalescesQueries (0.07s)
```
Full suite: `FAIL github.com/corescope/server 98.261s`, but the only
failures are `TestHandleNodePaths_PrefixCollision_1352`,
`TestHandleNodePaths_FallbackUniquePrefix_1352`, and
`TestHandleNodePaths_FallbackUnresolvableHop_1352`, all failing on a
`503 {"error":"index loading","retryAfter":5}` — an index-build race in
this container's timing, not this change. Confirmed by running the same
three against an unmodified, freshly-cloned `master` in the same
container: they fail there too (plus
`TestHandleNodePaths_PrefixCollision_1352_FallbackBranch`, which this
run happened not to hit). Nothing in this diff touches node-path
handling.

## Not done

The deeper query-plan issue flagged in #2029 (the outer scan is driven
by `payload_type`, not region, so a cold solo request still costs
several seconds regardless of concurrency) is filed separately as #2058,
with `EXPLAIN QUERY PLAN` output and row counts against production data.
Coalescing makes one slow query serve everybody; it doesn't make the
query itself fast.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: anieto <anieto@meshtexas.org>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 10:05:59 +02:00
liquidraverandClaude Opus 5 b695b979a2 fix(nodes): keep paginating past a page that post-LIMIT filtering shortened (#2061)
## Problem

`handleNodes` runs the geo-filter, `nodeBlacklist`, `hiddenNamePrefixes`
and area
passes **after** the SQL `LIMIT/OFFSET`, and rewrites `total` to the
filtered
length. A page that loses a row is therefore short **without being the
last
page**, and neither the page length nor `total` can tell a client
whether to ask
for another page.

#1606 added the pagination loop and chose the page length as the
canonical stop.
That is correct only where nothing is ever filtered. Everywhere else the
list
truncates at the first filtered page boundary and strands every node
behind it —
the #1598 symptom reached by a different route: a node that is relaying
right
now simply stops being in the list.

The comment at `app.js:240` rejects `total` for exactly the right
reason, then
picks the signal the same code path also breaks.

## Measured on a live 2346-node deployment

Page sizes for the query the map issues:

```
offset=0     returned=500     ← full, loop continues
offset=500   returned=499     ← one row filtered AFTER the LIMIT → loop STOPS
offset=1000  returned=500     ← never requested
offset=1500  returned=500     ← never requested
offset=2000  returned=345     ← never requested
```

| stop rule | requests | nodes reached |
|---|---:|---:|
| short page (master) | 2 | **999** |
| `has_more`, else empty page | 6 | **2344** |

**1341 nodes, 57%, unreachable through the UI.**

### One hidden node truncates the whole list

The deployment this came from has no `geoFilter`
(`/api/config/geo-filter`
returns `polygon: null`) and no `nodeBlacklist`. It has a single
`hiddenNamePrefixes` entry — a deliberate operator choice — and exactly
one node
whose name starts with it:

```
public_key   d4a46ea2…1054      (64 clean hex chars)
name         🚫🔥☀️
role         repeater
last_seen    2026-09-22T10:09:38Z
```

`handleNodes` drops that row in the `IsNameHidden` pass, which runs
after the SQL
`LIMIT`. The row is counted by the `LIMIT` and by `COUNT(*)`, so the
page it lands
in comes back exactly one short — and stops every client that treats a
short page
as the end.

Isolated against SQL on the same database, seconds apart:

```
SELECT lower(public_key) FROM nodes ORDER BY last_seen DESC LIMIT 500 OFFSET 500
  -> 500 rows
GET /api/nodes?limit=500&offset=500
  -> 499 rows
comm -23 sql.txt api.txt
  -> d4a46ea2e99cab132a3286ef3d9cce9099318790af7f25671fe83de453721054
```

Deterministic — `offset=500` returned 499 on three consecutive requests.
Not a
CDN artifact either: `cf-cache-status: DYNAMIC`, origin `cache-control:
no-store`, no `age` header, and four requests with deliberately unique
cache keys
all returned 499.

So **one deliberately hidden node makes 1341 of 2344 nodes
unreachable.** The
hiding feature does exactly what it was asked to do for that one node,
and takes
57% of the network with it, silently. A single `hiddenNamePrefixes`
entry is
enough; no geo-filter, blacklist or area filter is needed to reach this
state.

### The cutoff moves, which is why this reads as intermittent

The visible set is the sum of the pages up to and including the first
short one,
so the boundary sits wherever the unreturnable row currently sorts by
`last_seen`, and jumps a whole page as ingest reorders the list. Same
deployment, same code, same config, ~2h apart:

| dropped row's rank | first short page | nodes visible |
|---|---|---:|
| inside 0–499 | page 1 | 499 |
| inside 500–999 | page 2 | 999 |

A node is visible or invisible purely by where it lands relative to that
moving
line, so affected nodes appear to vanish and return on their own. Two
operators
on this deployment reported exactly that, independently, while I was
measuring.

### A named reproduction

`HU-ZA-Lentihegy` (`5287a33f…`), reported missing from the map by an
operator
whose companion had logged its advert at 04:20 local the same morning.

Ingest was fine. The row is in `nodes` with `last_seen`
`2026-09-22T02:20:35Z` — the same advert, to the second — valid GPS,
role
`repeater`, 1033 adverts, and `/api/nodes/search?q=lentihegy` returns
it.

```
rank by last_seen : 1081
cutoff at the time:  999
```

It missed by 82 positions. Walking the same live endpoint, same moment:

| stop rule | requests | nodes reached | Lentihegy |
|---|---:|---:|---|
| short page (master) | 2 | 999 | **not reached** |
| `has_more`, else empty page | 6 | 2340 | reached |

The practical shape of this on a busy mesh: 1081 nodes had been heard
more
recently than 9.4 hours, so on that deployment **anything last heard
more than
~9 hours ago was invisible**, alive or not.

`#/nodes` compounds it — its search box filters client-side over the
truncated
set, so the server-side `?search=` never runs and an operator cannot
find the
node by searching for it either, even though the endpoint would return
it.

## Change

**Server** — `NodeListResponse` gains `has_more`, computed from the raw
SQL page
against the real `COUNT(*)` before the filter passes run, so it survives
them:

```go
hasMore := offset+len(nodes) < total
```

Always emitted (no `omitempty`) so a client can tell `false` from an old
server.
No extra request in the fixed path: `has_more` ends the loop exactly,
where the
old rule needed a probe page.

**Clients** — `app.js` `fetchAllNodes`, `nodes.js` `loadNodes` and
`area-map.html`'s inline helper stop on `has_more`, falling back to a
zero-length
page against a server that predates it. An empty page always ends the
loop, so a
`has_more` against a concurrently-shrinking table cannot spin to
`safetyCap`.

Left alone: the three loops are still three copies. Collapsing them onto
`fetchAllNodes` is a bigger change than this fix needs, and `nodes.js`
has its
own inter-page progress UI. Happy to do it separately if you want it.

## Testing

- **Unit** (`tests/unit/test-fetch-all-nodes-pagination.js`): the
fixture now
models the real handler — a row counted by the LIMIT and by `COUNT(*)`,
then
removed from the page. Three new cases. Fails on the old rule at 499 of
1199.
- **E2E** (`tests/e2e/test-map-nodes-pagination-e2e.js`, already wired
into
  `deploy.yml`): the mock drops a page-1 row and emits `has_more`.
  Mutation-checked — restoring master's stop rule fails 3 of its steps.
- **Go** (`cmd/server/nodes_pagination_has_more_test.go`): asserts
`has_more`
stays true on a page filtering shortened. Mutation-checked — recomputing
it
  after the filter block fails the test.
- Full server suite `go test -race`: ok, 41.4s. `gofmt` clean, `go vet`
passes.
- **Against a real binary**, not just mocks: fixture DB migrated with
`corescope-migrate`, `hiddenNamePrefixes: ["SKCE"]`, `limit=3`. Page 1
returns
2 of 3 with `total` rewritten to 2 and `has_more=true`. Walking the real
server
with master's rule reaches 2 nodes; with `has_more`, all 199 visible of
200,
the hidden one still hidden. The real frontend against that server loads
199
  with no JS errors.

Two existing expectations changed, both deliberate:

1. `surfaces ALL nodes past the 500 server cap` — 3 → 4 requests. That
mock emits
no `has_more`, so the 200-row final page can no longer end the loop (a
short
page is exactly what a filtered page looks like) and a zero-length probe
   follows. Against a current server `has_more` still ends it at 3.
2. `rows missing public_key are NOT collapsed into one` — its stub
returned a
constant body, which would now be paged to `safetyCap`. It serves one
page
   then empties.

Local `test-all.sh` exits 1 on two XSS-gate self-tests
(`good-2-tested.js`, `good-4-tested.js`) via a `UnicodeEncodeError`
printing an
emoji under Windows cp1252. Identical on clean `origin/master` in a
scratch
worktree, so it is pre-existing and platform-local, not this branch.

There is a second identical filter block further down `routes.go` on
another list
endpoint. Likely the same class; not touched here.

If you would rather land your own version of this, say so and I will
close mine.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 09:45:18 +02:00
efitenandClaude Opus 5 980c5c4515 fix(nodes): stop the Heard By empty state claiming the node is out of range (#2063)
Follow-up to #2057, which is correct in what it does and overclaims in
one sentence.

## The sentence

When no observer heard the node directly, the card said:

> No observer is within radio range of this node.

#2057's own rule cannot establish that:

- every **direct route** is discarded, because the firmware removes the
sender from the path before retransmitting (`Mesh.cpp`,
`removeSelfFromPath`), so the packet cannot say who transmitted it.
#2057 measured these at **38% of transmissions over 7 days**.
- **28.8% of flood observations with a path** are dropped because the
last hop resolves to more than one candidate, and the rule
under-attributes rather than guesses.

So an empty direct list is missing evidence, not evidence of missing
coverage.

## Why it matters in practice

Sampled 40 repeaters on a production instance after #2057 shipped:

| | |
|---|---|
| at least one direct observer | 24 |
| empty card, "no observer is within radio range" | **16** |
| of those 16, with a non-zero `relayObserverCount` | **16** |

Every node showing "nobody is in radio range" also showed "Seen via
relay by N observers" two lines below. An operator reading that about a
working repeater concludes they have a coverage problem they do not
have.

## The change

Wording only, in both copies of the card (full page and side pane):

> No observation proves a direct reception here, which is not the same
as being out of range.

with the reason in a `title`, so the card stays one line:

> Only flood-routed transmissions identify who was heard: a direct route
removes the sender from the path before retransmitting (firmware
`Mesh.cpp`, `removeSelfFromPath`), and an ambiguous relay hop is left
unattributed rather than guessed. So an empty list is missing evidence,
not proof of missing coverage.

No API change. #2057's rule, shape and performance work are untouched —
I verified its firmware derivation against the clone at `0679dbef`
before writing this: `Packet.h:83`, the forwarder appending with
`packet->getPathHashSize()` at `Mesh.cpp:349`, and `removeSelfFromPath`
on the direct path at `Mesh.cpp:89-105` all read as described.

## Test

`tests/unit/test-direct-rf-heard-by.js` slices this template out of
`public/nodes.js`, so it pins the shipped markup. It now asserts the old
sentence is gone and the qualifier is present.

Worth recording how that assertion was reached: my first version banned
the phrase "out of range" from the card, and it failed — on the new
line, which contains that phrase precisely in order to deny it. A word
ban was the wrong instrument. Matching the old sentence and requiring
the new qualifier is the assertion that actually distinguishes the two
states.

8 of 8 in that suite, eslint clean.

## Not in this PR

The **Regions** line and **Region** column on the same card read
`o.iata`, which `HealthObserverRow` does not emit, so both have always
been dead. #2057 named this and left it; it is now **#2062** with the
file and line references, rather than a remark inside a merged
description.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 23:56:02 +02:00
efitenandClaude Opus 5 61f565c606 fix(node-health): credit zero-hop adverts as direct reception (#2064)
Reported by @dborup on #2057. The mechanism is real; the suspected scale
is not. Both parts measured below.

## The defect

`directHeardNode` rejects every non-flood route type before it looks at
the path, so the `hop == ""` branch that credits an advert's originator
can never run for a direct route. Zero-hop adverts — the clearest
direct-RF evidence the network produces — are discarded and the node is
listed under "Seen via relay" instead.

## Firmware

Read at `0679dbef` rather than taken on trust:

- `Mesh::sendZeroHop` sets `ROUTE_TYPE_DIRECT` and `path_len = 0`,
commented there as "path_len of zero means Zero Hop". The transport
overload does the same with `ROUTE_TYPE_TRANSPORT_DIRECT`.
- `examples/simple_repeater/MyMesh.cpp` sends the periodic **local
advert** through it (the `next_local_advert` branch), as does
`sendSelfAdvertisement` when `flood` is false.
- `examples/companion_radio/MyMesh.cpp` does the same for a companion's
own advert.

An ADVERT arriving on a direct route with an empty path therefore cannot
have been forwarded: the observer received the advertiser's own
transmission, and the advert carries its pubkey in the clear.

Every other direct case keeps #2057's rule. A non-empty path on a direct
route is the **remaining** route, because the forwarder ran
`removeSelfFromPath` before retransmitting, and `advertOriginPubkey`
already returns `""` for any payload type other than ADVERT — so the
payload guard costs nothing.

## Measured, 7-day window on a production instance

| | |
|---|---|
| zero-hop advert observations currently dropped | 8,166 |
| distinct nodes they evidence | 139 |
| node-observer pairs they evidence | 185 |
| pairs **not** already credited via an empty-path flood advert | **52**
|
| pairs currently credited from flood adverts | 217 |

So the fix restores 52 node-observer pairs of direct evidence that are
invisible today, roughly a quarter more advert-based direct evidence.

## What it does not explain

The report suspected this accounts for #2057's low headline numbers
("NL-BXE-RP01 | 433 → 0", 234 of 1,860 nodes with any direct observer).
The measurement does not support that:

- Of 30 sampled nodes with zero-hop advert evidence, **29 already show
at least one direct observer**, because they also send flood adverts
which #2057 credits.
- **NL-BXE-RP01 | 433 has zero zero-hop adverts** in the window. Its
empty list is not caused by this rule.

So this mostly enriches lists that are already non-empty, and flips few
cards from empty to populated. Worth doing on correctness grounds, not
as a fix for the counts.

## Tests

Four cases added to the `TestDirectHeardNode` table, which previously
covered direct routes only with `PayloadTXT_MSG`:

- direct + ADVERT + empty path credits the advertiser
- transport-direct + ADVERT + empty path credits the advertiser
- direct + empty path + **not** an advert credits nobody
- direct + ADVERT + **non-empty** path credits nobody

The last two matter as much as the first two: they pin the exception to
exactly the shape the firmware guarantees.

`gofmt` clean. Go tests not run locally (no cgo toolchain on this
machine since #1992, and per AGENTS.md `CGO_ENABLED=0` builds a stub
that proves nothing), so CI is their first run.

## Related

#2063 fixes the empty state's wording on the same card, which asserts
the node is out of range when the data cannot establish that. The two
are independent: this one adds evidence, that one stops overclaiming
when there is none.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 23:55:55 +02:00
efitenandClaude Opus 5 25f8426d32 test: make every suite in tests/e2e run, and tell the truth when it does (#2053)
Closes #2037.

Step 1 (#2045) wired in the nine suites that already passed. This is
steps 2 and 3: the four that ran and failed, and the five that could not
run at all. After this, the count of suites in `tests/e2e` invoked by
nothing goes from 18 to 0.

## Step 2 — the four that ran and failed

Triaged against a fixture server in CI with full output kept, not the
four-line tail the first probe saved.

| suite | verdict |
|---|---|
| `test-packets-scope-column.js` | passes on master today. It failed on
2026-09-18, so something between the two fixed it. Wired in unchanged
rather than investigated. |
| `test-node-reach-e2e.js` | test wrong. It waited for the reach map
whenever any *link* had GPS; `public/node-reach.js` builds the map only
when the *node* has coordinates. Guaranteed 10s timeout on a node with
positioned neighbours and no position of its own. |
| `test-channel-modal-e2e.js` | both failures test-side. The Add
button's visible label was shortened to "+ Add" with the accessible name
moved to `aria-label` (`public/channels.js:746`); and
`.ch-section-mychannels` is conditional on the visitor having added a
channel, which has not happened at that point in the suite. |
| `test-touch-targets.js` | three test-side, two a product finding. |

The three test-side touch-target failures were the harness measuring
controls that are not shown: `.compare-btn` (the CTA was removed in
#1646, and `style.css` says so), `.ch-back-btn` (`display:none` outside
the mobile channels layout), and `.filter-toggle-btn` (`display:none` on
mobile since #1461; the control shown is the navbar mirror, which
`mobile-page-actions.js:70` builds as a `.nav-btn`, so it was already
measured). All three are dropped from the table with the reason recorded
in the file.

The remaining two are **not** a test problem: `.nav-btn` and
`.ch-icon-btn` are each declared twice in `public/style.css`, 48px in
the touch-target block and 44px in their own component rule, and the
later one wins. Rather than lower the blanket or hide the failures, the
suite now has `DEFAULT_MIN = 48` plus a `MIN_OVERRIDES` table holding
those two at their effective 44, so a third selector dropping to 44
still fails the build. The contradiction is **#2052**, with both ways
out costed; the override entries should go when it is settled.

## Step 3 — the five that could not run

None is deleted. I checked each selector and seam against the product
before deciding, and every one still targets a surface that exists and
that nothing else covers.

Four were written against `@playwright/test`, a runner the project
neither installs nor uses anywhere else. Adopting a second runner for
twelve tests costs more than porting them, and the precedent is already
set: `test-path-inspector-coverage-e2e.js` exists, as its own header
says, because `test-path-inspector-e2e.js` could not run. So they are
ported to the plain-node Chromium pattern the other 109 suites use.

- **`test-issue-1522-trace-url-sync-e2e.js`** — the trace hash in the
URL, both directions. `test-e2e-playwright.js` covers that the page
loads and searches; it never looks at the URL, which is the whole of
#1522.
- **`test-marker-outline-weight.js`** — the canvas pulse ring never
thins below 2px. There is no CSS rule to read and axe cannot see inside
a canvas, so sampling the seam is the only way. Added a guard that the
ring was actually visible, so the weight check cannot pass vacuously on
a pulse that never rendered.
- **`test-pr-1490-live-map-gpu-animations-e2e.js`** — the queue drains,
the engine sleeps again, the fading trails stay under the cap of 5, and
the canvas sits on `animationsPane` rather than under the markers.
- **`test-path-inspector-e2e.js`** — reduced to what nothing else
covers: the map side pane, the `/#/traces/<hash>` redirect, the tools
landing. Its standalone-page test duplicated the wired coverage suite
and is dropped. Its "switching candidate clears prior polyline" case
ended after the click with a comment and no assertion, which is the same
green-but-empty problem this issue is about; it now compares path
counts, and skips loudly when the fixture yields too few candidates.

The fifth, **`test-table-sort.js`**, needed `jsdom`, which was declared
nowhere. It is a unit test of `public/table-sort.js` filed under
`tests/e2e`, so: `jsdom` is a devDependency (lockfile updated, `npm ci`
stays consistent), the file moved to `tests/unit/`, and the
`domIntegration` group in `scripts/non-unit-tests.json` is gone with its
only member. It runs 22 tests. 20 passed immediately; 2 had rotted,
because #1648 M2 replaced the up/down glyphs with Phosphor sprites and
the direction moved out of `textContent` into the `<use href>`. Those
two now read the sprite ref and the `aria-sort` value, so they also
guard the accessible announcement.

## Verification

`tests/unit/test-table-sort.js` 22/22 and `test-test-inventory.js` pass
locally; the E2E suites need a fixture server, which I cannot build here
(no cgo toolchain since #1992), so CI is their first run as committed.
The triage above was measured in CI, not assumed.

## Not done

The per-assertion skips named in the second comment on #2037 are
untouched: the two flaky packet-detail cases, the fixture-data ones, and
the two `clientRxCoverage` suites that skip wholesale while reporting
success. Those need a fixture deployment with coverage enabled, which is
its own change. I have not opened it.

Option 2 from the issue, making `test-test-inventory.js` require a
`deploy.yml` line for every `tests/e2e` file, is also not here. It is
the right guard and it is now enforceable, since the list is finally at
zero, but it belongs in its own change where a red build means what it
says.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 11:45:11 +02:00
efitenandClaude Opus 5 b614badb85 fix(packets): give each observation its own wire bytes in the detail API (#2055)
Closes #1999.

## The defect, measured on a production instance

Packet `96d716f18d885e78` on a live deployment, read from the deployed
build's own API:

| | |
|---|---|
| observations in the response | 60 |
| distinct `path_json` values | 50 |
| distinct `raw_hex` values returned | **1** |
| distinct frames actually stored in SQLite | **51** |

The contradiction the issue describes, from that same response:

```
obs 38410791  path ["58C0","1403","50C7"]  ->  hex 094258C01403AF37E39624E548FB7575F195A0BF
```

Three hops in the path, two path bytes in the frame. The bytes belong to
the 2-hop observation and are served for all 60.

Across the 3000 most recent transmissions on that database: 1977 have
more than one observation and **1844 of those (93%) hold genuinely
different frames**. 32659 of 36133 observations (90%) differ from their
transmission's canonical bytes. This is the normal case, not an edge
case.

## Cause

The store deliberately does not retain `obs.RawHex`. #881 dropped it as
a memory optimisation, ~98MB measured on a 1.7M-observation store, on
the assumption that one content hash implies one frame. The firmware
hashes payload and type independently of the relay path, so that
assumption is false.

Worth adding to the issue's diagnosis: all four load and ingest paths in
`cmd/server/store.go` still `SELECT o.raw_hex` and scan it into
`obsRawHex`, then use it nowhere — LoadAll, loadChunk, `IngestNewFromDB`
and `IngestNewObservations`. The bytes are read out of SQLite and
discarded, so a cold load pays the transfer for nothing.

## The fix

Keeps the memory saving and reads the bytes back only where a human is
looking at one packet.

- **`cmd/server/db.go`** gains `ObservationRawHexForHash`: one query
returning the stored frame per observation id. Two indexed lookups
regardless of observation count — `transmissions.hash` through the
prepared `stmtTxByHash` (`idx_transmissions_hash`), then
`observations.transmission_id` (`idx_observations_transmission_id`).
Guarded by `hasObsRawHex`, because #881 made the column optional and the
query would be a SQL error without it.
- **`cmd/server/routes.go`** backfills in `handlePacketDetail`: once per
request rather than once per observation, and after the store lock is
released. Bytes already present are never overwritten, and an
observation with no stored frame still falls back to the transmission's.

Against the acceptance list: observation bytes exposed with canonical as
fallback only ✓; the store's memory optimisation untouched ✓; bounded
indexed reads with no query per observation and no work under the store
lock ✓; startup-loaded, newly ingested and DB-fallback details all
covered, because both the store path and the DB path converge on this
one backfill and both key observations by an int `id` ✓.

**No frontend change is needed.** `public/packets.js` already spreads
the selected observation over the packet (`{...pkt, ...currentObs}`) and
already reasons about per-observation bytes: the comment there says
"post-#882 per-obs raw_hex with a different path length than the
top-level packet's raw_hex still gets accurate byte highlights". The
client was built for this and has been receiving 60 copies of one frame.

## Tests

`cmd/server/obs_raw_hex_test.go`:

- the per-id mapping, with three distinct frames and a fourth
observation storing none
- the `hasObsRawHex` guard, so a schema without the column is not
queried
- the handler regression: each observation carries its own frame, the
frameless one falls back to the canonical bytes, and at least three
distinct frames come back across four observations — the last assertion
so that a regression to repeating one frame fails, rather than passing
on shape

## Verification

`gofmt` clean. **Go tests were not run locally**: no cgo toolchain on
this machine since #1992, and per AGENTS.md `CGO_ENABLED=0` builds a
stub that proves nothing. CI is their first run.

Browser validation per AGENTS.md rule 2: I verified **the defect** in a
real browser and through the deployed API, with the numbers above. I
could **not** validate the fix in a browser, because the change is
server-side Go and is not deployed anywhere yet. Saying so rather than
claiming otherwise.

## Not done

- The four scan sites that fetch `o.raw_hex` and discard it are left
alone. Removing the column from those query builders would stop
transferring roughly ten frames per transmission on every cold load, but
it touches four builders and their `scanArgs` alignment and is not
needed for this defect.
- `fetchResolvedPathForObs`, immediately next to this code in
`enrichObsWithTx`, does run one query per observation. This change
deliberately does not copy that pattern, and does not fix it either.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 11:08:17 +02:00
efitenandClaude Opus 5 d1b615fc0d fix(node-health): list only observers that heard the node on air (#2057)
Closes #2056.

## What changes

The node detail "Heard By" card now lists only observers that received
the node's **own transmission off the air**, and reports the rest as a
count.

```
HEARD BY — DIRECT (8 OBSERVERS)
OBSERVER                REGION  PACKETS  AVG SNR   AVG RSSI
BE-DUF-SiSCD-01         —        16276    7.9 dB   -108 dBm
BE-BRU-Moris  repeater  —        13775   -5.6 dB   -122 dBm
...
Seen via relay by 29 observers. Those observers heard a repeater that
forwarded this node's traffic, not this node.
```

and for a node nothing hears:

```
HEARD BY — DIRECT (0 OBSERVERS)
No observer is within radio range of this node.
Seen via relay by 2 observers. …
```

## The rule, and where it comes from

Read out of the firmware rather than assumed:

| | |
|---|---|
| `Packet.h:83` | `setPathHashSizeAndCount(sz,n) { path_len =
((sz-1)<<6) \| (n&63); }` — hash size rides in the packet's `path_len`
byte |
| `Mesh.cpp:649,678` | only `sendFlood()` sets it, so the **originator**
decides; `CommonCLI.h:69` defaults `path_hash_mode = 0`, i.e. one byte |
| `Mesh.cpp:349` | a forwarding repeater appends its hash with the
packet's size — it cannot upgrade a packet, and the **last hop is who
was heard** |
| `Mesh.cpp:89,103` | on a direct route a forwarder matches the head of
the path and calls `removeSelfFromPath` before retransmitting, so the
path is the **remaining** route and the transmitter is not in it |

So an observation credits exactly one node:

1. Route type must be `ROUTE_TYPE_FLOOD` or
`ROUTE_TYPE_TRANSPORT_FLOOD`. Direct routes never qualify (38% of
transmissions over 7 days).
2. Empty path → the originator, known only for ADVERTs.
3. Otherwise the last hop.
4. The hop must resolve to exactly one candidate. Same gate
`resolvePathForObsColdLoad` already applies: under-attribute rather than
guess. It drops 418,530 of 1,455,721 flood observations with a path over
7 days (28.8%), and it is what stops the wrong-band credits.

## Measured effect

| node | before | after |
|---|---|---|
| BE-BRU-Moris | 36 observers | 3 |
| BE-KRO-RP01 \| ON1KW | 40 | 3 |
| BE-BRE-ON8AR | 38 | 2 |
| NL-BXE-RP01 \| 433 | 35 | 0 |

Network-wide over 7 days, 234 of 1,860 nodes have at least one direct
observer (161 have exactly one, maximum 8). The direct list is therefore
empty for most nodes, with the relay count below it. That is the correct
reading: no observer is in radio range of them.

Independent corroboration on staging: for BE-WIL-3EIK-01 the eight
direct observers are exactly the top eight entries of its Neighbors
table by score and observation count.

## Perf justification

`GetNodeHealth` is fast today precisely because it never walks
observations — it uses one representative observation per transmission.
Direct-RF needs the per-observation path, and that cannot be a
per-request walk: the reference store holds **232,928 transmissions /
2,887,861 observations**, one node's `byNode` slice alone holds **55,458
transmissions / 1,450,544 observations**, and
`/api/nodes/bulk-health?limit=200` would multiply that.

So the aggregate is rebuilt by a background recomputer on the existing
`newAnalyticsRecomputer` pattern, published into an `atomic.Value`.
Reads are `O(direct observers)`, which is **cheaper than before** — the
old code built per-observer sums over every transmission in `byNode` on
every request.

Proof, `BenchmarkBuildDirectHeardIndex`:

```
BenchmarkBuildDirectHeardIndex-12    1    63067900 ns/op
```

3,000,000 observations (60,000 transmissions × 50 observations, 8-hop
paths, 64 candidate repeaters) in **63 ms**, once per recompute
interval.

Per observation the walk does one route-type check, one backward scan of
`PathJSON` for the last quoted token (no allocation, no
`json.Unmarshal`), one prefix-map lookup and one counter update.

Rebuilding wholesale also means eviction needs no bookkeeping: a pass
simply does not see evicted transmissions. The alternative — a field on
`StoreObs` updated incrementally — would have needed the call at five
construction sites (`store.go:942,1264,2854,3179`,
`chunked_load.go:609`), which is the duplication that caused #1558, plus
matching decrements at eviction.

## API

Both `GetNodeHealth` and `GetBulkHealth` carried a near-identical copy
of the observer loop; they now share one builder.

- `observers` — direct-RF only. Same field names, so no client
migration. Rows are a named `HealthObserverRow` instead of
`map[string]interface{}` (one fewer occurrence in a touched file, per
the AGENTS.md ratchet).
- `relayObserverCount` — new integer, observers that saw traffic through
the node without hearing it. `stats.totalPackets` and `stats.avgHops`
still count relayed traffic, so without this number the card would
contradict the figures printed beside it.

`docs/api-spec.md` is updated for both endpoints. It also documented an
`iata` field on these rows that the endpoint has never emitted; removed.

## Tests

- `cmd/server/direct_heard_test.go` — table test over the rule: flood
with empty path and known originator, flood whose last hop is the node,
flood whose last hop is another node, direct and transport-direct routes
(never credit), ambiguous last-hop prefix, listener-only candidate,
1-byte and 2-byte hop sizes; plus aggregation and row-building.
- `cmd/server/node_health_direct_rf_test.go` — end-to-end through the
handler: an observer that only saw relayed traffic must not appear in
`observers` but must be counted in `relayObserverCount`. Plus the
benchmark.
- `tests/unit/test-direct-rf-heard-by.js` — slices the card template out
of `public/nodes.js` and evaluates it, so it tests the shipped markup
rather than a copy: heading, empty state, relay line, singular/plural,
signal columns, listener/repeater badge tri-state.
- `cmd/server/node_health_can_relay_case_1290_test.go` — updated to seed
a genuinely direct reception, since a relay-only observer no longer
carries a badge.
- `cmd/server/analytics_recompute_after_load_test.go` — recomputer count
10 → 11.

Verified locally: `cmd/server` suite green, `sh test-all.sh` green (180
suites), `tests/e2e/test-e2e-playwright.js` 131/134 passed with 3
skipped and 0 failures against the seeded fixture, plus
`test-issue-1147-section-order-e2e.js`,
`test-issue-1151-orphan-separators-e2e.js` and
`test-issue-1281-location-row-e2e.js`, which all assert on this card.
`gofmt` clean, `vet` clean across all modules.

Browser-validated on staging: both the full detail page and the side
pane, on a node with 8 direct observers and on the 433 MHz node with
none. No console errors.

## What this does not do

`prefixMap.resolveWithContext` still guesses on ambiguous hops, so
paths, neighbor edges and analytics keep their current attribution.
Making it abstain is a much larger change and needs its own issue.

The "Regions" line and Region column on this card read `o.iata`, which
this endpoint has never emitted, so both have always been dead. Left as
found rather than widened into this change.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 07:55:37 +02:00
efitenandClaude Opus 5 5c016a210a fix(analytics): show a building state on the distance index's 202, and stop caching it (#2051)
Closes #1997.

## What was wrong

`/api/analytics/distance` answers `202 {status:"building",
retry_after_seconds:5}` with no `summary` until the lazy index (#1011)
has been built.

1. `renderDistanceTab` read `data.summary.totalHops` straight away. The
TypeError is caught by the tab's own try/catch, so it never reaches
`window.onerror`: it is painted into the tab as `Failed to load distance
analytics: Cannot read properties of undefined (reading 'totalHops')`.
2. `api()` caches any `res.ok` body, and `res.ok` is true for 202, so
the placeholder was stored for the `analyticsRF` TTL. Even a correct
retry read the cached "building" body back. That is the half that made
the broken state outlast the index build.

## What changed

- `public/app.js:173` skips the cache write when `res.status === 202`. A
200 still caches, unchanged.
- `public/analytics.js` renders a building notice and retries itself,
honouring `retry_after_seconds` clamped to [1s, 30s]. The timer is
cleared on tab switch and in `destroy()`, and a new render supersedes a
pending retry, so two renders cannot write into the same tab.

The notice uses `.text-center`/`.text-muted` rather than the `.spinner`
class used at `analytics.js:1683`, because `.spinner` has no CSS
anywhere in the repo and renders nothing.

## Tests

`tests/unit/test-issue-1997-distance-building.js` (7 assertions, wired
into `test-all.sh`) pins the two pure decisions the renderer makes and
`api()`'s refusal to cache a 202 while still caching a 200. Red-run on
the unfixed sources: 6 of 7 fail, and the "a 200 is still cached"
control stays green.

`tests/e2e/test-issue-1997-distance-building-e2e.js` (classified in
`scripts/non-unit-tests.json`, invoked from `deploy.yml` with
`CHROMIUM_REQUIRE=1`) serves both responses by route interception, so it
does not depend on whether the server under test has an index built. It
asserts: the building state appears, the tab does not paint the error
text, a retry arrives with no interaction, the retry replaces the
placeholder once the server answers 200, and no retry fires after
leaving the tab.

Check (2) deliberately asserts on the rendered text and not on
`pageerror`: the TypeError is caught, so a `pageerror` assertion would
pass on the broken build too.

## Verification

No local cgo toolchain here since #1992, so I could not build a server
to run the E2E against. Instead I ran its five steps in Playwright
against a live instance with this branch's `public/app.js` and
`public/analytics.js` injected in place of the deployed ones (both files
are byte-identical between that instance and upstream master, so the
injection is faithful):

| check | deployed build | this branch |
|---|---|---|
| (1) building state shown | fail | pass |
| (2) not rendered as data | fail | pass |
| (3) retried on its own | fail (1 request) | pass (2 requests) |
| (4) real payload after retry | fail | pass |
| (5) no retry after leaving the tab | pass | pass |

(5) passes on the broken build too: it schedules no retry at all, so it
is a control and only means anything together with (3).

The committed E2E suite has not been run as committed. CI is its first
real run.

## Not done

- The server still recomputes on every 202 poll rather than signalling
readiness.
- No other analytics tab was audited for the same assume-a-summary
pattern.
- One pre-existing unit suite (`test-preflight-xss-gate.js`) fails on
this Windows machine with a cp1252 `UnicodeEncodeError` from its Python
helper, on a clean tree as well as with this change. Unrelated, and
green on Linux CI.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 17:04:37 +02:00
efitenandClaude Opus 5 7c6b95ea53 fix(store): merge background chunks in order instead of prepending them (#2050)
Fixes #2024.

`s.packets` is declared "sorted by first_seen ASC (oldest first; newest
at tail)" (`cmd/server/store.go:177`), and retention eviction depends on
it: `evictStaleInternal` walks from the head and stops at the first
transmission inside the window. A slice out of order is therefore
**under-evicted silently** rather than failing loudly.

## What breaks it

The background chunk loader. Chunks are windowed on `last_seen` (#1690),
so a transmission first heard weeks ago and heard again recently arrives
in a *recent* chunk carrying its old `first_seen`. The chunk was then
put in front of the slice:

```go
s.packets = append(localPackets, s.packets...)
```

and never re-sorted, so the next chunk, which covers an older window,
was prepended in front of it and left that ancient row sitting behind
newer ones. `LoadChunked` re-sorts after its own load; the background
merge did not. That asymmetry is the whole bug.

It is not a corner case. On a production database, of the **236080**
transmissions in a 14 day window, **2071** have a `first_seen` more than
a day older than their `last_seen`, and **1848** more than a week.

This matters more since #2035: with the accounting fixed, `maxMemoryMB`
actually triggers, and a walk that stops early works against it.

## The fix

`mergeChunkIntoPackets` merges the two sorted runs linearly. Re-sorting
the whole slice was not an option: this runs under `s.mu` once per
chunk, so it would sort hundreds of thousands of packets while ingest
waits for the lock. The chunk already arrives sorted, since the chunk
query ends in `ORDER BY t.first_seen ASC`, so the `sort.SliceIsSorted`
guard is a contract check costing one linear pass that never sorts in
production.

## Covered

- `TestMergeChunkIntoPackets_KeepsFirstSeenOrder` pins the merge against
an interleaving, deliberately unsorted chunk.
- `BenchmarkMergeChunkIntoPackets` guards the linear cost, against a
future simplification back into a sort.

The server suite runs under `-race` in CI and is green.

## Not covered, and I would rather say it than let the PR imply
otherwise

There is **no integration test driving `loadChunk` end to end**. I wrote
one and dropped it: a faithful seed database for that path needs more of
the schema and more of the loader's preconditions than the fix itself is
worth. Two CI rounds in, the seed was still loading zero packets (the
first attempt failed at `OpenDB` on a missing `nodes` table, the second
on the window). Both attempts are in this branch's history rather than
rewritten away.

So the end-to-end claim rests on the code path quoted above and on the
production measurement, not on a test that exercises it. The unit test
covers the function where the logic now lives, which is the part that
can regress.

Also not verified locally: `cmd/server` needs cgo for the #1992 driver
and this machine has no C toolchain, so CI is the check.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 15:12:36 +02:00
liquidraver 84eba41592 fix(live): anchor the legend toggle to .live-page so the VCR bar stops eating its clicks (#2049)
The PACKET TYPES legend could not be dismissed on Live: clicking the palette toggle did nothing, while document.querySelector('#legendToggleBtn').click() from the console worked. That asymmetry was the diagnosis, since a synthesised click skips hit testing.

Residual half of #1833. That fix corrected the button's offset to calc(var(--vcr-bar-height) + 10px) but left the anchor: .legend-toggle-btn was position: fixed, so the offset resolved against the viewport, while the .vcr-bar it clears is absolute inside .live-page, whose height subtracts --bottom-nav-reserve (56px + safe-area at <=768). The two anchors disagreed by exactly the reserve, dropping the button into the bar's band, and .vcr-bar at z-index 1000 against the button's 500 took every real click. The window is 641-768px: below 641 both buttons are display:none, above 768 the reserve is 0, which is why neither the desktop nor the phone layout test saw it.

Both buttons are now position: absolute, resolving against the same containing block as the bar. .feed-show-btn carried the identical defect but missed the bar by 16px because --legend-toggle-stack parks it a row higher, so that change prevents a break rather than repairs one.

Three review rounds, each verified rather than argued:
- The new E2E was classified in non-unit-tests.json but had no deploy.yml line, so it would never have run (#2037). Now wired in, and it covers both buttons.
- The '.feed-show-btn is display:none at <=768' claim was wrong (it is <=640), which the author measured and corrected.
- Its first real execution failed on its own setup: getComputedStyle().getPropertyValue() on an unregistered custom property returns the unevaluated declaration ('calc(56px + 0px)'), so the non-zero check could not work. It now asserts the measured gap instead, and deliberately does not pin 56 because the gap measures 58 on CI.

Final run: 14 passed, 0 failed. The discriminating assertion is (f), each button's offset net of the bar being identical at 720 and 1440 (11px and 123px slack); under the old anchor those differ by the reserve. (g) pins the desktop layout unmoved.

Merged by the interim maintainer without a second human reviewer: CI and the review above are the independent checks.

Fixes #1833
2026-09-20 14:00:45 +02:00
n30nex 7f4b357d89 fix(map): keep inspector controls clickable and ignore stale loads (#2048)
Fixes #1998 and #2030.

Map Controls overlapped the Path Inspector toggle on desktop and intercepted its clicks; the controls are now offset by the collapsed and expanded inspector widths, with the mobile layout unchanged.

Leaving the Map page while its node data was loading could log an invalidateSize error or let an old response overwrite a replacement map. Node-loading continuations and the deferred initialization callbacks are now tied to the Leaflet instance that started them, and obsolete results return before touching state or rendering.

Reviewed against the code: the identity check is sound because destroy() clears the reference (map.remove() then map = null, public/map.js:2428-2433), so every guard is false after teardown and after replacement. Each await in loadNodes is followed by a check, as are the two deferred callbacks, the error branch and the finally. Moving 'nodes = data.nodes' after the observers await is deliberate: shared state is only mutated once the load is known to still own the map.

The tests are the strong half. Overlap is asserted through elementFromPoint on the toggle's centre across five viewport widths and both pane states, rather than a forced click that would pass over an overlapping control, and 641/640 pin both sides of the breakpoint. Both extended suites are invoked by deploy.yml, one line each, which was checked rather than assumed. The commit chain is test-then-fix twice: e675100/7d736d1 and fd56ea2/823668b.

One nit left to the author: the CSS repeats the pane widths from style.css:4532-4533 four times, so changing the pane width would silently misalign the controls, caught only as a waitForFunction timeout.

Merged by the interim maintainer without a second human reviewer: CI and the review above are the independent checks.
2026-09-19 21:11:53 +02:00
efiten 29c0a3ba2f ci: name each E2E suite in the log before it runs (#2046)
The E2E steps invoke 105 suites in a row. GitHub echoes a step's whole script once at the top, so the per-suite output then runs together with nothing between it, and a line like 'node-reach-coverage E2E SKIP (clientRxCoverage disabled on this deployment)' cannot be attributed to its suite without reading the sources and guessing. That is how test-node-reach-coverage-e2e.js came to pass every build while asserting nothing, unnoticed until #2037.

Each invocation now prints '=== E2E SUITE: <file> ===' first, so the log is greppable per suite.

Safe by construction: the banners go to stdout only, never through tee, so e2e-output.txt is byte-identical and scripts/aggregate-e2e-pass.sh sees what it saw before (it keys on digits followed by 'passed', which a banner never matches). The change is mechanical and was checked as such: the diff removes no line and every added line is a banner, 105 for 105 invocations.

Verified on CI run 35347268407: all jobs green, and the banners make attribution work. Using them, exactly one wired suite skips wholesale (test-node-reach-coverage-e2e.js) and three skip individual assertions (test-e2e-playwright.js on a flaky case, test-touch-gestures-coverage-e2e.js and test-issue-1306-collisions-terminology-e2e.js on fixture gaps). That list was guesswork before.

Refs #2037. Merged by the interim maintainer without a second human reviewer: CI is the independent check.
2026-09-19 11:46:56 +02:00
efiten e01565737a feat(ingestor): store CoreDrive RX region answers, with position, clock and retention (#2047)
The Scope Audit page said a repeater's declared region list can come from CoreDrive RX while nothing the app sends ever reached it: the client-topic switch handled packets and rf only, so /regions was dropped without a log line, and node_declared_regions is read by region_keys.go and config.go but created by nothing in this tree. #2044 found that gap and proved it against a live instance.

This lands the implementation that has been carrying the feature in production on the ON8AR fork since 2026-09-06, the instance CoreDrive RX publishes to. Measured there: 1840 answers about 275 repeaters from 51 collectors, 2026-08-18 to 2026-09-19.

Beyond storing the answer it keeps three things the first version did not: position (lat, lon, pos_acc_m, filled on 1263 of 1840 rows, with acc_m dropped when the fix it qualifies was rejected), repeater_clock (filled on all 1840, so a wrong repeater clock cannot make an answer look newer than it is), and retention with per-collector history (pruneOldClientDeclaredRegionsAt bounds by age instead of keeping one row per target). On that dataset 138 of 275 repeaters have answers from more than one collector and 40 have collectors that disagree about the region list, which is the signal the Scope Audit exists to surface and which only survives while more than one answer does.

It gates on its own clientRegions block rather than riding on clientRxCoverage, so region answers can be accepted without GPS-tagged reception uploads, and AES block padding is trimmed from region names on ingest.

Taken from #2044 with the author credited as co-author: declaredRegionsTablePresent() and its test, a real bug this version lacked (supervisord starts both processes together, so a server that probes first ignores every answer until its next restart), and the docs/client-rx-coverage.md section.

Ported by cherry-picking the fork's nine commits rather than retyping, so this is the code that has been running. CI run 35434151026 is green: server ok 80.177s, ingestor ok 99.413s, race detector ok 112.406s, no --- FAIL lines.

Merged by the interim maintainer without a second human reviewer: CI and the production figures above are the independent checks.
2026-09-19 11:44:47 +02:00
efiten 20b0ee9742 ci: run the eight tests/e2e suites that had no runner at all (#2045)
Part of #2037. Of the 113 suites in tests/e2e, 18 were invoked by no deploy.yml line, no script and no workflow: written, classified in scripts/non-unit-tests.json, and never executed.

All 18 were run once on upstream master to find out which still work (CI run 35340835178, each tolerated and timed so one run produced the whole table). Nine passed; eight are wired in here, 45 seconds together: test-analytics-fluid-charts (2s), test-nodes-export-e2e (2s), test-1110-live-filter (3s), test-e2e-1267-mobile-vcr (6s), test-live-dedup (6s), test-issue-1274-legend-coverage (7s), test-issue-1648-m5-icons (9s), test-show-neighbors (10s).

The ninth, test-rx-coverage-mobile-nav-e2e.js, is deliberately left out: it exits 0 with 'SKIP (clientRxCoverage disabled on this deployment)' and coverage is off by default, so it would add the appearance of a guard rather than a guard.

Skip paths were checked rather than assumed: test-issue-1648-m5-icons honours CHROMIUM_REQUIRE and gets the flag, test-nodes-export-e2e's skip branch needs an empty dataset and the fixture holds 182 named positioned nodes of 200, and the other six have no skip path.

Verified on run 35343304933: all jobs green, each of the eight invoked exactly once in the Playwright job, and none of them printed a SKIP.

Left for #2037: the four that run and fail (test-channel-modal-e2e 12/2, test-packets-scope-column 4/3, test-touch-targets 5 assertions, test-node-reach-e2e TimeoutError) and the five that cannot run as node scripts (four import @playwright/test, one jsdom, neither is a dependency).

Merged by the interim maintainer without a second human reviewer: CI is the independent check.
2026-09-18 14:47:22 +02:00
efiten 6c4fa041de test(rx-coverage): budget the viewport assertions in pixels, not a flat 0.001° (#2040)
assertViewport allowed 0.001 degrees between the centre it asked for and the centre Leaflet reports back. Leaflet keeps the centre as a pixel coordinate, so a read-back centre is only accurate to the pixel it landed on, and one pixel spans 360/(256*2^zoom) degrees: 0.02197 at zoom 6, 0.00137 at zoom 10. The flat budget demanded 1/20 of a pixel at zoom 6 while its own comment said 'allow sub-pixel coordinate rounding'.

It passes on master because the container rounds favourably. It fails as soon as the layout changes: on the ON8AR fork, whose coverage page carries one extra toolbar row, the first assertion reported {lat:11.996338401936226, lng:33.99169921875001, zoom:6} against an expected 12/34 — 0.167 px of latitude and 0.378 px of longitude, so the view was right and the assertion was wrong. Any container-height change can trip it, fork or not.

The budget is now one pixel at the asserted zoom via a degPerPixel(zoom) helper, applied to both centre assertions and to the waitForFunction on the hash (zoom 10, where 0.001 was 0.73 px and merely lucky). For latitude the Mercator scale is the longitude figure times cos(lat), so the longitude figure is the looser bound there, deliberately: this asserts the view is where we asked, not a projection identity. A one-tile miss at zoom 6 is 5.6 degrees and still fails.

Verified by arithmetic rather than locally (the suite needs a served instance and a browser): the observed drift falls inside the new budget and a one-tile miss does not. CI run 35273657579 is green, with 'RX coverage viewport browser regressions OK' in the Playwright job.

Merged by the interim maintainer without a second human reviewer: CI is the independent check.
2026-09-18 10:51:34 +02:00
liquidraver bbf54cfe1b fix(rx-coverage): open at the configured map default, with its own saved viewport (#2033)
The coverage page opened at a hardcoded [51.0, 4.8] zoom 8 regardless of deployment, ignoring /api/config/map (#2032). It now follows the same precedence as the main map (URL hash, then saved position, then /api/config/map, then [37.6, -122.1] zoom 9) and persists its own position across visits, syncing lat/lon/zoom into the hash so a view is shareable.

The saved position lives under its own key, rx-coverage-view, and the page never reads or writes the main map's map-view: sharing the configured default was the bug, sharing the session position was not. Both suites assert map-view stays untouched after a pan, so reintroducing a shared write fails instead of passing quietly.

Also fixed here: selectedRx is now percent-encoded into the hash, and a generation counter stops a late /api/config/map response or a stale 150ms layout timer from building a map for a page that was already left.

Reviewed twice. Verified by mutation rather than by reading: writing map-view too, ignoring the saved coverage position, and dropping the /api/config/map fetch each make the unit suite exit 1, so it covers the feature, the fix and the rejected alternative. The deploy.yml invocation was confirmed to land inside Run Playwright E2E tests (fail-fast) by parsing the workflow, and CI run 35261689694 is the E2E suite's first real execution: Go prints 'RX coverage viewport regressions OK' and Playwright prints 'RX coverage viewport browser regressions OK'.

Worth recording for the next reviewer: that E2E asserted localStorage.getItem('map-view'), so wiring it into deploy.yml without updating the assertion would have turned the job red on its first ever run. It was registered in scripts/non-unit-tests.json but invoked by nothing, the gap tracked as #2037.

Merged by the interim maintainer without a second human reviewer: CI and the mutation checks above are the independent checks.

Fixes #2032
2026-09-17 22:02:36 +02:00
efiten aabeda0f2c test(server): set the #1239 lock-hold threshold from measurement, 150µs to 5ms (#2039)
TestComputeAnalyticsDistanceLockHoldDuration failed on two consecutive master commits (5430bc79 at 222µs, 89377333 at 156µs), both passing on a re-run of the identical tree, neither touching cmd/server runtime code. The flat 150µs limit sat inside the healthy band.

Measured, not assumed:

  healthy    156µs, 222µs, 402µs   three commits, 402µs from run 35252186369
  regressed  201203µs              fork run 35252430017, RLock deliberately
                                   held across the whole compute

A factor of 500 apart, so the limit only had to stop sitting inside the healthy band. 5ms is 12x above the worst healthy reading and 40x below the measured regression. The four numbers and the run IDs are in the doc comment.

The first attempt (a1767c77) calibrated against a control where readers churned a second store the writer never locks, on the assumption that their CPU load was slowing the writer. CI measured that control at 0µs: the readers cost the writer nothing, the variance is lock handoff, and the control could not see what it was meant to subtract. 5e457961 replaces it. Both commits are kept in this branch's history, and issue #2038 is corrected where it argued against raising the limit.

Methodology untouched: same eight readers, same 200 writer cycles, same 20000 hops and 200 paths.

Merged by the interim maintainer without a second human reviewer. Not run locally: cmd/server needs cgo for the #1992 driver and this machine has no C toolchain, so CI (run 35253209412) is the check, and the mutation run above is what proves the assertion still fails on a real regression.

Fixes #2038
2026-09-17 21:58:34 +02:00
Alex B e6323ec587 fix(store): account path, decode-cache and dedup-key bytes so maxMemoryMB eviction triggers (#2035)
trackedBytes undercounted the packet store by about 2.5x, so packetStore.maxMemoryMB never triggered: a tx was charged at creation, before pickBestObservation set its path, so the byPathHop and spTxIndex costs were never added, and eviction then re-estimated with the path known and subtracted more than had been added, drifting the total downwards. The ParsedDecoded cache, the obsKeys dedup key and several per-observation strings were not estimated at all.

StoreTx.accountedBytes now records what was charged, rechargeTx returns the delta after every pickBestObservation, and eviction subtracts accountedBytes instead of re-estimating. Measured by the author on a production database copy: trackedMB 151 against 402 MB of heap in use before, 351 against 396 MB after, with GC cycles dropping from ~0.71/s to ~0.011/s over 12 h on their instance.

Reviewed by auditing the accounting rather than the arithmetic: all six production pickBestObservation sites recharge, all three subtraction sites read accountedBytes, every recharge site holds s.mu (Load from :857, the two ingest paths at :2778 and :3142 with deferred unlocks), observations are charged only after acceptance, and no charged tx is discarded during the chunk merge. That lock audit is the independent check, because CI's race job covers the ingestor only.

Operator impact, both from the estimate growing rather than any limit moving: where maxMemoryMB is set, eviction now caps the real store size, and the cold load clamp drops about a third of the boot walk (124420 to 84374 packets at 650 MB). Where it is unset, which is the default, the change is inert. docs/go-migration.md claimed the Go server ignored the setting, which was never true, and is corrected here.

Merged by the interim maintainer without a second human reviewer: CI (run 35258839322) plus the review above are the independent checks.
2026-09-17 21:58:16 +02:00
Alex B 893773338e chore(tests): move root test-*.js into tests/unit and tests/e2e (#2036)
Moves 290 root test-*.js into tests/unit (177, listed in test-all.sh) and tests/e2e (113, classified in scripts/non-unit-tests.json), per #1981 and PR-D of #1385. Root goes from 348 entries to 48. test-all.sh and test-fixtures/ stay put. The inventory guard now fails if a test reappears in the root or sits in the wrong folder.

Verified independently of the diff: the invoked sets are unchanged (test-all.sh 177 before and after, deploy.yml 96 before and after, both identical as sets), and a full local run of test-all.sh on master and on the branch produced 4702 output lines each whose only differences are absolute paths, stack-trace line numbers shifted by the REPO_ROOT line, the inventory wording and two perf ratios. The guard was mutation-checked: a test back in the root, a unit suite in tests/e2e, and a suite dropped from test-all.sh each make it exit 1. CI run 35246304316 ran 97 suites from tests/e2e and is green.

Follow-up 9335c51d finished the instruction files: no bare root test command is left in AGENTS.md, the squad charters, .github or docs, and every tests/ path they name resolves.

Merged by the interim maintainer without a second human reviewer: CI and the local runs above are the independent checks.

Known and deliberately out of scope: 18 of the 113 files in tests/e2e are invoked by no runner at all, and one of them cannot run anywhere because it requires jsdom, which is not a declared dependency. Tracked separately.
2026-09-17 19:06:07 +02:00
efiten 5430bc7923 test(ingestor): anchor the RF-sample fixtures to now, not to a calendar date (#2034)
The three ClientRfDeltas tests seeded 2026-08-17T10:00:00.000Z and queried that window back. resolveRxTimeCore (cmd/ingestor/main.go:1527) replaces timestamps older than 30 days with the ingest time, so from 2026-09-16T10:00Z the seeds landed at time.Now() and every delta fell outside the queried window. Master and every open PR went red on it.

Fixtures now derive from a package-level base two hours in the past, computed once per test binary so two seeds cannot straddle a second boundary and break the exact WallMillis assertion.

Merged by the interim maintainer without a second human reviewer: CI is the only independent check (run 35222316327, ingestor tests ok in 97.033s, race detector ok, no --- FAIL). Fixed dates elsewhere in the ingestor tests are untouched, they assert row counts rather than querying by the seeded date.
2026-09-17 16:52:20 +02:00
efitenandClaude Opus 5 b8c8d98e61 fix(release): keep every platform when re-tagging :edge as a release (#2031)
## Problem

`crane mutate` works on one image, not on an index. Pointed at the
multi-arch `:edge` tag it silently resolves the default platform, so the
fast path published v3.11.0 as a single amd64 OCI manifest, and `crane
tag` then pointed `v3.11`, `v3` and `latest` at that same manifest.
`docker pull` on arm64 against any of those four tags fails.

Verified in the registry:

| tag | shape | arch |
|---|---|---|
| `v3.9.2`, `v3.10`, `v3.10.1`, `edge` | index, 4 children | multi-arch
|
| `v3.11.0`, `v3.11`, `v3`, `latest` | `oci.image.manifest.v1`, 17
layers | amd64 only, revision `a2ea18f7` |

Earlier releases are indexes, so this only hit v3.11.0. The GitHub
release and both `corescope-decrypt` binaries are unaffected.

## Change

- The fast path now reads the `:edge` manifest, mutates each runnable
platform child by digest (`/app/.image-version` plus the version label,
as #1807 intended) and reassembles an index with `crane index append`. A
single-platform `:edge` still takes the old single mutate path.
- Attestation manifests (`platform.architecture == "unknown"`) are not
carried over: they reference the pre-mutation digests, so copying them
would attest the wrong images.
- A new verification step compares the platform set of `vX.Y.Z`, `vX.Y`,
`vX` and `latest` against `:edge` and fails the run if any of them
differs. A release tag that resolves to one platform is worse than a
slow release, so this should break the build rather than ship.
- Scratch tags (`tmp-vX.Y.Z-linux-amd64`, ...) are deleted best-effort
afterwards; the index references the manifests by digest, so leaving
them behind is only untidy.
- Added `workflow_dispatch` with a `tag` input to republish the images
for an existing release. A dispatched run resolves the tagged commit
itself, because `github.sha` is then the ref the workflow file came
from, and it skips the `deploy.yml` dispatch: that release already
exists and releases here are immutable (the trap from #1955/#1956).

## Tests

None: this repository has no harness that executes workflow files, and
the CI jobs cannot reach a step that pushes to GHCR. What the change is
verified against instead:

- `crane index append` accepts `-m/--manifest` repeated plus `-t/--tag`,
with the base index optional, so building an index from scratch is
supported (crane docs for `index append`).
- The YAML parses and every `run:` block passes `bash -n`.
- The platform comparison was run by hand against the live registry:
`:edge` reports `linux/amd64,linux/arm64` and `v3.11.0` reports
`single`, which is exactly the case the new step must fail on.
- The real test is the dispatch on `v3.11.0` right after merge, which is
also the repair. If the verification step fails there, nothing is
published and the tags stay as they are.

## Not verified

- The scratch-tag delete needs `delete:packages`; `GITHUB_TOKEN` may not
have it. It cannot fail the run.
- Whether GHCR keeps the attestation manifests attached to `:edge`
reachable after the index is rebuilt for a release tag (they stay on
`:edge` itself, which is untouched).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:07:18 +02:00
Sylvain Rabot a2ea18f778 perf(sqlite): swap modernc.org/sqlite for mattn/go-sqlite3, cross-built with zig (#1992)
Swaps the SQLite driver from `modernc.org/sqlite` (pure Go, SQLite
3.46.0) to `github.com/mattn/go-sqlite3` (cgo, bundled SQLite 3.53.4),
and pays the resulting cross-compilation cost with `zig cc`.

Draft because the riskiest part of this deletes rows — see [Please
review this part first](#please-review-this-part-first) — and because
three things remain unverified at the bottom.

`modernc.org/sqlite` is a transpilation of the C amalgamation. This repo
is read-heavy: `cmd/server` chunk-loads a graph at startup and fans out
neighbour/topology/analytics queries per request, and it pays for that
transpilation on exactly those paths. Head-to-head on the same
120k-transmission / 240k-observation database, running our own hot-path
SQL under both drivers (Apple M4, `-count=5`, medians):

| workload | modernc | mattn | |
|---|---:|---:|---|
| chunk load (`chunked_load.go` v3 join, 20k tx) | 449ms | 196ms |
**2.3×** |
| aggregate scan (240k-row join + `GROUP BY`) | 276ms | 137ms | **2.0×**
|
| 1500 prepared-statement lookups | 512ms | 403ms | **1.3×** |

Allocations fall with it: 1.12M vs 1.64M allocs and 21MB vs 30MB on the
chunk load.

**Superseded by a production run.** @efiten measured both drivers on a
real instance — 11,077,038 observations, 9.7GB database, 4-core arm64 —
as server-only containers against the same live volume, one at a time,
with round 2 reversing the order so the page cache favours the old
driver:

| | audit 7d | audit 24h | background fill (13 chunks) | start →
/api/health |
|---|---:|---:|---:|---:|
| modernc, round 1 | 16.67s | 2.27s | 130.2s | 16.6s |
| mattn, round 1 | 7.87s | 1.34s | 93.8s | 13.5s |
| mattn, round 2 | 8.15s | 1.35s | 96.4s | 13.0s |
| modernc, round 2 | 13.46s | 2.29s | 137.8s | 15.5s |

Warm, the old driver improves to 13.46s on the 7d audit and still loses
by ~1.8×. Chunk load is ~1.4×. `/api/nodes?limit=500` is 0.039s against
0.037s — nothing.

**So the real gain is ~1.4–1.8× on the paths that matter, not 2–2.3×.**
The shape the harness predicted holds — scans and joins gain, small
lookups do not — which is more reassuring than the magnitude would have
been. Quote these numbers.

**The counterweight**, cold and native on that machine: a build goes
from **52s to 163s**. An instance that builds its own image pays that
per deploy.

## The build is cgo now, and one thing about that is a trap

**`CGO_ENABLED=0` still builds.** mattn links a stub, and the binary
dies on its first query with `go-sqlite3 requires cgo to work. This is a
stub`. A green build is not evidence of anything here, which is why
`AGENTS.md` now says so explicitly. `GOOS=linux go build` genuinely
cannot cross-compile any more.

A new root `Makefile` is the entry point. `make crossbuild` uses `zig cc
-target {x86_64,aarch64}-linux-musl` and links static, so each artifact
stays a single self-contained file and the `alpine:3.20` runtime no
longer depends on the base image's libc at all.

`-Wl,-s` is load-bearing: Go's own `-s -w` does not reach the musl
objects zig links in, and without it the server binary is 19.8MB instead
of 12.1MB.

The Dockerfile keeps its single `$BUILDPLATFORM` builder — still no QEMU
for compilation — and gains a checksum-pinned zig plus BuildKit cache
mounts. The mounts are not a nicety: without them an image build
recompiles the amalgamation from cold and takes over half an hour.

## Please review this part first

`internal/dbschema/dedup_index.go` **deletes observation rows**. It is
the one part of this change that can lose data, and it exists because
the migration exposed a real bug rather than causing one.

`stmtInsertObservation` resolves its `ON CONFLICT` against
`idx_observations_dedup`, which `cmd/ingestor/db.go` only ever created
inside the branch that creates the `observations` table for the first
time. Any database whose table predates that branch never got one, so
the UPSERT had no conflict target. modernc failed on the first insert;
mattn fails at `OpenStore`. Same bug, found earlier.

Creating the index unconditionally repairs it — but the index is what
was supposed to prevent duplicates, so a database that never had it can
already hold rows violating it. **`test-fixtures/e2e-fixture.db` in this
repo holds one.** So duplicates are collapsed first. Refusing is not the
safer option: without the index the ingestor cannot prepare its UPSERT,
so it cannot start at all.

Replaying that UPSERT faithfully is subtler than it looks, and a first
cut of this got it wrong twice:

- `COALESCE(excluded.x, x)` means the **incoming** value wins, so down a
group in id order the survivor keeps the **last** non-NULL value. Taking
the first silently discarded newer readings.
- The UPSERT names exactly five columns (`snr`, `rssi`, `score`,
`raw_hex`, `resolved_path`). Every other column must keep the surviving
row's own value; merging those too invents history the ingestor would
never have written.

Merge, delete and `CREATE UNIQUE INDEX` now share one transaction. Split
apart, a writer inserting a duplicate in the gap fails the index
creation while leaving the deletions committed — rows destroyed and no
index to show for it.

Cost, measured on 2.4M synthetic rows holding 5 duplicates: **4.1s**,
holding the write lock throughout, once, at ingestor startup before MQTT
subscribe. Materialising the duplicate-group scan once rather than per
column took that from 9.7s; the pathological case (400k of 600k rows
duplicated) is 5.7s, slightly worse than the 4.2s it was before that
change.

## Four more behavioural differences

Full detail in `docs/sqlite-driver-migration.md`. Briefly:

**Statement preparation is eager.** modernc's `newStmt` stored the SQL
and compiled lazily; mattn calls `sqlite3_prepare_v2` inside `Prepare`,
so SQL naming a missing table fails at *open*. 59 server tests failed on
this alone, all fixtures with partial schemas. `OpenDB` keeps failing
loudly (#1901; `main.go` gates on `dbschema.AssertReady` anyway) and the
fixtures now declare what they are prepared against via
`ensurePreparable`. This also exposed nine `nodes(pubkey …)`
declarations across seven files, where production has only ever had
`public_key` — lazy compilation had hidden the mismatch for as long as
it existed.

**`synchronous` silently dropped FULL → NORMAL.** mattn defaults it to
NORMAL and executes the pragma unconditionally, where SQLite's own
default (what modernc left alone) is FULL. In WAL mode that weakens
durability under power loss. Pinned in `dbschema.WriterDSN`, which both
writers now share — `cmd/migrate` kept a bare path at first and so
quietly wrote at NORMAL, which is what a second copy of a DSN buys you.

**The DSN dialects are mutually invisible.** modernc understood only
`_pragma=name(value)`, mattn only `_`-prefixed parameters, and neither
errors on the other's form — a driver-only rename would have dropped
every pragma in silence. `_journal_mode=WAL` is also gone from the
server's read handle: modernc ignored it, mattn honours it, and setting
`journal_mode` on a read-only connection is a write. Dropping
`_busy_timeout` with it costs nothing, since mattn already defaults to
5000ms — which means the read handle finally *gets* the busy timeout it
had silently lacked.

**`mode=ro` survives for a non-obvious reason.** mattn always passes
`READWRITE|CREATE` and its amalgamation has `SQLITE_USE_URI=0`; what
makes the URI work is its C wrapper ORing `SQLITE_OPEN_URI` in. So the
#1283/#1289 invariant holds with no build flags — but it depends on the
`file:` prefix. `cmd/decrypt` had been building its DSN without one, so
its `mode=ro` had never applied and a missing path was created
read-write. Fixed in passing; never a migration regression.

## What did not change

No modernc-specific API was in use: no `RegisterFunction`, no
`*sqlite.Conn`, no `sqlite/lib` error constants, no `sql.Register`. No
`time.Time` is ever bound as a query argument, so driver time handling
is not in play. Both drivers convert declared
`DATE`/`DATETIME`/`TIMESTAMP` columns to `time.Time`, so
`/api/dropped-packets` keeps emitting `dropped_at` as RFC3339 — an
earlier draft "fixed" that with a `CAST` and would have been the
regression.

## Tests and CI

New regression tests, each written because something got through without
it:

- `TestEnsureObservationsDedupIndexKeepsLatestValues` — the merge
ordering. The original test used complementary NULLs, which passes
whichever direction you pick, which is why the bug survived it.
- `TestCollapseDuplicatesAndIndexIsAtomic` — a failed index creation
must roll the deletions back.
- `TestOpenStorePragmas` / `TestWriterDSNPragmas` — every writer pragma,
read back through the store's own connection. A separate `sqlite3`
session or the startup log line would prove nothing.
- `TestOpenDBRefusesMissingDatabase` — the read-only invariant, which
now rests on a detail of the driver's C wrapper.
- `TestEnsurePreparableMatchesPrepareStatements` — fails when a new
prepared statement outgrows the fixture helper.

CI gains test execution for `cmd/migrate` and `internal/dbschema`, which
had none and both open the database. A PR-time two-arch build plus an
arm64 QEMU smoke gate is new: the GHCR push is push/tag-only, so without
it nothing on a PR would exercise zig, static musl linking or arm64, and
the first signal would arrive on master. `cache-dependency-path` widens
from 2 of the 5 tracked `go.sum` files to all of them.

`make test` passes across all 14 modules, `cmd/server` also under `-race
-count=2` with no failures and no races. `gofmt` and `go vet` clean.
Release-routing and Dockerfile COPY-invariant gates pass.

## Verified by running

- All 8 cross-builds static and correct-architecture; both arches of the
container image built, exported and run under QEMU, serving
`/api/health` and `/api/nodes` against a 2.9M-observation production
snapshot.
- The `migrate` binary repairing that snapshot's duplicate on bare
Alpine.
- `CGO_ENABLED=0` producing a binary that builds and then fails on first
query.

## Not verified

- ~~The 2–2.3× figures come from a standalone harness, not this load
under the old driver.~~ **Closed** by @efiten's production run above,
which also corrected the multiplier.
- SQLite 3.46.0 → 3.53.4 query-planner differences on queries with no
total `ORDER BY`.
- Sustained live ingest through the new writer DSN, and the duplicate
collapse against a database an ingestor is actively writing to. Verified
against a static snapshot only, and the collapse is measured at 4.1s on
2.4M synthetic rows with 5 duplicates — well short of an 11M-row
instance. @efiten has offered a staging instance taking real MQTT
traffic; **this is the item to close before the PR leaves draft.**

An earlier revision of this branch shipped the dedup merge in the wrong
direction with a green test suite, and review then found three more
things in the same file: the repair gated on an error string, a
non-atomic TEMP table drop aimed at the wrong connection, and a deletion
whose only record was a row count. All fixed in ac7e8d38. Passing tests
did not establish safety here, which is why the deletion path wanted a
second pair of eyes rather than a rubber stamp.
v3.11.0
2026-09-16 09:02:13 +02:00
Sylvain Rabot cb994d9f9a feat(home): link My Mesh node cards to the node page (#2027)
## What

The node cards in the **My Mesh** grid on the home page could open the
full
health panel or the node's packets, but there was no way to reach the
node
detail page from them — you had to go search for the node again.

Each card now leads with a **Node page →** button that navigates to
`#/nodes/<pubkey>`, the same route the channels, live and analytics
pages
already link to.

## Details

- Wired through the existing `.mnc-btn` click delegation, so it inherits
the
`stopPropagation()` that keeps the card's own click-to-health handler
from
  firing as well.
- The error-state card (health fetch failed, including the 404 *"waiting
for
first advert"* case) gets the button too, on its own actions row. *Full
health* and *View packets* stay off that card — the health fetch is
exactly
what failed, but the node page still resolves for a node that has so far
only
  been seen in channel messages.
- `.mnc-actions` now wraps, so three buttons don't overflow a narrow
card.

## Testing

`node --check public/home.js` passes. Lint and the Playwright suite were
not
run: the worktree this was written in has no `node_modules`. The
existing home
e2e test targets `.mnc-btn[data-action="health"]` specifically, so the
new
button does not disturb it.
2026-09-14 11:13:32 +02:00
efitenandClaude Opus 5 5d3af168b3 fix(live): keep the space the user is typing in the node filter (#2028)
## Problem

Typing a node name with a space slowly into the Live page node filter
glues the words together: "Dan's Local" ends up as `Dan'sLocal`, and
`/api/nodes/search` then returns no suggestions.

The debounced input handler commits the trimmed value (`public/live.js`
`applyFilterFromInput`, ~1791), so after "Dan's " the filter key is
`Dan's`. `setNodeFilter` calls `updateNodeFilterUI`, which wrote
`nodeFilterKeys.join(', ')` back into the input whenever it differed
from the raw input value (~2912). That dropped the trailing space the
user had just typed, and the next keystrokes were appended to `Dan's`.

Found while reviewing #2026.

## Change

`updateNodeFilterUI` no longer writes into the node filter field while
it has focus, and otherwise only when the trimmed input differs from the
keys. Besides the typing debounce, it runs for every matching live
packet, which could also eat a character typed inside the debounce, and
it replaced a picked suggestion's name with the pubkey. Restores from
`?node=` or localStorage (field not focused) still write the keys.

## Tests

- `test-live.js`: a trailing space typed into a focused field is kept
(fails on master); a focused field that differs from the keys is not
overwritten; an unfocused field with different text is; an unfocused
field that differs only by whitespace is not.
- Mutation: each of the three guards (focus, trimmed comparison, write
when different) fails one test on its own.
- `node test-live.js`: 100 passed. `sh test-all.sh`: all standalone
frontend suites pass. `npx eslint public/live.js`: 0 errors.

## Browser validation

Local server with the e2e fixture, headless Chromium, typing `Dan's `
(60 ms per key), a 500 ms pause, then `Local`:

| Build | Input | Stored filter | Suggestions |
|---|---|---|---|
| this branch | `Dan's Local` | `Dan's Local` | `Dan's Local Repeater` |
| master | `Dan'sLocal` | `Dan'sLocal` | none |

## Not verified

- Other browsers than Chromium, and mobile keyboards with autocorrect.
- A trailing space typed and left there: the filter key stays trimmed
while the input keeps the space, which is the intended difference.



## Review follow-up (commit `4e86171d`)

An independent review found the change correct but pointed at the same
bug on a second path, plus a weak test:

- **Packets during typing.** `updateNodeFilterUI` also runs for every
matching live packet (`public/live.js` ~3393). A packet arriving inside
the 200 ms debounce still rewrote the field: with filter `ab12` and the
user typing `3`, the `3` was lost. The write is now skipped while the
input has focus.
- **Picked suggestion.** The same write replaced the name
`selectSuggestion` had just put in the field with the node's
64-character pubkey. With the focus guard the field keeps the name; the
filter key is still the pubkey.
- **Tests.** The old "writes a different key" test started from an empty
input, so a check that only writes into an empty field passed too. The
tests now cover a focused input that differs (not overwritten), an
unfocused input with different text (overwritten) and an unfocused input
that differs only by whitespace (not overwritten). Each of the three
guards was mutated on its own and fails one test. `test-live.js`: 100
passed; `sh test-all.sh`: all standalone suites pass; eslint: 0 errors.

Browser, local server with the e2e fixture: typing `Dan's ` then `Local`
keeps `Dan's Local` with one suggestion; picking it shows `Dan's Local
Repeater` while the stored filter is the pubkey; reopening
`#/live?node=<pubkey>` shows the pubkey in the field, as before.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 23:13:27 +02:00
efitenandClaude Opus 5 52b9474d7a feat(map): filter repeaters by region name (#1862) (#2022)
Fixes #1862

## What

Adds a **Region Scope** picker to the map controls: pick `#be` and the
map keeps the nodes that declare `#be` or were seen carrying `#be`
traffic. It combines with the #2006 scope-state filter and persists in
localStorage the same way. While a region is picked, a small "Region:
#be · reset" chip sits on the map itself, so the filter stays visible
when the controls panel is collapsed (the default on phones) and can be
cleared from there.

## API

`/api/nodes` and `/api/nodes/{pubkey}` gain two fields on repeater/room
rows:

- `declared_regions`: named regions from the node's newest
declared-regions answer, split by the same function the Scope Audit now
uses for `declaredRegions` (`splitDeclaredRegions`,
`cmd/server/scope_config_state.go`), so both pages list a repeater under
the same names. `[]` means it answered and named no region. Absent means
it never answered, other roles, no declared-regions source, or the
declared-regions lookup failed.
- `declared_regions_truncated`: present, and `true`, only when that
answer was flagged as truncated, so the list is partial. Never `false`:
the `nodes.configured_scope` source does not record truncation, so
absence does not mean the list is complete.

The observed side reuses `transported_scopes`. Documented in
`docs/api-spec.md` and the served OpenAPI spec.

### Why no `?hashRegion=` query parameter
The observed side lives in the in-memory store. Filtering it after the
SQL `LIMIT`/`OFFSET` would corrupt `total` and paging, and the map pages
through `/api/nodes`. Same reasoning as the Data path section of #2001.

## Map behaviour

- Filtering is client-side over the nodes `fetchAllNodes` already
loaded: no new request. One pass over the loaded nodes to build the
picker counts, one Set lookup per node per render. The marker filter is
`nodePassesMapFilters` (`public/map.js:219`) and the observer stand-down
`observerLayerShown` (`:212`), both exported and tested.
- The picker and hint count only nodes with a map position, the same
test the marker filter applies first, so a count never promises markers
the map cannot draw.
- The observer layer stands down while a region is picked, for the same
reason it does for the scope-state filter.
- The popup lists declared and observed regions separately. A truncated
declared answer carries the same `truncated` badge the Scope Audit
shows.
- Absence is not read as a finding: the hint under the picker says a
node left off the map is not proof it lacks the region.

## Tests

- Go: `node_declared_regions_api_test.go` covers `declared_regions` on
list and detail endpoints, the no-source case, agreement with
`/api/scope-audit`, `declared_regions_truncated` (truncated,
truncated-empty, complete, configured_scope-only, newer untruncated
answer, companion) and the OpenAPI schema.
- JS: `test-issue-1862-map-region-filter.js` (30 tests) covers the pure
pieces (evidence, counts, options, hint, popup rows,
`nodePassesMapFilters`, `observerLayerShown`) and, at page level, runs
the registered map page through `init()` and `loadNodes()` in a vm
sandbox: picker built from loaded nodes, markers filtered by a stored
region, observer pins standing down, popup rows, the change handler
persisting, and the chip showing, resetting and rendering its text as
text.
- Mutation-checked: removing the region check, the observer stand-down,
the picker build in `loadNodes`, the popup rows, the persist on change,
the chip reset, or the truncated flag (Go or JS) each fails a test.
`test-issue-2001-map-scope-state.js` still passes.
- `go test ./...` in `cmd/server` passes; `check-css-vars` and
`check-xss-sinks --diff` are clean.

## Staging validation
Build `c646310f`, Chrome, no console errors:
- Picking `#be`: chip "Region: #be · reset" at the top of the map, hint
"221 nodes with a map position have evidence for #be: 129 declare it,
191 seen carrying its traffic. Absence here is not proof: ...".
- Clicking reset: picker back to "All regions", stored choice cleared,
chip hidden.
- First version, same instance: `declared_regions` on 136 of 500
`/api/nodes` rows; returning to "All regions" blocked the main thread
5.4 s against 3.2 s for the existing Status filter returning to "All".
The review measured the region filter's own added work at about 0.08 ms
per render plus popup rows that are already built for every marker, so
most of that time is the existing full re-render.

## Not verified
- The chip placement at phone width, and on narrow desktop widths where
it may sit under the expanded controls panel.
- The truncated badge with real truncated answers (staging has none
today).
- `test-map-clustering.js` has one failing test on `upstream/master`
too; untouched here.
- Marker badges from the original issue body are not implemented; the
popup rows are the per-node display.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:09:51 +02:00
efitenandClaude Opus 5 dda04ae918 fix(packets): empty the observer and type selections on Clear Filters (#2012) (#2015)
## What was wrong

On the Packets page, Clear Filters reset `filters.observer`,
`filters.type`, localStorage and every checkbox in both multi-select
menus, but not the Sets that hold the selection: `selectedObservers`
(`public/packets.js:1793`) and `selectedTypes`
(`public/packets.js:1845`). The next `change` event
(`public/packets.js:1826`, `:1876`) added to the stale Set, so select
observer A, Clear, select observer B wrote `A,B` to the URL and the
trigger read "2 Observers". Types behaved the same way. Clear also
unchecked the "All Observers" / "All Types" rows although no filter was
active.

## What changed

- `public/packets.js:1999-2006`: the Clear handler empties both Sets and
rebuilds both menus through `buildObserverMenu()` / `buildTypeMenu()`
plus `updateObsTrigger()` / `updateTypeTrigger()`, replacing the
hand-written checkbox and trigger resets. All four functions are in the
same scope as the handler.
- `test-issue-2012-clear-filters-selection.js` (new): runs the real
multi-select section and Clear handler from `packets.js` in one function
scope against a small fake DOM. For observers and types it drives select
A, Clear, select B and asserts that only B is selected (filters,
localStorage, trigger text) and that the All row is checked after Clear.
- `test-clear-filters.js`: the handler body now uses names this test did
not provide, so menu-agnostic cases get empty stand-ins. The old
"unchecks every checkbox" case expected the All row to be unchecked (the
bug), so it now asserts that the Sets are emptied and both menus
rebuilt.
- Registered the new test in `test-all.sh` and the unit-test step in
`.github/workflows/deploy.yml`.

## Tests

- `node test-issue-2012-clear-filters-selection.js`: 4 passed. With the
`packets.js` change reverted: 0 passed, 4 failed (`obsA,obsB`, `4,5`,
All row `false`).
- `node test-clear-filters.js`: 6 passed, 2 failed, the same counts as
on master. The two failing `updatePacketsUrl` cases fail on master with
`location is not defined`. This file is not run by `test-all.sh` or CI.
- Also green: `test-packets-local-channels.js`,
`test-issue-1415-packets-layout.js`, `test-frontend-helpers.js`,
`test-observer-iata-1188.js`, `test-packet-filter.js`,
`test-packet-filter-ux.js`, `test-xss-escape-sinks.js`,
`scripts/check-css-vars.js`.

## Browser validation

Deployed together with #2013's, #1851's and #1868's branches to a
staging instance with live traffic (build `e84d2da6`), in Chrome:

- Observers: select BE-BRU-Moris, Clear (URL back to `#/packets`,
trigger "All Observers", All row checked), select BE-BRU-Bécodok: URL
holds only Bécodok's key, trigger shows Bécodok.
- Types: select one type, Clear (All row checked, "All Types"), select
another: stored filter holds only the second.

## Not verified

- Other ways of resetting filters (navigating back to `#/packets`
without params) were not checked for the same stale Sets.
- The full `test-all.sh` run was not done locally.

Fixes #2012



## Review follow-up (commit `4b14f54e`)

An independent review of this PR reproduced the bug on master and
confirmed the fix in headless Chromium against a real server. It found
one related pre-existing problem, now fixed here:

- The type multi-select change handler never called
`updatePacketsUrl()`, which is what shows or hides the Clear button
(`public/packets.js` ~779-783). With only a type selected the button
stayed hidden, so the Types half of #2012 could not be reached by a
click. The handler now makes that call, like the observer handler does
(`public/packets.js:1887`). Type is not part of the URL, so the call
only toggles the button.
- The regression test now runs the real `updatePacketsUrl()` and adds
two cases: picking only a type, and only an observer, shows the Clear
button, and Clear hides it again. The type case fails without the new
call.
- The test's fake DOM also provides `#observerList` and
`#observerSearchInput`, so it keeps working if #1884 merges first.
Checked against a local merge of #1884: 6 of 6 pass. That merge has one
conflict in the Clear handler; whichever PR lands second must keep both
the Set clear and the search box reset.
- Not mentioned before: the fix also resets the type trigger's `title`
tooltip through `updateTypeTrigger()`, which the old handler left stale.

Validated on staging together with the other follow-ups (build
`c646310f`). Not verified: mobile viewport, full Playwright suite.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:09:31 +02:00
efitenandClaude Opus 5 efdb3ea0b3 feat(analytics): retransmission pressure over time (#1699) (#2023)
## Summary
Adds `GET /api/analytics/retransmissions` and a "Retransmission Pressure
(proxy)" chart on the Analytics Topology tab, implementing the metric
agreed in #1699: for each flood, the number of distinct repeaters in the
union of the paths of all its observations (`[A]`, `[A,B,C]`, `[A,D]`
gives 4), averaged per time bucket.

Topology is the tab that already shows hop counts and repeaters in
paths, so the chart sits there instead of in a new tab.

## Definition
- Flood routes only (`route_type` 0/1); TRACE excluded. Direct routes
carry the route still to travel (firmware
`src/Mesh.cpp:78-106,334-342`), zero-hop sends are direct
(`src/Mesh.cpp:717-737`), TRACE path bytes are SNR values
(`src/Mesh.cpp:59-61`, refused by `sendFlood` at
`src/Mesh.cpp:637-641`). Firmware commit 0679dbef.
- **Flood events, not hashes.** `transmissions.hash` is UNIQUE and the
packet hash excludes the path (`src/Packet.cpp:41-50`), so when the same
bytes flood again the observations land on the same transmission.
Observations are sorted by time and split into events wherever two
consecutive observations are more than 5 minutes apart. Each event is
counted on its own and bucketed by its first observation.
- Why 5 minutes: a node holds a flood for at most 32 s
(`src/Dispatcher.cpp:11,243-251`) plus a random retransmit delay. On
live over 7 days, 72,806 of 74,347 flood transmissions span 60 s or
less, and of 1,372,283 consecutive observation gaps, 52 fall between 60
s and 300 s against 1,823 above 300 s.
- Events that start before the store retention floor (now minus
`retentionHours`) are left out for every request shape. The store keeps
older observations only for hashes heard again recently, so they do not
represent that period. Eviction of those transmissions is tracked in
#2024.
- A flood event heard only with an empty path counts as 0 repeaters.
- **Prefixes are not resolved to nodes, and a prefix counts once per
event**, whether it repeats across observations or inside one path. On
live (7 days), a repeated 2-byte prefix inside one path occurs in 1.08%
of flood transmissions and 6,178 of 6,596 such repeats match exactly one
known node; for 3-byte it is 0.69% and 104 of 104. That is one node
forwarding again after its 160-slot cyclic duplicate filter dropped the
hash (`src/helpers/SimpleMeshTables.h:9,52-57`). A repeated 1-byte
prefix (44.9% of 1-byte transmissions) is mostly two nodes; counting it
once keeps the value a lower bound. `summary.one_byte_packets` reports
how many events that affects.
- Observations are stored once per observer and path per hash, so a
later event of the same hash only holds pairs not stored before; its
count is a lower bound too. On live these are 1,466 of 75,356 events
(1.9%), and they are kept in the average.
- Resolution was not used: on live, 1-byte observations nearly all have
`resolved_path` NULL, and cold load refuses context-based resolution of
history (`cmd/server/neighbor_persist.go:155-168`).
- Buckets `5m|15m|1h|6h|1d`. `region` filters on observers like
`/api/analytics/rf`, after the event split; a region with no known
observers is not filtered, the same as the other analytics endpoints.
`area` is not supported.

## Implementation
- `cmd/server/retransmission_pressure.go:255` `addPath`: scans path JSON
directly into a generation-stamped hash set, no allocation per
observation.
- `cmd/server/retransmission_pressure.go:367`
`computeRetransmissionPressure`: one pass under `s.mu.RLock`. Per flood
transmission it sorts the observations by cached parsed time into a
reused scratch slice, splits events and counts each in `addEvent`
(`:319`). O(T + O log k + H).
- `cmd/server/retransmission_pressure.go:470`
`GetRetransmissionPressure`: default shape from the recomputer (#1659
warm-up gate). Other shapes come from a typed TTL cache (max 64 entries)
cleared on new paths and eviction (`cmd/server/store.go:2289,2336`);
concurrent misses on one key share one compute through singleflight
(`store.go:204`).
- `cmd/server/retransmission_pressure.go:515` handler,
`cmd/server/routes.go:331`, `cmd/server/openapi.go:108`,
`docs/api-spec.md:1283`.
- `public/analytics.js:761` card, `:855` `renderRetransmissionChart`
(CSS variables only, lines break at missing buckets, caption states it
is a proxy, names the observer coverage bias, the once-per-flood prefix
rule and the 5 minute event split), `:829` loader with stale-response
guard.

## Performance
- `BenchmarkComputeRetransmissionPressure`, 50k transmissions x 20
observations, `-cpu 1`, i5-1335U: median 161 ms/op (132 ms/op before the
event split); first pass after startup with timestamps not yet parsed
192 ms/op. About 22 KB and 281 allocations per op.
- On staging the default shape is served from the recomputer in 0.3 s;
the post-load recompute of this recomputer took 994 ms on a
121k-transmission store (log line quoted in #2025). A 336h store would
be about twice that, every recompute interval, under the store read
lock.

## Tests
- `cmd/server/retransmission_pressure_test.go`: union counting (reporter
example, overlaps, once per event for 1/2/3-byte, width/case, growth);
route/TRACE/zero-hop filter, bucketing, window by event start; event
split and the 5 minute settle gap (boundary, chained steps, unsorted
input), retention floor; region filter, region applied after the split,
unknown region, 1-byte share; recomputer read, TTL cache invalidation on
new paths and on eviction, cache expiry, singleflight, recomputer gate
wiring, handler, warm-up gate.
- `test-issue-1699-retransmission-chart.js` (33 tests, registered in
`test-all.sh` and `deploy.yml`).
- Mutation-checked: 18 mutations of the event split, floor, prefix rule,
bucketing, region order, cache clears, expiry, gate wiring and
singleflight, all killed.
- `go test ./...` in cmd/server passes; `scripts/check-css-vars.js` OK.

## Staging validation
Build `c646310f` (this rework plus #2025 and the other review
follow-ups), after a container restart and full load: default shape
74,974 flood events, average 27.08 repeaters, 169 hourly buckets from
2026-09-06 16:00 (the 168h floor) to the current hour, highest hourly
average 53.3. Before the rework the same instance showed buckets back to
2026-07-18, averages up to 148, and for the first minutes after a
restart only 5,911 packets.

## Merge order with #2025
#2025 fixes the recomputer startup for all analytics endpoints (the
stale first snapshot seen here). Whichever of the two merges second has
to add `recompRetransmissions` to `analyticsRecomputersLocked`, wire it
to that PR's `loadedGate` instead of `LoadComplete`, bump the recomputer
count in `TestAnalyticsRecomputers_PostLoadOrder` from 9 to 10, and make
`TestStartAnalyticsRecomputers_RetransmissionsGatedOnLoadComplete` call
`signalStartupLoadDone()`. That resolution is what ran on staging.

## Not verified
- Recompute timing on a production-size (336h) store; only extrapolated.
- Phone-width layout and dark theme of the reworked chart.
- E2E Playwright suite.

Fixes #1699

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:37:14 +02:00
914bd4cf0e feat(channels): show each message's region (#1851) (#2018)
## Show each channel message's region scope

Each message in the Channels view now shows the region scope it was sent
with, as a small chip in the meta line: the region name (for example
`#be`), `unknown scope`, or nothing.

This is the channel-message part of #1852 by @dborup, extracted as a
focused change. #1852 was closed unmerged because it had grown to the
whole fork diff. The implementation follows dborup's commits c686ae3f,
350bf7ee and a94d57ed on dborup/CoreScope, adapted to current master.
dborup is co-author on the commit.

### What changed
- `cmd/server/db.go:2099,2195`: `GetChannelMessages` selects
`t.scope_name` when the column exists and returns it as `scope_name`.
- `cmd/server/store.go:5654`: the in-memory `GetChannelMessages` returns
`scope_name`, so the field does not depend on which path serves the
endpoint.
- `cmd/server/store.go:2966,3244`: both WebSocket broadcast builders
carry `scope_name`, so a live message shows its region immediately.
- `public/channels.js:342,2278`: `messageScopeChipHtml` renders the chip
with the existing `.sa-chip-declared` / `.sa-chip-unmatched` styles from
`scope-audit.css`. The name goes through `escapeHtml`. No new CSS.
- `public/channels.js:678,695,1435,1487`: the decrypt path and the
WebSocket path keep `scope_name` on the message.

### Differences from #1852
- The field is `scope_name`, the name `/api/packets` already uses.
- No `routeType` field. `transmissions.scope_name` already tells the
states apart: NULL means no transport code, an empty string means a
transport code that no configured region key matched. The frontend uses
`??`, not `||`, so the empty string is kept.
- A chip instead of `scope: <name>` text. The area label from later
#1852 commits is not included.

### Perf
One extra column per observation row in the page query (at most `limit`
transmissions), and one extra map entry per broadcast observation. No
new queries, loops or API calls.

### Tests
- `cmd/server/channel_message_scope_name_test.go`: the three states
through the DB query, the store, `/api/channels/{hash}/messages` over
both paths, a schema without the column, and both broadcast builders. 5
of its 6 tests fail without the change; the sixth guards the
missing-column case and passes either way.
- `test-issue-1851-channel-message-scope.js`: the REST, WebSocket and
client-side decrypt paths, escaping, and the name / unknown / none
render. 4/4 fail without the change. Registered in `test-all.sh` and the
unit step of `deploy.yml`.
- Mutation checks: returning `nil` for `scope_name` in the DB path fails
the DB and endpoint tests; `||` instead of `??` in the WebSocket path
fails the WebSocket test.
- `go test ./...` in `cmd/server`: ok. gofmt and go vet clean.

### Browser validation
On a staging instance with live traffic (build `e84d2da6`), in Chrome:
- `/api/channels/{hash}/messages` carries the `scope_name` key on every
message in the 19 channels whose results I read. `#hamradio`, latest 50:
34 named, 1 empty string, 15 NULL.
- Opening `#hamradio` renders 104 chips: `#nl` 53, `#be` 32, `#de` 11,
`#eu` 6, `#bx` 1 and `unknown scope` 1, and no chip on unscoped
messages. Chip text `rgb(26, 26, 46)` on `rgb(238, 242, 255)` in the
light theme.

### Not verified
- Dark theme not checked.
- The real-decrypt branch of `decryptCandidates` has no test and was not
exercised in the browser; the already-decrypted branch is tested.
- Messages already in the client decrypt cache show no chip until they
are decrypted again.
- `go test -race` and the Playwright E2E suite were not run locally.

Fixes #1851



## Review follow-up (commit `50346589`)

An independent review found no correctness or XSS problem and confirmed
DB, store and WebSocket agree on the value. Changed:

- **Real decrypt branch tested.** A new test runs the real AES+HMAC
decrypt branch in `decryptCandidates` with one packet per scope state;
deleting `scope_name` there now fails 2 of 7 tests.
- **Tooltip wording.** The unknown-scope tooltip now says the scope
"could not be matched to a single region on this instance"
(`public/channels.js:338-347`). The ingestor stores an empty name both
when no key matches and when several match without exactly one
operator-configured key (`cmd/ingestor/region_keys.go:364-393`), so
"matches none of the configured keys" was wrong for the second case.
- **Old decrypt cache.** Decrypted messages cached before this change
had no `scope_name` key and stayed chipless as long as the candidate
count did not change. A cached message missing the key now forces one
full decrypt; a cache that has it still takes the delta path. A test
covers each case.
- **Docs.** `docs/api-spec.md` documents `scope_name` on the channel
messages response, with the null / empty string / name semantics.

Corrections to the description:

- **Test counts.** With the `db.go` and `store.go` changes reverted, 4
of the 5 top-level Go tests fail (6 of 7 counting subtests); only the
missing-column test passes.
- **Broadcast payload.** `scope_name` is added to `pkt`, which is copied
into `broadcastMap` and also nested as `packet` (`store.go` ~2974-2980,
~3252-3257), so the key appears twice per observation: 36 bytes for
`null`, 48 bytes for `"#belgium"`.
- **Side effect on the Packets page.** The live table reads
`m.data.packet` (`packets.js` ~1316-1318), so flat rows and expanded
group children now show Scope for live packets. In grouped mode a new
group copies a fixed field list without `scope_name` (~1384-1395) and
shows the empty placeholder until reload. Before this PR every live row
showed that placeholder, so this is not a regression.

The three copies of the three-state scope rendering (`app.js`,
`packets.js`, `channels.js`) are left as they are.

---------

Co-authored-by: dborup <3627142+dborup@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:36:34 +02:00
2c6d7bb6ae feat(packets): add filter to All Observer dropdown (#1884)
Lets users type a prefix to filter the observer checkbox list, instead
of scrolling a long list.

---------

Co-authored-by: efiten <erwin.fiten@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:36:13 +02:00
efitenandClaude Opus 5 a059299588 feat(node-analytics): hop-count statistics per node (#1812) (#2021)
## Summary
Adds per-node hop-count statistics so repeater operators can choose
`flood.max`, `flood.max.unscoped` and `flood.max.advert` from what their
node actually sees.

- New endpoint `GET /api/nodes/{pubkey}/hop_analytics?days=N`
(`cmd/server/routes.go:299`, `cmd/server/node_hop_analytics.go:312`),
separate from `/analytics` as requested in the issue.
- New card "Hop Count at This Node" on the node analytics page
(`public/node-hop-analytics.js`, wired at
`public/node-analytics.js:130,174`): histogram of hop counts with a box
plot on the same x axis, filters `flood.max` (default),
`flood.max.advert`, `flood.max.unscoped`, driven by the existing range
picker.
- The existing "Hop Distribution" chart is unchanged: it shows path
length at the observer, a different quantity.
- No `direct` tag, although the issue lists one: for DIRECT packets the
path is the remaining route and no flood limit applies, so there is no
hop count to report.

## Hop count definition (firmware 0679dbef)
- `src/helpers/RoutingPolicy.h:15-21`: limits compare
`getPathHashCount()`; `.unscoped` applies to route type FLOOD, `.advert`
to adverts.
- `src/Mesh.cpp:344-350`: `routeRecvPacket` checks with n hashes in the
path, then writes its own hash at index n. So hops = the node's
zero-based index in the path, no +1.
- `src/Mesh.cpp:265-285`: a node forwards a flood once;
`src/Mesh.cpp:651,680`: an originator never forwards its own flood.
- DIRECT packets are excluded: their path is the remaining route
(`src/Mesh.cpp:78-103,334-341`).

Response: `{timeRange, packets: [{hash, timestamp, hops, tags}],
ambiguous}`. Tags: `flood`, `scoped` or `unscoped`, `advert`. Documented
in `docs/api-spec.md:679` and `cmd/server/openapi.go:90`.

## Attribution
`cmd/server/node_hop_analytics.go:198-309`. The result depends only on
the observed paths, the prefix map and the neighbor graph, so it is the
same after a restart as after live ingest.

- Every observation of every flood packet in the window is read.
`byNode` holds the server resolver's pick at ingest and other picks
after a cold load; `byPathHop` indexes only each packet's longest path,
which for a busy relay often runs through another branch of the flood.
- A packet counts when the node's prefix sits at exactly one index
across its observations, and either the node is the only relay candidate
for that prefix (`prefixMap.relayCandidates`,
`cmd/server/store.go:6795`), or the hop resolves to the node under the
ingestor's strict rule (`cmd/ingestor/path_resolver.go:143-214`) in at
least one observation and to another node in none. Strict rule: earlier
hops identified without a tiebreak, exactly one candidate adjacent in
`neighbor_edges` to the previous hop (the originator for hop 0 of an
advert), nodes already on the path excluded.
- The server resolver's tiebreaks (affinity, GPS distance, advert count,
pubkey order) are not used.
- Everything else with the node's prefix goes to `ambiguous`. In
practice that is most packets with a colliding 1-byte path hash.

On a read-only 7-day dump of a 1,669-node mesh DB, for one busy
repeater: 23,081 packets attributed, 11,437 ambiguous. Taking candidates
from `byPathHop` instead gave 9,995 attributed, with the histogram mode
moved from 2 to 3-5 hops.

## Performance
Scans `s.packets` under the read lock, no SQL per packet. Per
observation: one substring test for the node's first prefix byte; the
hop scan only for observations containing it; the strict walk only for
colliding prefixes, with per-request caches for candidates and
adjacency. `BenchmarkNodeHopPackets` models one 7-day request at that
scale (73,782 flood packets, 1,430,280 observations): 44-87 ms/op, 13.4
MB, 40 allocs on a throttling laptop.

Response size for that repeater over 7 days: about 23k entries, 2.3 MB
JSON, 375 KB gzipped. `hash` and `timestamp` are 61% of the raw and 91%
of the gzipped bytes; they stay because the issue asks for them so a
client can join entries to packets and bin by time.

## Tests
- Go: `cmd/server/node_hop_analytics_test.go`: 12 unit tests, a
live-ingest test through `IngestNewFromDB` (a colliding prefix without
independent attribution goes to `ambiguous`, not to the node the
resolver picked), live ingest versus cold load of the same DB, route
test, benchmark. 15 mutations of the attribution logic each fail a test.
- JS: `test-node-hop-analytics.js` (filters, histogram, quartiles and
whiskers with a fixture that separates 1.5 IQR from 3 IQR, render),
registered in `test-all.sh` and `.github/workflows/deploy.yml`.
- `gofmt`, `go vet ./...`, `go test ./...` in `cmd/server`,
`scripts/check-css-vars.js` pass.

## Staging validation
Build `c646310f` (this PR's review follow-up together with the other
open follow-ups), after a container restart and full load, on a busy
Belgian repeater:

- `hop_analytics?days=7`: 23,302 packets, 11,548 ambiguous, median 4,
adverts never above hop 7 (matching the firmware default
`flood_max_advert = 8`, `examples/simple_repeater/MyMesh.cpp:922`), 1.2
s. The first version reported 23,035 packets and 86 ambiguous in 534 ms,
because it trusted the resolver's pick for colliding prefixes.
- The card rendered on the first version with no console errors; the
rework does not touch the frontend beyond a test fixture.

## Not verified
- Response time and lock hold for 30 days on the busiest node on a
14-day store.
- Server relay candidates exclude companions and listeners while the
ingestor's prefix index does not, so a few strict attributions can
differ from the ingestor's persisted `resolved_path`.
- Identical numbers across a second container restart were shown in a Go
test, not repeated on staging.
- Dark theme, phone width, and switching the range picker in the
browser.
- Filter state is not reflected in the URL hash (the range picker is not
either).

Fixes #1812

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:59:56 +02:00
efitenandClaude Opus 5 2d4019f719 fix(analytics): recompute once the store has fully loaded (#2025)
Refs #2023, #1659, #1724

### Problem
`main.go:258` waits only for the first load chunk, then `main.go:402`
starts the analytics recomputers. `Start()` computes immediately on that
chunk (`analytics_recomputer.go:86` on master) and the next compute
waits a full interval (`:93`, 5 min default). The chunk loader walks by
ascending id, so that chunk holds the oldest transmissions.

- RF, topology, channels: the #1659 gate checked `LoadComplete()` after
the compute (`analytics_warmup_1659.go:122`). `LoadComplete` flips at
the end of the hot window (`chunked_load.go:489`), before the background
fill (`store.go:1455`), so the gate could open on a snapshot without the
background fill, and the 60 s force timeout (`:73`, `:162`) opened it on
the first-chunk snapshot. In an end-to-end test on master,
`/api/analytics/rf` returned 200 with 8 of 100 packets before the
background fill ran.
- Distance, hash-collisions, hash-sizes, roles, observers-clock-skew,
nodes-clock-skew: no gate, partial snapshot served from the start.
- Distance additionally served a snapshot from the previous index for up
to one interval after each lazy index build.

On a staging instance, a default analytics request returned 5,911
packets with hours-old last buckets until the next recompute (about
74k).

### Change
- `StartupLoadDone()` (`chunked_load.go:108`): closed when
`RunStartupLoad` returns, on every path (`chunked_load.go:202`). Closing
it drops the hash-size info cache (15 s TTL) and the clock-skew engine
throttle (30 s, `clock_skew.go:225`), both read by the post-load
computes.
- `recomputeWhenLoaded` (`analytics_recomputer.go:172`): on that signal,
recompute each recomputer once, sequentially, via `RecomputeNow`
(`:154`), which runs on the recomputer's own loop and restarts its
ticker (`:106`). Order (`:255`): rf, topology, channels, distance,
hash-collisions, hash-sizes, observers-clock-skew, nodes-clock-skew,
roles (roles reads the nodes-clock-skew snapshot). Logs one line with
per-recomputer durations.
- Warm-up gate: now the same signal (`:343-354`), sampled before the
compute starts (`:129`), so a pass that began on partial data never
opens it. 503 + `Retry-After: 5` and the force timeout are unchanged; a
forced-open snapshot is replaced by the post-load recompute.
- Ungated endpoints: no new 503s (their API has none); snapshot replaced
right after the load.
- Distance: the lazy index build refreshes the distance recomputer
before reporting built (`store.go:4476`).
- Recompute intervals and config unchanged.

### Performance
One extra compute per recomputer per process start, run sequentially so
they do not all hold the store read lock at once. Ticker phases
afterwards are offset by the cumulative post-load compute durations
instead of all starting within the first-chunk compute window (relevant
to #1724; the effect on lock waves is not measured).

### Tests
`analytics_recompute_after_load_test.go`: signal open during background
fill, closed after success and failure; cache drops; immediate and
ordered post-load recompute; gate not opened by a pass started before
the load; forced-open snapshot replaced on load; ticker restart;
distance refresh before 202 ends; end to end with recomputers started
before the background fill (RF 503 until load, then `totalTransmissions`
equals the full store; six ungated endpoints 200 during load; all nine
recomputed after load). 9 of these failed on master with stubs; 8
single-line mutations each caught. `go test ./...` in `cmd/server`
passes.

### Staging validation
Deployed together with the review follow-ups of #2015-#2023 (build
`c646310f`), container restart:

```
16:35:20 [store] first chunk ready (chunkSize=10000)
16:35:25 [store] LoadChunked complete ... starting background fill loader
16:36:58 [store] background load complete: 121120/121282 packets in memory (coverage=99.9%)
16:37:03 [analytics-recompute] startup load done: recomputed 10 snapshots in 5.155s (rf=955ms topology=1.684s channels=43ms distance=49ms hash-collisions=30ms hash-sizes=338ms observers-clock-skew=369ms nodes-clock-skew=692ms roles=2ms retransmissions=994ms)
```

Right after that line, `/api/analytics/rf` reported `totalTransmissions`
121,121 against 120,700 packets in memory, and the retransmissions
default shape from #2023 covered the full 7 days. Before this change
both waited for the next 5 minute tick.

### Merge order with #2023
#2023 adds a tenth recomputer. Whichever of the two merges second has to
add `recompRetransmissions` to `analyticsRecomputersLocked`, wire it to
`loadedGate` instead of `LoadComplete`, and change 9 to 10 in
`TestAnalyticsRecomputers_PostLoadOrder`; the retransmissions gate test
then calls `signalStartupLoadDone()` instead of setting `loadComplete`.
That resolution is what ran on staging above.

### Not verified
- Repeater-enrich recomputer and the region/window TTL caches
(hash-collisions region results have a 1 h TTL) may also keep partial
results after the load; not changed here.
- Recompute order is tested structurally, not with roles/clock-skew
data.
- Whether this reduces the #1724 stalls; not measured.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:59:37 +02:00
efitenandClaude Opus 5 7f8f71e773 fix(live): reconnect when the websocket goes silent (#1074) (#2020)
## Problem

#1074 reports that after a proxy dropped the WebSocket, live updates
only came back 8 to 10 minutes later.

The client only reconnects from `onclose` (`public/app.js:791` on
master). A half-open connection (a proxy or NAT dropping state without a
FIN reaching the browser, a laptop that slept) can keep a WebSocket OPEN
for minutes without `onclose`, and nothing retries in the meantime.

The server does ping every 30s (`cmd/server/websocket.go:252` on
master), but ping frames are answered by the browser below the page and
JS cannot observe them. The only app-level frames are packet broadcasts
(`websocket.go:337, 343, 366`), which stop on a quiet mesh. So the
client had no signal to tell a quiet mesh from a dead socket.

## Change

Server (`cmd/server/websocket.go`):
- On the existing ping tick, `writePump` also writes the text frame
`{"type":"heartbeat"}` (`:283`, bytes at `:80`). One 20-byte frame per
client per 30s, from the goroutine that already writes the ping, no hub
lock.
- The interval moves to `Hub.pingInterval` (default 30s, `:118`) so a
test can shorten it.

Client (`public/app.js`):
- Every frame refreshes `wsLastMessageAt` (`:841`). One timer
(`checkWSLiveness`, `:806`) fires at last message + `WS_STALE_MS` (75s,
one late or lost heartbeat of slack) and replaces the socket if it is
still silent. It is armed at socket creation, so a stuck handshake is
covered too.
- `dropWS` (`:796`) detaches the old socket's handlers before `close()`,
so a late close event cannot schedule a second connection.
- `connectWS` (`:818`) cancels a pending reconnect and drops the
previous socket, so the watchdog, resume checks, `onclose` and
pull-to-reconnect cannot stack sockets. Before this, `pullReconnect` on
a non-open socket left a third socket 3s later. The 3s `WS_RECONNECT_MS`
delay after `onclose` is unchanged (`:837`).
- `visibilitychange` (to visible) and `online` run the check immediately
(`:861`), because a hidden or sleeping tab's timers can run late.
- Heartbeat frames are matched by exact bytes (`:842`) and are not
pulsed or dispatched to `onWS` listeners.

Compatibility: tabs loaded before the deploy dispatch heartbeats to
their listeners until reloaded. Every current listener filters on
`msg.type`, so the visible effect is a logo pulse and a `/stats` cache
refresh every 30s.

Perf: one `Date.now()` and one string compare per WS message on the
client; one extra 20-byte write per client per 30s on the server.

## Tests

- `test-ws-stale-watchdog-1074.js`: real `app.js` in a vm with a fake
clock, timers and WebSocket. 12 tests: silence past the threshold
replaces the socket exactly once; a handshake that never opens is
replaced; heartbeats and packet traffic keep the socket; heartbeats are
not dispatched; resume and `online` after silence reconnect immediately,
with recent traffic they do not, and hiding does not trigger a check;
repeated resume events open one socket; after `onclose` only the
reconnect timer is pending; pull-to-reconnect leaves one socket. 9 of 12
fail on master. 12 of 12 source mutations (threshold, reconnect path,
detaching, timer cleanup, resume wiring, heartbeat filter) are caught.
Registered in `test-all.sh` and the deploy.yml unit step.
- `TestWritePumpSendsAppHeartbeat`: fails with a read timeout without
the heartbeat, even with pings every 20ms. `TestHubDefaultPingInterval`
pins the 30s interval that `WS_STALE_MS` assumes.
- `go test ./...` in `cmd/server` passes; gofmt and go vet are clean.

## Browser validation

On a staging instance (build `139e484e`, together with #1979's branch),
in Chrome, no console errors:

- A `{"type":"heartbeat"}` frame arrived on the open socket within the
observation window.
- Silent socket: after `ws.onmessage = null`, the page replaced the
socket after 76.2s (threshold 75s plus a 250ms poll); the old socket
ended in CLOSED, the new one OPEN.
- Normal close: `ws.close()` led to a new OPEN socket after 4.1s, and
exactly one new `WebSocket` was constructed.

## Not verified

- The reporter's proxy setup was not reproduced; that their delay was a
half-open socket is a hypothesis consistent with the symptom. Hence
`Refs`, not `Fixes`.
- Laptop sleep and the `visibilitychange` / `online` resume path were
only covered by the unit test, not in a browser.
- Behaviour under Chrome's intensive background-timer throttling and
mobile tab freezing was not measured; a frozen but healthy tab may do
one unnecessary reconnect on resume.
- Go tests were run without `-race`.

Refs #1074



## Review follow-up (commit `72e5e906`)

An independent review found no blocking bug: all data writes stay on the
write goroutine, pong-based dead-client detection still works, and no
ordering of onclose, watchdog, resume checks and pull ends with two live
sockets or none. It reproduced the silent-socket case in headless
Chromium through a blackholing TCP proxy (replacement 75.0 s after the
last frame). Changed:

1. **Startup wiring tested.** The first resume test now boots through
the page's real `DOMContentLoaded` listeners, so removing
`setupWSResumeCheck()` from startup makes it fail.
2. **Wall clock stepping back.** If the clock steps back after the last
message, the watchdog no longer re-arms for the size of the step (a 1 h
step used to delay detection by about an hour). A negative silence
reading is treated as stale, so the socket is replaced within
`WS_STALE_MS` of the step (`public/app.js:810-814`). A step in either
direction costs at most one extra reconnect on a healthy socket.
`Date.now()` stays the clock so a tab resumed after sleep is still
checked against real elapsed time.
3. **Pull-to-reconnect at once.** On an OPEN socket, pull-to-reconnect
now replaces it through `connectWS()` instead of closing it and waiting
for onclose, which took 63 s on a half-open connection in the review's
measurement (`public/app.js:927-934`). This was slow on master too; it
is safe now that `connectWS()` detaches the old socket.

Tests: 12 to 16 in `test-ws-stale-watchdog-1074.js`;
`test-pull-to-reconnect.js`, `test-pull-to-reconnect-1091.js` and
`test-live.js` pass.

Correction to the compatibility note: tabs opened before the deploy
treat the heartbeat like any other message. Besides the logo pulse,
`app.js` runs `updateNavStats` on every message and invalidates the
cached `/stats` and `/nodes` responses 5 s later; `packets.js` also
pushes every message into `pauseBuffer` unfiltered (~1310-1313), so an
old tab with Packets paused sees its counter rise by 2 per minute.
Cosmetic: heartbeats are filtered out on replay, and a reload ends it.

Not verified: real hidden-tab or mobile freeze behaviour, Firefox and
Safari, and the reporter's proxy setup.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:59:18 +02:00
efitenandClaude Opus 5 0fea3f2a75 feat(analytics): break scope adverts down by node role (#1979) (#2019)
## Summary

Adds a breakdown of flood adverts by sender role to `/api/scope-stats`
and the Scopes tab, in the descriptive shape agreed in #1979: per node
role, how many flood adverts were unscoped, scoped with an unnamed
region, or scoped with a named region. It reports what was sent, not
why.

## Changes

- `cmd/server/db.go:3114-3144`: one grouped query in `GetScopeStats`.
ADVERT packets on flood routes (TRANSPORT_FLOOD 0, FLOOD 1) in the
window, `LEFT JOIN nodes` on `from_pubkey`, split by the three
`scope_name` states (NULL, empty string, name). A missing or empty role
becomes `"unknown"`. Ordered by total descending, then role. Zero-hop
adverts are excluded because firmware sends them as
DIRECT/TRANSPORT_DIRECT (`src/Mesh.cpp:717-730`, `Mesh::sendZeroHop`),
so they would inflate "unscoped".
- `cmd/server/types.go:116-133`: `ScopeAdvertRoleCount` and
`ScopeStatsResponse.AdvertsByRole` (`advertsByRole`, always an array).
- `public/analytics.js:4760`: `scopeAdvertsByRoleHtml` renders a table
under the time-series chart with the count per state and its share of
the row. Role text is escaped. It reuses the existing `/scope-stats`
response, so there is no extra request.
- `docs/api-spec.md:1763-1786`: documents the new field.
`/api/scope-stats` is in `openapi_known_gaps.json`, so there is no
`openapi.go` entry to update.

API addition (existing fields unchanged):

    "advertsByRole": [
{ "role": "repeater", "unscoped": 7741, "unknownScope": 7, "named": 1562
}
    ]

## Performance

The query runs inside `GetScopeStats`, so it shares the existing 30s
cache per window. The unary `+` on `payload_type` keeps SQLite on the
`first_seen` range index. Without it the planner picked the
`payload_type` index and walked every stored advert whatever the window.
Read-only timing on a production DB (1,063,345 transmissions, 161,634
adverts, sqlite3 CLI 3.45.1):

| Window | payload_type index | first_seen index (this PR) |
|---|---|---|
| 7d | 0.231s | 0.059s |
| 24h | 0.214s | 0.008s |
| 1h | 0.213s | 0.001s |

## Tests

- `TestGetScopeStatsAdvertsByRole` (`cmd/server/db_test.go:2295`): the
three states, flood-only routes, non-advert and out-of-window exclusion,
`unknown` for a missing node row, an empty role and a NULL
`from_pubkey`, and ordering. Mutation checked: widening to routes 0-3
and dropping the empty-role fallback both fail it.
- `TestGetScopeStatsAdvertsByRoleEmpty` (`:2372`): empty result is `[]`,
not null.
- `test-issue-1979-scope-adverts-by-role.js`: renders the real
`analytics.js` helper in a vm sandbox. Covers row order, totals and
shares, columns, escaping (mutation checked), the empty state, and the
non-causal caption. Registered in `test-all.sh` and the deploy.yml unit
step.
- `go test ./...` in `cmd/server` passes, gofmt and go vet are clean,
`check-css-vars.js` OK, `check-xss-sinks.sh --diff` exits 0.

## Browser validation

On a staging instance with live traffic (build `139e484e`, together with
#1074's branch), in Chrome, no console errors: `/#/analytics?tab=scopes`
shows "Flood adverts by node role" under the time-series chart with its
caption, and a table of 6 roles for the default window, for example
`repeater 1.361 | 1.109 (81.5%) | 2 (0.1%) | 250 (18.4%)`. Shares in
each row add up to 100%. `/api/scope-stats?window=7d` returns
`advertsByRole` with the same six roles.

## Not verified

- Timings are from the sqlite3 CLI, not the modernc driver in the server
process.
- Role is the sender's current `nodes.role`. A node whose advert type
changed within the window is counted under its latest role. The data
also contains a raw `type-13` role, shown as is.
- Switching the window in the browser was not exercised; the API was
checked for 7d.
- No Playwright E2E added.

Fixes #1979



## Review follow-up (commit `43af46b8`)

An independent review found no correctness, security or performance
problem: counts are per transmission, the window matches the rest of the
Scopes tab, zero-hop adverts are DIRECT per firmware, and the
`+t.payload_type` hint holds on modernc SQLite 3.46.0. It found two test
gaps and a docs gap. Changed:

- **Ordering.** The Go fixture gave companion, repeater and unknown 3
adverts each, so ordering by total was never checked (`ORDER BY COUNT(*)
ASC` still passed). The fixture now has repeater 4, unknown 3, companion
2, sensor 2, so the expected order differs from alphabetical and the
companion/sensor tie checks the role-name tie-break. Reversing the count
order, dropping it, or reversing the tie-break now fails the test.
- **Column positions.** The JS test only checked that each cell string
appeared somewhere in the row, so swapping two columns passed. It now
compares each row's cells and the header cells by position; both swaps
fail.
- **Docs.** `docs/api-spec.md` and the `ScopeAdvertRoleCount` comment
now name every source of `unknown`: a NULL `from_pubkey` (legacy rows
the #1143 backfill has not reached), a sender with no `nodes` row,
including one moved to `inactive_nodes` by node retention (inside the 7d
window only with `retention.nodeDays` below 7), and an empty role.

Not added: a query-plan test pinning the `+t.payload_type` hint. The SQL
is inline in `GetScopeStats`, so the test would have to copy it or the
query would have to move into a constant; left for a follow-up if
wanted. The raw `type-13` role in the table is the ingestor's
placeholder for reserved advert types 5-15
(`cmd/ingestor/decoder.go:1229`, #1279), shown as is.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:59:00 +02:00
3ca176a676 fix(packets): decode CONTROL discover fields for humans (#1868) (#2017)
## What

CONTROL `DISCOVER_REQ` / `DISCOVER_RESP` packets on the Packets page now
show their fields in a readable form:

- Node type and the `DISCOVER_REQ` type filter bitmask render as
Companion / Repeater / Room Server / Sensor instead of a number or hex.
- SNR renders in dB (wire value / 4, sign kept: `-11` shows `-2.75 dB`)
instead of the raw wire byte.
- The responder pubkey renders as the node name (detail header, and the
row preview once the node index is loaded) or a link to the node (detail
field table) when the node is known, and as the first 8 hex chars when
it is not. The full key no longer appears in the row.
- The detail field table has Subtype / Type Filter / Tag / Since and
Node Type / SNR / Tag / Public Key rows with byte offsets, instead of
one generic `Raw` row. Unknown sub-types keep a `Raw` row.

## Credit

Port of @dborup's fix in dborup/CoreScope@ad1680ca (merged on their fork
as 601f1e4), adapted to current master. dborup is co-author on the
commit.

Differences from that commit:
- Name lookup uses the existing bulk node index (`HopResolver`) instead
of `/api/nodes/{pubkey}` on every opened packet, per AGENTS.md rule 10.
New `HopResolver.nodeForKey` (`public/hop-resolver.js:357`) does an O(1)
lookup by full key or 8-byte prefix; an ambiguous prefix returns null.
- `renderDetail` loads that index for CONTROL packets with a pubkey
(`public/packets.js:3293`), because zero-hop packets have no path to
trigger it.
- The ADVERT app-flags row and CONTROL share one type label map
(`advTypeLabel`, `public/packets.js:2992`).

## Firmware references (meshcore-dev/MeshCore @ 0679dbef)

- `docs/payloads.md:259-282`: field layout; DISCOVER_RESP `snr` is
"signed, SNR*4"
- `examples/simple_repeater/MyMesh.cpp:798-799`: sub-type values `0x80`
/ `0x90`
- `examples/simple_repeater/MyMesh.cpp:817`: `filter & (1 <<
ADV_TYPE_REPEATER)`
- `examples/simple_repeater/MyMesh.cpp:820-821`: node type in low
nibble, `data[1] = packet->_snr`
- `src/Dispatcher.cpp:206`: `_snr = getLastSNR() * 4.0f`
- `src/Packet.h:51,92`: `int8_t _snr`, `getSNR()` returns `_snr / 4.0f`
- `src/helpers/AdvertDataHelpers.h:7-11`: `ADV_TYPE_*` values

`cmd/ingestor/decoder.go:750-803` (`decodeControl`) already emits
`ctrlSNR` as `int(int8(buf[1]))` and `ctrlPubKey` as 64 or 16 hex chars,
so this is display-only. No backend change.

## Performance

Row preview adds one object lookup per CONTROL row (O(1)). The 8-byte
prefix index adds one entry per node at `HopResolver.init`. No new API
calls.

## Tests

- New `test-issue-1868-control-decode.js` (14 cases: type names, filter
bitmask, SNR 16 and -21, known full key, known 8-byte prefix, unknown
key, escaping of node names). Registered in `test-all.sh` and the unit
step of `.github/workflows/deploy.yml`.
- Written first: 14/14 failed before the change, 14/14 pass after.
- Mutation checks: removing `/ 4` fails 3 cases; treating the byte as
unsigned fails 2; disabling the prefix index fails 1.
- `test-packets.js`: CONTROL DISCOVER_RESP assertion updated from full
pubkey to 8-char prefix. Its other 13 failures are unchanged from
master.

## Browser validation

On a staging instance with live DISCOVER traffic (build `e84d2da6`), in
Chrome, no console errors:

- `4d67731357dac2ce`, raw payload `92 39 80 6C B8 41 CC 14 67 5D ...`:
the field table shows Subtype `DISCOVER_RESP` (flags `0x92`), Node Type
`Repeater`, SNR `14.25 dB` (wire `57 / 4`), Tag `0x41B86C80`, Public Key
linked as `BE-JBE-Permeke`.
- `99c6121e3fca3782`: SNR `-2.75 dB` (wire `-11 / 4`), Public Key
`BE-RIN-SPECTRUM-ESP-01`, link target `#/nodes/6b8c65b7...`, which
matches that node's key in `/api/nodes`.

## Not verified

- The row preview shows the 8-char prefix until the node index has
loaded, which a page listing only zero-hop packets does not trigger.
Seen on staging: the row kept `pubkey=cc14675d` while the detail panel
showed the name.
- `DISCOVER_REQ` packets were not opened in the browser; only the unit
test covers them.
- eslint and the full `test-all.sh` were not run locally.

Fixes #1868



## Review follow-up (commit `ee1a3887`)

An independent review re-derived every DISCOVER field and offset from
the firmware and found them correct, but found untested offsets and
three small display gaps. Changed:

- **Bytes past the last decoded field are kept.** The CONTROL field
table now adds a Raw row at the real offset for any payload bytes after
the last decoded field (`public/packets.js:3802`, `:3844-3846`). Before,
a truncated DISCOVER_REQ such as `80 04 12`, or a DISCOVER_RESP with a
partial or oversized key, dropped those bytes, while master showed them
in a generic Raw row. UNKNOWN subtypes use the same path now.
- **Unknown filter bits are visible.** Type filter bits outside ADV_TYPE
1..4 (ADV_TYPE_NONE and the reserved 5..15 range,
`src/helpers/AdvertDataHelpers.h:7-12`) are shown as hex next to the
names: `filter=Repeater+0x20` in the row, `Requesting: Repeater +0x21`
in the detail.
- **prefix_only and since=0.** The DISCOVER_REQ Subtype row decodes
`prefix_only` (flags bit 0, `docs/payloads.md:270`,
`examples/simple_repeater/MyMesh.cpp:818`), and `since=0` renders as "0
(no filter)" in the detail (`MyMesh.cpp:811-817`).
- **Unknown responder key in full.** A key with no known node now
appears in full (64 or 16 hex) in the field table; the row preview keeps
the 8-character prefix #1868 asked for.
- **Tests: 14 to 35.** They check the Offset column of every REQ and
RESP row, the since, prefix_only and key-length labels, the UNKNOWN and
truncated or oversized payload cases, an ambiguous 8-byte prefix in
`HopResolver.nodeForKey`, and that `renderDetail` loads the node index
for a zero-hop DISCOVER_RESP. Every mutation listed in the review
(offsets, Since row, Raw row, labels, resolver loading, ambiguous
prefix) now fails the suite, and so do 7 more on the new code.

Firmware references re-checked at `0679dbef`. Not verified: how the
wrapped 64-hex key looks in the detail pane.

---------

Co-authored-by: dborup <3627142+dborup@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:57:55 +02:00
efitenandClaude Opus 5 5efe61eef2 fix(ingestor): set an explicit MQTT ClientID per source (#2013) (#2016)
## What

`buildMQTTOpts` (`cmd/ingestor/main.go:591`) never called `SetClientID`,
so with paho.mqtt.golang v1.5.0 every ingestor connected with a
zero-length ClientID and `CleanSession=true`. The session identity then
depended on the broker.

This PR:

- adds an optional `clientId` per `mqttSources` entry
(`cmd/ingestor/config.go:29`)
- when unset, uses `corescope-<name>-<6 hex chars>`
(`cmd/ingestor/main.go:653`). The name is reduced to `[0-9A-Za-z-]`,
with the broker host as fallback when the name is empty. The suffix
comes from `crypto/rand` and changes on every ingestor start.
- sets the ID once per source in `buildMQTTOpts`
(`cmd/ingestor/main.go:616`). paho copies the options into the client
and reuses them for every reconnect, and the watchdog force-reconnect
reuses the same client, so the ID is stable for the life of the process.
- logs the ID on connect: `MQTT [tag] connected to <broker> as client
<id>` (`cmd/ingestor/main.go:150`)
- documents the key in `config.example.json:210` as a
`_comment_clientId` entry rather than a value, because
`docker/entrypoint.sh:6` copies that file as a live config and a literal
value would give every default deployment the same ID. Also listed in
`cmd/ingestor/README.md:94`.

## paho behaviour

- No client-side length limit. `SetClientID` only stores the value; the
65535 check in `packets/connect.go:156` is in `Validate()`, which the
client never calls.
- The default ID is longer than the MQTT 3.1 limit of 23 characters for
most source names. paho falls back to MQTT 3.1 after any refused CONNACK
when no protocol version is set (`client.go:412`), so on a broker that
refuses the first 3.1.1 attempt, the retry may hit that limit. I did not
cap the length because paho does not require it and the 3.1.1 path
accepts it (see below).

## Tests

`cmd/ingestor/mqtt_opts_test.go:53-107`:

- default ID is non-empty, has the sanitized name prefix, and contains
only `[0-9A-Za-z-]`
- broker host is used when the name is empty
- configured `clientId` is used verbatim
- two unconfigured sources with the same name get different IDs
- the client built from the options reports the same ID

Mutation checks: removing the random bytes fails the "different IDs"
test; removing sanitization fails the prefix and character-set tests.

`go test ./...` in `cmd/ingestor` passes except
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails locally on
Windows for a symlink privilege reason. `gofmt` and `go vet` are clean.

## Validation against a real broker

On a staging instance (build `e84d2da6`) connecting to a Mosquitto
bridge:

```
MQTT [lincomatic] connection attempt #1 to tcp://mosquitto-bridge:1883
MQTT [lincomatic] connected to tcp://mosquitto-bridge:1883 as client corescope-lincomatic-71a6eb
MQTT [lincomatic] subscribed to meshcore/#
```

The 27-character default was accepted on the first attempt and packets
kept arriving afterwards.

## Not verified

- Only one broker type (Mosquitto) was tried.
- `-race` was not run locally (no cgo toolchain on the test machine).
- The case where both the source name and the broker host are empty (ID
becomes `corescope-<hex>`) has no test.

Fixes #2013



## Review follow-up (commit `a1d6709e`)

An independent review found no bug in the ID handling, but the tests
covered less than their names said. Changed, tests only
(`cmd/ingestor/mqtt_opts_test.go`):

- `TestBuildMQTTOpts_ClientIDSurvivesReconnects` replaces the old
stability test, which only checked that paho copies the options. Against
a loopback fake broker built on paho's `packets` codec, the first
CONNECT, paho's auto-reconnect after the broker drops the socket, and
the watchdog force-reconnect (`buildForceReconnectFn`) must all carry
the same non-empty ID. It runs in about 0.01 s and passed `-count=30
-cpu 1,2,8`.
- `TestBuildMQTTOpts_ClientIDDefaultShape` asserts full IDs:
`^corescope-local-feed-1-[0-9a-f]{6}$`,
`^corescope-mqtt-example-com-[0-9a-f]{6}$` for the broker host fallback
(no port), and `^corescope-[0-9a-f]{6}$` with neither a name nor a host.
- Mutations now caught: `SetClientID` removed, `u.Host` instead of
`u.Hostname()`, a 1-byte suffix, the name guard dropped, sanitization
removed. The "as client" log line has no test because it is logged from
a closure inside `main()`.

Corrections to the description:

- **Fallback to MQTT 3.1.** paho falls back after any failed handshake
once the socket is open, not only after a refused CONNACK: also a read
error or timeout before any CONNACK, or a first packet that is not a
CONNACK (`client.go:401-416`, `net.go:83-97`). A failed dial does not
trigger it (`client.go:387-391`), and after the first successful connect
the protocol version is locked in (`client.go:422-424`). Without this PR
the 3.1 retry sent an empty ID, which MQTT 3.1 forbids as well, so
nothing gets worse.
- **Broker side.** On the EMQX broker we run, authorization has
per-username and all-client rules and no client-ID rules (checked
through its REST API). Two per-username rules use `${clientid}` in a
topic, but both are publish rules and the ingestor only subscribes, so
no rule can match it. A broker that caps IDs at 23 characters but
accepted the empty ID before would now reject the default ID for source
names of 7 or more characters; I have no evidence such a broker is in
use.
- The "Not verified" item about the empty name and host case no longer
applies.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:57:37 +02:00
n30nex cd9b4c04d0 test: unify frontend test runs and prevent inventory drift (#1965)
Make `test-all.sh` the authoritative standalone frontend runner for npm
and CI. Restore stale assertions and reject missing, duplicate, removed
or undocumented inventory entries.

Fixes #1858.

Rebased onto `a2f039d4`. Retains release-routing, map scope-state and
Scope Audit stylesheet tests, adds `test-packets-local-channels.js` to
the sorted runner, and classifies `test-neighbor-map-btn-clip-e2e.js`
under browser. Its separate CI browser step is preserved. Inventory: 280
root suites, 167 standalone, 113 requiring separate setup.

The icon repair fixes two suites red on master:
`test-issue-1648-m2-emoji-scan.js` and
`test-issue-1648-m6-final-sweep.js`. Node/live configured-scope
confirmations now use the existing accessible Phosphor check sprite.
Values, visibility conditions and scanner assertions are preserved.

- Red evidence: `89e45a9` inventory assertions; `4e255df` accessible
confirmation assertion. The latest rebase also reproduced both
unclassified-file failures before adding their entries. This follow-up
only changes runner/classification configuration and counts; no test
files modified.
- Local validation: all 167 standalone suites; 27 M2 browser checks and
16 neighbor geometry checks in Chromium. Syntax, whitespace, PII,
CSS-variable and XSS checks passed.
- Browser coverage includes populated/empty/null configured scopes in
node and live views.
- No new dependencies, requests, application settings or Go changes.
Workflow outside the unit step matches master.
- Windows validation uses process-local UTF-8 settings. Encoding and
node-reach confirmation follow-ups remain separate, as requested.

## Preflight override

External `run-all.sh` is unavailable; applicable repository checks were
run directly.
2026-09-13 19:19:39 +02:00