3030 Commits
Author SHA1 Message Date
Alex BandClaude Opus 5.5 dc7ebcd6f5 perf(store): check observation duplicates without per-transmission maps (#2181)
## Problem
Every `StoreTx` carried two maps for observation de-duplication:
- `obsKeys`, keyed by `observerID + "|" + pathJSON` (a concatenated
string allocated for every observation);
- `observerSet`, for `UniqueObserverCount`.

On a busy regional network's 7-day window (64k transmissions, 400k
observations), these maps and their concatenated keys held about 60 MB
of heap. Most transmissions are heard a handful of times; the busiest in
that week was heard 40 times. The maps cost a few hundred bytes per
transmission to de-duplicate lists that are usually under ten long.

## Change
- `hasObservation` and `hasObserver` scan `tx.Observations`, comparing
the observer ID and path strings the observations already hold. These
are the same two values the old map key was built from, so the answers
are identical.
- Once a transmission passes 64 observations (`obsIndexThreshold`), it
builds the maps. They're keyed by a struct that shares the observations'
strings, so no key bytes are allocated, and busy transmissions keep O(1)
checks (the case #543 / #355 addressed).
- The five places that appended an observation now call
`addObservation`, which keeps `UniqueObserverCount` and
`ObservationCount` in step. Those three lines were previously copied at
each site.
- The memory estimate no longer charges for the per-transmission maps or
the concatenated keys.

## Perf justification (AGENTS.md rule 0)
- **Complexity.**
- Below the threshold, a duplicate check is O(k) for a transmission with
k < 64 observations, so building it costs at most 63·64/2 = 2016
string-header comparisons.
- At and above 64 observations, checks are O(1) map lookups, as before.
- Nothing is O(n²) in the number of transmissions or observations
overall: the scan is bounded per transmission.
- **Scale.** Real data peaks at 40 observations per transmission. The
scan compares string headers (length first, then bytes), and
64-character observer IDs differ early.

`BenchmarkObsDedup` (a miss, the common case on ingest):

| observations on tx | check used | time |
|---|---|---|
| 4 | scan | 15 ns |
| 16 | scan | 70 ns |
| 40 | scan | 151 ns |
| 63 | scan | 233 ns |
| 64+ | map (struct key) | 15 ns |

For comparison, the old map lookup was about 15 ns too, plus building
and allocating the `observerID|path` key string on every check.

At 40 observations that's about 3 µs extra to build a whole
transmission, against roughly 2 KB of maps no longer allocated for it.

**Measured on a copy of that 7-day window** (`GOMEMLIMIT=750MiB`, same
queries):

| | before | after |
|---|---|---|
| in-use heap after load | 612 MB | 568 MB |
| HeapAlloc after load | 660 MB | 564 MB |
| peak RSS | 1077 MB | 1009 MB |
| startup to ready | 31 s | 21 s |

It has also run on a staging instance (live MQTT ingest) for about nine
hours alongside other store changes, with no change in observation or
observer counts.

## Tests
- `obs_dedup_test.go`:
- duplicates (same observer and path) are rejected, while a new path or
a new observer is kept;
  - a transmission under the threshold carries no maps;
  - across the threshold the maps take over with the same answers;
  - empty observer IDs aren't counted as observers;
  - `BenchmarkObsDedup`.
- Red commit first (stubs, failing on assertions), then green.
- `cd cmd/server && go test ./...` passes.
- SonarQube: quality gate passed, no issues on added lines.

Backend only: no UI, API shape or config change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 14:31:12 +02:00
Alex BandClaude Opus 5.5 955386b63f test(map): settle the router's focus frame before the teardown snapshot (#2179)
## Summary

Fixes an intermittent failure in
`tests/e2e/test-map-nodes-pagination-e2e.js`, the "late response after
map teardown" cases added for #2030 in #2048:

```
✗ late /observers response after map teardown: old response must not change the destination page
```

It failed on #2164 ([run
38080812509](https://github.com/Kpa-clawbot/CoreScope/actions/runs/38080812509),
shard 2/3), a Go-only change, and passed on the rerun.

## Cause

`checkMapTeardown` compares `#app`'s HTML before and after releasing the
old map's pending response. When it failed, the only difference was on
the destination page's heading:

```
before: <div class="tools-landing"><h2>Tools</h2>…
after:  <div class="tools-landing"><h2 tabindex="-1">Tools</h2>…
```

That attribute comes from the router, not the old response. Since #630
(`app.js`, "SPA focus management") it focuses the new page's first
heading in a `requestAnimationFrame` after the page renders. The
before-snapshot was taken as soon as `#leaflet-map` detached, sometimes
ahead of that frame.

## Fix

The test now waits two frames before the before-snapshot, the same
settle it already does before the after-snapshot. The router's frame is
queued by then, so it always lands first. The change is test-only: no
app code changes, and every assertion stays as it was.

## Tests

Run locally with the e2e fixture (`tools/freshen-fixture.sh`, migrated)
served by this branch's server, headless Chromium:

| | runs | failed |
|---|---|---|
| test on `master` (`352a0ccd`) | 20 | 2 |
| this branch | 50 | 0 |

- **Before:** the two failures on `master` were "old response must not
change the destination page" (teardown, `/config/regions` and `nodes`),
plus one 30 s timeout in an "after map replacement" case in the same
runs.
- **After:** none of these failures or the timeout came back in 50 runs.

## Checks

- **Lint:** ESLint passes with the repo config and with Node globals.
SonarQube rules on the file report nothing on the added lines; its four
existing minor findings are on lines this PR doesn't touch.
- **Repo rules:** one logical change, explicit `git add`, no private
data, no new dependencies. This is a test-only fix, so there's no
red/green pair: the "red" is the intermittent failure above, measured as
2 in 20.

Related: #2030, #2048 (where the test was added).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 13:28:25 +02:00
CoderNemesisandCoderNemesis d9342bbebd (feat): UserManagement, Add Support for Postal Email Delivery Provider (#2154)
This PR extends email support to add another email delivery provider
called Postal for user management. It is a, open source, self hosted
email delivery and reputation management service.

It adds 2 new additional configuration settings to the
userManagement.mail object which are "postalBaseUrl" and "postalApiKey".
It also adds "postal" as a supported value for the "provider" setting.
All existing functionality including email delivery notification is
supported by Postal.

Included test cases to support the new functionality and all pipelines
pass.

---------

Co-authored-by: CoderNemesis <rfedor@comchan.net>
2026-10-11 11:28:20 +02:00
Alex BandClaude Opus 5.5 a4ee3a7328 feat(scope-audit): observer region filter for both views (#2171)
## Summary

Builds on the transport view from #2167 and #2169 (both merged); rebased
onto `master` (`352a0ccd`), so the diff is this PR's six commits only.

The Scope Audit was the one analysis page without the region selector.
This adds the shared region filter to it, for both the declared and the
transport view. `?region=SFO,SJC` means what it means on every other
endpoint: count only forwarding **heard by observers in those regions**,
not repeaters located there.

- **API** (`GET /api/scope-audit`, both modes):
- The region codes are resolved to their observers first (a handful of
rows), then the audit's hop scan gets `AND o.observer_idx IN (…)` on v3,
or `o.observer_id IN (…)` on v2. The 3.5M-row hop scan never joins
`observers`.
- `All` means no filter (shared `normalizeRegionCodes`). Codes with no
observer match nothing, instead of silently widening to every observer.
- Responses echo the normalised codes as `region`. The declared and
transport caches and their singleflight keys now include the region, so
each region set is cached separately with the existing TTLs.
- Declared-region verification (`regionEvidence`) is deliberately left
unfiltered: it corroborates what a repeater declares from its own
traffic, which doesn't depend on where a hop was heard.
- **Page:** the shared `RegionFilter` sits with the view and window
controls; a change reloads the view, and the selection persists like
every other page. The transport view's text box is relabelled **Scope**
(it already sends `?scope=` since the rename in #2167), so the page no
longer has two controls called Region.
- **Docs:** `docs/api-spec.md` (parameter, both response shapes) and the
OpenAPI parameters for the route (`mode`, `scope` and `region` were all
missing).

Partial fix for #2142 (its region filter for rollout tracking), together
with #2167 and #2169.

## Performance

The unfiltered path must not get slower: it's the same query with an
empty clause. `BenchmarkComputeScopeTransport` (from #2167, 1,200
repeaters), the PR it stacks on vs this branch, `-count 8`:

| | before | after | |
|---|---|---|---|
| sec/op | 220.3 ms | 218.0 ms | ~ (p=0.065) |
| B/op | 43.77 MiB | 43.70 MiB | ~ |
| allocs/op | 833.5k | 833.5k | ~ |

**Filtered requests, measured on staging** (a copy of a production
database): the first version drove `o.observer_idx IN (…)` from
`idx_observations_observer_idx`, reading every observation the region's
observers ever made. Fixed in `80414242` with a unary `+` (the same
treatment as #2166), so the scan starts from the `first_seen` window and
filters per row. Red test `67e063b1` asserts the query plan against a
fixture with the real indexes.

| scan, one region over 24h (13 observers) | plan | time | hop rows |
|---|---|---|---|
| `o.observer_idx IN (…)` | observer index | 13.44 s | 102,114 |
| `+o.observer_idx IN (…)` | `first_seen` window | **0.25 s** | 102,114
|
| unfiltered | `first_seen` window | 0.35 s | 225,983 |

Cold end-to-end requests on staging after the fix: 24h declared,
unfiltered 0.84 s · one region 0.48 s (was 16.7 s) · two regions 0.22 s
· transport, one region 0.42 s · transport, six regions 0.75 s · 7d, one
region 3.46 s. 24h transport, live data: 383 repeaters unfiltered, 204
heard by one country's observers.

## Tests

- Red `629e4f36` (Go): extends the audit fixture with an `observers`
table and `observations.observer_idx`. Covers the declared view ignoring
a scope heard only by another region's observer, the transport view
listing only repeaters heard in the region, `region=All`, per-region
caching, a region with no observers, and the echoed codes. Five fail on
assertions; `region=All` passes because the stub ignores the filter,
which is the behaviour it pins.
- Green `6f1b1e1c`: the implementation, plus a v2-schema case
(`observer_id`), so both branches of the clause are covered.
- Red `b4933d68` / green `08ecc172` (frontend): `apiPath` carries the
selection in both views and adds nothing when none is selected. The
region control on the page is a net-new UI surface.
- `go test` (cmd/server) passes. Frontend suites: 193 of 194. The one
failure, `test-issue-1956-release-routing.js`, needs CI's
`GITHUB_REF_NAME` and fails the same way on `master` locally.

## Screenshot

The same transport view with the shared region filter set to **SJC**:
only forwarding heard by San Jose observers counts, so 7 repeaters are
listed instead of 44, with their carried scopes counted from those
observers alone. Taken against `test-fixtures/e2e-fixture.db` with
fictional scopes.

![Scope Audit with the observer region filter on
SJC](https://raw.githubusercontent.com/A13xB0/CoreScope/pr-screenshots/2026-10/scope-audit-region-filter.png)

## Browser validation

The branch's server was run locally with 14 configured regions (dropdown
layout), checked in headless Chromium at 1280, 900, 700 and 390 px
(touch):

| check | 1280 | 900 | 700 | 390 touch |
|---|---|---|---|---|
| ticking SJC re-requests with `&region=SJC` | yes | yes | yes | yes |
| menu fully on screen | yes | yes | yes | yes |
| page wider than viewport | no | no | no | no |
| page errors | none | none | none | none |

Found and fixed this way: the filter is the right-most control on
desktop, and the menu (which opens rightwards) ran 66 px off the page.
It now opens leftwards above 640 px. On a phone the controls wrap, the
filter starts a line, and the default applies.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **S3776 on `ScopeAuditForwardingHeardIn`:** complexity 55. The
function it generalises, `ScopeAuditForwarding`, is 54 on `master`; the
region clause adds one point.
- **S3776 on `computeScopeAudit`:** complexity 65, unchanged from
`master`.
- Both pre-date this PR. Splitting them is better as a refactor PR of
its own than mixed into a feature.
- The page code's findings were fixed in the transport-view PR it stacks
on.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 11:17:12 +02:00
efitenandClaude Opus 5.5 55dd28164d test(e2e): assert the configured-scope row after the theme-refresh re-render (#2180)
## What

`test-issue-1648-m2-icons-e2e.js` now asserts the configured-scope row
only after the page's `theme-refresh` re-render has fired and the detail
has rendered again.

## Why

On 2026-10-10 this assertion failed CI on four PRs (#2169, #2171, #2154,
#2164), each time at the first live-map case: `confirmed scope row must
be visible, including an empty scope`. A re-run of the failed jobs
passed every time.

The code path:
- `customize-v2.js:814` dispatches `theme-changed` once
`/api/config/theme` resolves, on every page load.
- `app.js:1309-1315` turns that into `theme-refresh` 300 ms later.
- `live.js:4713` re-runs `showNodeDetail(activeNodeDetailKey)` on
`theme-refresh`, and `showNodeDetail` first sets `#nodeDetailContent` to
"Loading…" (`live.js:2598`). This is the only caller besides the marker
click (`live.js:3090`).
- The test waited for `#nodeDetailContent table` and then called
`row.isVisible()`, which does not wait. A refresh landing between the
two fails the check.

The same race could let the null-scope case pass on the "Loading…"
panel, because `row.count() === 0` holds there too.

The fix uses the same `addInitScript` flag pattern as
`test-e2e-playwright.js:83`.

## Testing

- Locally against a server on the E2E fixture: 5 of 5 runs pass, 27 of
27 checks each.
- **Not reproduced locally.** The unchanged test also passed every local
run: 5 plain, 3 with `/api/config/theme` delayed by 1 s, 5 with 8x CPU
throttling. So the reasoning above comes from the code path and the four
CI failures, not from a local red-then-green.
- Test-only change, no product code touched.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 10:58:20 +02:00
Alex BandClaude Opus 5.5 4e91c5cf0b perf(analytics): opt-in pauseWhenIdle skips recomputes nobody reads (#2164)
## Summary

The analytics recomputers (#1240) recompute eleven snapshots (topology,
RF, distance, channels, hash collisions, hash sizes, roles, two
clock-skew views, retransmissions, direct-heard) on a fixed interval,
300 s by default, whether or not anyone reads them. That is the right
trade for a busy instance: reads always hit a warm cache. On a quiet one
it means most of the server's allocation goes on results nobody looks
at.

Measured on one production instance (2 vCPU / 2 GB):

- topology, RF and hash-sizes were each **read 23 times in 45 hours**
while being recomputed every 5 minutes;
- `analyticsRecomputer.runOnce` accounted for **30% of all bytes
allocated** (121 GB of 391 GB in 45 h);
- garbage collection was 87% of the process's CPU (30 s CPU profile:
`gcDrain` 39.9 s of 45.9 s sampled), with an 805 MB live heap against a
1 GiB `GOMEMLIMIT`, so allocation rate translates directly into
collector time.

This PR adds an **opt-in** `analytics.pauseWhenIdle`:

- a recomputer's tick is skipped when nothing has read its snapshot
since the previous compute, and the recomputer is marked paused;
- the next read still returns immediately with the snapshot it holds
(`Load` stays non-blocking, as #1240 requires), and makes a non-blocking
send that wakes the loop to recompute; reads after that get the fresh
snapshot and the ticker resumes from there.

**Default is off**, so behaviour is unchanged unless an operator turns
it on. The trade-off for operators who do: the first read after a quiet
spell can show a snapshot older than the interval (refreshed seconds
later). It is documented in `config.example.json`.

No existing issue covers this.

## Config example

Off by default. To turn it on, add one key to the existing `analytics`
block in `config.json`:

```json
"analytics": {
  "defaultIntervalSeconds": 300,
  "pauseWhenIdle": true
}
```

Nothing else changes: per-endpoint `recomputeIntervalSeconds` still
apply while an endpoint is being read. There is no UI change, so no
screenshot.

## Performance

Staging A/B: the same 30 minutes of real production API traffic (5,640
requests from a day's nginx log, replayed at 2x against a copy of the
production database with the live MQTT feed), on `master` and on this
branch with `pauseWhenIdle: true`, each after a restart and warm-up.
Allocation is the delta of `/debug/pprof/allocs` across the replay:

| during the replay | master | this PR |
|---|---|---|
| allocated by `analyticsRecomputer.runOnce` | 459.7 MB | **53.4 MB
(-88%)** |
| total allocated by the process | 3,704 MB | 3,372 MB (-9%) |
| GC cycles | 52 | 41 |
| requests: p50 / p95 / errors | 9.9 ms / 57.0 ms / 0 | 10.4 ms / 60.1
ms / 0 |

The replayed traffic included real analytics reads, so the paused
recomputers were woken by them; the remaining 53 MB is that. Over 15
minutes each recomputer only ticks about three times, so on a longer
quiet period the saving grows with the idle time. GC count and CPU on
this shared host vary between runs on their own, so the allocation
figures are the claim, not the GC count.

There is no Go benchmark for this change: the behaviour is "how many
computes happen in a period of no reads", which the tests below assert
directly (16 computes on master vs at most 2 here over the same idle
period).

## Tests

- Red commit `24056c57`: config field, accessor stub and an unused
recomputer flag, plus tests. Three fail on the stub:
- `TestAnalyticsRecomputer_PauseWhenIdleSkipsUnreadTicks`: an unread
recomputer with a 20 ms interval computes at most twice in 300 ms (16
times on the stub);
- `TestConfig_AnalyticsPauseWhenIdle`: `pauseWhenIdle: true` enables it,
absent or nil config leaves it off;
- `TestStartAnalyticsRecomputers_AppliesPauseWhenIdle`: all eleven
recomputers receive the flag.
- Also added, passing on both commits on purpose: a read on a paused
recomputer returns the existing snapshot and triggers a refresh; a
recomputer that is being read keeps ticking; the default (off) keeps
ticking unread.
- Green commit `19de3636`: the behaviour, `config.example.json` entry
with `_comment_pauseWhenIdle`, and the `Load` doc comment.
- `go test -race` on the recomputer, warm-up (#1659) and config tests,
three runs: pass. `make test`: 16 of 16 modules pass. The frontend tests
that read `config.example.json` pass.

## Review notes

- `roles` reads the nodes-clock-skew snapshot inside its own compute,
which counts as a read, so that dependency keeps its source fresh while
roles is in use.
- The #1659 warm-up gate is untouched: a skipped tick never marks a
first pass done.
- Config documented per WORKFLOW.md section 7. Customizer (AGENTS.md
rule 8): it is an operator setting, not a display value; no customizer
entry proposed.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- No issues on the added lines.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 09:51:48 +02:00
Alex BandClaude Opus 5.5 352a0ccd19 feat(region-filter): configurable region quick picks (#2170)
## Summary

**Independent of every other open PR.** Choosing a pick saves the
selection, redraws the control and notifies the page's listeners itself,
exactly as a click in the control does, so it re-queries straight away
whether or not #2168 (`setSelected` notifies) is merged; with both in,
nothing fires twice.

Region quick picks are named groups of region (IATA) codes, offered as
one-tap choices inside the region filter, e.g. "Northern California" for
a dozen airport codes instead of ticking each box. One deployment has
been injecting these into the page through an nginx `sub_filter` and a
script that wraps `window.RegionFilter`; this makes them a configured
feature.

```jsonc
"regionQuickPicks": [
  { "name": "Northern California", "description": "Bay Area and Sacramento observers.", "regions": ["SJC", "SFO", "OAK", "SMF"] }
]
```

- **Server:** `GET /api/config/region-quick-picks` returns the groups
cleaned up: names and codes trimmed, codes upper-cased and de-duplicated
in their configured order, groups without a name or codes dropped, at
most 20 groups of 200 codes, always a list. `/api/config/regions` is
unchanged.
- **Region filter:** loads the groups with the regions and offers every
group with at least one observer today, inside the region control:
- dropdown layout: rows of the same kind as **All** (checkbox and bold
name), directly under All with a divider before the single regions; the
row is ticked while that pick is the selection, ticking it selects the
pick and closes the menu, unticking goes back to all;
- pill layout: the picks lead the same bar, set off from the single
regions by a thin divider.
- A tap selects the group's codes that have an observer and re-queries
straight away.
- The selector then reads the group's name ("South Bay ▾") instead of "2
Regions", and the button shows as pressed. Selecting the same codes by
hand does the same.
  - Tapping the active group again goes back to all regions.
- A failed fetch, or no configured groups, leaves the filter exactly as
before.
- **Mobile:** the rows take the menu's 44 px row height under `(pointer:
coarse)`; the pill-bar picks reuse `.region-pill` and its touch target.
Inside `.filter-bar` (Packets) the existing touch rule for inputs made
every region checkbox a 44 px square; the checkbox is now 20 px and the
44 px row is the tap target.
- **Semantics:** like every region filter, a group means "heard by
observers in these regions", not "located there"; `config.example.json`
says so.

No existing issue covers this.

## Screenshot

Nodes page with two picks configured (example config below), after
tapping **South Bay**: the pick and its codes (SJC, MRY) are selected,
and the list re-queried for them. Taken against
`test-fixtures/e2e-fixture.db`.

![Region quick picks on the Nodes
page](https://raw.githubusercontent.com/A13xB0/CoreScope/pr-screenshots/2026-10/region-quick-picks.png)

## Config example

```json
"regionQuickPicks": [
  { "name": "South Bay", "description": "San Jose and Monterey observers", "regions": ["SJC", "MRY"] },
  { "name": "East and North Bay", "description": "Oakland and San Francisco observers", "regions": ["OAK", "SFO"] }
]
```

Each pick needs a `name` and at least one code; `description` is
optional (shown as the tooltip). Codes are trimmed, upper-cased and
de-duplicated; a pick with no code that has an observer today is not
shown. At most 20 picks of 200 codes.

## Browser validation

The branch's server was run locally against a migrated copy of
`test-fixtures/e2e-fixture.db` (observers in MRY, OAK, SFO, SJC) with
three groups configured: "South Bay" (SJC, MRY, SCZ), "East and North
Bay" (OAK, SFO, sts) and "Nowhere yet" (XXX). Checked in headless
Chromium:

| check | desktop 1280 px | phone 390 px (touch) |
|---|---|---|
| groups shown | South Bay, East and North Bay ("Nowhere yet" hidden) |
same |
| tap South Bay selects | SJC, MRY (SCZ has no observer) | same |
| next `/api/packets` request | `region=SJC,MRY` | `region=SJC,MRY` |
| selector label | "South Bay ▾" | "South Bay ▾" |
| second tap | back to all regions | same |
| button height | 24 px | 44 px |
| page wider than viewport | no | no |
| page errors | none | none |

Checked on Packets (dropdown layout) and Nodes (pill layout). Three
layout problems found this way on the phone were fixed before opening: a
sideways-scrolling row clipped at the screen edge, the selector squeezed
into a narrow column (the filter group is a non-wrapping flex row), and
the two wrapped rows spread apart on Nodes (`align-content`).

**Revision 1:** the first version put the picks in a separate row beside
the selector; they moved inside the dropdown. Re-checked with a
14-region list and three picks on Packets: desktop 1280 px dark and
phone 390 px light both keep the menu on screen (menu 340 px wide, page
not wider than the viewport), the pick closes the menu and the trigger
names the pick.

**Revision 2:** picks are checkbox rows like **All** instead of chips.
Re-checked on Packets at 1280 px dark and 390 px light: rows render bold
under All, the divider separates the single regions, ticking a pick
ticks its codes and names the trigger, no sideways scroll.

Unrelated and not changed: on Nodes at 390 px the region row (existing
pills included) starts flush with the screen edge.

## Tests

- Red commit `81ea67d6`: the config type, a stub normaliser and endpoint
returning nothing, and tests.
- Go: cleaning (trim, upper-case, de-duplicate, drop incomplete groups),
caps, nil config, and the endpoint's JSON.
- Frontend, `tests/e2e/test-region-quick-picks.js` (moved from
`tests/unit/` in `fced69a7`; it needs jsdom, so it runs in the E2E job
like `test-table-sort.js`): the real `region-filter.js` in **jsdom**
against faked config endpoints. Nine cases: rendering and hiding groups
with no observer, selecting and re-querying, the selector label,
toggling back, a hand-made selection marked active, the pill layout, no
groups, a failed fetch, HTML escaping. Seven fail on the stub, all on
assertions.
- Green commit `0e33210e`: the implementation, styles,
`config.example.json` (`regionQuickPicks` +
`_comment_regionQuickPicks`), `docs/api-spec.md`, and the OpenAPI route
list (`TestOpenAPICompleteness` requires it). The jsdom test also had a
cross-realm comparison bug fixed in this commit (`Array.from` on arrays
coming out of the jsdom window).
- `make test`: 16 of 16 modules pass. Frontend suites: 195 of 196; the
one failure, `test-issue-1956-release-routing.js`, needs CI's
`GITHUB_REF_NAME` and fails the same way on `master` locally.

- Red commit `f991060d` (revision): picks live inside the dropdown menu
before the checkboxes, choosing one closes the menu, and in the pill
layout they share the region bar; two fail on assertions against the old
layout. Green commit `1f720e35`. Frontend 195 of 196 (same
environment-only failure).

- Red `134f5256` / green `da0adb50` (second revision): picks are
checkbox rows between All and the single regions with a divider, no
chips in the menu; ticking and unticking a pick row. The green commit
also repairs `style.css`: `1f720e35` had prepended its block to the top
of the file (its edit matched an earlier `.region-pill {`) and left the
previous rules in place. The file now differs from `master` by 13 lines.

## Review notes

- Groups are offered only when at least one of their codes appears in
`/api/config/regions`, and only those codes are selected. The stored
selection is pruned to known codes on load anyway, so selecting absent
codes would not survive a reload.
- Customizer (AGENTS.md rule 8): groups are operator configuration; a
customizer editor could come later.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **Fixed** in `b84ace4e`:
- `normalizeQuickPickCodes` split out of the config normaliser (S3776,
17 > 15);
  - a comment on why a failed quick-pick fetch is ignored (S2486);
  - `for…of` in `activePick` (S4138);
- in the tests: `node:` specifiers, `includes()`, and an explicit string
comparator for sorted selections (S7772, S7765, S2871).
- **Kept on purpose:**
- S2486 is still reported on that `catch`. Swallowing the error is the
intended behaviour: the filter works as before, just without picks.
  - S1523 on the `vm` test loader, as above.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:23:05 +02:00
Alex BandClaude Opus 5.5 99efaed33d fix(region-filter): setSelected notifies change listeners on a real change (#2168)
## Summary

`RegionFilter.setSelected()` saved the selection and redrew the control
but never ran the `onChange` listeners; only a click (`toggleRegion`)
did. A programmatic selection therefore updated the dropdown while the
page kept showing data for the old selection until the next click.

Affected today: the Packets page applies a `?region=` URL parameter
through `setSelected`. It registers its listener afterwards, so its
first load is fine, but any other page's listener that is still
registered sees nothing. And anything that offers one-click region
groups has to wrap `setSelected` to make it re-query (an nginx-injected
overlay on one deployment does exactly that). This PR stands alone;
#2170 (configurable quick picks) does not need it.

`setSelected` now compares the new selection with the current one
(order-insensitive; empty and null both mean "all") and, only when they
differ, runs the listeners after saving and re-rendering, exactly as a
click does. Re-applying the current selection stays silent, so a caller
restoring state does not trigger a reload.

No existing issue covers this.

## Tests

- Red commit `75569b78`: `tests/unit/test-region-filter-setselected.js`
loads the real `public/region-filter.js` in a `vm` sandbox. Two cases
fail on `master` (a programmatic selection notifies once with the new
codes; a clear notifies once with `null`). Two pass on both commits on
purpose: re-applying the current selection in another order, and
clearing an empty selection, must not notify.
- Green commit `8a7ddfe1`: the fix.
- Listed in `test-all.sh` in sorted order; `test-test-inventory.js`
passes. `test-1108-region-hide-nodes.js` (the other region-filter suite)
passes. Frontend suites: 194 of 195; the one failure,
`test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and
fails the same way on `master` locally.

## Review notes

- `packets.js` "clear all filters" calls `setSelected([])` and then
reloads itself. When a region was selected, that path now also triggers
the page's region listener, so the packets list is requested twice. It
is a cosmetic double fetch on an explicit user action; left as is to
keep this PR to the component.
- No performance claim; this is a correctness fix.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **Fixed** in `9e732706`: `node:` specifiers (S7772).
- **Kept on purpose:** S1523 on `vm.runInContext`, which loads the
repo's own `public/region-filter.js` as the other frontend unit tests
do.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:22:54 +02:00
Alex BandClaude Opus 5.5 deb55dc82e feat(scope-audit): transport view on the Scope Audit page (#2142, page) (#2169)
> **Requires #2167 to be merged first.** This page calls the transport
view API that #2167 adds (`/api/scope-audit?mode=transport`), so it
cannot be split from it; until #2167 merges this PR also shows its
commits.

## Summary

Page half of #2142. **Stacked on #2167** (the API); review that first.
#2171 adds the observer region filter on top. Partial fix for #2142
until all three are in.

Adds a **Declared / Transport** toggle to `#/scope-audit`. The transport
view calls `?mode=transport` and lists every repeater seen forwarding in
the window:

- **Seen carrying:** one chip per region scope with its packet count
(first/last seen in the tooltip);
- **Declared:** the repeater's declared regions coloured by whether they
were seen carried (the declared view's own chips), or *not asked* when
no declared-regions answer exists. Never asked is worded differently
from declaring nothing;
- **Other traffic:** unscoped, unmatched and ambiguous-hop counts.

A **Region** box (suggestions from the scopes seen) adds a *Carries X*
column, lists the repeaters not yet carrying it first, and summarises
carriers against active non-carriers: the rollout report the issue asks
for. Window, view and region are deep-linked
(`#/scope-audit?window=7d&mode=transport&scope=be`) and survive a
reload. The page says plainly that *transported* means seen carrying,
not configured, and that over a short window a quiet region looks the
same as a missing one. The intro text changes with the view.

## Screenshot

The transport view with scope `be` typed in: every repeater seen
forwarding in the last 24 h, whether it carries `be` (non-carriers
first), the scopes it was seen carrying with counts, and its declared
regions or *not asked*. Taken against `test-fixtures/e2e-fixture.db`
with fictional `be` / `be-van` / `nl` scopes.

![Scope Audit transport
view](https://raw.githubusercontent.com/A13xB0/CoreScope/pr-screenshots/2026-10/scope-audit-transport-view.png)

## Browser validation (staging, real data)

Headless Chromium against a staging copy of a production database with
its live feed:

- `#/scope-audit?window=24h`: the declared view, 16 rows, as before.
- Clicking **Transport**: 385 rows; the hash gains `mode=transport`.
- Typing a scope as `#XX-YY` in capitals: normalised to `xx-yy`; the
hash gains it (now `scope=xx-yy`); summary "385 repeaters seen
forwarding in the last 24h · 16 with a declared-regions answer · 51
carried xx-yy · 334 active but not seen carrying it"; the first rows are
the non-carriers.
- Reloading that hash restores view and region. Switching back to
Declared hides the Region box and restores the declared intro. No page
errors.
- Checked at 1280 px and 390 px wide.

Two layout bugs found this way were fixed before opening: the region box
inherited `width: 100%` from `.nodes-search`, and `display: flex`
overrode the `hidden` attribute on the region bar.

## Tests

`tests/unit/test-frontend-helpers.js` gains 10 cases on rendered markup
and state, using the page's exported internals like the existing
scope-audit tests: hash and API paths (and that a region never leaks
into the declared view's link), region normalisation (`#`, capitals,
empty), *not asked* vs declared chips (with the `*` wildcard),
carried-scope chips with counts and the filtered region marked, the
carries column only with a region, HTML escaping of names, the summary
counts, region suggestions, and the per-view intro. 768 passing in that
file.

Net-new UI surface, so per TDD.md the tests land in this PR rather than
as a preceding red commit. Frontend suites on this branch: 193 of 194;
the one failure, `test-issue-1956-release-routing.js`, needs CI's
`GITHUB_REF_NAME` and fails identically on `master` locally (#2149 and
#2160 address that).

## Review notes

- Colours only through existing CSS variables (no hex literals); the
filtered region is marked with the link colour, not a new status colour.
- The window note's wording is in the same spirit as the declared view's
`windowHonestyNote`.
- No API calls per row: one request per view/window/region, cached
client-side for 30 s like the declared view.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **Fixed** in `df53e770`:
- one `modeQuery` helper for the page hash and the API path (S3358; the
two copies were identical);
  - the carries-scope cell built before the row (S3358);
- `localeCompare` for the region suggestions (S2871), `for…of` and
`.dataset` (S4138, S7761), optional chaining (S6582);
  - object spread in the helper tests (S6661).
- Optional chaining and spread are already used throughout `public/`. No
behaviour change.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:22:37 +02:00
Alex BandClaude Opus 5.5 696906315a feat(scope-audit): transport view of every repeater seen forwarding (#2142, API) (#2167)
## Summary

Backend half of #2142 (the page is the stacked PR #2169).

`GET /api/scope-audit` lists only repeaters that have answered a
declared-regions request: 17 of almost 200 nodes on the instance the
request came from. This adds a transport view, `GET
/api/scope-audit?mode=transport&window=…[&scope=X]`, with one row per
repeater **seen forwarding** in the window:

- `transported`: the named region scopes it was seen carrying, each with
packets and first/last seen (most packets first), plus unscoped,
unmatched and ambiguous-hop counts;
- `asked`, `declaredRegions`, `notObserved`, `configState`, `declaredAt`
where a declared answer exists; `asked: false` with null declared fields
where it does not ("not asked" is not "declares nothing");
- with `scope=X` (spelled `be`, `#be` or `BE`): `carriesScope` on every
row and `carrying` / `notCarrying` counts, which is the rollout report
the issue asks for.

`mode=declared` (or no `mode`) is unchanged. An unknown `mode` is a 400.

Partial fix for #2142: this PR provides the API; the page is #2169 and
the observer region filter #2171. The issue should only be considered
done once all three are in.

## How it follows the issue's notes

- **The scan does not change.** It is the same single
`ScopeAuditForwarding` pass, with every repeater and room in `nodes`
(plus any declared target) as a target instead of only the declared
ones. Same attribution rules: FLOOD-family route types only,
`minForwarderHopHexLen` (no 1-byte hops), and a hop whose prefix matches
more than one repeater is credited to none of them.
- **Identity lookup with 1,000+ pubkeys:** `scopeAuditNodeIdentities`
runs one `IN` query; exercised at 1,200 keys in the benchmark below.
- **Row count and response size:** measured below.
- **Cache:** the transport view is cached per window with the declared
view's TTLs (30 s; 5 min for 7d), under the same singleflight group with
a `transport|` key. The scope filter is applied to a copy of the cached
result, so one scan serves every scope. (`scope=`, not `region=`: across
the API `region=` is the observer IATA filter.)
- **Wording:** "transported" is documented as *seen carrying*, not
configured for; the declared side stays separate.

## Performance

`BenchmarkComputeScopeTransport`: 1,200 repeaters (#1975 quotes an
instance with 1,179), 30k flood transmissions in the last 24h on 4-hop
paths of 3-byte prefixes, a third region-scoped:

```
BenchmarkComputeScopeTransport-12   5   219253727 ns/op   425.4 response-KB   1200 rows
```

219 ms per compute, served from cache for 30 s (24h) or 5 min (7d); a
425 KB uncompressed response for 1,200 rows. This is new code with no
"before"; the declared view's numbers are untouched (same scan, same
cache).

### On real data (a staging copy of a production database and its live
feed)

| request | rows | result | time |
|---|---|---|---|
| declared view, 24h | 16 | (unchanged) | |
| `mode=transport`, 24h | **385** repeaters seen forwarding, 16 of them
with a declared answer | | 1.5 s cold |
| `mode=transport&scope=<scope A>`, 7d | 664 | **51 carrying it, 613
not** | 14.4 s cold |
| `mode=transport&scope=%23<SCOPE B>` (`#` and capitals), 7d, same
window now cached | 664 | 237 carrying, 427 not | 0.35 s |

The 7d cold compute is the same full-window hop scan the declared view
already pays for 7d (its own TTL comment measures 16.7 s on a
live-shaped database); like the declared view it is then cached for 5
minutes, and every scope filter is served from that one scan.
- Commit `ceaeb56b` renamed the filter from `region=` to `scope=` after
opening; the real-data figures above were taken before the rename and
are otherwise the same code.

## Tests

- Red commit `2a9b1286`: response types and a `mode=transport` stub that
returns no rows, plus tests. Four fail on the stub:
- repeaters that never answered are listed with `asked: false` and null
declared fields (and the answered one with its declared list and
`notObserved`);
- `scope=fr` / `%23fr` / `FR` marks carriers and non-carriers and counts
them; no scope means no `carriesScope` or counts;
  - an unknown mode is a 400;
- the transport and declared views are cached separately (requesting one
does not serve the other).
- Also added, passing on the stub by construction: a repeater with no
traffic in the window is not listed; the node blacklist applies.
- Green commit `60c669f4`: the implementation, `docs/api-spec.md` (new
`mode` and `scope` parameters, a Transport view section with the
response shape), and the benchmark.
- `make test`: 16 of 16 modules pass, including every existing
scope-audit test.

## Review notes

- Hidden-name and blacklist filtering use the same rules as the declared
view.
- `newTestStoreWithDB` now takes `testing.TB` so the benchmark can use
it (callers unchanged).
- No new `map[string]interface{}`; the response is a named struct.
- No config added.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **Fixed** in `060668b6`: S3776 (cognitive complexity 41 > 15) in the
new `computeScopeTransport`. It is split into target, visibility, row
and sort helpers; the response is the same and the transport tests are
unchanged.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:22:20 +02:00
Alex BandClaude Opus 5.5 8a3b923e81 perf(scope-stats): keep the per-region count on the first_seen index (#2166)
## Summary

`/api/scope-stats` averaged **2.6 s on one production instance (max 17.7
s)**, whatever the window. Four of its five queries take 10-100 ms on
that database. The fifth, the per-region count, took **2.5-2.9 s for
every window, 1h included**: SQLite chose `idx_tx_scope_name`
(`scope_name > ''`) and walked every scoped transmission ever stored (45
days of them), filtering by `first_seen` afterwards.

```
SEARCH transmissions USING INDEX idx_tx_scope_name (scope_name>?) | USE TEMP B-TREE FOR ORDER BY
```

A unary `+` on every `scope_name` reference in that query (`WHERE …
+scope_name …`, `GROUP BY +scope_name`) stops the planner using that
index. The query then searches the window on
`idx_transmissions_first_seen` and groups in a temp b-tree, the same
idiom the `advertsByRole` query in the same function already uses for
`payload_type`. Putting `+` in the `WHERE` alone is not enough: the
planner then walks `idx_tx_scope_name` to avoid sorting for the `GROUP
BY`, which is just as slow; the `GROUP BY +scope_name` is what fixes it.

```
SEARCH transmissions USING INDEX idx_transmissions_first_seen (first_seen>?) | USE TEMP B-TREE FOR GROUP BY | USE TEMP B-TREE FOR ORDER BY
```

Related to #2146, which lists `/api/scope-stats` (2.3 s average on one
instance) and records the 24 h summary as 7 ms warm there. On the
instance measured here the per-region query planned onto
`idx_tx_scope_name` and took 2.5-2.9 s for every window; this PR fixes
that plan. Not marked as a fix: #2146 is about pool and lock waits.

## Results are identical

On a copy of the production database, the old and new query run inside
**one read transaction** (same snapshot) return exactly the same rows
for 1h, 24h, 7d and 45d (5, 10, 12 and 18 regions; 73, 2,024, 13,822 and
76,083 transmissions). `TestScopeStatsByRegion_CountsMatchWindow` also
checks the per-region counts against a direct count of the window.

## Performance

### The query alone, on a copy of the production database

311k transmissions, 76k of them region-scoped over 45 days; best of
three:

| window | before | after |
|---|---|---|
| 1h | 2,530 ms | 0.1 ms |
| 24h | 2,570 ms | 11.6 ms |
| 7d | 2,876 ms | 124 ms |

### The endpoint, on staging

Cold `GET /api/scope-stats` (the response cache is 30 s, so each sample
is a full compute) on a staging host with a copy of a production
database and its live feed; two samples each, 32 s apart:

| window | master | this PR |
|---|---|---|
| 1h | 3.08 s, 2.83 s | **16 ms, 12 ms** |
| 24h | 2.69 s, 2.79 s | **131 ms, 101 ms** |
| 7d | 3.35 s, 4.53 s | **626 ms, 663 ms** |

### Go benchmark

`BenchmarkGetScopeStats24h`: all five scope-stats queries for 24h, on
300k transmissions spread over 45 days (a quarter region-scoped), with
production's tables and both competing indexes and `ANALYZE` run as
production does (#2072). 10 runs:

```
cpu: AMD Ryzen 5 3600XT 6-Core Processor
                    │   master    │             this PR              │
GetScopeStats24h-12   33.186m ± 3%   8.058m ± 6%  -75.72% (p=0.000 n=10)
```

The fixture is smaller and better cached than production, which is why
the absolute times are lower than above; the plan change is the same.

## Tests

- Red commit `a7dcf98f`: moves the query into `scopeStatsByRegionQuery`
(unchanged), adds the fixture, the benchmark,
`TestScopeStatsByRegion_UsesFirstSeenIndex` (`EXPLAIN QUERY PLAN` must
use `idx_transmissions_first_seen` and not `idx_tx_scope_name`; fails on
`master` with exactly the production plan shown above) and
`TestScopeStatsByRegion_CountsMatchWindow` (passes on both commits;
guards correctness).
- Green commit `f9ba986b`: the `+` hints and a comment explaining them.
- All existing scope-stats tests pass; `make test`: 16 of 16 modules
pass.

## Review notes

- Only the per-region query changes; the summary, non-transport,
time-series and adverts-by-role queries already search by `first_seen`.
- If `idx_transmissions_first_seen` were ever missing, the query would
still run correctly (full scan); `+` only removes `idx_tx_scope_name`
from consideration, unlike `INDEXED BY`, which would make the query
fail.
- `/api/scope-stats` still has no singleflight around its 30 s cache
(unlike stats, channels and the scope audit). With the query now well
under a second it matters much less; left for a separate change.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- No issues on the added lines.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:22:02 +02:00
Alex BandClaude Opus 5.5 95672ac2be perf(server): share /api/nodes responses for identical queries for 15 s (#2165)
## Summary

`/api/nodes` scans every node row into a map, enriches it (hash size,
multi-byte capability, relay info, usefulness scores, declared scope)
and encodes the page on **every** request. On one production instance (2
vCPU / 2 GB) that was **95 GB of allocation in 45 hours, about 6.7 MB
per request and a quarter of all bytes allocated**, on a process where
garbage collection was 87% of CPU.

The queries repeat heavily. In that instance's nginx log for one day,
three pages of one query (`limit=500&offset=0|500|1000&lastHeard=30d`)
were 10,000 of the 14,053 `/api/nodes` requests.

This PR keeps the encoded response for 15 s per distinct query:

- `handleNodes` serves the cached bytes when it can; otherwise it builds
through `buildNodesResponse` (the old handler body, unchanged apart from
returning instead of writing) inside a singleflight, so concurrent
identical misses share one build.
- **Cache key:** blacklist generation | hidden-prefix generation |
sorted query string. Parameter order does not split the cache, and any
change to either filter list misses it at once. This is the pattern
`/api/nodes/{pubkey}/reach` already uses (#1629).
- `setGeoFilter` drops the cache, so a geo-filter change through the
admin endpoint shows immediately.
- **Bounded twice:** by `maxCacheEntries` (256) like the other keyed
caches, and by bytes, since a full-list response is several MB. Expired
entries are dropped on every insert, the total is capped at 32 MB, and a
single response over 8 MB is not cached.
- **15 s** matches the relay/usefulness caches the response is built
from (#1257); the client already keeps its own node list for 90 s.

Related to #2146: `/api/nodes` is in its endpoint table (144 s max on
one instance). This PR reduces how often the handler runs and allocates,
not the SQL wait the issue is about, so it is not a fix for it.

## Performance

### Go benchmarks

1,100 nodes, half repeaters, 10 runs each:

```
cpu: AMD Ryzen 5 3600XT 6-Core Processor
                       │    master     │              this PR                 │
                       │    sec/op     │    sec/op     vs base                │
HandleNodesRepeated-12   5991.14µ ± 3%    41.21µ ± 7%  -99.31% (p=0.000 n=10)
HandleNodesUnique-12       6.010m ± 2%    5.967m ± 0%        ~ (p=0.393 n=10)

                       │     B/op      │     B/op      vs base                │
HandleNodesRepeated-12   1611.0Ki ± 1%   240.9Ki ± 0%  -85.05% (p=0.000 n=10)
HandleNodesUnique-12      1.566Mi ± 1%   1.753Mi ± 1%  +11.96% (p=0.000 n=10)
```

`Repeated` is one query asked over and over (the production pattern).
`Unique` makes every request a different query, so nothing is served
from cache: the time is unchanged, and it allocates 12% more because the
body is now built in memory so it can be kept.

### Expected hit rate

Replaying that day's 14,053 `/api/nodes` requests through a cache of
each lifetime:

| TTL | served from cache |
|---|---|
| 5 s | 39% |
| 10 s | 47% |
| **15 s** | **52%** |
| 30 s | 58% |
| 60 s | 69% |

### Staging A/B

The same 30 minutes of real production API traffic (5,640 requests, 279
of them `/api/nodes`), replayed at 2x on a staging host against a copy
of a production database, on `master` and on this branch:

| `/api/nodes` | master | this PR |
|---|---|---|
| p50 | 47.1 ms | **31.0 ms** |
| p95 | 165.4 ms | **114.4 ms** |
| max | 759.9 ms | 321.7 ms |
| allocated by `handleNodes` during the replay | 1,000 MB | **594 MB
(-41%)** |
| total allocated by the process | 3,704 MB | 3,349 MB (-10%) |

The -41% matches the expectation from the hit rate and the miss-path
cost: about 0.48 × 1.12 + 0.52 × 0.15 ≈ 0.61 of the old allocation.

## Tests

- Red commit `ab33bc29`: a test-only build hook, the tests and both
benchmarks. Three fail on `master`: identical queries are built once,
reordered parameters share an entry, ten concurrent misses share one
build. Three more guard against over-caching and pass on both commits:
different queries are built separately, an expired entry is rebuilt, a
geo-filter change drops entries.
- Green commit `ea0d1cf2`: the cache, plus
`TestHandleNodes_FilterListChangesMissCache` (hidden-prefix and
blacklist changes miss the cache) and `TestNodesCache_BoundedByBytes`
(stays under 32 MB; an oversized response is not cached).
- One existing test changed: `TestHandleNodesRegionUsesStore` already
clears the store's own 30 s region cache by hand before its second
request; it now clears this cache on the same line.
`TestHiddenNamePrefix_1181_NodesList`, which changes the hidden-prefix
list between requests, passes unchanged because the generation is in the
key.
- `go test -race` on the new tests passes. `make test`: 16 of 16 modules
pass.

## Review notes

- The response depends only on the query string and server state, not on
who is asking (no auth or header-dependent output in the handler), so a
shared cache is safe.
- `writeJSONBytes` writes the cached bytes with the same `Content-Type`
and status as `writeJSON`; the body has the trailing newline
`json.Encoder` adds, so output is byte-identical.
- No new `map[string]interface{}` added (the existing node maps are
unchanged; converting them to a struct is a separate, larger job).
- The TTL and byte budget are constants. They could be operator settings
later (AGENTS.md rule 8); not added here.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **S3776 on `buildNodesResponse`:** complexity 91. That is upstream's
existing handler body moved unchanged (it is also 91 on `master`); this
PR makes the handler itself simpler.
- **Fixed** in `25d40950`: the test hook moved to the cache-miss path,
so the builder stays identical to the old body rather than one point
above it. Splitting that function is a refactor of its own.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:21:50 +02:00
Alex BandClaude Opus 5.5 323b8040d4 perf(frontend): throttle nav stats to one refresh per 15 s, none in hidden tabs (#2162)
## Summary

The nav bar's packet / node / observer counts were refreshed three ways
at once: a 15 s timer, every batch of live WebSocket packets
(`debouncedOnWS`, 250 ms), and the `/stats` client cache being cleared
every 5 s while packets arrive. None of them checked whether the tab was
visible. So every open tab, including hidden ones, asked `/api/stats`
every few seconds for as long as it stayed open.

On one production instance's nginx log for a single day, `/api/stats`
was **135,654 of 187,411** CoreScope API requests (72%), about 20,000 a
day for each browser that kept a tab open. Each response is cheap thanks
to the 10 s server-side cache (#1910), so this is about request volume,
bandwidth and log noise, not server CPU.

This PR adds `createVisibleThrottle()` to `app.js` and routes all three
triggers through one throttled refresh:

- at most one `/api/stats` request per 15 s from a visible tab (the old
timer's own period);
- none from a hidden tab;
- one immediate refresh when a hidden tab becomes visible again.

No existing issue covers this.

## Performance

Measured with headless Chromium, counting `/api/stats` requests from a
real tab: 60 s visible, then 60 s with the page hidden
(`document.hidden` and a `visibilitychange` event), then back to
visible.

| `/api/stats` requests | production (current release) | staging,
`master` | staging, this PR |
|---|---|---|---|
| visible tab, per minute | 9 | 4 | 4 (at +10, +25, +40, +55 s) |
| hidden tab, per minute | 6 | 5 | **0** |
| on return to visible | n/a | n/a | 1, immediately |

The same check confirmed the nav bar still renders its counts (`… pkts ·
… nodes · … obs`).

## Tests

- Red commit `2c81a3e4`: `tests/unit/test-nav-stats-throttle.js` loads
the real `public/app.js` in a `vm` context. `createVisibleThrottle`
lands as a pass-through stub, so the test fails on its first assertion
(40 refreshes where 1 is expected).
- Green commit `f0ba92ab`: the throttle, and the nav stats timer, WS
hook and a new `visibilitychange` listener calling one throttled
trigger. All 6 cases pass: one refresh per burst, again after the
interval, at most 241 an hour, none while hidden, one on return, and a
source check that `updateNavStats` is wrapped and not called by the
timer directly.
- Added to `test-all.sh` in sorted order; `test-test-inventory.js`
passes.
- Frontend suites: 194 of 195 pass. The one failure,
`test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and
fails the same way on `master` locally.

## Review notes

- `home.js` keeps its own one-shot `/stats` load on the home page; only
the nav bar's polling changed.
- A tab opened in the background loads its counts when it is first
shown.
- The 15 s interval is a constant at the call site. Could be exposed in
the customizer later (AGENTS.md rule 8); not added here.
- No Playwright E2E added: the change is timing behaviour, covered by
the unit test against the real file and by the browser measurement
above. Happy to add an E2E if wanted.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **Fixed** in `9eb94ea1`: `node:` specifiers (S7772).
- **Kept on purpose:**
- S7773 wants `Number.isFinite`/`Number.parseInt` for the `isFinite` and
`parseInt` handed to the `vm` sandbox, but `app.js` calls the global
functions, so those are the ones it needs.
- S1523 flags `vm.runInContext`, which is how the frontend unit tests
load the repo's own `public/app.js`.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:21:30 +02:00
Alex BandClaude Opus 5.5 be6aeffe45 perf(store): drop the duplicate subpath count map (#2163)
## Summary

The subpath index keeps every contiguous run of 2-8 hops from every
stored transmission's path, in two maps:

- `spIndex`: key → count
- `spTxIndex`: key → transmissions containing that subpath, one entry
per occurrence

Both were always updated together (the only production callers were
`addTxToSubpathIndexFull` / `removeTxFromSubpathIndexFull`, which touch
both), so every count equalled `len(spTxIndex[key])`. The existing
`TestSubpathTxIndexPopulated` asserted exactly that. The index therefore
held its 1M+ keys twice.

This PR removes `spIndex`. The two analytics consumers read `len(txs)`,
and the add/remove helpers take only the transaction map. No API
response changes.

No existing upstream issue covers this; it came out of profiling a
production instance where garbage collection was 87% of CPU (30 s CPU
profile: `gcDrain` 39.9 s of 45.9 s sampled) with an 805 MB live heap
against a 1 GiB `GOMEMLIMIT`. That heap profile put 186 MB in this
index; on that instance's 7-day window there are 1.2M distinct subpath
keys (79% seen once) and 3.07M entries.

## Performance

### Go benchmark

`BenchmarkBuildSubpathIndex` builds the index for a synthetic,
production-shaped window: 41k transmissions with path lengths drawn from
the production hop-count histogram, and hops from random walks over a
400-node graph with 3-byte prefixes. That gives 1.48M keys, 85% seen
once (production: 1.2M, 79%). No real data is used. 10 runs each:

```
cpu: AMD Ryzen 5 3600XT 6-Core Processor
                     │    master    │              this PR               │
                     │    sec/op    │   sec/op     vs base               │
BuildSubpathIndex-12    2.157 ± 1%    1.610 ± 5%  -25.33% (p=0.000 n=10)

                     │  retained-B  │  retained-B   vs base              │
BuildSubpathIndex-12   238.5Mi ± 0%   187.7Mi ± 0%  -21.31% (p=0.000 n=10)

                     │     B/op     │     B/op      vs base              │
BuildSubpathIndex-12   572.0Mi ± 0%   465.6Mi ± 0%  -18.61% (p=0.000 n=10)
```

(`retained-B` is the live heap the index keeps after a forced GC: 250.1
MB → 196.8 MB.)

### Staging heap

Live heap after a forced GC (`/debug/pprof/heap?gc=1`), on a staging
host running a copy of a production database and its live MQTT feed,
attributed to the subpath index builder:

| | master | this PR |
|---|---|---|
| subpath index (`buildSubpathIndex`, cum) | 172.4 MB | 116.1 MB |
| packets in memory at the time | 27,521 | 24,086 |
| per packet | 6.27 KB | 4.82 KB (**-23%**) |

The two snapshots were taken minutes apart with different numbers of
packets loaded, hence the per-packet figure; it agrees with the
benchmark's -21%.

The same A/B replay used for the other PRs (30 minutes of real
production API traffic) showed no latency change (p50 10.3 ms vs 10.7
ms, p95 59.5 ms vs 52.9 ms, 0 errors). GC count and CPU were lower in
this run (44 → 28 collections, 13.0% → 9.6% CPU), but runs on this
shared host vary by that much on their own, so no claim is made from
them.

## Tests

- Red commit `ded359ff`: the fixture, `TestSubpathFixtureShape` (keeps
the fixture production-like: 60-95% singleton keys),
`BenchmarkBuildSubpathIndex`, and `TestSubpathIndexRetainedHeap`, which
allows at most 225 MB for the 41k fixture. It fails on `master` with
250.1 MB.
- Green commit `e399c561`: the change. The heap test passes at about 197
MB.
- `make test`: 16 of 16 modules pass; `go test -race` over the subpath,
eviction, path-hop and resolved-index tests passes.

Test changes in the green commit are mechanical, as required for a
structural change:

- struct literals stop setting `spIndex`; where a literal did not set
`spTxIndex` it now does, because appending to a nil map panics where the
old code skipped the transaction map when passed `nil`;
- count assertions read `len(spTxIndex[key])`;
- the old "spIndex and spTxIndex agree" check becomes "no key is kept
with an empty list", which is what keeps `len()` a correct count.

## Review notes

- Removal still takes out one occurrence per subpath occurrence and
deletes a key whose list becomes empty, so `len()` stays equal to the
old count through eviction.
- The store's memory estimate (`perSubpathEntryBytes`) is unchanged
here; accounting accuracy is a separate item.
- No new `map[string]interface{}`.

## SonarLint

SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.

- **Fixed** in `8af9582f`: S3776 (cognitive complexity 16 > 15) in the
test fixture. Path generation is split into graph and hop-count helpers;
the random draws stay in the same order, so the fixture is unchanged.

## Checked against the repo rules

- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 21:21:13 +02:00
n30nex 7f2c999266 fix(ingestor): record cancelled background migrations separately (#2161)
Refs #1739. Depends on the small local-validation fix in #2160; merge
that first. This branch contains its prerequisite commit plus one
migration-status commit.

An interrupted background migration currently records `failed`, making
an operator shutdown look like a migration fault. Errors matching
`context.Canceled`, including wrapped errors, now record `cancelled` and
retain the error and completion time without logging `FAILED`. Ordinary
errors, deadline expiry and panics remain failures even when the context
is cancelled; a successful callback still records `done`.

Cancelled work follows the existing retry path on the next run.
Regression tests cover classification and clearing stale
completion/error fields before retrying, then successful completion. The
migration runbook documents the distinction.

This covers the existing ingestor bookkeeping portion of #1739. The
issue's server status mapper and banner depend on closed, unmerged #1735
and are absent from current master; this PR does not add that API/UI and
leaves #1739 open for the remaining scope. No new configuration,
customizer setting or schema migration is required.

Validation: direct and wrapped cancellation failed the new regression
before the change; focused tests and the race check pass afterward.
Native full server/ingestor suites and all 194 standalone frontend
suites pass with #2160. The browser smoke suite passed 148 checks, with
three fixture-dependent skips and two pre-existing disabled cases.
Windows-only filesystem capability skips are documented in #2160; Linux
CI retains the security assertions.

Performance impact: one error-chain check when a background migration
finishes, using the existing single status update. No packet-ingestion
loop, batching, polling or retention behavior changes, and no speedup is
claimed.
2026-10-10 21:20:57 +02:00
n30nex a5b60ca594 test: make local validation independent of runner environment (#2160)
Fixes #2159. Related to #1995.

Standalone release tests failed when `GITHUB_REF_NAME` was absent and
could use an unrelated runner tag when it was present. Supply the
simulated tag explicitly; exercise the real release-notes step with both
missing and unrelated host values, including present and absent notes
files.

Two filesystem fixtures now report unsupported Windows capabilities
explicitly: POSIX permission bits, and symlink creation denied with
error 1314. Unix mode checks and all symlink safety assertions remain
intact; other symlink setup errors still fail. Production code and
workflows are unchanged. This is 29 added lines across three test files.

Validation: reproduced all three failures before fixing them; focused
regressions, the full standalone frontend suite, and native Go
server/ingestor suites pass. The POSIX permission and unprivileged
Windows symlink fixtures are explicitly skipped on this host and remain
enabled in Linux CI. The browser smoke suite passed 148 checks, with
three fixture-dependent skips and two pre-existing disabled cases. No
configuration or customizer changes.
2026-10-10 21:20:40 +02:00
efitenandClaude Opus 5.5 0737f10a5d feat(notify): optional external node alerts feed and the node.external event (#2177) (#2178)
Closes #2177.

## What

Node notifications can mail alerts computed outside CoreScope. With
`userManagement.notifications.externalAlerts` configured, users can
choose a new event, `node.external`, which is off by default and only
offered when configured:

```json
"notifications": {
  "enabled": true,
  "externalAlerts": { "url": "http://coverage-job/alerts.json", "maxAgeHours": 48, "label": "Coverage" }
}
```

The feed:

```json
{ "generatedAt": "2026-10-11T01:35:12Z",
  "alerts": [ { "pubkey": "<64 hex>", "key": "SE", "text": "never heard to the south-east", "url": "https://example.org/coverage" } ] }
```

Mail lines: `<name>: Coverage: never heard to the south-east, <time>`
when an alert appears, `<name>: Coverage resolved (SE), <time>` when it
is gone from a fresh feed.

## How

- `cmd/server/notify_external.go`: feed types, `parseExternalFeed` (RFC
3339 `generatedAt`, age check, per-alert validation),
`fetchExternalFeed` (200 only, 1 MB cap, 10 s timeout).
- `notify_eval.go`: `evalExternal`. Subjects are `<pubkey>/<key>`; a
`<pubkey>/*` row is the per-node baseline, so the first fresh feed after
watching stores the current alerts without a mail (like `foreign.new`),
and later alerts mail at once. `External == nil` (not configured, nobody
chose it, or stale) leaves every `node.external` state as it is.
- `notify_service.go`: the feed is fetched only when configured and at
least one user chose the event. Each new reason the feed is unusable is
logged once, and so is the recovery. Mail text and link: the alert's
`url`, otherwise the node page.
- `notify_handlers.go`: `availableEvents` and the new `externalLabel`
include the event only when configured; PUT answers 400 for
`node.external` otherwise.
- `internal/users/notify.go`: event constant, `OptionalNodeNotifyEvents`
(canonical order: node, optional, admin), `RemoveWatch` also deletes the
node's `<pubkey>/...` rows.
- Account page label, OpenAPI descriptions, `config.example.json`
comment, `docs/user-guide/accounts.md` section.

Decisions on the three open questions in #2177: the generic name
`node.external`; the resolved mail names the key instead of repeating
the text, so there is **no `users.db` schema change** (a bump would also
make older binaries refuse the file); one feed.

## Invariants

- Read/write separation (#1283): the server only reads the feed over
HTTP. State rows go to `users.db` through `internal/users`, like the
other events.
- A dead, garbled or old feed never mails "resolved":
`TestEvaluateExternalStaleFeedChangesNothing` and
`TestNotifierMailsExternalAlertsAndSkipsAStaleFeed` (HTTP 502 and a 72 h
old feed in a row, then a fresh empty feed resolves).

## Performance

The evaluator adds one pass over the stored states (already read every
tick) plus, per user with the event on and per watched node, the node's
current alerts. Bounded by `maxWatchesPerUser` (50) times the alerts per
node. The feed is fetched once per tick (default 5 minutes) only when
someone chose the event, capped at 1 MB. No packet-store lock is taken
for it. Nothing changes for instances without `externalAlerts`.

## Tests

- `cmd/server`: `go test` green locally (Windows;
`TestSaveGeoFilterPreservesFileMode` skipped as it needs Unix mode
bits). New: `notify_external_test.go` (feed parsing and limits, HTTP
fetch 200/404/oversized, config resolution, evaluator
baseline/new/resolved/returning/stale/not chosen/notifications off,
notifier end to end with a fake feed, account routes).
- `internal/users`: green, 2 new tests (optional event and canonical
order, `RemoveWatch` deletes external subjects only for that node).
- Frontend: `test-all.sh` green, 1 new test in
`test-notifications-ui.js` (label from `externalLabel`, escaped).
- `scripts/check-xss-sinks.sh --file public/notifications.js`: clean.

## Not done

- No E2E test: the account page change is one label, covered by the unit
test.
- Not run against a live instance yet. The first feed in use is the
coverage job of on8ar.eu (`alerts.json`, 29 alerts at the moment).

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 20:54:35 +02:00
efitenandClaude Opus 5.5 c2a590b741 feat(companions): linked-only ingest, account Companions and My coverage (#2176)
Part 3 of 3 for #2173 (companion linking): the optional linked-only
ingest filter, the account page and the "My coverage" toggle, plus the
docs. Builds on #2174 and #2175 (both merged); `upstream/master` is
merged in, so the diff shows only this part.

## Changes in this PR

**Ingestor: linked-only ingest** (`cmd/ingestor/linked_companions.go`)
- With `clientRxCoverage.requireLinkedCompanion` on (default off) and
user management enabled, every `meshcore/client/<pubkey>/...` message
whose pubkey is not in `companion_links` is dropped before any client
handler runs, and counted (`client_unlinked_dropped` in the stats file).
- The ingestor opens `users.db` read-only (`mode=ro`, raw SQL, no
`internal/users` import), like the approved-channels reader, so the
#1283 write invariant holds. `TestLinkedCompanionSetIsReadOnly` pins it.
- The set refreshes every 60 s. A miss re-reads at most once per 5 s, so
a companion linked a moment ago is accepted within seconds while a flood
of unknown pubkeys costs one read per 5 s.
- It passes everything until its first successful read, and when
`users.db` has no `companion_links` table yet (a server older than
this). After the first read it keeps the last set on a read error. It is
a filter, not a security boundary: anyone with the shared broker
credential can still publish under a linked pubkey.

**Frontend**
- Account page: device tokens show under Devices as "CoreDrive RX" with
their label; a new Companions section lists linked companions (name,
short key, linked since, last seen) with an Unlink button.
- Coverage page (`#/rx-coverage`): a "My coverage" toggle for logged-in
users (`mine=1`). Only the newest coverage reply is drawn, and a refused
`mine=1` turns the toggle off. Under "My coverage" the page does not ask
for the "nothing received" cells from #2172, and their legend entry is
hidden.

**Docs**: user guide (accounts: devices, companions),
`docs/client-rx-coverage.md` (linking, linked-only ingest, the CORS
exception), `config.example.json`, and the spec status.

## Performance

The filter runs on every client message. Benchmarks with 1000 linked
companions (12-core dev box, `go test -bench`, 3 runs each):

| Path | Time | Allocations |
|---|---|---|
| linked pubkey (lock-free map lookup) | 54 ns/op | 0 |
| linked pubkey, parallel | 11-14 ns/op | 0 |
| unlinked pubkey inside the 5 s miss gap | 85-95 ns/op | 0 |
| one `users.db` re-read | 0.38 ms | 3030 |

The unlinked benchmark asserts that no extra `users.db` read happens. A
re-read happens at most once per 5 s on a miss and once a minute from
the refresh loop.

## Verification

- `cmd/ingestor`: 15 new tests (filter off changes nothing, linked
passes, unlinked dropped before every handler, miss refresh accepts a
new link, the 5 s cap holds under a flood, unlink honoured after the
periodic refresh, missing table and missing file, fail-closed only after
the first read, read-only open, startup, stats counter) and 4
benchmarks. Full ingestor suite green.
- Frontend unit tests for the account page (device tokens, companions
list escaping, empty state, Unlink confirm and cancel, load errors) and
`tests/unit/test-rx-coverage-mine.js`.
- E2E `tests/e2e/test-user-management-e2e.js` gains a companion step
(device token under Devices, a companion linked with a real Ed25519
signature under Companions, Unlink). Run locally against the e2etest
build: 28/28. The coverage suites (`test-rx-coverage-gaps-e2e.js`,
`-noise-`, `-viewport-`) still pass.
- Frontend unit suite (`test-all.sh`) green.

## Not in this PR

- Narrowing the "nothing received" cells to the caller's companions
under "My coverage" (the track query takes one companion today).
- Per-user broker credentials or ACLs, which would turn the filter into
an access control (non-goal in the spec).

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-10 18:38:27 +02:00
efitenandClaude Opus 5.5 b2cda21a27 feat(server): companion linking routes, device tokens and bearer auth (#2175)
Part 2 of 3 for #2173 (companion linking): the server routes, bearer
auth and the link handshake. Builds on #2174 (merged); `upstream/master`
is merged in, so the diff shows only this part. Everything here is
behind `userManagement.enabled` and answers 404 with it off.

## Changes in this PR

**Device tokens and bearer auth**
- `POST /api/auth/device-token`: email + password returns a bearer
token, stored as a `kind = 'device'` session with a label. Same failure
answers and the same rate limit as login.
- Server-enforced scope for a bearer token: `/api/auth/me`,
`/api/auth/logout`, `/api/account/companions*`, `/api/account/settings`.
Everything else answers 403.
- A bearer header only counts when there is no session cookie, so an
auth proxy that adds its own `Authorization` header does not hide a
logged-in browser. Cookies never accept a device token.
- The sessions list on the account API carries `kind` and `label`;
logout with a bearer token revokes that device.

**Linking a companion**
- `POST /api/account/companions/challenge` `{pubkey}` returns
`{challenge, expiresAt, host}`: 32 random bytes, single use, 5 minutes,
bound to the caller and the pubkey.
- `POST /api/account/companions` `{pubkey, challenge, signature, name}`:
the signature is the companion's Ed25519 signature over
`corescope-link:<host>:<challenge>`
(`internal/sigvalidate.VerifyMessage`). The challenge is consumed in
every case.
- A companion linked to another account moves to the caller (the newest
proof wins). Both users get a `companion.transfer` audit row and the
previous owner a mail. That mail is a security notice: it is sent
whatever the node notification settings are.
- `GET /api/account/companions` (with `lastSeenAt` from
`client_receptions`) and `DELETE /api/account/companions/{pubkey}`.
- Linking does not touch the synced `meshcore-my-nodes`: a companion is
not a node to monitor.
- Admin user detail lists linked companions; the account data export
includes companions and device sessions.

**Coverage and config**
- `GET /api/rx-coverage?mine=1`: only coverage of the caller's linked
companions (401 when logged out, 403 for a device token). With `gaps=1`
(#2172) the reply carries no `gaps` member under `mine=1`, because the
track query takes one companion and cannot narrow to a set yet
(`TestRxCoverageMineOmitsGaps`).
- `clientRxCoverage.requireLinkedCompanion` (default off) is read and
reported in `/api/config/client`, which also announces
`userManagement.companionLinking`, so CoreDrive RX shows its login only
where linking exists. The ingestor side is part 3.
- CORS: for origins in `corsAllowedOrigins`, the device-token routes
alone also allow POST/PUT/DELETE and the `Authorization` and
`Content-Type` headers. No credentialed CORS. Documented in
`config.example.json`.
- OpenAPI and `docs/api-spec.md` cover the new routes.

## Verification

- 26 new tests in `cmd/server` and `internal/sigvalidate`, among them:
`TestCompanionLinkFlow`, `TestCompanionLinkRejections` (signature by
another key, malformed signature, challenge bound to another pubkey,
unknown challenge; reuse and expiry are covered by the store tests in
#2174), `TestCompanionTransfer`,
`TestCompanionTransferMailIgnoresNotificationSettings`,
`TestBearerScopedRoutes`, `TestBearerIgnoredWithCookie`,
`TestDeviceTokenSharesLoginRateLimit`, `TestCORSBearerRoutes`,
`TestCompanionRoutesAbsentWhenOff`, `TestRxCoverageMine`,
`TestRxCoverageMineOmitsGaps` (fails without the `mine == nil` guard),
`TestVerifyMessage`.
- `cmd/server`, `internal/sigvalidate` and `internal/users`: `go vet`
and `go test ./...` green, `gofmt` clean.
- The linking flow runs on analyzer.on8ar.eu since 9 October 2026 with
CoreDrive RX v1.19.0 (a real companion linked through the signed
challenge).

## Not in this PR

- The ingestor's linked-only filter, the account page (Devices,
Companions), the "My coverage" toggle and the user docs: part 3.
- Per-user broker credentials and per-user API keys (non-goals in the
spec).

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-10 18:20:49 +02:00
efitenandClaude Opus 5.5 9897d52cf6 feat(users): companion linking store (schema v6, device sessions, link challenges) (#2174)
Part 1 of 3 for #2173 (companion linking). This PR adds only the
`users.db` side: the schema and store methods the server (part 2) and
ingestor (part 3) build on. No route, no UI and no behaviour change for
a running server.

## Changes in this PR

- **Design spec** `docs/specs/2026-10-08-companion-linking-design.md`,
including the optional linked-only ingest.
- **Schema v6** (`internal/users/schema.go`):
- `sessions` gains `kind` (`'web'` or `'device'`), `label` and `scopes`
(`''` = full, for web sessions).
- `companion_links (pubkey PRIMARY KEY, user_id, name, linked_at)`,
cascading on user delete.
- `link_challenges`: single-use challenges bound to a user and a pubkey,
with an expiry.
- **Device sessions** (`sessions.go`): create a device session with a
label and scopes, list sessions with kind and label.
- **Challenges** (`challenges.go`): issue and consume a challenge
exactly once; consumed, expired, or bound to another user or pubkey all
fail.
- **Companion links** (`companions.go`): link, list, unlink. Linking a
pubkey that belongs to another user moves it and reports the previous
owner (transfer detection), so the server can audit and mail.
- **Janitor**: the existing janitor also prunes expired challenges
(`cmd/server/auth_service.go`, 3 lines).

## Upgrade note

The store migrates `users.db` to v6 on start. A binary that only knows
v5 refuses to open a v6 file (`users: database schema version 6 is newer
than this binary supports (5)`), so rolling back after this lands needs
the pre-upgrade `users.db` backup.

## Verification

- `internal/users`: new tests for the v6 migration from v5, device
sessions, challenges (single use, expiry, binding) and companion links
(transfer, cascade on user delete). `go vet` and `go test ./...` green.
- `cmd/server` full suite green against the v6 store.

## Not in this PR

- Routes, bearer auth and the link handshake: part 2.
- Ingestor filter, account page and "My coverage": part 3.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-10 18:08:51 +02:00
efitenandClaude Opus 5.5 f3004d977e feat(rx-coverage): show where a companion drove and received nothing (#2172)
The Mobile RX coverage page (`#/rx-coverage`) shows only cells where a
companion received something. Where a companion drove and received
nothing is just as useful for operators, and the data is already there:
the RF sample track (`client_rf_samples`, #1905/#1906) and the
receptions (`client_receptions`) use the same hex grid (`hexCellAt`,
`zoomToHexRes`). This PR adds those cells as a hatched layer under the
reception cells.

## Changes in this PR

**Server**
- `GET /api/rx-coverage?...&gaps=1` adds a `gaps` member to the
FeatureCollection: cells with moving (non-stationary) track samples and
no reception in the same `days` window and `bbox`, each with `samples`.
`gaps_truncated` is set when more than 5000 cells qualify (the
most-sampled are kept, same pattern as `coverageFeatureCap` and
`rfNoiseFeatureCap`).
- Gaps are only returned when `clientRfSamples` is on and there is no
`node` filter: with `?node=` the question would be "this node was not
heard", which these cells do not answer. `?rx=` limits both the track
and the receptions to that companion.
- Without `gaps=1` the payload is unchanged: the `gaps` member is
omitted.
- New code lives in `cmd/server/rx_coverage_gaps.go` (pure
`aggregateGaps` plus `queryTrackPoints`); the handler change in
`rx_dashboard.go` is 14 lines.
- `/api/rx-coverage` is now documented in `openapi.go` and removed from
`openapi_known_gaps.json` (#1670).

**Frontend** (`public/rx-coverage.js`, `public/node-reach-coverage.css`)
- A "Nothing received" toggle on the signal layer, on by default, only
rendered with `MC_CLIENT_RF_SAMPLES`. Hidden on the noise layer.
- Gap cells use a 45-degree hatch (an SVG pattern in a dedicated
renderer, so they sit under the reception cells). The hatch is denser
for cells driven through more often (1 / 2-3 / 4+ samples). Tooltip:
sample count. Legend entry: "driven, nothing received".
- Off state is deep-linkable as `gaps=0`.
- New token `--nq-cov-gap` for light and dark. Plain grey
(`--nq-cov-grey`) already means "received, no SNR", so the hatch keeps
the two apart.

## Threshold: one sample

Measured on a live network (analyzer.on8ar.eu, Belgium, September to
October 2026), with cells of about 250 m: a cell with track samples and
no reception in one half of a month, driven again in the other half. How
often did it receive something the second time?

| Minimum samples | Cells "nothing received" | Driven again | Received
something later |
|---|---|---|---|
| 1 | 3971 | 1263 | 29% |
| 2 | 1565 | 592 | 32% |
| 3 | 619 | 243 | 35% |
| 5 | 142 | 63 | 40% |

A higher threshold hides cells without making the rest more reliable, so
the threshold is 1 and the wording is "nothing received in this period",
never "no coverage". Using every decoded packet as reception evidence
(including 1-byte path hashes) gave 25% to 32%, so the capture rule for
1-byte hashes is not what drives this. The median cell has 2 samples.

## Verification

- Go: 7 new tests in `rx_coverage_gaps_test.go` (gap/no gap, stationary
excluded, single sample, cap keeps the most-sampled, endpoint, `?rx=`
filter, no `gaps` member when not asked / with `?node=` / with RF
samples off). The endpoint tests fail without the handler change. Full
`cmd/server` suite green locally, `go vet` and `gofmt` clean.
- E2E: `tests/e2e/test-rx-coverage-gaps-e2e.js` (4 steps: no toggle with
RF samples off, on by default with hatched cells under the reception
cells, toggle off with `gaps=0` and deep link, noise layer hides it). 2
steps fail without the frontend change. Existing
`test-rx-coverage-noise-e2e.js` and `test-rx-coverage-viewport-e2e.js`
still pass. Listed in `scripts/non-unit-tests.json`.
- Frontend unit suite (`test-all.sh`) green.
- Real data on a staging instance (7 days, bbox 50.8,4.9,51.3,5.8,
z=12): 478 reception cells and 379 gap cells, no cell in both, reception
features identical with and without `gaps=1`, 0.48 s to 0.52 s per
request both ways.
- Browser: checked on staging with real data (light theme), and light,
dark and 390 px wide with mocked data.

## Not in this PR

- A per-node variant on the Reach page ("this repeater transmitted while
a companion was nearby, and was not heard"). That needs the transmit
moments from `observations` and is a heavier query.
- `deploy.yml` does not get a line for the new E2E file, like the other
entries that are only in `non-unit-tests.json` (#2037).

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 17:39:33 +02:00
efitenandClaude Opus 5.5 673442ccb9 perf(ingestor): refresh planner statistics 24h after the last refresh, not after every start (#2158)
Refs #2146.

## Situation

The daily planner statistics refresh (#2058) runs its first `ANALYZE` 2
minutes after every ingestor start. On a cold page cache that bounded
`ANALYZE` (analysis_limit 10000) takes minutes on a 10-11 GB database,
and it holds the ingestor's single write connection the whole time:

| Where | Logged after a restart |
|---|---|
| production instance (11.5 GB) | `component=analyze
duration=469855.2ms`, plus `[neighbor-build] SLOW tick: 7m49.98s` behind
it |
| staging instance, 12 restarts logged on 2026-10-09 | 4m11s to 9m11s
each; 7 of the 12 took over 7 minutes |

`EnsurePlannerStats` already documents what that costs: ingest stalls
and buffers for those minutes. The server's reads compete with the scan
for the disk during that window.

Every restart pays this for statistics that are at most a day old.

## My change in this PR

- `RefreshPlannerStats` records when it succeeded, in a file next to the
database (`<db>.planner-stats-refreshed`, RFC 3339).
- The refresh loop waits until that time is 24 hours old: at least the
existing 2-minute stagger, at most 24 hours (`nextPlannerStatsRefresh`).
A restart within a day of the last refresh runs no `ANALYZE`.
- No recorded time (the first start after upgrading, or a new database)
refreshes after 2 minutes as before, and records it.
`EnsurePlannerStats` records its build too, so a new database no longer
runs a second `ANALYZE` 2 minutes after the first.
- A failed refresh keeps the daily rhythm instead of retrying every 2
minutes, the same as the old ticker.

## Verification

On staging (`/var/lib/docker/volumes/.../meshcore.db`, 168 h window):

```
first start:
  [analyze] next planner statistics refresh in 2m0s
  [analyze] planner statistics refreshed in 7m26.086s (analysis_limit=10000)
  [analyze] next planner statistics refresh in 24h0m0s
  meshcore.db.planner-stats-refreshed: 2026-10-09T21:08:34Z
docker restart, then:
  [analyze] next planner statistics refresh in 24h0m0s
  (no ANALYZE and no db-slow-writer line in the following 6 minutes)
```

Tests:

- `TestNextPlannerStatsRefresh`: never refreshed, 1 hour ago, 23h59m
ago, two days ago, a recorded time in the future.
- `TestRefreshPlannerStatsRecordsWhen`: a refresh records about now, and
a reopened store reads it back as a wait of about 24 hours.
- `TestRefreshPlannerStatsDisabledRecordsNothing`.
- The existing #2058 planner-statistics tests pass unchanged.
`cmd/ingestor`: `go test ./...` green.

## Not in this PR

- The time lives in a file, not in the database: the schema has no
metadata table, and a new table would need a migration for one
timestamp. If the file is lost, the only cost is one refresh 2 minutes
after the next start, as today.
- The cold `ANALYZE` itself is unchanged; it now runs once a day instead
of after every restart. Its time of day is wherever the last refresh
landed.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 11:04:08 +02:00
efitenandClaude Opus 5.5 a2b5a4931e perf(server): battery history via the primary key, cached /api/perf row counts (#2156)
Refs #2146 (items: `/api/nodes/:pubkey/battery`, `/api/perf`).

## Situation

Two endpoints ran SQL that reads more than it needs, on every request.
Timings are read-only on a production database (18.3M observations, 283k
`observer_metrics` rows), first run / second run:

| Query | Now | With this PR |
|---|---|---|
| battery history, `WHERE LOWER(observer_id) = ?` | 174 / 29 ms (plan:
`idx_observer_metrics_timestamp`) | 5.9 / 5.5 ms (plan: primary key
`(observer_id, timestamp)`) |
| `/api/perf`: `SELECT COUNT(*) FROM observations` | 118 / 116 ms on
every call | once per 60 s |

## My changes in this PR

1. **Battery history** (`5bdd3c7f`). `LOWER(observer_id) = ?` keeps
SQLite off the primary-key index, so it walked the timestamp index over
every observer's metrics in the window. The query now matches
`observer_id IN (?, ?)` with the pubkey lowercased and uppercased, the
two casings observer IDs are stored in.
2. **`/api/perf` row counts** (`e24c24cf`). The four `COUNT(*)`
(transmissions, observations, nodes, observers) are cached for 60 s.
File, WAL and freelist sizes are still read on every call.

## Tests

- `TestNodeBatteryQueryUsesObserverIndex` loads the `sqlite_stat1`
values of that database and checks the plan uses
`sqlite_autoindex_observer_metrics_1`. With the `LOWER()` form it fails
with `SEARCH observer_metrics USING INDEX idx_observer_metrics_timestamp
(timestamp>?)`.
- `TestGetNodeBatteryHistory_LowercaseStoredID`: a lowercase-stored
observer ID is still found.
- `TestGetDBSizeStatsTypedCachesRowCounts`: the cached count is served
within the TTL and refreshed after it.
- `cmd/server`: `go test ./...` green.

## Not in this PR

Other per-request queries listed in #2146 measured too small to change
on the same database: `COUNT(*) FROM dropped_packets` 0.0 ms, `SELECT
DISTINCT iata FROM observers` 0.2 ms, the observers packet-count query
1.3 ms, the 7-day active node count 0.1 ms. An observer ID stored in
mixed case would no longer match the battery query; in that database all
41 distinct `observer_metrics.observer_id` values are uppercase.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 11:03:56 +02:00
efitenandClaude Opus 5.5 be9f382526 perf(server): /paths reads canonical paths in batches and skips redundant collision checks (#2157)
Refs #2146 (item: `/api/nodes/:pubkey/paths`).

## Situation

Per-step timing in a staging build (3 most-connected nodes, 2 calls
each) showed where a `/paths` request spends its time:

| Step | Time |
|---|---|
| build candidates under the read lock | 73-115 ms |
| collision check: one `INSTR(LOWER(resolved_path), ?)` query per index
hit, 13,367-16,790 per call | 1.3-2.7 s (first call after start: 13.5 s)
|
| canonical resolved_path per survivor (`bestResolvedPath`, one or two
queries each) | 1.7-2.9 s |
| aggregation under the read lock | 82-120 ms |

So 3.2-4.1 s of a 3.4-5.1 s request was SQL.

On a production database (busiest node, 27k transmissions, read-only):
the collision check costs 1.85 s one by one and still 1.69 s batched,
because SQLite reads and lowercases every stored path either way.
Reading the canonical path (the longest observation's stored path) costs
0.18 s one by one and 0.09 s batched by id.

## My changes in this PR

- **The collision check runs only where it decides something.** The
aggregation keeps a transmission with a canonical path only if that path
names the node (`if !containsTarget { continue }`). If no stored path
names the node, the canonical one does not either. So for index hits
with a canonical path the check changes nothing, and it now runs only
for index hits without one, as before.
- **Canonical paths are read in batches.** `loadCanonicalResolvedPaths`
reads the stored path of each candidate's longest observation, one query
per 500 transmissions. When that path is missing, empty or unparseable
it calls `bestResolvedPath` for that transmission, so the result is the
same as before.
- The snapshot of observation IDs and `path_json` strings is taken in
the existing candidate read lock (string copies, no parsing).

The PR has two commits: a first version that batched both queries and
did not get faster on staging, and the change described above, whose
message has the measurements. Squash-merging is fine.

## Measurements

Staging, after the ingestor's startup `ANALYZE`, same probe for both
builds (ms):

| | without this PR | with this PR |
|---|---|---|
| node 1 (degree 55), call 1 / call 2 | 22,237 / 3,364 | 1,793 / 1,686 |
| node 2 (degree 47) | 4,427 / 3,615 | 1,828 / 1,809 |
| node 3 (degree 46) | 5,171 / 4,035 | 3,050 / 2,408 |
| 4 concurrent clients over 40 nodes, 120 s: calls / p50 / p95 | 73 /
6,717 / 15,160 | 125 / 2,861 / 10,630 |

Both staging builds also contain #2147 and #2155.

## Tests

- `TestLoadCanonicalResolvedPathsMatchesBestResolvedPath` compares the
batched result with `bestResolvedPath` on seven cases: the #810 case
(longest observation without a stored path), equal lengths, an empty
stored path, case, an observation missing from the snapshot, longest
wins, nothing stored.
- `TestLoadCanonicalResolvedPathsBatchesQueries`: 1,001 transmissions
take 3 batched queries and no per-transmission lookups.
- `TestHandleNodePaths_CanonicalPathDecidesIndexHits`: the longest
observation names A, a shorter one names B; both are index hits. A gets
the transmission, B does not, with no per-transmission query. Run
against master, the A/B results are the same and only the query-count
check fails (1 there, 0 here).
- The existing `/paths` tests pass unchanged (#810, #1278, #1352
collision and fallback cases, anchor bias).
- `cmd/server`: `go test ./...` and `go test -race` green.

## Not in this PR

- `nodeInResolvedPathViaIndex` still calls the per-transmission check;
it has no production caller, only tests.
- The concurrent p95 is still 10.6 s on staging. What is left there is
the 4-connection read pool shared with other endpoints; a staging A/B
with 8 connections gave no gain (`/paths` throughput dropped from
133-146 to 81-103 calls per 120 s), so the pool stays at 4.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 10:58:33 +02:00
efitenandClaude Opus 5.5 3e21b179ec fix(server): drive the reach scan from the observations time window (#2155)
Refs #2146 (item: `/api/nodes/:pubkey/reach`).

## Situation

The reach scan joins `observations` to `transmissions` and filters on
`o.timestamp >= ?` plus one `o.path_json LIKE '%"XX"%'` per token. With
the planner statistics `ANALYZE` produces on a large database, SQLite
runs it as `SCAN t`: every transmission, with an index probe into
`observations` (`idx_observations_tx_ts`) per row. The cost is then the
size of `transmissions`, whatever node is asked for and whatever the
window.

## My change in this PR

`scanReachRows` uses `CROSS JOIN` for the transmissions join. In SQLite
that fixes the join order: `observations` is the outer loop, read
through `idx_observations_timestamp` for the window, and each row looks
up its transmission by primary key. Results are the same rows; it is
still an inner join.

The query text moves into `reachScanSQL(nTokens)`, so a test can inspect
the plan.

## Measurements

Read-only on a database with 18.3M observations and 1.5M transmissions
(Python `sqlite3` 3.45.1 on the host), 7-day window:

| node (rows returned) | `JOIN`, first run | `JOIN`, warm | `CROSS
JOIN`, both runs |
|---|---|---|---|
| busiest relay (54,387) | 10.29 s | 3.27 s | 1.95 s |
| quartile (3,913) | 3.16 s | 3.15 s | 1.81 s |
| median (1,310) | 3.17 s | 3.18 s | 1.82 s |
| a repeater reported as timing out in the browser (34,209) | 14.11 s |
3.25 s | 1.93 s |

On that instance `/api/nodes/:pubkey/reach` reports p95 188.6 s and max
192.6 s over 34 calls in `/api/perf`; a single uncached call from the
host took 3.4 to 3.6 s, so most of that is waiting for the 2 build slots
and the read pool.

On a staging instance (smaller database), with unique cache keys and
`days` 6 or 7, one cycle per build after the ingestor's startup
`ANALYZE` had finished:

| | without this PR | with this PR |
|---|---|---|
| sequential, 12 calls, p50 | 3402-3801 ms | 2078-2467 ms |
| 2 concurrent clients (= `reachMaxConcurrentBuilds`), p50 / p95 |
3575-4087 / 4049-5291 ms | 2271-2462 / 2814-2928 ms |
| calls completed in 120 s | 58-68 | 96-101 |

The staging builds also contain #2147.

## Verification

- `TestReachScanStartsFromObservationsTimeIndex` creates the two
production indexes, loads the `sqlite_stat1` values from the
18.3M-observation database, and checks that the plan starts with `SEARCH
o USING INDEX idx_observations_timestamp`. With the old `JOIN` it fails
with the production plan (`SCAN t`, then `SEARCH o USING INDEX
idx_observations_tx_ts`).
- The existing reach tests (`TestScanReachRows_DecodesRows`,
`TestScanReachRows_CapTruncates`, attribution tests) pass unchanged.
- `cmd/server`: `go test ./...` green.

## Not in this PR

- The scan still reads every observation in the window and applies the
`LIKE` per row, so its cost grows with the window and with traffic.
Removing the scan needs an index on hop tokens (for example a table the
ingestor maintains). I have not measured what that would cost in
storage.
- The build-slot limit (2) and the 5-minute cache are unchanged.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 21:20:34 +02:00
efitenandClaude Opus 5.5 f50312767f fix(server): no SQL while holding PacketStore.mu (#2147)
Refs #2146 (first item: "Stop issuing SQL while holding `store.mu`").

## Situation

Several store functions ran SQLite queries while holding
`PacketStore.mu`. The server's read pool has 4 connections
(`db.go:153`), so a store-lock holder could wait for a pool connection.
`mu` is a `sync.RWMutex`: once the ingest poller asks for the write
lock, every new reader queues behind it, including `/api/healthz`. A
test that scans the package (below) finds 16 such call sites on `master`
at 9dbc2875.

## My changes in this PR

One commit per part:

1. **Caches serve stale while one refresh runs** (`4d9a818e`).
`getCachedNodesAndPM`, `resolveRegionObservers` and `resolveAreaNodes`
ran their SQL inline on a 30 s TTL miss, also from ingest under the
write lock. Past the TTL they now return the cached value and start one
background refresh per key (`cache_refresh.go`).
- A cache that was never built, or was invalidated with
`InvalidateNodeCache`, still loads inline. A generation counter stops a
refresh that started before an invalidation from marking the cache
fresh.
- The region cache no longer holds `regionObsMu` across its query. Its
refresh re-queries only the regions looked up since the previous
refresh. That keeps it bounded, because the region comes from the query
string.
2. **resolved_path after the lock** (`750dd951`). `QueryPackets`,
`QueryMultiNodePackets`, `GetTransmissionByID`, `GetPacketByHash`,
`GetPacketByID`, `GetObservationsForHash` and `GetNodeHealth` record the
lookups under the lock (`rpJob`, with a snapshot of observation IDs and
path JSON) and run them after releasing it. The number of queries per
row is unchanged.
3. **Single queries before the lock** (`bf85e63e`): the `node=` filter
(`resolveQueryNode`), multi-node pubkey resolution, the can_relay sets
in node health, the observers read for hash-size home regions
(`observerIATAs`), and the stats read in topology.
4. **Bulk health** (`93b7e8ec`): the nodes-table read runs before the
lock. The region filter, which needs the in-memory indexes, is applied
to those rows under the lock.
5. **`/paths` observations race** (`c30e9b43`, `fbf0d7ef`).
`handleNodePaths` read each candidate's canonical resolved_path after
releasing the lock, and that lookup walked `tx.Observations`, which
ingest appends to under the write lock. `-race` reports it in the new
test. The handler now copies the survivors' observation IDs and path
JSON in a short read lock, and parses after releasing it.
`fetchResolvedPathForTxBest` had no caller left and is replaced by
`bestResolvedPath`.
6. **Topology node count** (`4094d215`). Topology used `db.GetStats()`
for one field, the 7-day active node count. `GetStats` also runs
`COUNT(*)` over transmissions and observations. That count is now
`CountActiveNodes`, one query on `nodes`, and `GetStats` uses it too.

## Tests

- `TestNoSQLUnderStoreLock` parses the package and fails on any call
inside a `store.mu` section that reaches SQL, directly or through other
package functions. It allows the three caches above as callees (they
only query when cold) and the startup loaders as holders.
`TestNoSQLUnderStoreLockDetects` runs the scan on a small synthetic
source, so a scan that stops matching cannot pass silently. On master:
16 call sites. On this branch: 0.
- `TestSlowResolvedPathSQLDoesNotStallReaders` reproduces the convoy. A
test hook makes the resolved_path query take 300 ms; `/api/packets` sits
in it, the ingest writer queues, then a reader takes `RLock`. With the
lookup back under the lock (as on master) the reader waited 590 ms. On
this branch it did not wait measurably.
- `TestResolvedPathSQLRunsWithoutStoreLock` checks all seven functions
from part 2. `cache_swr_test.go` covers stale serving, one refresh per
key, inline rebuild after invalidation, and a cold region not blocking a
cached one. `TestNodePathsDoesNotReadObservationsUnlocked` fails under
`-race` with the old read. `TestCountActiveNodes` pins the 7-day count
against `GetStats`.

## Verification

- `cmd/server`: `go test ./...` green, and `go test -race` on the full
package green with 0 races, both on the final commit.
- E2E against a local server on the CI fixture: `test-e2e-playwright.js`
148/151 (3 skipped by the suite itself), plus
`test-issue-1486-collapse-reopens-detail-e2e.js`,
`test-issue-1692-packets-init-parallel-e2e.js` and
`test-channel-issue-1087-e2e.js`, all passing. The other E2E suites were
not run locally.
- Browser: packets, node detail, analytics hash sizes and topology
render. In the API responses, resolved_path is present on 44 of 50
packets, on packet detail, and on health `recentPackets`. No console
errors.

### Staging measurement

A staging instance with 129k transmissions in memory. Load ran on the
same host: 4 threads looping `/paths`,
`/api/packets?limit=200&expand=observations`, node health,
`hash-sizes?region=`, `topology?region=` and `/api/packets?node=`, while
one thread sampled `/api/healthz` every 200 ms. Each build had one
cycle: 360 s during the ingestor's startup `ANALYZE` (heavy disk reads
for 7 to 9 minutes after every restart), then 2 runs of 120 s after it
finished. Values in ms.

| | without this PR | with this PR |
|---|---|---|
| **During ANALYZE** | | |
| `healthz` p95 / p99 / max | 156 / 1105 / 5681 | 42 / 905 / 4280 |
| `topology?region` p99 / max | 2537 / 3061 | 1942 / 4159 |
| `hash-sizes?region` max | 5809 | 2866 |
| node health max | 5719 | 1110 |
| `packets?expand` max | 3912 | 2570 |
| calls over 5 s, excluding `/paths` | 2 | 0 |
| **After ANALYZE** (two runs) | | |
| `healthz` p95 / max | 444-667 / 2139-4356 | 189-335 / 1377-3238 |
| `topology?region` max | 3401-5363 | 1440-1927 |
| `/paths` p50 | 1466-4736 | 1125-2648 |
| calls over 5 s | 3 | 0 |

The two builds were measured about 50 minutes apart, so the background
traffic was not identical.

An earlier version of this branch, without `fbf0d7ef`, raised `healthz`
p95 to 2.7-2.9 s in the same setup, because it parsed every `/paths`
candidate's path JSON under the read lock. Without `4094d215`,
`topology?region` reached 8-17 s during `ANALYZE` in three cycles. The
table is for the final branch.

## Not in this PR

- The query count of `/paths` (1 to 3 per candidate,
`resolved_index.go:160`). It still dominates that endpoint, at p95 8 to
15 s on staging. It is its own item in #2146.
- Batching the resolved_path queries.
- The other per-endpoint items in #2146 (reach, channels, perf,
observers, config/regions, battery, neighbors).
- A cache's first lookup still queries inline, under the lock where the
caller holds it. That happens once per key per process, by design, and
the invariant test allows it explicitly.
- `cmd/ingestor` is unchanged. No measurement on an instance the size of
the ones in #2146.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 19:58:42 +02:00
efitenandClaude Opus 5.5 9dbc287579 docs: release notes for v3.14.0 (#2143)
Release notes and the CHANGELOG entry for v3.14.0. The release workflow
reads `docs/release-notes/v3.14.0.md` from the tagged commit as the
release body, so this lands before the tag.

23 commits since v3.13.1: 13 fix, 8 feat, 1 ci, and this docs commit.
The headline is optional user accounts (#2128). "Read this before
upgrading" covers the new limits on unauthenticated endpoints, the
`traffic_share_score` drop (#2117), the region filter window (#2114) and
the config file mode (#2126).

After merge: tag this commit as `v3.14.0` once its `:edge` image is
built, then dispatch the CI/CD pipeline on the tag.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
v3.14.0
2026-10-08 23:19:04 +02:00
nullrouten0andClaude Mythos 5.1 57395e2b28 fix(ui): escape observer iata, name and id, and node names, in eight render sinks (#2118)
Anyone who can publish to an MQTT broker that CoreScope reads controls
three strings: the observer `iata` (a topic segment), the observer
`name` (the JSON `origin` field) and the observer `id` (a topic
segment). Anyone with a radio controls the node `name` in an ADVERT.

Eight places wrote these into the page with `innerHTML` without
escaping, so a crafted value runs as script in every visitor's browser:

| file | what |
|---|---|
| `packets.js` | Region badge in the group, child and single packet rows
|
| `observers.js` | Region badge in the observers table and in the
slide-over |
| `live.js` | "Heard By — Regions" header in the node detail |
| `nodes.js` | observer name in the clock-skew evidence panel |
| `analytics.js` | `observer_id` inside `data-observer=` and `id=`
attributes |
| `app.js` | node name in the favorites dropdown |
| `customize-v2.js` | node name in the geofilter prune preview (admin
page) |

**Fix:** each value now goes through the existing `escapeHtml()` /
`esc()` helper. No behaviour change for normal values.

**Tests:** one test per sink added to
`tests/unit/test-xss-escape-sinks.js`, following the file's existing
pattern (pull the template out of the source, render it with a hostile
value, check the output is escaped). Reverting any single fix makes its
test fail. 41/41 pass.

**Notes for reviewers:**
- The ingestor uppercases `iata` before storing it. That is not a
defence: tag and attribute names are case-insensitive, and script can be
written with numeric character references.
- An operator who has set an `observerIATAWhitelist` is protected from
the `iata` sinks but not from the name/id sinks.
- Follow-ups worth a separate PR: validate `iata` and observer `id` to a
safe character set at ingest, cap the `origin` length, and add a
`Content-Security-Policy` header. This PR is escaping only so it is easy
to review and safe to ship.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 22:59:25 +02:00
n30nex 29095658e9 fix(nodes): explain packet and observation counts (#2132)
Red commit: `d8a6d42` ([CI
assertion](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37695889461)).
Review-regression red: `f73643b` ([CI
assertion](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37703995494)).

Fixes #2131. Partial fix for #1154: U2 only; other audit items remain
open.

Both node views explain distinct transmissions versus recorded
observations. The shared help is reachable by keyboard and touch, stays
readable at either scroll edge, and remains open when the pointer enters
its text. Escape dismisses help while preserving node context and focus;
a second Escape keeps the existing navigation behavior.

Count expressions and API requests are preserved. One controller per
nodes page handles one active help element and one queued animation
frame. Scroll/resize work is coalesced; listeners are removed on page
teardown. Work is constant relative to packet/node counts, with no
dataset scans or new settings/dependencies.

E2E assertion added: `tests/e2e/test-e2e-playwright.js:114`. Thirteen
cases cover exact counts, accessible description, keyboard/touch, hover
transfer, Escape, and viewport/scroll bounds. Existing metric assertions
are unchanged.

Browser verified: http://localhost:63508 using Chromium, controlled
fixtures and actual API data. Desktop/mobile screenshots retained
locally; additional checks exercised resizing and repeated page cleanup.

Local validation: 189 standalone suites; core browser 148 passed, 3
fixture-dependent skips (existing manual skips unchanged);
location/metric suite 6 passed; syntax, theme variables and lint passed
(0 errors, 92 existing warnings).

## Preflight overrides
- OpenClaw preflight/profile unavailable; repository checks and local
Chromium used.
- Windows ingestor symlink fixture lacks permission. Backend files are
unchanged; full Linux Go CI is required.
2026-10-08 22:37:15 +02:00
efitenandClaude Opus 5.5 2456762179 feat: account data export and users.db backup (#2141)
A small follow-up to parts A to E of #2128: users can download all data
the instance holds about them, and the server backs up `users.db`.

Stacked on #2140. Review the commits after that branch.

## The situation

- An instance with accounts holds personal data, but a user has no way
to get a copy (GDPR articles 15 and 20). Self-delete already exists.
- `users.db` is the only file the server writes, and nothing backs it
up: `GET /api/backup` snapshots the analyzer database only.

## What this PR adds

**Data export.** `GET /api/account/export` and a "Download my data"
button on the account page give one JSON file with:
- the profile, including activation and a pending email change;
- sessions (times and user agent);
- synced settings, own proposals, notification prefs and watches;
- audit rows where the user is actor or target, with other accounts as
an id only;
- the mail log.

Password hash, session and link tokens and the unsubscribe token are
left out: they are credentials. Each export is audited as `user.export`.

**users.db backup**
- Daily `VACUUM INTO` snapshots through `internal/users`, on by default
when user management is on. They go to `backups/` next to `users.db`,
the 7 newest are kept, the directory is 0700 and the files 0600.
Rotation only touches `users-<timestamp>.db` files and never the
snapshot it just wrote. Config: `userManagement.backup {enabled, dir,
keep}`.
- `GET /api/admin/users-backup`: an admin downloads a fresh snapshot to
keep off the server; audited as `user.backup`.
- `docs/user-guide/accounts.md` gains a restore procedure. It covers
what a restore brings back (accounts deleted after the snapshot, old
passwords, revoked sessions) and how to clean that up afterwards.

Spec:
[`docs/specs/2026-10-08-account-export-and-users-backup-design.md`](https://github.com/efiten/CoreScope/blob/feat/account-export-backup/docs/specs/2026-10-08-account-export-and-users-backup-design.md).

## Verification

- `internal/users` with `-race` and the `cmd/server` suite pass locally,
30 new Go tests. The file-mode tests skip on Windows.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 27 of 27 steps locally.
- On our production instance since 8 October 2026; the first snapshot
there was 143 kB.

## Not in this PR

- Importing an export into another instance.
- Off-site upload of the snapshots. The admin download covers that by
hand.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 22:16:49 +02:00
efitenandClaude Opus 5.5 de364a8bb6 feat: mail notifications for watched nodes (user management part E) (#2140)
Part E of #2128: a logged-in user watches nodes and gets one mail when a
watched node goes offline, comes back, or reports a low battery. Admins
can add instance events: a new foreign node, and an observer going
offline or back. This covers the per-user mail part of #775.

Stacked on #2139 (part D, which builds on #2138). Review the commits
after that branch.

## The situation

A sysop learns that a repeater went silent, or that its battery is
running down, only by opening CoreScope. #775 asks for notifications.
Accounts (part A) now give a verified address to mail and a mailer with
delivery status.

## What this PR adds

**Events.** They use the instance thresholds, so a mail does not
disagree with the node page:
- `node.offline`: silent for the role's `healthThresholds` window. Last
heard comes from the packet store, and repeaters and rooms count relayed
traffic (#1598).
- `node.battery`: advert telemetry below `batteryThresholds.lowMv`,
recovered at `lowMv + 100`.
- Admins only: `foreign.new` (once per node) and `observer.offline`.

**Store.** Schema v5 with `notification_prefs`, `notification_watches`
and `notification_state`.

**Server** (only with `userManagement.notifications.enabled`)
- A notifier next to the janitor, every `intervalMinutes` (5). The
evaluator is a pure function of watches, prefs, stored states and node
snapshots.
- No burst after a restart or a new watch: a subject's first evaluation
stores its state without mailing. The loop waits for the startup load.
While ingest is stale (newest packet older than 30 minutes) the offline
checks pause, and after recovery they wait one silent window.
- One mail per user per check. Limits: 20 per user and 100 per instance
per rolling 24 hours, 50 watches per user. A change over a limit is
recorded and never mailed later.
- Every mail has a one-click unsubscribe (`List-Unsubscribe` and
`List-Unsubscribe-Post`). The GET only redirects to a confirm page, so
link scanners cannot unsubscribe anyone. Node names are cleaned of
control and bidi characters before they go into a mail.
- Routes: `GET`/`PUT /api/account/notifications`, `PUT`/`DELETE
/api/account/notifications/watches/{pubkey}`, `POST
/api/account/notifications/watch-my-nodes` (copies the synced "my nodes"
list), `GET`/`POST /api/notifications/unsubscribe`. `mailer.Message`
gains `Headers`.

**Frontend.** A "Notify me" toggle on the node pages, a Notifications
section on the account page (`#/account?section=notifications`), an
unsubscribe page, and the notification mail count on the admin overview.

Spec:
[`docs/specs/2026-10-07-node-notifications-design.md`](https://github.com/efiten/CoreScope/blob/feat/notifications/docs/specs/2026-10-07-node-notifications-design.md).

## Performance

Each check reads watches, prefs and states from `users.db` (one query
each), the watched nodes from the analyzer DB in chunks of 500, and
last-heard times from the packet store.

| Measurement | Result |
|---|---|
| One check on our staging instance, 130,000 packets in memory, 1
watcher | 6.9 ms, of which 0.93 ms holding the store read lock |
| Evaluator benchmark, 100 users with 50 watches each, 2,000 nodes |
1.75 ms |

No work is added to ingest, broadcast or any request path.

## Verification

- `internal/users` and `internal/mailer` with `-race`; the `cmd/server`
suite plus `-race` on the notifier tests; 74 new Go tests. The same
Windows-only failure as noted in #2138 applies.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 26 of 26 steps locally. With notifications off
the new steps fail, so they do test the flag.
- On our staging and production instance since 7 October 2026.

## Not in this PR

- Other channels (Discord, Telegram, webhooks). Detection is separate
from delivery, so one can be added.
- The topology, RF and anomaly alerts of #775.
- If the whole server starts after a feed outage that already ended, the
grace window is unknown and a watcher can get one wrong offline mail.
The user guide says so.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 22:02:39 +02:00
efitenandClaude Opus 5.5 abbe6846ff feat: propose and approve hashtag channels (user management part D) (#2139)
Part D of #2128: logged-in users propose a hashtag channel, an admin
approves, rejects or later revokes it, and the ingestor decrypts
approved channels without a restart. It is the first consumer of a
generic proposal table, and answers the request in #2092.

Stacked on #2138 (part C). Review the commits after that branch; GitHub
shows C's commits here until it is merged.

## The situation

- An instance decrypts only the hashtag channels listed in
`hashChannels`. Users who want a city or club channel shown have to ask
the operator, who edits `config.json` and restarts the ingestor.
- #2092 asked for a suggest-and-approve flow; a downstream fork
(dborup/CoreScope#99) runs one with anonymous suggestions and a file
queue.

## What this PR adds

**Name rules (`internal/channel`).** `ValidateHashtagName`, mirrored in
the browser:
- at most 31 UTF-8 bytes including `#` (firmware
`ChannelDetails.name[32]`), case preserved, `Public` refused in any
case;
- control, bidi, separator and format characters refused except ZWJ,
plus invisible fillers (`Other_Default_Ignorable_Code_Point`, U+2800,
non-ASCII spaces).

Go and JS give the same verdict for every code point.

**Store.** Schema v4, one generic `proposals` table (kind, subject,
status pending/approved/rejected/revoked, proposer, reviewer, note). One
row per (kind, subject); limits are checked in the same transaction as
the change.

**Server** (only with `userManagement.channelProposals.enabled`)
- `POST /api/proposals`, `GET /api/account/proposals`, `GET
/api/admin/proposals`, `POST
/api/admin/proposals/{id}/approve|reject|revoke`.
- 400 for an invalid name. 409 for a duplicate, a rejected name, or a
name already in `hashChannels`/`channelKeys` ("this channel is already
decrypted on this instance"). 429 over a limit (`maxPending` 100,
`maxApproved` 128, `perUserPerDay` 5).
- `GET /api/channels` gains `approvedChannels`, so approved channels are
listed before they have traffic. With the feature off the response is
byte-identical.

**Ingestor.** Opens `users.db` with `mode=ro` every 60 seconds,
re-validates the names, and swaps its key map through an
`atomic.Pointer`. Configured channels always win, and a failed read
keeps the last good set. The ingestor never writes `users.db`; AGENTS.md
now says so.

**Frontend.** "Propose for everyone" in the add-channel dialog, an
"approved" marker in the channel list, "My proposals" on the account
page, and an admin "Proposals" tab with a warning on approve. The packet
detail pane now escapes the channel name: until now only the operator
chose those names.

Spec:
[`docs/specs/2026-10-07-channel-proposals-design.md`](https://github.com/efiten/CoreScope/blob/feat/channel-proposals/docs/specs/2026-10-07-channel-proposals-design.md).

## Performance

Every key is tried on each GRP_TXT that no key opens (`decodeGrpTxt`).
Benchmark of that worst case:

| Keys | Time per packet |
|---|---|
| 320 configured | 204 to 216 µs |
| 320 configured + 128 approved | 272 to 320 µs |

That is about +33%, linear in the key count and capped by `maxApproved`.
With the feature off the decoder gets the identical map and no ticker
runs.

## Verification

- `internal/channel`, `internal/users` (with `-race`), `cmd/server` and
`cmd/ingestor` pass locally, 53 new Go tests. The ingestor's
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced` needs symlink rights on
Windows and fails on `master` too.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E: 21 of 21 steps locally (propose, approve, the
channel appears, revoke removes it).
- On our staging and production instance since 7 October 2026.

## Not in this PR

- Private (PSK) channels: their keys stay in the browser (#725).
- Anonymous proposals.
- Names that need ZWNJ (some Persian and Urdu spellings) and subdivision
flags are refused, because ZWNJ and tag characters are format
characters. Allowing them later only widens the rule, so no approved
name gets stranded.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 21:51:42 +02:00
efitenandClaude Opus 5.5 18a634c115 feat: admin dashboard with overview and audit log (user management part C) (#2138)
Part C of #2128: an admin area with an overview, the user table and a
global audit log. Off by default like the rest of user management: with
`userManagement` off nothing changes, and non-admins get no new UI.

This is the first of 4 stacked PRs (C, D, E, then a small export/backup
follow-up). Each later one contains this branch; review them in order.

## The situation

- An admin could manage users one by one in `#/admin/users`, but nothing
showed whether the instance needs attention: accounts stuck in
activation, bouncing mail, someone guessing a password, an MQTT source
that dropped.
- The audit log existed (`internal/users/audit.go`) but could only be
read per user, and logins were not recorded.

## What this PR adds

**Store (`internal/users`)**
- Schema v3: an index on `audit_log(at)`.
- `AuditList` with filters (action or group prefix like `user.login.*`,
user as actor or target, period) and keyset pagination; `PruneAudit`;
`Stats` for the user figures.

**Server**
- `GET /api/admin/stats` (typed struct): accounts by status, admins,
registrations and active users over 7 and 30 days, logins and failed
logins in 24 hours, mail by final status, and the attention items
computed server-side.
- `GET /api/admin/audit`: filtered, newest first, `next` cursor.
- `GET /api/admin/users` gains `bouncing=1`.
- Logins are audited as `user.login` and `user.login.failed` (reason
`wrong_password`, `pending` or `disabled`). An unknown address writes no
row. The writes are asynchronous, so the login response does not wait on
`users.db`. Login rows are pruned after 90 days.

**Frontend**
- `#/admin?tab=overview|users|audit`: `admin.js` (tab shell),
`admin-overview.js` ("Needs attention", Users card, System card from the
existing health, MQTT and observer endpoints), `admin-audit.js` (filters
in the URL, "Load more"). `admin-users.js` becomes the Users tab; the
old `#/admin/users` link rewrites to it.
- `/api/healthz` is read on open and on Refresh only, not on the
60-second timer, because it walks every packet under a read lock.

Spec:
[`docs/specs/2026-10-07-admin-dashboard-design.md`](https://github.com/efiten/CoreScope/blob/feat/admin-dashboard/docs/specs/2026-10-07-admin-dashboard-design.md).

## Verification

- `internal/users` and `cmd/server`: `go vet` and `go test` pass
locally, 26 new Go tests. One upstream test,
`TestSaveGeoFilterPreservesFileMode`, also fails on Windows on plain
`master` (file modes) and is unrelated.
- `sh test-all.sh` exits 0; the XSS gate in diff mode passes.
- User-management E2E with the `e2etest` build: 13 of 13 steps locally.
- Running on our staging and production instance since 7 October 2026.

## Not in this PR

- Restricting existing pages (perf, MQTT status) to logged-in users or
admins. The admin area adds no access mechanism of its own, so that
stays a later router rule plus endpoint check.
- Server-enforced customizer tab restrictions (#1508).
- The 24-hour, 5-attempt and 10-minute attention thresholds are
constants for now (AGENTS.md rule 8: customizer later).

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 21:38:30 +02:00
efitenandClaude Opus 5.5 a4445b4985 ci: run E2E, Go tests and the image build side by side (#2135)
Closes #2134.

## What changes

The test jobs and the image build now run side by side. Only the GHCR
push waits for all of them.

```
changes ─┬─ go-test ─────────────────────────┐
         ├─ race-test (ingestor changes only)  │
         ├─ e2e-shard ×3 ── e2e-test (gate) ───┼─ build-and-publish → deploy, publish
         └─ image-check (two-arch build) ──────┘
```

- **`e2e-shard` (new, matrix of 3).** It needs only `changes`. The suite
list stays in the workflow: each line now reads `suite <shard> <file>
[VAR=value ...]`, with the same environment variables as before.
- Shards are balanced on the measured suite times, about 6 minutes each.
- A shard number outside 1..3 fails the step, so a typo cannot drop a
suite silently.
- `test-issue-1648-m4-icons-e2e.js` stays in the same shard as, and
after, the a11y suite. Its distance-tab check only counts once the lazy
distance index (#1011) has been built, and the a11y suite is what first
requests it.
- **`e2e-test`** is now a gate under the old check name. It runs only if
all three shards passed, checks that three results arrived, and builds
the same `e2e-badges` artifact as before.
- **`image-check` (new).** It runs the two-arch build (no push) and the
arm64 QEMU smoke on every event, beside the tests. It also computes the
build metadata once, which the push reuses.
- **`build-and-publish`** now needs `go-test`, `e2e-test` and
`image-check`. On a PR every step is skipped. On push and tag refs it
only pushes, with the same build args as `image-check`, so every layer
should be a cache hit.
- **The unused `docker compose build` of the staging image is gone.**
`docker compose config --quiet` still validates the file.
- `AXE_SCREENSHOT_DIR` now actually reaches the a11y suite.
- `tests/unit/test-issue-1956-release-routing.js` follows the new jobs,
and checks more than before:
- a failure in `go-test`, `e2e-shard`, `e2e-test` or `image-check` skips
`build-and-publish`;
- the only non-publishing build lives in `image-check` and runs on every
event;
  - the push takes its build args from `image-check`.

## Measurements

The old numbers are from upstream; the new ones are from PR runs on my
fork (a branch on top of this commit, plus one temporary commit that
also triggers the workflow for PRs into a test base branch).

| | Before (PR run 37766553643) | After (fork runs 37812660655 /
37814554422) |
|---|---|---|
| Go Build & Test | 4.9 min | 7.2 / 7.5 min, in parallel |
| E2E | 18.4 min | 3 shards of 7.1–9.0 min, in parallel |
| Image build | 6.9 min, after E2E | 5.9 / 5.4 min, in parallel |
| **Total** | **30.5 min** | **9.6 / 10.6 min** |

**Same tests:** I compared the second fork run with the E2E log of the
upstream run, suite by suite. All 118 suites ran, and each one passed
the same number of checks (1161 in total).

**Cost:** more runner minutes, because each shard repeats about 2.5
minutes of setup. Standard runners are free for a public repository.

## Not verified here

- **A master push.** On a fork the GHCR push cannot run. The expectation
is that `Build and push to GHCR` gets cache hits from `image-check` and
finishes in a minute or two instead of 4–4.5. The first master run after
merge will show whether it does.
- **Branch protection.** If it requires other job names than `🎭
Playwright E2E Tests` and `🏗️ Build & Publish Docker Image`, those may
need adding. I cannot read the protection settings, so I could not
check.

## Tests run locally

- `sh test-all.sh` passes (on Windows with `PYTHONUTF8=1`).
- `test-issue-1956-release-routing.js` passes, and fails as expected
when `go-test` is removed from the `needs` of `build-and-publish`.
- actionlint 1.7.7 reports only the existing self-hosted runner label.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-08 20:13:09 +02:00
nullrouten0andClaude Mythos 5.1 caa66356bd fix(api): cap concurrent cold reach scans at 2, answer 429 beyond that (#2127)
`GET /api/nodes/{pubkey}/reach` is unauthenticated and a cold-cache
request runs a full scan. `singleflight` only collapses identical keys,
so distinct `(pubkey, days)` requests each start a scan, and they share
the 4-connection SQLite pool with every other handler. A handful of
requests for different nodes can hold every reader and stall the whole
API.

**Fix:** a small semaphore allows two cold scans at once. A cold request
that finds both slots busy gets `429` with `Retry-After: 5` instead of
queueing. Cached answers are unaffected, and the two in-flight scans
still complete normally.

**Tests:** `node_reach_concurrency_test.go` — with both slots held a
cold request returns 429 with `Retry-After`; after release the same
request runs normally. Full server suite passes.

**Note:** our instance runs a fork that already bounds reach work
differently (an async job queue), so this exact change is not what we
run in production. The test suite is the verification here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 16:26:15 +02:00
nullrouten0andClaude Mythos 5.1 0a6870c508 fix(api): cap request bodies on /api/decode and /api/packets/observations (#2119)
`POST /api/decode` and `POST /api/packets/observations` are
unauthenticated and decoded the request body with no size limit.
`json.Decoder` buffers the whole value in memory, and `/api/decode`
hex-decodes the string before the 184-byte payload check runs. On a test
instance one 100 MB request raised RSS by about 180 MB. Many requests in
parallel can push a container past its memory limit, and a restart
reloads the dataset for several minutes.

**Fix:** `http.MaxBytesReader` on both handlers, the same way the
geo-filter PUT and `/api/paths/inspect` already do it.

| endpoint | cap | why |
|---|---|---|
| `/api/decode` | 4 KB | a packet is at most ~512 hex chars |
| `/api/packets/observations` | 64 KB | 200 hashes of 64 chars is ~14 KB
|

Oversized bodies get HTTP 413 instead of a generic 400.

**Tests:** `body_limits_test.go` — normal bodies still return 200,
oversized bodies return 413, for both endpoints. Full `go test ./...` in
`cmd/server` passes.

**Note:** deployments behind a reverse proxy with its own body limit
were already safe. The shipped Caddy config has no body limit, so
default installs were not.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 16:26:10 +02:00
nullrouten0andClaude Mythos 5.1 2e7a4ebcfe fix(ingestor): bound the unauthenticated /neighbors report (#2122)
`handleNeighborsReport` trusted whatever an observer published on the
`/neighbors` topic. The sender chose `origin_id` (whose "self" it is)
and could list any pubkey as a responded neighbor; each got its
`configured_scope` written with any scope string, stamped with the
sender's own timestamp. The store is last-write-wins on that timestamp
and `normalizeReportTS` accepted any RFC3339 time, so one report dated
years ahead was written once and then blocked every genuine later report
for that node until someone edited the database. The value is shown on
the reach page as the confirmed scope and feeds `/api/scope-audit`.

**Fix (three guards):**
- pubkeys must be 64 hex chars — anything else cannot match a node
anyway, so it is dropped instead of running UPDATEs that never match
- the normalised scope list is capped at 256 bytes
- a report stamped more than 5 minutes ahead of our clock is dropped, so
a far-future timestamp can no longer lock the node

**Tests:** `neighbors_guard_test.go` covers each guard, including that a
genuine report still lands after a future-stamped one was rejected and
that ordinary clock skew is still accepted. Full ingestor suite passes.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:58:45 +02:00
nullrouten0andClaude Mythos 5.1 a981420d21 fix(ingestor): cap observer-supplied string lengths (#2123)
Observer `id` and `iata` come from the MQTT topic; `origin` (name),
`model`, `firmware`, `client_version` and `radio` come from the status
JSON. Any publisher controls them and nothing bounded their length, so
one message could store a 64 KB observer id or name. Each new id is also
a new `observers` row, and that table is joined by most packet queries.

**Fix:** a small `clampObserverField` helper strips control characters
and truncates: ids to 128 runes, text fields to 128, IATA to 16. Applied
on the status path, the packet path and in `extractObserverMeta`. Values
are truncated rather than rejected, so a legitimate observer with a long
name still appears.

**Tests:** `observer_fields_test.go` — short values untouched, long
values cut at 128 runes (not bytes, so multi-byte names are not split),
control characters removed, `extractObserverMeta` caps all string
fields. Full ingestor suite passes.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:58:13 +02:00
nullrouten0andClaude Mythos 5.1 548e4cd4a4 fix(api): cap the nodes= list on /api/packets at 50 entries (#2120)
Each entry in the comma-separated `nodes=` list on `GET /api/packets`
costs one SQLite lookup (`resolveNodePubkey`) while the packet store's
read lock is held. A 1 MB URL fits about 15,000 entries. On a test
instance 12,000 entries took 1.3 s per request, against 0.9 ms for one
entry, and the lock stalls the poller's writes for that long. A few
parallel clients can keep the site busy and the live feed stale.

**Fix:** lists longer than 50 entries get HTTP 400 with a clear message.
No UI page sends more than a handful.

**Tests:** `multi_node_cap_test.go` — 50 entries return 200, 51 return
400. Full `go test ./...` in `cmd/server` passes.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:58:07 +02:00
nullrouten0andClaude Mythos 5.1 4ce6d9f6d4 fix(ui): pin CDN script versions and add integrity hashes (#2121)
`leaflet.heat` and `chart.js` in `index.html`, and swagger-ui on
`/api/docs`, were loaded from unpkg with no `integrity` attribute.
`chart.js@4` and `swagger-ui-dist@5` also floated on a major version, so
a new release would load unreviewed. A compromised CDN or package would
run as our own code on every page. Leaflet itself already had a hash.

**Fix:** pin `chart.js@4.5.1`, `leaflet.heat@0.2.0`,
`swagger-ui-dist@5.33.1`, each with a sha384 hash and
`crossorigin="anonymous"`.

**How the hashes were made:** download each pinned file, `openssl dgst
-sha384 -binary | base64`, then download again and check the hash
matches.

**Note:** bumping chart.js or swagger-ui now means updating the hash
too. Vendoring them into `public/vendor/` (as markercluster already is)
would remove the CDN dependency entirely; happy to do that instead if
preferred.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:56:48 +02:00
nullrouten0andClaude Mythos 5.1 8b64439635 fix(server): keep config.json's file mode when saving the geo-filter (#2126)
`SaveGeoFilter` rewrote `config.json` through a temp file created with
mode 0644, so a config an operator had made 0600 (it holds the API key
and broker passwords) became world-readable after the first geo-filter
save.

**Fix:** stat the original and reuse its mode for the temp file. Falls
back to 0644 when the stat fails.

**Tests:** `config_mode_test.go` — a 0600 config stays 0600 after a
save; a 0644 config stays 0644.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:56:42 +02:00
nullrouten0andClaude Mythos 5.1 523a5edfa1 fix(ui): escape the route string on the unknown-route page (#2124)
`app.js` wrote the URL fragment into `innerHTML` as the heading of the
"Page not yet implemented" view for routes it does not know. Browsers
percent-encode `<`, `>` and `"` in fragments, so this is unlikely to be
exploitable today, but it is a plain `innerHTML` sink on a
URL-controlled string and costs one `escapeHtml()` to close.

**Tests:** one test added to `tests/unit/test-xss-escape-sinks.js` in
the file's existing style.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:56:30 +02:00
dborupandClaude Opus 5.5 7309bdb5a9 fix(store): index a transmission once per relay key in byPathHop (#2117)
Relates to #2108

## Problem

`traffic_share_score` grows with server uptime until it is far above
reality, and many relays end up clamped at 1.0.

The score is the number of non-advert entries in `byPathHop[pubkey]`
divided by the number of non-advert transmissions.
`indexResolvedPathHops` runs once per **observation**, on the
live-ingest path (`IngestNewFromDB`) and on the late-observation path
(`IngestNewObservations`). `addResolvedPubkeysToPathHopIndex` only
de-duplicates within one call, so every further observation through the
same relay appends the transmission to that relay's bucket again. The
denominator counts each transmission once.

After a restart the values look right, because `buildPathHopIndex` →
`retainResolvedPathHops` de-duplicates by `*StoreTx`. They then drift
upwards as live observations arrive.

None of the affected functions has changed on `master` since the
diagnosis against `415362c`. This branch is based on `000d9ab`.

## Fix

1. **Idempotent insert, per (transmission, relay key).**
`addResolvedPubkeysToPathHopIndex` keeps a side map `pathHopResolved
map[*StoreTx][]string`. It holds the resolved keys each transmission is
already indexed under and skips those.
- Keys are compared **exactly**, so no collision can ever drop an entry.
- A later observation through a **new** relay still adds that relay,
once.
   - `StoreTx` is unchanged.
2. **Interned keys.** `pathHopKeys map[string]string` keeps one shared
copy of each resolved key. The record then holds a 16-byte string header
per entry, and not the string each observation's resolve allocated: a
fresh `json.Unmarshal` string per persisted observation on `Load`, and a
`strings.ToLower` result on the live paths. The `byPathHop` key is the
same shared copy.
3. **Eviction and rebuild.**
- `evictStaleInternal` drops the record of evicted transmissions, plus
any interned key whose bucket it deletes.
- `retainResolvedPathHops` keeps the records of live transmissions only
and drops interned keys whose bucket was not carried over. When a
rebuild starts from an empty index, it clears both.
4. **Defence in depth.** The single (`GetRepeaterUsefulnessScore`),
batch (`GetRepeaterNodeStatsBatch`) and bulk
(`computeRepeaterUsefulnessScoreMap`) scores count **distinct**
non-advert transmissions per bucket, via `countDistinctNonAdvert`.
   - IDs are collected in a reused slice.
- A bucket in ascending ID order is already distinct; only out-of-order
buckets are sorted.
- The bulk pass does this for full-pubkey keys only. Raw-hop buckets get
one entry per transmission from `addTxToPathHopIndex`.
5. **Cache side effect.** A repeated observation through known relays no
longer mutates `byPathHop`, so it no longer drops the batch relay-stats
cache, as the cache contract from `1164` intends. An observation through
a new relay still drops it.

The other `byPathHop` consumers already de-duplicate by `tx.ID` or read
only raw prefix keys: `GetNodeHopAnalytics`,
`computeMultiByteCapability`, `handleNodePaths`,
`computeRepeaterRelayInfoMap` and the relay info in
`GetRepeaterNodeStatsBatch`. Two tests pin that they ignore duplicate
entries.

## Exact keys vs. a 64-bit fingerprint

The first version of this fix stored a 64-bit FNV-1a hash per key. As
requested, I measured both, plus exact keys without interning, on the
real `addResolvedPubkeysToPathHopIndex`.

**Keys per transmission.** In the e2e fixture, the union of resolved
relay keys over all observations of one transmission has a mean of 5.4
and a maximum of 17. The protocol limit is 64 path bytes
(`MAX_PATH_SIZE`), i.e. 64 / hash size hops.

**Memory** retained by the record, per transmission.
`BenchmarkPathHopRecordMemory_2108`: 100K transmissions, keys from 4,000
relays, each transmission heard twice with freshly allocated keys. The
record map and the interned keys are included.

| keys per tx | exact, interned (this PR) | 64-bit fingerprint | exact,
not interned |
|---|---|---|---|
| 2 | 88 B | 68 B | 207 B |
| 5 | 136 B | 101 B | 447 B |
| 17 | 344 B | 197 B | 1,423 B |

- At realistic key counts, exact keys cost 20–35 B per transmission more
than the fingerprint. That is about 3.5 MB per 100K transmissions, and
small next to what the store already charges per transmission
(`storeTxBaseBytes` alone is 384 B).
- Without interning, exact keys would cost 3–4× more, and most of that
would be held during every `Load`. Interning is what makes exact keys
affordable.

**Time.** I ran the variants interleaved, 3 rounds each, median ns/op on
an Apple M4. The record step is not measurable inside the full call. The
isolated record step was faster with exact keys in a separate
micro-benchmark, because comparing a handful of 64-char strings is
cheaper than hashing each one.

| benchmark | exact, interned | fingerprint | exact, not interned |
|---|---|---|---|
| late observation, 20K txs | 312 | 358 | 331 |
| late observation, 100K txs | 760 | 836 | 758 |
| `IngestNewObservations` (SQL + resolve + index) | 531 µs | 497 µs |
527 µs |

**Collision probability of the fingerprint.** For k distinct keys in one
transmission it is about k(k−1)/2 / 2^64:

| k | per transmission | per 10^9 transmissions |
|---|---|---|
| 2 | 5.4e-20 | 5.4e-11 |
| 5 | 5.4e-19 | 5.4e-10 |
| 17 | 7.4e-18 | 7.4e-9 |
| 64 | 1.1e-16 | 1.1e-7 |

For random keys this is negligible. FNV-1a is unkeyed, though, so two
relay keys that collide could be found deliberately (a birthday search
over about 2^32 key pairs). A collision would leave a transmission out
of one relay's bucket.

**Decision.** Exact keys are cheap enough once interned: the same speed
within noise, and a few tens of bytes per transmission. They remove the
question of collision-driven omissions entirely, so this PR uses exact
keys and drops the fingerprint.

The two new maps (`pathHopResolved`, `pathHopKeys`) are not added to
`trackedBytes`, so the `maxMemoryMB` trigger undercounts by roughly the
per-transmission figures above. Both are bounded by live state (eviction
and rebuild prune them), so this is a steady proportional undercount
rather than a leak.

## Tests

**New: `cmd/server/pathhop_dedupe_2108_test.go`.** Behaviour tests that
use only existing API. Each one runs against the real SQLite schema
through `Load`, `IngestNewFromDB` and `IngestNewObservations` where
noted.

| test | covers | on `master` |
|---|---|---|
| `TestPathHopIndexOncePerTx_LateObservations_2108` | 1 + 10 late
observations through the same relays: one entry per relay | **fails**
(11 entries) |
| `TestPathHopIndexOncePerTx_LiveIngestBatch_2108` | 11 observations in
one `IngestNewFromDB` batch | **fails** (11) |
| `TestPathHopIndexOncePerTx_AfterLoad_2108` | `Load` + rebuild, then
one live observation | **fails** (2) |
| `TestPathHopIndexAddsNewRelayFromLaterObservation_2108` | a later
observation through a **new** relay adds it exactly once, and its share
becomes correct | **fails** (2) |
| `TestTrafficShareStableAcrossLateObservations_2108` | single, batch
and bulk scores agree, equal the definition, stay put over rounds of
late observations, and equal a fresh `Load` of the same data | **fails**
(all four relays at 1.0, want 0.4–0.5) |
| `TestPathHopIndexSizeBoundedByTransmissions_2108` | the index size
does not grow with observations | **fails** |
| `TestRelayStatsCacheAcrossRepeatedObservations_2108` | a repeated
observation keeps the relay-stats cache, and the kept cache equals a
fresh compute; a new relay drops it, and the next read sees the relay |
**fails** |
| `TestAddResolvedPubkeysToPathHopIndex_PerRelayIdempotent_2108` | the
helper's return value and cache invalidation per (tx, relay), in any key
order | **fails** |
| `TestTrafficShareCountsDistinctTransmissions_2108` | all three scores
count distinct transmissions in a duplicated bucket | **fails** |
| `TestPathHopConsumersIgnoreDuplicateEntries_2108` | pin of the
consumer audit | passes (by design) |
| `TestMultiByteCapabilityIgnoresDuplicateEntries_2108` | pin of the
consumer audit | passes (by design) |

**New: `cmd/server/pathhop_record_2108_test.go`.** These tests use the
new symbols, so they do not compile against `master`.
- `TestCountDistinctNonAdvert_2108`: ascending, descending, interleaved
duplicates, adverts, nils, untyped.
- `TestPathHopResolvedRecordBoundedAndEvicted_2108`: the record stays at
2 keys after 11 observations. Eviction removes the record and every
entry. A survivor's next observation adds nothing. Evicting everything
leaves no record, no interned key and no bucket.
- `TestPathHopResolvedRecordAcrossRebuild_2108`: a rebuild keeps the
records of live transmissions, so the next observation adds nothing. It
drops the record of a removed transmission and the interned key only
that transmission used.
- `TestPathHopResolvedRecordClearedWithEmptyIndex_2108`: a rebuild from
an empty index clears the record and the interned keys, and the next
observation puts the transmission back.
- `TestPathHopResolvedRecordInternsKeys_2108`: record entries and the
`byPathHop` key share one copy, even though every observation passes
freshly allocated keys.

**Changed: `pathhop_eviction_1908_test.go`.** It built its duplicate
entries by repeating `indexResolvedPathHops`, which no longer
duplicates. It now seeds the duplicates directly, so the `1908` sweep is
still tested against buckets that hold one transmission several times.
The expected buckets are unchanged.

**Changed: `db_test.go`.** `setupTestDB` takes `testing.TB`, so the SQL
benchmark can use it.

**Mutants.** I applied each mutant on its own to this branch. All 14 are
killed:

| mutant | killed by |
|---|---|
| record never consulted | the OncePerTx, stable-share, size, cache and
record tests |
| dedupe per transmission instead of per relay |
`AddsNewRelayFromLaterObservation`,
`RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent` |
| eviction keeps the record | `RecordBoundedAndEvicted` |
| rebuild keeps records of removed transmissions | `RecordAcrossRebuild`
|
| rebuild from an empty index keeps the record |
`RecordClearedWithEmptyIndex` |
| cache dropped on every call |
`RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`,
existing `NoMutation_PreservesCache` |
| cache kept although `byPathHop` changed |
`RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`,
existing `InvalidatesRelayStatsCache` |
| single score counts entries |
`TrafficShareCountsDistinctTransmissions` |
| batch score counts entries | `TrafficShareCountsDistinctTransmissions`
|
| bulk score counts entries | `TrafficShareCountsDistinctTransmissions`
|
| distinct count trusts any bucket order |
`TrafficShareCountsDistinctTransmissions`, `CountDistinctNonAdvert` |
| keys not interned | `RecordInternsKeys` |
| eviction keeps interned keys | `RecordBoundedAndEvicted` |
| rebuild keeps interned keys of dropped buckets | `RecordAcrossRebuild`
|

**Commands run:**
- `gofmt -l` on all tracked Go files: clean.
- `go vet ./...` in all 14 modules: clean.
- `cd cmd/server && go test -race ./...`: pass.
- `cd cmd/ingestor && go test ./...`: pass.
- `sh test-all.sh`: 184 of 186 suites pass locally. The other two,
`test-issue-1956-release-routing.js` and `test-preflight-xss-gate.js`,
shell out to scripts that need bash ≥ 4 (`mapfile`). They fail under
macOS's bash 3.2 regardless of this change. This PR touches no frontend
or script files.

## Benchmark: `master` vs. this branch

I ran `master`'s sources (`000d9ab`) and this branch interleaved, 5
rounds, with the same benchmark files. Medians on an Apple M4.

| benchmark | `master` | this PR | change |
|---|---|---|---|
| late observation through known relays, 20K txs | 446 ns, 38 B/op | 441
ns, 0 B/op | within noise |
| late observation through known relays, 100K txs | 710 ns, 69 B/op |
755 ns, 0 B/op | within noise (runs overlap) |
| … index size afterwards, entries/tx (20K / 100K) | 24.99 / 8.99, still
growing | 4.99 / 4.99 | bounded |
| `IngestNewObservations`, one observation for each of 20 txs | 532 µs |
513 µs | within noise |
| bulk score pass, clean index, 20K, ingest order | 246 µs | 316 µs |
+28 % |
| bulk score pass, clean index, 100K, ingest order | 3.03 ms | 4.30 ms |
+42 % |
| bulk score pass, clean index, 20K, reversed buckets | 257 µs | 467 µs
| +82 % |
| bulk score pass, clean index, 100K, reversed buckets | 3.38 ms | 4.79
ms | +42 % |
| bulk score pass after 11 observations per tx, 20K | 663 µs | 284 µs |
−57 % |
| bulk score pass after 11 observations per tx, 100K | 4.72 ms | 3.19 ms
| −32 % |

- On an index that is clean on both builds, the distinct count makes the
bulk pass slower. That pass runs on the cache-miss path, which the
background recomputer refreshes every 5 minutes by default.
- On the index a live server actually holds without this fix, `master`
walks every accumulated duplicate, so the real-world pass is faster
after the fix. The gap grows with uptime.
- The late-observation step no longer allocates. On `master` its index
grows with every observation.

Benchmarks: `BenchmarkLateObservationIndex_2108`,
`BenchmarkIngestNewObservations_2108`,
`BenchmarkTrafficShareScoreMap_2108` and
`BenchmarkPathHopRecordMemory_2108`.

## Production motivation

We have run this fix on two production instances. Before the fix, the
summed `traffic_share_score` grew past 180 and many relays sat at the
1.0 clamp; with the fix no node reaches `≥ 0.999`.

A controlled 12 h A/B run makes the drift explicit. Two servers read one
identical database — one on this `master` base, one with the fix —
alongside a reference server that freshly loads the same database (a
fresh load is correct, because the rebuild dedupes). The unfixed server
drifted to 9 relays at the 1.0 clamp and up to 0.95 absolute error per
node against the reference; the fixed server stayed within 0.05 of the
reference for every node, with no relay at the clamp.

A smaller residual rise remains and is a separate cause (startup-vs-live
hop resolution); it is deliberately left for a follow-up so its effect
stays measurable.

## Out of scope

That remaining slow rise comes from a separate cause: live ingest
resolves some hops that the startup load does not. This PR deliberately
leaves that drift alone, so its effect stays measurable. The fix will
follow in a separate PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 14:34:31 +02:00
efitenandClaude Opus 5.5 f3ae8ac25a feat: sync a logged-in user's settings across devices (part B) (#2130)
Part B of #2128: a logged-in user's settings follow them across devices.
Log in on a phone and your own nodes, favorites, customizer and filters
are there; a change on one device reaches the others within a minute or
when you return to the tab.

**This PR builds on #2129.** Until that one is merged, the diff here
includes it. The commits for this part start at `docs(specs): settings
sync for optional user management (sub-project B)`.

## The situation

Everything a visitor sets up lives in one browser's `localStorage`
(about 100 keys in `public/`). A second device or a cleared cache starts
from zero (#895).

## What this PR adds

**Storage.** `users.db` schema v2: one JSON document per user in
`user_settings`, with a revision number and a generation id. A write
succeeds only when the client's revision and generation match the stored
ones, so two devices cannot overwrite each other silently.

**Server.** `GET`, `PUT` and `DELETE /api/account/settings`, behind the
same session and CSRF checks as the account routes.
- The server owns the list of synced keys (61 keys,
[`settings_allowlist.go`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/cmd/server/settings_allowlist.go))
and sends it to the client, so the two cannot drift.
- A hard denylist, checked first, refuses `meshcore-api-key`, every
`corescope_channel_*` key and `live-channel-colors` (#725). The colour
map is keyed by channel hash, and for a user-added channel that hash is
`user:<name>`, which would expose hashtag channel names.
- Documents are capped at 256 KiB, measured like `JSON.stringify`. PUT
is limited to 60 requests per hour per user. A stale revision gets 409
with the current document.

**Client**
([`settings-sync.js`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/public/settings-sync.js)).
Inert unless the feature is on and someone is logged in.
- It wraps `localStorage.setItem` and `removeItem` for allowlisted keys
only and pushes 2 seconds after the last change.
- It pulls on login, page load, tab focus and every 60 seconds while the
tab is visible.
- **Merge:** three-way, against a per-device baseline that belongs to
one user and one document generation. Lists (own nodes, favorites, saved
filters) merge per item, so an item added anywhere is kept and an item
removed on one device does not come back from another. Single values:
the profile wins unless only this device changed it.
- Remote changes are written without a push, theme and colour-blind
preset are re-applied, and the current page re-renders (skipped on
account pages and while the geofilter editor is open).

**UI.**
- Logout asks: keep my settings on this device (default), remove them
from this device, or cancel. Channel keys are never removed: no copy
exists anywhere else.
- The account page gets a "Settings sync" section: last synced time,
"Sync now", what is and is not synced, and "Delete synced settings from
my account".

## Not synced

Layout and device state (panel and column widths, collapsed panels, map
positions, geofilter drafts), channel data (#725), the API key, and all
`sessionStorage`. The full list is in the
[spec](https://github.com/efiten/CoreScope/blob/feat/settings-sync/docs/specs/2026-10-06-user-settings-sync-design.md).

## Performance

- One GET per page load, tab focus and minute while visible; one
debounced PUT per burst of changes.
- The `setItem` wrapper costs one Set lookup per write for non-synced
keys. A synced write reads one small revision key, not the stored
document.
- The server reads or writes one row per request.

## Verification

- `internal/users` and `cmd/server`: `go vet` and `go test` pass locally
(22 new Go tests), including a test that every allowlisted key still
occurs in `public/`, and denylist tests.
- `tests/unit/test-settings-sync.js`: 79 passing (vm, real module). The
cases cover the merge table, two tabs sharing one storage, stale answers
after a push, delete while a push is in flight, and logout while the
final push fails.
- `sh test-all.sh` exits 0.
- `tests/e2e/test-user-management-e2e.js` (10 steps, 4 of them new)
passed locally with two browser contexts as two devices: a favorite and
the packet time window travel from device 1 to device 2, a removal does
not come back, and "remove from this device" clears the synced keys
while a channel key stays.
- Checked by hand on a staging instance with a desktop and a phone on
one account.

## Not in this PR

- On a shared browser where the previous user chose "keep", the next
user's first login merges those settings into their own account. The
user guide says to choose "remove" on shared computers.
- Saved filter expressions are synced as typed, including any channel
names written in them. The guide says so.
- Realtime push between devices; the minute pull is the sync interval.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 11:20:44 +02:00
efitenandClaude Opus 5.5 d232072f85 feat: optional user accounts (part A: foundation) (#2129)
Part A of #2128: optional, off-by-default user accounts. With the
feature off nothing changes; with it on, visitors can register and log
in, and admins manage users and use the operator actions without the API
key.

PR #2130 (settings sync) builds on this one. The two are meant to be
merged together.

## The situation

- Operator actions (geofilter save and prune, backup, perf reset) need
the shared `apiKey`. There is no per-person right.
- Nothing in CoreScope knows who a visitor is, so the requests in #2128
that need that (#1835, #2092, #1508, #730) have nothing to build on.

## What this PR adds

**Two new Go modules**
- `internal/users`: a separate `users.db` (SQLite through
`modernc.org/sqlite`) with users, sessions, single-use tokens, an audit
log and a mail log. Passwords use argon2id.
- `internal/mailer`: a `Mailer` interface with a Brevo client (send,
delivery events, webhook parsing) and an in-memory fake for tests.

**Server (`cmd/server`)**, active only with `userManagement.enabled`
- 24 routes, all documented in OpenAPI under the `users` tag
([`auth_routes.go`](https://github.com/efiten/CoreScope/blob/feat/user-management/cmd/server/auth_routes.go)):
  - auth: register, activate, login, logout, me, forgot, reset;
- account: profile, password, email change with confirmation, sessions,
self-delete;
- admin: list, detail, disable, enable, delete, role, resend activation,
manual activation, mail status refresh;
  - a Brevo webhook, registered only when `mail.webhookSecret` is set.
- `requireAdmin` replaces `requireAPIKey` at the 7 operator call sites:
the API key **or** an admin session. With the feature off it is the old
API-key gate (`TestRequireAdminWithoutUserManagementIsAPIKeyGate`).
- `/api/config/client` gets `userManagement: {enabled: true}` only when
the service started; with the feature off the response is
byte-identical.

**Frontend**
- `auth.js` (header account control, request helper that adds the CSRF
header), `account.js` (login, register, activate, forgot, reset, confirm
email, my account), `admin-users.js` (`#/admin/users`, deep-linked
filters), `account.css` (theme tokens only).
- On phones the top-bar control is hidden, so a conditional entry goes
into the bottom-nav "More" sheet and the nav drawer.
- The customizer geofilter tab and the Perf "Reset stats" button use the
admin session when there is one.

**Config.** A `userManagement` block (`config.example.json`,
[`docs/user-guide/accounts.md`](https://github.com/efiten/CoreScope/blob/feat/user-management/docs/user-guide/accounts.md)).
The Brevo key can come from `CORESCOPE_BREVO_API_KEY`. The server
refuses to start when the block is enabled but incomplete.

## Security choices

- Session cookie `cs_session`: HttpOnly, SameSite=Lax, Secure when
`publicBaseUrl` is https. Every cookie-authenticated state change needs
the `X-CS-CSRF` header and a matching Origin.
- Activation needs the token **and** the account password. Without the
password, an attacker who keeps re-registering a known address could get
the owner to activate an account that carries the attacker's password.
- Register, forgot and email change answer identically for known and
unknown addresses. A password reset ends all sessions, a password change
ends all other sessions, and both end outstanding email-change links.
- Rate limits: login 10 per 15 minutes, register and forgot 5 per hour,
per IP and per address. The bucket count is capped. `trustedProxies`
makes the per-IP limits see real client IPs behind a proxy.
- Server logs carry `#<user id>`, never addresses, tokens or passwords;
mail-provider error texts are redacted before logging.

## Performance

No change to an existing hot path with the feature off. With it on:
- One `users.db` lookup per authenticated request (session by token
hash).
- The admin user table rebuilds its `tbody` on each filter change.
`users.List` caps the result at 1000 rows (`internal/users/users.go`),
which bounds the rebuild.
- `map[string]interface{}` in `openapi.go`: 79 before, 78 after.

## Verification

- `internal/users`, `internal/mailer` and `cmd/server`: `go vet` and `go
test -race` pass locally. 121 new Go tests.
- `cmd/server` with `-tags e2etest`: vet and the e2e hook tests pass.
- `sh test-all.sh` exits 0. `tests/unit/test-user-management-ui.js`: 67
passing (vm, real modules).
- `tests/e2e/test-user-management-e2e.js` (6 steps) passed locally
against an `e2etest` build with the fake mailer and against a
feature-off build. CI builds the `e2etest` binary and runs the suite on
a second server (`deploy.yml`).
- On a staging instance with a real Brevo key: register, activation mail
delivered, activate, admin table, "Refresh status" showing sent,
deferred, delivered, opened and clicked.

## Not in this PR

- Settings sync (#2130), the admin dashboard, approval flows and
notifications (parts B to E of #2128).
- A `requireReadAuth` mode (#1835). Sessions from this PR are what such
a mode would accept.
- Binary size and build time with `modernc.org/sqlite` linked next to
`mattn/go-sqlite3` were not measured. Their driver names do not collide.
#1992 discusses the driver choice.
- No Brevo webhook was configured on staging; delivery status there came
from "Refresh status".

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 09:25:29 +02:00
efitenandClaude Opus 5.5 4363403495 fix(nodes): match the region node filter on observer ID, not ingest-time IATA (#2115)
Follow-up to #2114, from its review.

`RegionNodePubkeys` compared the IATA copied onto each observation at
ingest (`StoreObs.ObserverIATA`). When an operator changes an observer's
code, those copies stay as they were until a restart or until the
observations age out. `/api/nodes?region=` then kept matching the old
region, while the other store region filters already used the new one.
Those filters resolve observers by ID through `resolveRegionObservers`,
for example the packets query and `computeNodeHomeRegions`. The SQL
subquery that #2114 replaced also read `observers.iata` at query time,
so this restores that behaviour.

## Also fixes a regression from #2114

Found in review: `loadChunk`, the background history loader
(`cmd/server/store.go`, the chunk SELECT and the `StoreObs` it builds),
sets `ObserverID` on each observation but never `ObserverIATA`. With
#2114 matching on that IATA, **every node whose adverts came only from
background-loaded history dropped out of `/api/nodes?region=`** on
current master. Matching on observer ID fixes it.
`TestRegionNodePubkeysMatchesChunkLoadedObservations` pins it (an
observation with the observer ID and an empty IATA) and fails on
master's `region_nodes.go`.

## Change

- Resolve the region to observer IDs with `resolveRegionObservers` (own
mutex, 30 s cache) and match observations by `ObserverID`. Lock order:
`regionNodesMu` is released before it, and `s.mu` is taken after it;
none of the three is held together.
- Without a database there is nothing to resolve, so `RegionNodePubkeys`
reports no set and the handler keeps the SQL path.

## Tests

- The region tests now seed an observers table and leave each
observation's IATA at a stale value, so they can only pass through the
table.
- New `TestRegionNodePubkeysFollowsObserverIATAChange`: an observer that
moved from SJC to SFO matches SFO and not SJC. It fails on the previous
code (`got [pk_moved]` for SJC).
- `TestRegionNodePubkeysMatchesChunkLoadedObservations` (second commit),
see above.
- `setupTestDB` takes `testing.TB` so the benchmark can use it; every
existing caller passes `*testing.T` unchanged.
- Full `cmd/server` suite passes locally, the region tests also with
`-race`.

## Performance

`BenchmarkRegionNodePubkeys` (220k adverts × 8 observations): 37 ms to
30 ms per uncached scan, a map lookup per observation instead of a
string normalisation. The observer lookup is one query on the small
`observers` table, cached for 30 s.

## Not done

- No singleflight on a cold cache, and the 64-entry cache still resets
when full. Both were non-blocking in the #2114 review.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 23:42:07 +02:00
efitenandClaude Opus 5.5 d5e6d1b2a6 fix(nodes): resolve the /api/nodes region filter from the packet store (#2114)
Fixes #2101.

`/api/nodes?region=` filtered nodes with a subquery that joins every
advert to all of its observations and observers, with no time bound
(`cmd/server/db.go`, `GetNodes`). It ran twice per request, once for
`COUNT(*)` and once for the page. A client paging through nodes repeats
both per page.

## Reproduced

On our staging database (10.3 GB, 16.5M observations, 219,750 adverts),
read-only `sqlite3`, region `BRU`:

| | real | user | sys |
|---|---|---|---|
| one regional `COUNT(*)`, as `GetNodes` builds it | **154.8 s** | 1.9 s
| 5.8 s |

Almost all of it is waiting on disk. The plan walks
`idx_transmissions_payload_type` for every advert and then
`idx_observations_dedup` for each one's observations. Two of those per
request, a few requests at once, and the reader pool is gone, which
matches the report of all four database workers sitting in `GetNodes`.

## Fix

- **`PacketStore.RegionNodePubkeys(region)`** (new
`cmd/server/region_nodes.go`) walks the store's in-memory adverts
(`byPayloadType[ADVERT]`) once. It keeps the pubkey of every node with
an advert heard by an observer in the region, using the `ObserverIATA`
each observation already carries. It is cached for 30 s per region, and
the cache is bounded to 64 entries because its keys come from the
client's parameter. Lock discipline follows `resolveAreaNodes`: the
cache mutex and `s.mu` are never held together, and the ordering note in
`store.go` lists it.
- **`GetNodes` takes a `NodeQuery` struct** with a `RegionPubkeys`
field, passed as one `json_each` parameter: `public_key IN (SELECT value
FROM json_each(?))`, a primary-key lookup. An empty set matches nothing.
The SQL subquery stays for a server without a store (tests, tooling).
- The advert-pubkey lookup that `trackAdvertPubkey`,
`untrackAdvertPubkey` and `computeNodeHomeRegions` each copied is now
one helper, `advertPubkey`.

The struct instead of a second `GetNodes` variant keeps the
`map[string]interface{}` count unchanged in `db.go` (75) and `routes.go`
(59).

## Behaviour change

The region filter now covers the adverts the store holds
(`packetStore.retentionHours`), which is the window the rest of the UI
shows, instead of all database history. A node heard in a region only
before that window no longer matches the filter. While the store is
still loading after a restart, the set grows as history loads.

## Performance

| | before | after |
|---|---|---|
| region set, uncached | 154.8 s (SQL count, staging) | 37 ms
(`BenchmarkRegionNodePubkeys`: 220k adverts × 8 observations, 4,000
nodes) |
| region set, within 30 s | same again | cache hit |
| node count over the set | (included above) | 2 ms on staging (1,200
keys) |
| 500-row page over the set | same scan again | 3 ms on staging |

The scan holds `s.mu` for reading for those 37 ms, at most once per
region every 30 s.

## Tests

- `TestRegionNodePubkeys`: the in-region advert, an advert heard in two
regions, case and whitespace in codes, a comma list, an unknown region
giving an empty set, a non-advert never counting, a blank region giving
no filter.
- `TestRegionNodePubkeysCached`, `TestRegionNodePubkeysCacheIsBounded`
(1,000 distinct regions stay within 64 entries).
- `TestGetNodesRegionPubkeys`: the set combines with the role filter and
counts correctly, and an empty set returns nothing even with `Region`
set.
- `TestHandleNodesRegionUsesStore`: an advert that only the store knows
about shows up through `/api/nodes?region=`, so the handler is proven
not to ask SQL.
- Existing region tests (`TestGetNodesRegionFilterV2` and the
`db_test.go` region cases) pass unchanged through the SQL fallback. Full
`cmd/server` suite passes locally; the new tests also pass with `-race`.

## Not done

- No request context on these queries, also raised in the issue. With
the scan gone they take milliseconds, so I left that out of this change.
- The region semantics stay "heard by an observer in the region". #1879
argues for the node's home region instead; that is a separate decision.
- No frontend change. I did not check this in a browser; the nodes page
and map call the same endpoint with the same parameters.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 21:45:02 +02:00
Sylvain Rabot b5b230e884 feat(channels): show the sender's path hash size on each message (#2089)
Each channel message now shows the hash size its sender's path uses, read
from bits 7-6 of the path byte that the originator writes and repeaters keep.
The server sends it as path_hash_size (packetpath.HashSize); the frontend
helper pathHashSize() applies the same rule and returns 0 for unknown.

One rule on both pages: the packet detail Hash Size row and the hex
breakdown call pathHashSize() too, so a 0-hop flood message reports the
same size on Channels and on its packet page. A direct packet with no hops
left reports no size, matching cmd/server/decoder.go. The cases live in
test-fixtures/path-hash-size-cases.json, read by the Go and JS tests.
2026-10-05 20:32:40 +02:00
efitenandClaude Opus 5.5 000d9ab030 feat(coverage): RF noise-floor layer on the Mobile RX coverage page (#2113)
The ingestor already stores the noise floor that CoreDrive RX companions
report with each GPS fix (`client_rf_samples`, opt-in through
`clientRfSamples`, `cmd/ingestor/client_rf_sample.go:62`), but nothing
reads it back. An operator who enables it collects the data and cannot
see it. This adds the read side: `GET /api/rf-noise` and a Signal/Noise
toggle on the Mobile RX coverage page.

## What it does

- **`GET /api/rf-noise?bbox=&z=&days=`** returns a GeoJSON hex grid with
the median, quietest and noisiest noise floor per cell, in the same
shape as `/api/rx-coverage`. It is registered always and 404s unless
`clientRfSamples.enabled` is true, the same pattern as the coverage
routes.
- **Stationary samples are excluded.** A parked companion logs hundreds
of readings at one point, which would otherwise define its cell.
- **Coverage page:** a Signal/Noise toggle, rendered only when
`/api/config/client` reports `clientRfSamples: true`. The noise layer
reuses the coverage colour tokens with the axis inverted, because a
lower dBm is quieter. Tiers are at -115 and -108 dBm.
- **Deep link:** `#/rx-coverage?layer=noise` opens on the noise layer.
- **Empty and failed answers** are labelled on the map ("No RF samples
in this view yet", or a retry hint), so a blank map never reads as
"feature off".

No new configuration key: it reads the existing `clientRfSamples`
section that the ingestor already uses. Default off, so nothing changes
for an instance that has not opted in.

## Where it comes from

Ported from the efiten/CoreScope fork, where it has run on
analyzer.on8ar.eu since September (fork commits `42d09f0e`, `80581c43`,
`cde95078`). The cherry-picks conflicted with upstream's newer
`routes.go`, `types.go` and `rx-coverage.js`, so the final state was
ported by hand. The fork-only `/scopes` route that sat next to it in the
same hunk is deliberately left out.

## Performance

The query is bounded by `sampled_at` (indexed, `idx_crf_prune`) and the
bbox, aggregation is one pass plus a per-cell sort, and the response is
capped at 5000 cells. Measured on analyzer.on8ar.eu, all of Belgium
(`bbox=49.4,2.4,51.6,6.5&z=9`), from a client in Belgium, so network
time included:

| window | samples | cells | response time |
|---|---|---|---|
| 7 days | 10,835 | 241 | 0.37 s |
| 30 days | 43,303 | 429 | 0.39 s |

That table holds 45,270 rows in total. It is only read when someone
opens the noise layer, never on ingest or WebSocket paths.

## Tests

- `cmd/server/rf_noise_test.go`: 8 tests for the aggregation (median and
extremes, stationary exclusion, the cell cap, empty input) and the gate
(404 when off, even with data present). All pass, and the full
`cmd/server` suite passes locally.
- `tests/unit/test-rx-coverage-noise.js` (new, in `test-all.sh`): the
colour axis runs the right way, including both tier bounds and a string
median from the API. Slices the real function out of `rx-coverage.js`
and fails on master's copy, which has no thresholds.
- `tests/e2e/test-rx-coverage-noise-e2e.js` (new, wired in `deploy.yml`
and `scripts/non-unit-tests.json`): no toggle when the flag is off;
Noise fetches `/api/rf-noise`, draws the cells, swaps legend and
subtitle, puts `layer=noise` in the hash; a deep link opens on the noise
layer and an empty answer shows the message. 3/3 locally; against
master's `rx-coverage.js` and `roles.js`, 1/3 (only the "flag off" case
passes, as it should).
- `test-rx-coverage-viewport-e2e.js` still passes. ESLint 8 reports 0
errors on the changed files.

## Not done

- The thresholds (-115 / -108 dBm) are fitted to the fork's own data
(1,241 moving samples at the time) and are constants in
`rx-coverage.js`. Per AGENTS.md rule 8 they belong in the customizer
eventually; not in this PR.
- The E2E suite mocks the API. The real endpoint is covered by the Go
tests and by the measurements above, not by a fixture with seeded
samples.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 22:11:34 +02:00
efitenandClaude Opus 5.5 9c3d76e14f docs(release): notes and changelog for v3.13.1 (#2112)
Release notes and CHANGELOG entry for v3.13.1, a patch release that
ships #2111 (node-discover replies count as coverage).

Once merged, the tag goes on the merge commit so `deploy.yml` picks up
`docs/release-notes/v3.13.1.md` as the release body.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
v3.13.1
2026-10-04 17:22:56 +02:00