Commit Graph
2994 Commits
Author SHA1 Message Date
nullrouten0andClaude Mythos 5.1 2e7a4ebcfe fix(ingestor): bound the unauthenticated /neighbors report (#2122)
`handleNeighborsReport` trusted whatever an observer published on the
`/neighbors` topic. The sender chose `origin_id` (whose "self" it is)
and could list any pubkey as a responded neighbor; each got its
`configured_scope` written with any scope string, stamped with the
sender's own timestamp. The store is last-write-wins on that timestamp
and `normalizeReportTS` accepted any RFC3339 time, so one report dated
years ahead was written once and then blocked every genuine later report
for that node until someone edited the database. The value is shown on
the reach page as the confirmed scope and feeds `/api/scope-audit`.

**Fix (three guards):**
- pubkeys must be 64 hex chars — anything else cannot match a node
anyway, so it is dropped instead of running UPDATEs that never match
- the normalised scope list is capped at 256 bytes
- a report stamped more than 5 minutes ahead of our clock is dropped, so
a far-future timestamp can no longer lock the node

**Tests:** `neighbors_guard_test.go` covers each guard, including that a
genuine report still lands after a future-stamped one was rejected and
that ordinary clock skew is still accepted. Full ingestor suite passes.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:58:45 +02:00
nullrouten0andClaude Mythos 5.1 a981420d21 fix(ingestor): cap observer-supplied string lengths (#2123)
Observer `id` and `iata` come from the MQTT topic; `origin` (name),
`model`, `firmware`, `client_version` and `radio` come from the status
JSON. Any publisher controls them and nothing bounded their length, so
one message could store a 64 KB observer id or name. Each new id is also
a new `observers` row, and that table is joined by most packet queries.

**Fix:** a small `clampObserverField` helper strips control characters
and truncates: ids to 128 runes, text fields to 128, IATA to 16. Applied
on the status path, the packet path and in `extractObserverMeta`. Values
are truncated rather than rejected, so a legitimate observer with a long
name still appears.

**Tests:** `observer_fields_test.go` — short values untouched, long
values cut at 128 runes (not bytes, so multi-byte names are not split),
control characters removed, `extractObserverMeta` caps all string
fields. Full ingestor suite passes.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:58:13 +02:00
nullrouten0andClaude Mythos 5.1 548e4cd4a4 fix(api): cap the nodes= list on /api/packets at 50 entries (#2120)
Each entry in the comma-separated `nodes=` list on `GET /api/packets`
costs one SQLite lookup (`resolveNodePubkey`) while the packet store's
read lock is held. A 1 MB URL fits about 15,000 entries. On a test
instance 12,000 entries took 1.3 s per request, against 0.9 ms for one
entry, and the lock stalls the poller's writes for that long. A few
parallel clients can keep the site busy and the live feed stale.

**Fix:** lists longer than 50 entries get HTTP 400 with a clear message.
No UI page sends more than a handful.

**Tests:** `multi_node_cap_test.go` — 50 entries return 200, 51 return
400. Full `go test ./...` in `cmd/server` passes.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:58:07 +02:00
nullrouten0andClaude Mythos 5.1 4ce6d9f6d4 fix(ui): pin CDN script versions and add integrity hashes (#2121)
`leaflet.heat` and `chart.js` in `index.html`, and swagger-ui on
`/api/docs`, were loaded from unpkg with no `integrity` attribute.
`chart.js@4` and `swagger-ui-dist@5` also floated on a major version, so
a new release would load unreviewed. A compromised CDN or package would
run as our own code on every page. Leaflet itself already had a hash.

**Fix:** pin `chart.js@4.5.1`, `leaflet.heat@0.2.0`,
`swagger-ui-dist@5.33.1`, each with a sha384 hash and
`crossorigin="anonymous"`.

**How the hashes were made:** download each pinned file, `openssl dgst
-sha384 -binary | base64`, then download again and check the hash
matches.

**Note:** bumping chart.js or swagger-ui now means updating the hash
too. Vendoring them into `public/vendor/` (as markercluster already is)
would remove the CDN dependency entirely; happy to do that instead if
preferred.

Running in production on our instance since 2026-10-07.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:56:48 +02:00
nullrouten0andClaude Mythos 5.1 8b64439635 fix(server): keep config.json's file mode when saving the geo-filter (#2126)
`SaveGeoFilter` rewrote `config.json` through a temp file created with
mode 0644, so a config an operator had made 0600 (it holds the API key
and broker passwords) became world-readable after the first geo-filter
save.

**Fix:** stat the original and reuse its mode for the temp file. Falls
back to 0644 when the stat fails.

**Tests:** `config_mode_test.go` — a 0600 config stays 0600 after a
save; a 0644 config stays 0644.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:56:42 +02:00
nullrouten0andClaude Mythos 5.1 523a5edfa1 fix(ui): escape the route string on the unknown-route page (#2124)
`app.js` wrote the URL fragment into `innerHTML` as the heading of the
"Page not yet implemented" view for routes it does not know. Browsers
percent-encode `<`, `>` and `"` in fragments, so this is unlikely to be
exploitable today, but it is a plain `innerHTML` sink on a
URL-controlled string and costs one `escapeHtml()` to close.

**Tests:** one test added to `tests/unit/test-xss-escape-sinks.js` in
the file's existing style.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
2026-10-08 15:56:30 +02:00
dborupandClaude Opus 5.5 7309bdb5a9 fix(store): index a transmission once per relay key in byPathHop (#2117)
Relates to #2108

## Problem

`traffic_share_score` grows with server uptime until it is far above
reality, and many relays end up clamped at 1.0.

The score is the number of non-advert entries in `byPathHop[pubkey]`
divided by the number of non-advert transmissions.
`indexResolvedPathHops` runs once per **observation**, on the
live-ingest path (`IngestNewFromDB`) and on the late-observation path
(`IngestNewObservations`). `addResolvedPubkeysToPathHopIndex` only
de-duplicates within one call, so every further observation through the
same relay appends the transmission to that relay's bucket again. The
denominator counts each transmission once.

After a restart the values look right, because `buildPathHopIndex` →
`retainResolvedPathHops` de-duplicates by `*StoreTx`. They then drift
upwards as live observations arrive.

None of the affected functions has changed on `master` since the
diagnosis against `415362c`. This branch is based on `000d9ab`.

## Fix

1. **Idempotent insert, per (transmission, relay key).**
`addResolvedPubkeysToPathHopIndex` keeps a side map `pathHopResolved
map[*StoreTx][]string`. It holds the resolved keys each transmission is
already indexed under and skips those.
- Keys are compared **exactly**, so no collision can ever drop an entry.
- A later observation through a **new** relay still adds that relay,
once.
   - `StoreTx` is unchanged.
2. **Interned keys.** `pathHopKeys map[string]string` keeps one shared
copy of each resolved key. The record then holds a 16-byte string header
per entry, and not the string each observation's resolve allocated: a
fresh `json.Unmarshal` string per persisted observation on `Load`, and a
`strings.ToLower` result on the live paths. The `byPathHop` key is the
same shared copy.
3. **Eviction and rebuild.**
- `evictStaleInternal` drops the record of evicted transmissions, plus
any interned key whose bucket it deletes.
- `retainResolvedPathHops` keeps the records of live transmissions only
and drops interned keys whose bucket was not carried over. When a
rebuild starts from an empty index, it clears both.
4. **Defence in depth.** The single (`GetRepeaterUsefulnessScore`),
batch (`GetRepeaterNodeStatsBatch`) and bulk
(`computeRepeaterUsefulnessScoreMap`) scores count **distinct**
non-advert transmissions per bucket, via `countDistinctNonAdvert`.
   - IDs are collected in a reused slice.
- A bucket in ascending ID order is already distinct; only out-of-order
buckets are sorted.
- The bulk pass does this for full-pubkey keys only. Raw-hop buckets get
one entry per transmission from `addTxToPathHopIndex`.
5. **Cache side effect.** A repeated observation through known relays no
longer mutates `byPathHop`, so it no longer drops the batch relay-stats
cache, as the cache contract from `1164` intends. An observation through
a new relay still drops it.

The other `byPathHop` consumers already de-duplicate by `tx.ID` or read
only raw prefix keys: `GetNodeHopAnalytics`,
`computeMultiByteCapability`, `handleNodePaths`,
`computeRepeaterRelayInfoMap` and the relay info in
`GetRepeaterNodeStatsBatch`. Two tests pin that they ignore duplicate
entries.

## Exact keys vs. a 64-bit fingerprint

The first version of this fix stored a 64-bit FNV-1a hash per key. As
requested, I measured both, plus exact keys without interning, on the
real `addResolvedPubkeysToPathHopIndex`.

**Keys per transmission.** In the e2e fixture, the union of resolved
relay keys over all observations of one transmission has a mean of 5.4
and a maximum of 17. The protocol limit is 64 path bytes
(`MAX_PATH_SIZE`), i.e. 64 / hash size hops.

**Memory** retained by the record, per transmission.
`BenchmarkPathHopRecordMemory_2108`: 100K transmissions, keys from 4,000
relays, each transmission heard twice with freshly allocated keys. The
record map and the interned keys are included.

| keys per tx | exact, interned (this PR) | 64-bit fingerprint | exact,
not interned |
|---|---|---|---|
| 2 | 88 B | 68 B | 207 B |
| 5 | 136 B | 101 B | 447 B |
| 17 | 344 B | 197 B | 1,423 B |

- At realistic key counts, exact keys cost 20–35 B per transmission more
than the fingerprint. That is about 3.5 MB per 100K transmissions, and
small next to what the store already charges per transmission
(`storeTxBaseBytes` alone is 384 B).
- Without interning, exact keys would cost 3–4× more, and most of that
would be held during every `Load`. Interning is what makes exact keys
affordable.

**Time.** I ran the variants interleaved, 3 rounds each, median ns/op on
an Apple M4. The record step is not measurable inside the full call. The
isolated record step was faster with exact keys in a separate
micro-benchmark, because comparing a handful of 64-char strings is
cheaper than hashing each one.

| benchmark | exact, interned | fingerprint | exact, not interned |
|---|---|---|---|
| late observation, 20K txs | 312 | 358 | 331 |
| late observation, 100K txs | 760 | 836 | 758 |
| `IngestNewObservations` (SQL + resolve + index) | 531 µs | 497 µs |
527 µs |

**Collision probability of the fingerprint.** For k distinct keys in one
transmission it is about k(k−1)/2 / 2^64:

| k | per transmission | per 10^9 transmissions |
|---|---|---|
| 2 | 5.4e-20 | 5.4e-11 |
| 5 | 5.4e-19 | 5.4e-10 |
| 17 | 7.4e-18 | 7.4e-9 |
| 64 | 1.1e-16 | 1.1e-7 |

For random keys this is negligible. FNV-1a is unkeyed, though, so two
relay keys that collide could be found deliberately (a birthday search
over about 2^32 key pairs). A collision would leave a transmission out
of one relay's bucket.

**Decision.** Exact keys are cheap enough once interned: the same speed
within noise, and a few tens of bytes per transmission. They remove the
question of collision-driven omissions entirely, so this PR uses exact
keys and drops the fingerprint.

The two new maps (`pathHopResolved`, `pathHopKeys`) are not added to
`trackedBytes`, so the `maxMemoryMB` trigger undercounts by roughly the
per-transmission figures above. Both are bounded by live state (eviction
and rebuild prune them), so this is a steady proportional undercount
rather than a leak.

## Tests

**New: `cmd/server/pathhop_dedupe_2108_test.go`.** Behaviour tests that
use only existing API. Each one runs against the real SQLite schema
through `Load`, `IngestNewFromDB` and `IngestNewObservations` where
noted.

| test | covers | on `master` |
|---|---|---|
| `TestPathHopIndexOncePerTx_LateObservations_2108` | 1 + 10 late
observations through the same relays: one entry per relay | **fails**
(11 entries) |
| `TestPathHopIndexOncePerTx_LiveIngestBatch_2108` | 11 observations in
one `IngestNewFromDB` batch | **fails** (11) |
| `TestPathHopIndexOncePerTx_AfterLoad_2108` | `Load` + rebuild, then
one live observation | **fails** (2) |
| `TestPathHopIndexAddsNewRelayFromLaterObservation_2108` | a later
observation through a **new** relay adds it exactly once, and its share
becomes correct | **fails** (2) |
| `TestTrafficShareStableAcrossLateObservations_2108` | single, batch
and bulk scores agree, equal the definition, stay put over rounds of
late observations, and equal a fresh `Load` of the same data | **fails**
(all four relays at 1.0, want 0.4–0.5) |
| `TestPathHopIndexSizeBoundedByTransmissions_2108` | the index size
does not grow with observations | **fails** |
| `TestRelayStatsCacheAcrossRepeatedObservations_2108` | a repeated
observation keeps the relay-stats cache, and the kept cache equals a
fresh compute; a new relay drops it, and the next read sees the relay |
**fails** |
| `TestAddResolvedPubkeysToPathHopIndex_PerRelayIdempotent_2108` | the
helper's return value and cache invalidation per (tx, relay), in any key
order | **fails** |
| `TestTrafficShareCountsDistinctTransmissions_2108` | all three scores
count distinct transmissions in a duplicated bucket | **fails** |
| `TestPathHopConsumersIgnoreDuplicateEntries_2108` | pin of the
consumer audit | passes (by design) |
| `TestMultiByteCapabilityIgnoresDuplicateEntries_2108` | pin of the
consumer audit | passes (by design) |

**New: `cmd/server/pathhop_record_2108_test.go`.** These tests use the
new symbols, so they do not compile against `master`.
- `TestCountDistinctNonAdvert_2108`: ascending, descending, interleaved
duplicates, adverts, nils, untyped.
- `TestPathHopResolvedRecordBoundedAndEvicted_2108`: the record stays at
2 keys after 11 observations. Eviction removes the record and every
entry. A survivor's next observation adds nothing. Evicting everything
leaves no record, no interned key and no bucket.
- `TestPathHopResolvedRecordAcrossRebuild_2108`: a rebuild keeps the
records of live transmissions, so the next observation adds nothing. It
drops the record of a removed transmission and the interned key only
that transmission used.
- `TestPathHopResolvedRecordClearedWithEmptyIndex_2108`: a rebuild from
an empty index clears the record and the interned keys, and the next
observation puts the transmission back.
- `TestPathHopResolvedRecordInternsKeys_2108`: record entries and the
`byPathHop` key share one copy, even though every observation passes
freshly allocated keys.

**Changed: `pathhop_eviction_1908_test.go`.** It built its duplicate
entries by repeating `indexResolvedPathHops`, which no longer
duplicates. It now seeds the duplicates directly, so the `1908` sweep is
still tested against buckets that hold one transmission several times.
The expected buckets are unchanged.

**Changed: `db_test.go`.** `setupTestDB` takes `testing.TB`, so the SQL
benchmark can use it.

**Mutants.** I applied each mutant on its own to this branch. All 14 are
killed:

| mutant | killed by |
|---|---|
| record never consulted | the OncePerTx, stable-share, size, cache and
record tests |
| dedupe per transmission instead of per relay |
`AddsNewRelayFromLaterObservation`,
`RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent` |
| eviction keeps the record | `RecordBoundedAndEvicted` |
| rebuild keeps records of removed transmissions | `RecordAcrossRebuild`
|
| rebuild from an empty index keeps the record |
`RecordClearedWithEmptyIndex` |
| cache dropped on every call |
`RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`,
existing `NoMutation_PreservesCache` |
| cache kept although `byPathHop` changed |
`RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`,
existing `InvalidatesRelayStatsCache` |
| single score counts entries |
`TrafficShareCountsDistinctTransmissions` |
| batch score counts entries | `TrafficShareCountsDistinctTransmissions`
|
| bulk score counts entries | `TrafficShareCountsDistinctTransmissions`
|
| distinct count trusts any bucket order |
`TrafficShareCountsDistinctTransmissions`, `CountDistinctNonAdvert` |
| keys not interned | `RecordInternsKeys` |
| eviction keeps interned keys | `RecordBoundedAndEvicted` |
| rebuild keeps interned keys of dropped buckets | `RecordAcrossRebuild`
|

**Commands run:**
- `gofmt -l` on all tracked Go files: clean.
- `go vet ./...` in all 14 modules: clean.
- `cd cmd/server && go test -race ./...`: pass.
- `cd cmd/ingestor && go test ./...`: pass.
- `sh test-all.sh`: 184 of 186 suites pass locally. The other two,
`test-issue-1956-release-routing.js` and `test-preflight-xss-gate.js`,
shell out to scripts that need bash ≥ 4 (`mapfile`). They fail under
macOS's bash 3.2 regardless of this change. This PR touches no frontend
or script files.

## Benchmark: `master` vs. this branch

I ran `master`'s sources (`000d9ab`) and this branch interleaved, 5
rounds, with the same benchmark files. Medians on an Apple M4.

| benchmark | `master` | this PR | change |
|---|---|---|---|
| late observation through known relays, 20K txs | 446 ns, 38 B/op | 441
ns, 0 B/op | within noise |
| late observation through known relays, 100K txs | 710 ns, 69 B/op |
755 ns, 0 B/op | within noise (runs overlap) |
| … index size afterwards, entries/tx (20K / 100K) | 24.99 / 8.99, still
growing | 4.99 / 4.99 | bounded |
| `IngestNewObservations`, one observation for each of 20 txs | 532 µs |
513 µs | within noise |
| bulk score pass, clean index, 20K, ingest order | 246 µs | 316 µs |
+28 % |
| bulk score pass, clean index, 100K, ingest order | 3.03 ms | 4.30 ms |
+42 % |
| bulk score pass, clean index, 20K, reversed buckets | 257 µs | 467 µs
| +82 % |
| bulk score pass, clean index, 100K, reversed buckets | 3.38 ms | 4.79
ms | +42 % |
| bulk score pass after 11 observations per tx, 20K | 663 µs | 284 µs |
−57 % |
| bulk score pass after 11 observations per tx, 100K | 4.72 ms | 3.19 ms
| −32 % |

- On an index that is clean on both builds, the distinct count makes the
bulk pass slower. That pass runs on the cache-miss path, which the
background recomputer refreshes every 5 minutes by default.
- On the index a live server actually holds without this fix, `master`
walks every accumulated duplicate, so the real-world pass is faster
after the fix. The gap grows with uptime.
- The late-observation step no longer allocates. On `master` its index
grows with every observation.

Benchmarks: `BenchmarkLateObservationIndex_2108`,
`BenchmarkIngestNewObservations_2108`,
`BenchmarkTrafficShareScoreMap_2108` and
`BenchmarkPathHopRecordMemory_2108`.

## Production motivation

We have run this fix on two production instances. Before the fix, the
summed `traffic_share_score` grew past 180 and many relays sat at the
1.0 clamp; with the fix no node reaches `≥ 0.999`.

A controlled 12 h A/B run makes the drift explicit. Two servers read one
identical database — one on this `master` base, one with the fix —
alongside a reference server that freshly loads the same database (a
fresh load is correct, because the rebuild dedupes). The unfixed server
drifted to 9 relays at the 1.0 clamp and up to 0.95 absolute error per
node against the reference; the fixed server stayed within 0.05 of the
reference for every node, with no relay at the clamp.

A smaller residual rise remains and is a separate cause (startup-vs-live
hop resolution); it is deliberately left for a follow-up so its effect
stays measurable.

## Out of scope

That remaining slow rise comes from a separate cause: live ingest
resolves some hops that the startup load does not. This PR deliberately
leaves that drift alone, so its effect stays measurable. The fix will
follow in a separate PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 14:34:31 +02:00
efitenandClaude Opus 5.5 f3ae8ac25a feat: sync a logged-in user's settings across devices (part B) (#2130)
Part B of #2128: a logged-in user's settings follow them across devices.
Log in on a phone and your own nodes, favorites, customizer and filters
are there; a change on one device reaches the others within a minute or
when you return to the tab.

**This PR builds on #2129.** Until that one is merged, the diff here
includes it. The commits for this part start at `docs(specs): settings
sync for optional user management (sub-project B)`.

## The situation

Everything a visitor sets up lives in one browser's `localStorage`
(about 100 keys in `public/`). A second device or a cleared cache starts
from zero (#895).

## What this PR adds

**Storage.** `users.db` schema v2: one JSON document per user in
`user_settings`, with a revision number and a generation id. A write
succeeds only when the client's revision and generation match the stored
ones, so two devices cannot overwrite each other silently.

**Server.** `GET`, `PUT` and `DELETE /api/account/settings`, behind the
same session and CSRF checks as the account routes.
- The server owns the list of synced keys (61 keys,
[`settings_allowlist.go`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/cmd/server/settings_allowlist.go))
and sends it to the client, so the two cannot drift.
- A hard denylist, checked first, refuses `meshcore-api-key`, every
`corescope_channel_*` key and `live-channel-colors` (#725). The colour
map is keyed by channel hash, and for a user-added channel that hash is
`user:<name>`, which would expose hashtag channel names.
- Documents are capped at 256 KiB, measured like `JSON.stringify`. PUT
is limited to 60 requests per hour per user. A stale revision gets 409
with the current document.

**Client**
([`settings-sync.js`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/public/settings-sync.js)).
Inert unless the feature is on and someone is logged in.
- It wraps `localStorage.setItem` and `removeItem` for allowlisted keys
only and pushes 2 seconds after the last change.
- It pulls on login, page load, tab focus and every 60 seconds while the
tab is visible.
- **Merge:** three-way, against a per-device baseline that belongs to
one user and one document generation. Lists (own nodes, favorites, saved
filters) merge per item, so an item added anywhere is kept and an item
removed on one device does not come back from another. Single values:
the profile wins unless only this device changed it.
- Remote changes are written without a push, theme and colour-blind
preset are re-applied, and the current page re-renders (skipped on
account pages and while the geofilter editor is open).

**UI.**
- Logout asks: keep my settings on this device (default), remove them
from this device, or cancel. Channel keys are never removed: no copy
exists anywhere else.
- The account page gets a "Settings sync" section: last synced time,
"Sync now", what is and is not synced, and "Delete synced settings from
my account".

## Not synced

Layout and device state (panel and column widths, collapsed panels, map
positions, geofilter drafts), channel data (#725), the API key, and all
`sessionStorage`. The full list is in the
[spec](https://github.com/efiten/CoreScope/blob/feat/settings-sync/docs/specs/2026-10-06-user-settings-sync-design.md).

## Performance

- One GET per page load, tab focus and minute while visible; one
debounced PUT per burst of changes.
- The `setItem` wrapper costs one Set lookup per write for non-synced
keys. A synced write reads one small revision key, not the stored
document.
- The server reads or writes one row per request.

## Verification

- `internal/users` and `cmd/server`: `go vet` and `go test` pass locally
(22 new Go tests), including a test that every allowlisted key still
occurs in `public/`, and denylist tests.
- `tests/unit/test-settings-sync.js`: 79 passing (vm, real module). The
cases cover the merge table, two tabs sharing one storage, stale answers
after a push, delete while a push is in flight, and logout while the
final push fails.
- `sh test-all.sh` exits 0.
- `tests/e2e/test-user-management-e2e.js` (10 steps, 4 of them new)
passed locally with two browser contexts as two devices: a favorite and
the packet time window travel from device 1 to device 2, a removal does
not come back, and "remove from this device" clears the synced keys
while a channel key stays.
- Checked by hand on a staging instance with a desktop and a phone on
one account.

## Not in this PR

- On a shared browser where the previous user chose "keep", the next
user's first login merges those settings into their own account. The
user guide says to choose "remove" on shared computers.
- Saved filter expressions are synced as typed, including any channel
names written in them. The guide says so.
- Realtime push between devices; the minute pull is the sync interval.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 11:20:44 +02:00
efitenandClaude Opus 5.5 d232072f85 feat: optional user accounts (part A: foundation) (#2129)
Part A of #2128: optional, off-by-default user accounts. With the
feature off nothing changes; with it on, visitors can register and log
in, and admins manage users and use the operator actions without the API
key.

PR #2130 (settings sync) builds on this one. The two are meant to be
merged together.

## The situation

- Operator actions (geofilter save and prune, backup, perf reset) need
the shared `apiKey`. There is no per-person right.
- Nothing in CoreScope knows who a visitor is, so the requests in #2128
that need that (#1835, #2092, #1508, #730) have nothing to build on.

## What this PR adds

**Two new Go modules**
- `internal/users`: a separate `users.db` (SQLite through
`modernc.org/sqlite`) with users, sessions, single-use tokens, an audit
log and a mail log. Passwords use argon2id.
- `internal/mailer`: a `Mailer` interface with a Brevo client (send,
delivery events, webhook parsing) and an in-memory fake for tests.

**Server (`cmd/server`)**, active only with `userManagement.enabled`
- 24 routes, all documented in OpenAPI under the `users` tag
([`auth_routes.go`](https://github.com/efiten/CoreScope/blob/feat/user-management/cmd/server/auth_routes.go)):
  - auth: register, activate, login, logout, me, forgot, reset;
- account: profile, password, email change with confirmation, sessions,
self-delete;
- admin: list, detail, disable, enable, delete, role, resend activation,
manual activation, mail status refresh;
  - a Brevo webhook, registered only when `mail.webhookSecret` is set.
- `requireAdmin` replaces `requireAPIKey` at the 7 operator call sites:
the API key **or** an admin session. With the feature off it is the old
API-key gate (`TestRequireAdminWithoutUserManagementIsAPIKeyGate`).
- `/api/config/client` gets `userManagement: {enabled: true}` only when
the service started; with the feature off the response is
byte-identical.

**Frontend**
- `auth.js` (header account control, request helper that adds the CSRF
header), `account.js` (login, register, activate, forgot, reset, confirm
email, my account), `admin-users.js` (`#/admin/users`, deep-linked
filters), `account.css` (theme tokens only).
- On phones the top-bar control is hidden, so a conditional entry goes
into the bottom-nav "More" sheet and the nav drawer.
- The customizer geofilter tab and the Perf "Reset stats" button use the
admin session when there is one.

**Config.** A `userManagement` block (`config.example.json`,
[`docs/user-guide/accounts.md`](https://github.com/efiten/CoreScope/blob/feat/user-management/docs/user-guide/accounts.md)).
The Brevo key can come from `CORESCOPE_BREVO_API_KEY`. The server
refuses to start when the block is enabled but incomplete.

## Security choices

- Session cookie `cs_session`: HttpOnly, SameSite=Lax, Secure when
`publicBaseUrl` is https. Every cookie-authenticated state change needs
the `X-CS-CSRF` header and a matching Origin.
- Activation needs the token **and** the account password. Without the
password, an attacker who keeps re-registering a known address could get
the owner to activate an account that carries the attacker's password.
- Register, forgot and email change answer identically for known and
unknown addresses. A password reset ends all sessions, a password change
ends all other sessions, and both end outstanding email-change links.
- Rate limits: login 10 per 15 minutes, register and forgot 5 per hour,
per IP and per address. The bucket count is capped. `trustedProxies`
makes the per-IP limits see real client IPs behind a proxy.
- Server logs carry `#<user id>`, never addresses, tokens or passwords;
mail-provider error texts are redacted before logging.

## Performance

No change to an existing hot path with the feature off. With it on:
- One `users.db` lookup per authenticated request (session by token
hash).
- The admin user table rebuilds its `tbody` on each filter change.
`users.List` caps the result at 1000 rows (`internal/users/users.go`),
which bounds the rebuild.
- `map[string]interface{}` in `openapi.go`: 79 before, 78 after.

## Verification

- `internal/users`, `internal/mailer` and `cmd/server`: `go vet` and `go
test -race` pass locally. 121 new Go tests.
- `cmd/server` with `-tags e2etest`: vet and the e2e hook tests pass.
- `sh test-all.sh` exits 0. `tests/unit/test-user-management-ui.js`: 67
passing (vm, real modules).
- `tests/e2e/test-user-management-e2e.js` (6 steps) passed locally
against an `e2etest` build with the fake mailer and against a
feature-off build. CI builds the `e2etest` binary and runs the suite on
a second server (`deploy.yml`).
- On a staging instance with a real Brevo key: register, activation mail
delivered, activate, admin table, "Refresh status" showing sent,
deferred, delivered, opened and clicked.

## Not in this PR

- Settings sync (#2130), the admin dashboard, approval flows and
notifications (parts B to E of #2128).
- A `requireReadAuth` mode (#1835). Sessions from this PR are what such
a mode would accept.
- Binary size and build time with `modernc.org/sqlite` linked next to
`mattn/go-sqlite3` were not measured. Their driver names do not collide.
#1992 discusses the driver choice.
- No Brevo webhook was configured on staging; delivery status there came
from "Refresh status".

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 09:25:29 +02:00
efitenandClaude Opus 5.5 4363403495 fix(nodes): match the region node filter on observer ID, not ingest-time IATA (#2115)
Follow-up to #2114, from its review.

`RegionNodePubkeys` compared the IATA copied onto each observation at
ingest (`StoreObs.ObserverIATA`). When an operator changes an observer's
code, those copies stay as they were until a restart or until the
observations age out. `/api/nodes?region=` then kept matching the old
region, while the other store region filters already used the new one.
Those filters resolve observers by ID through `resolveRegionObservers`,
for example the packets query and `computeNodeHomeRegions`. The SQL
subquery that #2114 replaced also read `observers.iata` at query time,
so this restores that behaviour.

## Also fixes a regression from #2114

Found in review: `loadChunk`, the background history loader
(`cmd/server/store.go`, the chunk SELECT and the `StoreObs` it builds),
sets `ObserverID` on each observation but never `ObserverIATA`. With
#2114 matching on that IATA, **every node whose adverts came only from
background-loaded history dropped out of `/api/nodes?region=`** on
current master. Matching on observer ID fixes it.
`TestRegionNodePubkeysMatchesChunkLoadedObservations` pins it (an
observation with the observer ID and an empty IATA) and fails on
master's `region_nodes.go`.

## Change

- Resolve the region to observer IDs with `resolveRegionObservers` (own
mutex, 30 s cache) and match observations by `ObserverID`. Lock order:
`regionNodesMu` is released before it, and `s.mu` is taken after it;
none of the three is held together.
- Without a database there is nothing to resolve, so `RegionNodePubkeys`
reports no set and the handler keeps the SQL path.

## Tests

- The region tests now seed an observers table and leave each
observation's IATA at a stale value, so they can only pass through the
table.
- New `TestRegionNodePubkeysFollowsObserverIATAChange`: an observer that
moved from SJC to SFO matches SFO and not SJC. It fails on the previous
code (`got [pk_moved]` for SJC).
- `TestRegionNodePubkeysMatchesChunkLoadedObservations` (second commit),
see above.
- `setupTestDB` takes `testing.TB` so the benchmark can use it; every
existing caller passes `*testing.T` unchanged.
- Full `cmd/server` suite passes locally, the region tests also with
`-race`.

## Performance

`BenchmarkRegionNodePubkeys` (220k adverts × 8 observations): 37 ms to
30 ms per uncached scan, a map lookup per observation instead of a
string normalisation. The observer lookup is one query on the small
`observers` table, cached for 30 s.

## Not done

- No singleflight on a cold cache, and the 64-entry cache still resets
when full. Both were non-blocking in the #2114 review.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 23:42:07 +02:00
efitenandClaude Opus 5.5 d5e6d1b2a6 fix(nodes): resolve the /api/nodes region filter from the packet store (#2114)
Fixes #2101.

`/api/nodes?region=` filtered nodes with a subquery that joins every
advert to all of its observations and observers, with no time bound
(`cmd/server/db.go`, `GetNodes`). It ran twice per request, once for
`COUNT(*)` and once for the page. A client paging through nodes repeats
both per page.

## Reproduced

On our staging database (10.3 GB, 16.5M observations, 219,750 adverts),
read-only `sqlite3`, region `BRU`:

| | real | user | sys |
|---|---|---|---|
| one regional `COUNT(*)`, as `GetNodes` builds it | **154.8 s** | 1.9 s
| 5.8 s |

Almost all of it is waiting on disk. The plan walks
`idx_transmissions_payload_type` for every advert and then
`idx_observations_dedup` for each one's observations. Two of those per
request, a few requests at once, and the reader pool is gone, which
matches the report of all four database workers sitting in `GetNodes`.

## Fix

- **`PacketStore.RegionNodePubkeys(region)`** (new
`cmd/server/region_nodes.go`) walks the store's in-memory adverts
(`byPayloadType[ADVERT]`) once. It keeps the pubkey of every node with
an advert heard by an observer in the region, using the `ObserverIATA`
each observation already carries. It is cached for 30 s per region, and
the cache is bounded to 64 entries because its keys come from the
client's parameter. Lock discipline follows `resolveAreaNodes`: the
cache mutex and `s.mu` are never held together, and the ordering note in
`store.go` lists it.
- **`GetNodes` takes a `NodeQuery` struct** with a `RegionPubkeys`
field, passed as one `json_each` parameter: `public_key IN (SELECT value
FROM json_each(?))`, a primary-key lookup. An empty set matches nothing.
The SQL subquery stays for a server without a store (tests, tooling).
- The advert-pubkey lookup that `trackAdvertPubkey`,
`untrackAdvertPubkey` and `computeNodeHomeRegions` each copied is now
one helper, `advertPubkey`.

The struct instead of a second `GetNodes` variant keeps the
`map[string]interface{}` count unchanged in `db.go` (75) and `routes.go`
(59).

## Behaviour change

The region filter now covers the adverts the store holds
(`packetStore.retentionHours`), which is the window the rest of the UI
shows, instead of all database history. A node heard in a region only
before that window no longer matches the filter. While the store is
still loading after a restart, the set grows as history loads.

## Performance

| | before | after |
|---|---|---|
| region set, uncached | 154.8 s (SQL count, staging) | 37 ms
(`BenchmarkRegionNodePubkeys`: 220k adverts × 8 observations, 4,000
nodes) |
| region set, within 30 s | same again | cache hit |
| node count over the set | (included above) | 2 ms on staging (1,200
keys) |
| 500-row page over the set | same scan again | 3 ms on staging |

The scan holds `s.mu` for reading for those 37 ms, at most once per
region every 30 s.

## Tests

- `TestRegionNodePubkeys`: the in-region advert, an advert heard in two
regions, case and whitespace in codes, a comma list, an unknown region
giving an empty set, a non-advert never counting, a blank region giving
no filter.
- `TestRegionNodePubkeysCached`, `TestRegionNodePubkeysCacheIsBounded`
(1,000 distinct regions stay within 64 entries).
- `TestGetNodesRegionPubkeys`: the set combines with the role filter and
counts correctly, and an empty set returns nothing even with `Region`
set.
- `TestHandleNodesRegionUsesStore`: an advert that only the store knows
about shows up through `/api/nodes?region=`, so the handler is proven
not to ask SQL.
- Existing region tests (`TestGetNodesRegionFilterV2` and the
`db_test.go` region cases) pass unchanged through the SQL fallback. Full
`cmd/server` suite passes locally; the new tests also pass with `-race`.

## Not done

- No request context on these queries, also raised in the issue. With
the scan gone they take milliseconds, so I left that out of this change.
- The region semantics stay "heard by an observer in the region". #1879
argues for the node's home region instead; that is a separate decision.
- No frontend change. I did not check this in a browser; the nodes page
and map call the same endpoint with the same parameters.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 21:45:02 +02:00
Sylvain Rabot b5b230e884 feat(channels): show the sender's path hash size on each message (#2089)
Each channel message now shows the hash size its sender's path uses, read
from bits 7-6 of the path byte that the originator writes and repeaters keep.
The server sends it as path_hash_size (packetpath.HashSize); the frontend
helper pathHashSize() applies the same rule and returns 0 for unknown.

One rule on both pages: the packet detail Hash Size row and the hex
breakdown call pathHashSize() too, so a 0-hop flood message reports the
same size on Channels and on its packet page. A direct packet with no hops
left reports no size, matching cmd/server/decoder.go. The cases live in
test-fixtures/path-hash-size-cases.json, read by the Go and JS tests.
2026-10-05 20:32:40 +02:00
efitenandClaude Opus 5.5 000d9ab030 feat(coverage): RF noise-floor layer on the Mobile RX coverage page (#2113)
The ingestor already stores the noise floor that CoreDrive RX companions
report with each GPS fix (`client_rf_samples`, opt-in through
`clientRfSamples`, `cmd/ingestor/client_rf_sample.go:62`), but nothing
reads it back. An operator who enables it collects the data and cannot
see it. This adds the read side: `GET /api/rf-noise` and a Signal/Noise
toggle on the Mobile RX coverage page.

## What it does

- **`GET /api/rf-noise?bbox=&z=&days=`** returns a GeoJSON hex grid with
the median, quietest and noisiest noise floor per cell, in the same
shape as `/api/rx-coverage`. It is registered always and 404s unless
`clientRfSamples.enabled` is true, the same pattern as the coverage
routes.
- **Stationary samples are excluded.** A parked companion logs hundreds
of readings at one point, which would otherwise define its cell.
- **Coverage page:** a Signal/Noise toggle, rendered only when
`/api/config/client` reports `clientRfSamples: true`. The noise layer
reuses the coverage colour tokens with the axis inverted, because a
lower dBm is quieter. Tiers are at -115 and -108 dBm.
- **Deep link:** `#/rx-coverage?layer=noise` opens on the noise layer.
- **Empty and failed answers** are labelled on the map ("No RF samples
in this view yet", or a retry hint), so a blank map never reads as
"feature off".

No new configuration key: it reads the existing `clientRfSamples`
section that the ingestor already uses. Default off, so nothing changes
for an instance that has not opted in.

## Where it comes from

Ported from the efiten/CoreScope fork, where it has run on
analyzer.on8ar.eu since September (fork commits `42d09f0e`, `80581c43`,
`cde95078`). The cherry-picks conflicted with upstream's newer
`routes.go`, `types.go` and `rx-coverage.js`, so the final state was
ported by hand. The fork-only `/scopes` route that sat next to it in the
same hunk is deliberately left out.

## Performance

The query is bounded by `sampled_at` (indexed, `idx_crf_prune`) and the
bbox, aggregation is one pass plus a per-cell sort, and the response is
capped at 5000 cells. Measured on analyzer.on8ar.eu, all of Belgium
(`bbox=49.4,2.4,51.6,6.5&z=9`), from a client in Belgium, so network
time included:

| window | samples | cells | response time |
|---|---|---|---|
| 7 days | 10,835 | 241 | 0.37 s |
| 30 days | 43,303 | 429 | 0.39 s |

That table holds 45,270 rows in total. It is only read when someone
opens the noise layer, never on ingest or WebSocket paths.

## Tests

- `cmd/server/rf_noise_test.go`: 8 tests for the aggregation (median and
extremes, stationary exclusion, the cell cap, empty input) and the gate
(404 when off, even with data present). All pass, and the full
`cmd/server` suite passes locally.
- `tests/unit/test-rx-coverage-noise.js` (new, in `test-all.sh`): the
colour axis runs the right way, including both tier bounds and a string
median from the API. Slices the real function out of `rx-coverage.js`
and fails on master's copy, which has no thresholds.
- `tests/e2e/test-rx-coverage-noise-e2e.js` (new, wired in `deploy.yml`
and `scripts/non-unit-tests.json`): no toggle when the flag is off;
Noise fetches `/api/rf-noise`, draws the cells, swaps legend and
subtitle, puts `layer=noise` in the hash; a deep link opens on the noise
layer and an empty answer shows the message. 3/3 locally; against
master's `rx-coverage.js` and `roles.js`, 1/3 (only the "flag off" case
passes, as it should).
- `test-rx-coverage-viewport-e2e.js` still passes. ESLint 8 reports 0
errors on the changed files.

## Not done

- The thresholds (-115 / -108 dBm) are fitted to the fork's own data
(1,241 moving samples at the time) and are constants in
`rx-coverage.js`. Per AGENTS.md rule 8 they belong in the customizer
eventually; not in this PR.
- The E2E suite mocks the API. The real endpoint is covered by the Go
tests and by the measurements above, not by a fixture with seeded
samples.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 22:11:34 +02:00
efitenandClaude Opus 5.5 9c3d76e14f docs(release): notes and changelog for v3.13.1 (#2112)
Release notes and CHANGELOG entry for v3.13.1, a patch release that
ships #2111 (node-discover replies count as coverage).

Once merged, the tag goes on the merge commit so `deploy.yml` picks up
`docs/release-notes/v3.13.1.md` as the release body.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
v3.13.1
2026-10-04 17:22:56 +02:00
efitenandClaude Opus 5 00ee4f9e23 fix(coverage): attribute node-discover replies to the responder (#2111)
## Problem

A CoreDrive RX companion that runs node-discover gets no coverage on an
upstream instance. Reported today by an operator on v3.13.0: the app
logged 32 discover replies from one repeater in 16 minutes (`heard
751a49f0c5dadc70 (8B, discover)`), all published, and
`client_receptions` stayed at 0 rows for that companion while
`client_rx_observations` held 93.

Two gaps, both on the upstream side:

1. **Ingestor.** `deriveHeardKey` attributes a FLOOD `path[last]` and a
0-hop advert, and drops a 0-hop `CONTROL/DISCOVER_RESP` (firmware
`CTL_TYPE_NODE_DISCOVER_RESP`). The reply carries the responder's own
pubkey at offset 6 of the payload, which `decoder.go` already parses
into `CtrlPubKey` (#1802), but the coverage path never used it.
2. **Server.** CoreDrive RX asks for `DISCOVER_PREFIX_ONLY`, so the
firmware answers with an 8-byte prefix
(`simple_repeater/MyMesh.cpp:799-805`). `coverageHeardKeyCandidates`
only built the 64, 6 and 4 hex candidates, so a 16-hex `heard_key` would
be stored and then matched by no per-node coverage query.

## Change

- `cmd/ingestor/client_reception.go`: a third branch in `deriveHeardKey`
for a discover response with no hops. Accepts exactly 8 or 32 bytes,
nothing truncated, stored with `src='discover'`.
- `cmd/server/rx_coverage.go`: adds the 16-hex prefix to
`coverageHeardKeyCandidates`.
- `docs/client-rx-coverage.md`: documents the `discover` source, the
8-byte keylen and the four-candidate lookup.

The leaderboard and `/api/rx-coverage` read `client_receptions` without
a key filter, so they pick the rows up without a change. Name resolution
goes through `batchResolveHeardKeys`, which is a prefix lookup and
handles 16 hex as is.

## Evidence from a deployment that has had this since 2026-08-19

On analyzer.on8ar.eu, `client_receptions` over the last 7 days by `src`:
discover 4232, rxlog 4909, geo 548, advert 34. Discover replies are 44%
of all coverage rows there (4232 of 9723); on an upstream instance those
rows are not written. They cannot be backfilled afterwards either:
`client_rx_observations` keeps no raw bytes, so the responder pubkey is
gone.

## Tests

- `TestDeriveHeardKey` and `TestBuildClientReception` gain discover
cases: 8-byte and 32-byte keys accepted (32-byte uppercase input
lowercased), a 3-byte and an empty key rejected, a non-discover CONTROL
rejected, a discover response with hops not attributed.
- `TestHandleClientPacketDiscoverRespWritesReception`: end to end, a raw
0-hop DISCOVER_RESP on the client topic writes one `client_receptions`
row with `src='discover'`.
- `TestCoverageHeardKeyCandidatesIncludesDiscoverPrefix`: the 16-hex
prefix is among the per-node candidates.
- Ran locally on Windows: `go test ./...` in `cmd/server` passes; in
`cmd/ingestor` everything passes except
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which needs the symlink
privilege on Windows and fails on clean master too.

Not done: no browser validation, as the change is ingestor and server
only and the frontend reads the same endpoints. Not included: geographic
resolution of 1-byte hops (`src='geo'`) and the RF noise layer, which
are separate changes.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:15:50 +02:00
efitenandClaude Opus 5.5 c2a9500060 docs(release): notes and changelog for v3.13.0 (#2110)
Release notes and CHANGELOG entry for v3.13.0, in the shape used for
v3.12.0. Once merged, the tag goes on the merge commit so `deploy.yml`
picks up `docs/release-notes/v3.13.0.md` as the release body (#2076).

- 16 commits since v3.12.0: 10 fix, 4 test, 2 feat, so a minor bump.
- "Read this before upgrading" covers the one-off advert route evidence
backfill from #2088. Durations are measured on two instances (live 24m
1s on 16.8M observations, staging 36m 23s on 16.5M), and ingest during
the backfill window is measured as unaffected on live.
- Issues closed come from each PR's closing references: 14 from 16 PRs.

Docs only. #2089 is deliberately not in this release.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
v3.13.0
2026-10-04 13:23:25 +02:00
efitenandClaude Opus 5.5 ef12535364 test(map): wait for the new route sidebar before measuring it (#2109)
Step (9) of `tests/e2e/test-path-inspector-e2e.js` fails intermittently
on master with `(3-9) fixture route coverage: 0 !== 700`: 3 of the 18
runs that executed it since 2026-09-30, including master runs
[37120155356](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37120155356)
and
[37156362836](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37156362836).
A red master run skips the image build, so `:edge` is not rebuilt.

## Cause

The step clicks a candidate and then waits for any `.mc-rt-sidebar`. The
previous route's sidebar survives the hash-only `goto` to `#/map`, and
`drawPacketRoute` (`public/map.js:1080`) replaces it only after
`/api/resolve-hops` answers. So the wait matched the old sidebar, and
the checks that follow measured either the old sidebar or, if the
replacement landed between the locator resolving and the evaluation, a
detached element, whose `getBoundingClientRect()` width is 0.

This is a test defect. The app keeps the previous route on screen while
the next one loads, which is intended.

## Fix

Remember the sidebar before the click and wait until a different one is
in the DOM. Test-only, one file.

## Evidence

Locally against the fixture, with `/api/resolve-hops` delayed through
`page.route` in a throwaway copy of the test, measuring which sidebar
the step-9 assertion sees:

| Delay | Before | After |
|---|---|---|
| 0 ms | new sidebar, pass | new sidebar, pass |
| 80 to 100 ms | old sidebar, `isConnected: false`, width 0 (3 of 5
runs) | new sidebar, pass (5 of 5) |
| 1500 ms | old sidebar, pass: the step checked the wrong page | new
sidebar, pass |

The unmodified fixed test then passed 3 of 3 runs.

## Not done

- I did not look for the same wait pattern in other suites.
- The delay probe is not committed. It needs a timing window to be
useful, and that window would be a fixed sleep in CI.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:27:02 +02:00
n30nex 9e49db9725 feat(nodes): group adverts by observed route evidence (#2085)
Red commit: `ac60248` ([two intended label
failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37137429333)).
Original grouping red: `767da9b`
([CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36353796690)).

Fixes #2073.

#2088 is merged. Rebased onto master `415362c`; all eight patches are
unchanged, including the test-first history.

Recent Adverts groups the bounded sample by authoritative `advert_kind`:
Flood, Mixed flood / direct (empty path), Direct (empty path), and Other
/ unknown. Missing evidence stays unknown. Each advert appears once,
preserving reception metadata, order, packet links, counts and
origin/history explanations. Wording describes observed remaining paths
without inferring original send mode or RF distance.

Validation at `6317baf`:
- All 185 standalone suites passed; syntax, XSS and lint passed (zero
errors, 92 existing warnings).
- Browser verified: `http://127.0.0.1:51827` (temporary Go fixture
server, now stopped), desktop pane/full view and mobile. Mixed evidence,
exact links and caveats checked again after pushing.
- The expanded fixture exposes the existing #2104 failure in the
unmodified core harness. Supplemental integration with #2105's exact
test correction passes 135 checks, with three existing skips. Both PRs
remain separate.
- [Fresh green
CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37138366726)
passed; three independent automated reviews found no must-fix issues.

E2E assertion added: `tests/e2e/test-e2e-playwright.js:142`.

Grouping is O(n) over the existing 20-advert sample. No new requests,
settings, colors or dependencies.

## Preflight overrides
External OpenClaw runner/profile unavailable; repository checks and
actual local Chromium were used.
2026-10-03 23:49:08 +02:00
n30nex 99a326a6b3 test(packets): verify observation hex after rerenders (#2105)
Red commit: `ca553e4` ([expected assertion
failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37137426691)).

Fixes #2104.

The observation hex test now opens the existing grouped fixture
directly, selects A → B → A with stable-ID locators, and asserts the
selected ID and exact raw bytes after each render. Missing fixture data
fails explicitly. Distinct observation frames are seeded after schema
migration; this test's fixed sleeps and silent skips are removed.

Rebased onto master `415362c` after #2090 merged. Both patches are
unchanged. The previous #1122 layout blocker now passes against the
actual current server and assets.

Validation at `5db8017`:
- All 185 standalone frontend suites passed.
- Core browser: 133 passed, three existing skips; #1122 layout 6/6,
filter 11/11, grouped-collapse 4/4.
- Syntax, YAML, fixture SQL, XSS and lint passed (zero errors, 92
existing warnings).
- Fresh red CI reached the intended detached-observation assertion.
[Green
CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37138187496)
passed; three independent automated reviews found no must-fix issues.

Only test code and fixture setup change; no configuration/customizer
implications.

## Preflight overrides
External OpenClaw runner/profile unavailable; repository checks and
actual local Chromium were used.
2026-10-03 23:48:32 +02:00
efitenandClaude Opus 5.5 415362c5af fix(packets): resolve path hops with the observer that heard them (#2097) (#2099)
Fixes #2097.

Every 1-byte path prefix on this network is shared. Measured on the live
deployment: 254 prefixes cover all 2043 nodes, nine of them on a single
prefix, and not one node has a prefix to itself. The packets page
printed one of those nine as a certainty.

It could already have said otherwise. `hop-resolver.js` marks such a hop
`ambiguous` and returns its full candidate list; `hop-display.js`
renders a warning badge with a count and a candidate popover;
`renderHop(h, observerId)` reads a per-observer cache key. None of it
fired, because `HopResolver.resolve()` takes six parameters and
`packets.js` passed one.

Without `observerId`, `packetIata` is null, `nodeInRegion()` never runs,
no candidate is flagged `regional`, `globalFallback` stays false, and
`hop-display.js` computes a badge count of 0. The whole chain stayed
silent while the data said it was a guess.

## What this changes

**Resolution carries the observer.** `resolveHops()` takes it and writes
the per-observer cache key that `renderHop()` has always read and
nothing ever wrote. `resolveHopsForPackets()` groups a multi-packet call
by observer so each group is filtered on its own. One site was the bug
in miniature: it wrote its result under `hopNameCache[k + ':' +
pkt.observer_id]` after resolving without that observer, so the key
promised something the value was not.

**The observer's position is the anchor.** `resolve()` has always
accepted `observerLat`/`observerLon` as the anchor at the receiving end:
when nothing later in the path is resolved, the observer is the next
known position, which is what `pickByAffinity` needs.
`observerPosition()` is the single lookup, used by both resolve paths.

This matters more than the IATA route, which turns out to be dead on
this network. `nodeInRegion()` looks the observer's code up in
`/api/iata-coords`; **0 of 42 observers have their code in that table**
(`ANR BRU GNE HEP KJK LGG MST NRW OBL OST` are all missing, while its 54
entries are `AMS`, `APC`, `ATL` and the like). So the 300 km filter has
never excluded anything here. 27 of 42 observers report lat/lon
directly, and that works. Filed separately as #2098, since it is a
network-wide lookup that resolves nothing and hop resolution may not be
its only consumer.

**Each badge belongs to its pill.** The badge is a sibling after the
pill, so a path read `A [8] → B [6] → C` with nothing to say which name
the 8 qualified. Pill and badge now render inside one `white-space:
nowrap` `.hop-group`, leaving the separator outside the pair. Only a hop
that carries a badge is wrapped; wrapping every hop would change the
layout of every path for nothing. That markup predates this work, but
these commits are what made it visible — before them the badge
essentially never appeared on this page.

**The list summarises, the detail pane does not.** Every 1-byte hop
getting its own badge turned a table row into a line of warning
triangles. `renderPath()` takes `{ summary: true }` for the three list
call sites: names with no per-hop badge, and one indicator reading "N of
M hops have more than one candidate". The two detail-pane sites are
unchanged and keep a badge per hop, because that is where the candidates
actually get read. `HopDisplay` gains `opts.badge === false` for it, and
the hop keeps its `hop-ambiguous` class either way so CSS can still mark
it.

## Scope

Deliberately not included: chaining the next resolved hop to narrow a
candidate set to one. `pickByAffinity` already scores on graph edges and
would do it; it needs neighbouring context this call does not yet
supply, and that belongs in its own change to `hop-resolver.js`. What is
here makes the uncertainty visible and picks a defensible candidate
rather than the first in index order.

## Verification

30 unit tests in `tests/unit/test-issue-2097-hop-ambiguity-badge.js`,
registered in `test-all.sh`. Full frontend suite exits 0.

Four of them are structural guards over `packets.js`, and each was
mutation-checked by reverting the change it protects:

- no `HopResolver.resolve()` call passes the hops alone
- no call passes an observer id and drops its position
- one helper does the observer lookup
- the detail pane calls `renderPath` without `summary`

The behavioural test for the anchor runs with `iataCoords` deliberately
empty, which is the live state, and lists the distant candidate first:
without an anchor the resolver keeps candidate order and picks it, with
one it does not, and the hop stays reported as ambiguous with all
candidates listed.

**Browser validated** on a staging deployment of this branch, against
live data:

| | list | detail pane |
|---|---|---|
| per-hop badges | 0 | 24 |
| path indicators | 1 | 0 |
| badges outside a `.hop-group` | — | 0 |

On the packet that prompted the issue, the first hop now resolves to a
repeater 10.6 km from the transmitter instead of one 126 km away, and
the second hop resolves to the only candidate that has a recorded
neighbour edge matching the next hop in the path.

## Note for reviewers

The last commit exists because staging disagreed with the tests. After
the anchor landed, the detail pane still picked the distant candidate:
it resolves through its own `HopResolver.resolve` call, which had the
observer id but a null position. The guard that now fails on exactly
that shape is what would have caught it before the deploy.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:40 +02:00
efitenandClaude Opus 5.5 8933e2c703 fix(channels): keep client-only state across a channel-list refresh (#2095) (#2096)
Fixes #2095.

`loadChannels()` replaced the `channels` array with the server snapshot
and carried nothing across, and `mergeUserChannels()` only ever ran from
`init()`. So any refresh destroyed every field that exists only in the
tab.

**No race needed for the worst of it.** Change the region filter or
toggle show-encrypted with a PSK channel open, and:

- the **My Channels** section disappears — `renderChannelList()` derives
it as `channels.filter(c => c.userAdded === true)`, and the server
returns neither `userAdded` nor a `user:`-prefixed hash
- every unread badge resets to 0
- the user's own labels vanish
- `reconcileSelectionAfterChannelRefresh()` cannot find the `user:*`
hash in the snapshot, so it nulls `selectedHash`, sets `messages = []`,
rewrites the URL to `#/channels` and replaces the open conversation with
"Choose a channel from the sidebar"

## What this does

`mergeClientChannelState(fresh, prev)` mirrors
`mergeWsAppendedIntoRest()`, which already does exactly this job for
`messages` (#1498). Same shape: pure, takes both arrays as parameters,
returns a fresh array, never aliases or mutates an input.

It carries `unread`, `userAdded` and `userLabel` by hash, and keeps
`lastActivityMs` / `lastSender` / `lastMessage` / `messageCount` when a
WebSocket batch landed while the request was in flight and is therefore
newer than the snapshot. Those four move together: a sender without its
message reads as a different message.

`mergeUserChannels()` now runs inside `loadChannels()`, before the
render and before the reconcile, so all three call sites get it instead
of `init()` alone.

## What it deliberately does not do

**It does not resurrect a channel the snapshot left out.** A
region-filter change legitimately narrows the list, so carrying
survivors over would defeat the filter, which is a worse bug than the
one being fixed. The helper only enriches rows already present in the
fresh snapshot.

That leaves half of finding 2 in the issue unfixed: a channel pushed by
the WebSocket handler during the initial in-flight window is still
dropped. That one self-heals on the channel's next packet, and the
reverted preview line is overwritten by the next WS batch. Fixing it
properly needs a way to tell "dropped because the snapshot is stale"
from "dropped because the filter excludes it", which is a larger change
than this.

## Tests

`tests/unit/test-issue-2095-channels-client-state.js`, 11 cases,
registered in `test-all.sh`.

Part 1 exercises the helper directly. Part 2 drives the real
`loadChannels()` through the existing `_channelsLoadChannelsForTest`
hook with a stubbed `api()`, which is what proves the helper is wired in
rather than merely defined.

**Verified by mutation.** With the wiring removed from `loadChannels()`
but the helper left in place, the three reproduction cases fail with
exactly the reported symptoms:

```
FAIL  a refresh keeps the My Channels rows
      My Channels lost on refresh (got ["public1"])
FAIL  a refresh keeps unread badges
      the unread badge reset to 0 on refresh   undefined !== 7
FAIL  a refresh does not close an open PSK conversation
      the open PSK channel was deselected      + null  - 'user:MyPSK'
```

The fourth case, "a refresh still drops a channel the server filtered
out", stays green throughout, so the fix cannot be defeating the region
filter.

Full frontend suite green: `sh test-all.sh` exits 0.

## Not done

**No browser validation.** This is frontend JS covered by unit tests
that call the production function through its own hook, but I did not
run Playwright against a server, so I am not claiming a browser check I
did not do.

`init()` still calls `mergeUserChannels()` and `renderChannelList()`
after `loadChannels()` resolves, which is now redundant and costs one
extra full sidebar render per page load. Left alone deliberately: it is
idempotent, and the comment there records the regression it was added to
fix (`test-channel-issue-1111-e2e.js`, case 2). Worth removing
separately with that e2e test watched.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 12:55:39 +02:00
n30nex 3b415fb007 fix(live): enable node filtering after its handlers are ready (#2103)
Red commit: `8974616` ([CI assertion
failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37064461392)).

Fixes #2094.

The Live node-filter input now starts natively disabled and becomes
editable after its input, keyboard and blur handlers attach. Previously
it accepted text while initial node loading was pending, then never
processed that text without another keystroke.

The change adds one HTML attribute and one initialization assignment.
Failed initial loading still reaches the existing error handler and
installs the controls. Saved filters, URL precedence, autocomplete and
keyboard selection keep their existing behavior. No new settings,
styling, dependencies, requests or packet-processing work.

## Validation

- The existing CI-wired Live browser suite delays or aborts the real
initial node request. It verifies rejected premature typing, later
editability, actual suggestions, keyboard selection and URL state. Both
new cases failed on the red commit; all six cases pass locally on
`11fab46`.
- All 183 standalone frontend suites passed, plus lint and XSS gates.
Separate browser checks passed saved-filter restoration and URL
override; desktop/mobile loading/ready screenshots were inspected.
- A broader local core run exposed the unchanged stale
observation-handle test tracked in #2104. It is outside this Live fix.
Full
[CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37067337557)
passed at `11fab46`, including Go, browsers and both container
architectures; three independent reviews are clear.

E2E assertion added: `tests/e2e/test-1110-live-filter.js:69`.

Preflight overrides: external script/OpenClaw profile unavailable;
repository checks and local Chromium used.
2026-10-03 12:18:33 +02:00
n30nex b9fb3d244c feat(analytics): split adverts using recorded route evidence (#2088)
Red commits: `dc7dcbe`, `76e3217` ([assertion
failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37059977385)).

Fixes #2041; supplies #2085's backend contract.

Records at most two route bits per retained advert. Node flood counts
now use the same evidence, including mixed and transport-flood adverts.
Analytics failures log/count errors while core observation, path and
liveness writes continue. The unrelated recent-advert limit clamp is
removed.

Direct adverts (empty path) describes the observed remaining path across
valid hash widths. Firmware can forward an advert down to an empty path;
this cannot establish original send mode or RF distance. The `zero_hop`
API spelling remains compatible.

Successful post-upgrade evidence writes survive observation replacement
and restart. Older overwritten frames and failed writes can leave gaps.
The UI explains automatic background backfill.

## Measured validation

Synthetic Windows workloads; timings are not production guarantees:

- 300K retained adverts, one-hour window: median 64.35→39.84 ms; 13.98
MB→145 KB allocated. Seven-day allocations unchanged; five paired
samples.
- 1M transmissions/11M observations: backfill 5m43s. Simple 10 Hz
writer: all 3,423 observations persisted; p99 28 ms, maximum 323 ms.
- Simulated pre-backfill cursor: retained evidence ready 4m21s; full
catch-up 6m41s, 122 cache invalidations. Normal completed-migration
startup bulk-loads current evidence.

Portable opt-in scale tests and window benchmarks are included. All 183
frontend suites, full server tests, targeted races, lint and real-API
desktop/mobile checks passed. Full [Linux
CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37063210229)
passed at `36def9c`; three independent reviews are clear. Local Windows
testing hit the unchanged symlink-privilege test. No configuration
changes.

Preflight overrides: external script/OpenClaw profile unavailable;
repository checks and Chromium used.
2026-10-03 12:16:55 +02:00
Sylvain Rabot a4a447acea fix(packets): pack short columns, one expand arrow, Full Names toggle (#2090)
The packets table wastes most of its width on short columns, draws two
arrows on grouped rows, and cuts every observer and hop name short. This
fixes all three and adds a way to read the full path when it does not
fit.

## Why the columns were wide

`makeColumnsResizable` measured each column once on load and stored the
widths as **percentages** of the table. Percentages grow with the
screen: on a 2000px display "17s ago" sat in a 180px column. The
measurement also read the virtual-scroll spacer row — a single `colspan`
cell as wide as the table — as column 0, so the expand column came out
160px wide locally and about 950px on a live instance with grouped rows.

Measured on the fixture at 2000px:

| Column | Before | After |
|---|---|---|
| Expand | 162–230px | 32px |
| Time | 111px | 67px |
| Size | 64px | 46px |
| HB | 44px | 29px |
| Scope | 80px | 48px |
| Path / Details | ~180px each | ~650px each |

## The change

**Column packing.** A new `fitColumnsToContent` (`app.js`) sizes every
column except Path and Details to its widest cell, in px; those two
split what is left. Every visible column gets an explicit width and the
table gets their sum: a spacer row keeps a slot for each hidden column,
and leaving any column `auto` hands those slots a share of the slack as
a blank strip at the right edge.

When space is short (detail panel open, narrow window) the widest fixed
columns give way first, then Path and Details shrink to 60px, and only
then does the wrapper scroll. The split itself is a pure function,
`distributeColumnWidths`, with unit tests.

It re-lays out when the container resizes (detail panel), when
TableResponsive hides or reveals columns (a new `table-columns-changed`
event), and after sorting or a column toggle. After each render `grow()`
widens columns for the rows just inserted and never narrows them, so
nothing jitters while scrolling. Drag handles store px; double-clicking
one resets it. Path and Details have no handle.

`makeColumnsResizable`, still used by the nodes, observers and analytics
tables, now skips `colspan` rows when measuring; other rows are measured
as before.

**One expand arrow.** A CSS `::before` "▶" sat next to the SVG caret,
which pointed up when collapsed. The glyph is gone and the caret points
right, then down.

**Full Names.** A toolbar toggle next to Hex Paths shows observer and
hop names untruncated. It is saved per browser; a link can carry
`fullNames=1`, which applies to that page without overwriting the
visitor's saved preference. Paths stay on one line: the virtual scroller
positions rows at one fixed height, so wrapping would make scrolling
jump.

**The "+N" path pill, and a full-path popover.** Hops past the column
edge were meant to go behind a "+N" pill (#1124), but the pill was
appended after the overflowing chips and so sat past the clipped edge:
it never showed. It is now sticky at the right edge, opaque on hover,
and counts hop chips only (it used to count warning icons too). Hovering
or focusing it shows the full path, one hop per line and numbered;
click, Enter or Space pins it until an outside click or Escape, and it
closes when the page changes or its row scrolls away. Clicking the pill
no longer selects the row underneath.

**Two related fixes.**
- Rows drawn before `/api/observers` resolves show raw 64-char pubkeys.
`loadObservers` now re-renders and re-measures, instead of leaving
Observer at the pubkey's width (~600px under Full Names).
- Group child rows indented every cell by 20px, which misaligned them
with the headers and widened every column on expand. Only the Time cell
is indented now.

Anyone who had dragged packet columns gets the new defaults once: the
old percentage key is dropped.

## Performance

`grow()` runs after every `renderVisibleRows`, so it is on the hot path.
Measured in Chromium on the fixture, 300 iterations, three runs:

| Case | Cost |
|---|---|
| Full render, 63 rows measured | +0.88–0.93ms over the forced reflow
alone |
| Scroll step (incremental render) | ~73 `Range` rect reads, only the
inserted rows |
| Large jump / re-render | ~569 rect reads |

Before narrowing `grow()` to inserted rows, one scroll step with 92
rendered rows made 5,472 rect reads. The DOM row count is bounded by the
virtual-scroll window, so the cost does not grow with the packet count.
`distributeColumnWidths` is a binary search over at most ~10 columns.

## Tests

- `tests/e2e/test-packets-compact-columns-e2e.js`, 16 cases: short
columns stay under fixed px bounds at 1920px; Path and Details take the
freed width; hiding both still fills the wrapper; no blank strip and no
horizontal scroll, with and without the detail panel, at 1920, 1100,
1000 and 375px; responsive reveal and re-hide refit; stale percentage
widths dropped; drag follows the pointer and double-click resets;
expanding a group leaves widths alone; Full Names untruncates, keeps
paths on one line, persists and round-trips through the URL; the
Observer column shrinks back when observer names arrive late; the pill
popover lists every hop vertically, closes when the pointer leaves, pins
on click without selecting the row, closes on Escape and on navigation,
and opens on keyboard focus. Registered in `scripts/non-unit-tests.json`
and `deploy.yml`.
- On `master` the suite fails where it should, and restoring each fix's
old behaviour (child-row indent, the late-observer refit) turns its case
red.
- `tests/unit/test-packets.js`: `distributeColumnWidths` (room to spare,
rounding, widest-first capping, floors, dragged columns, the 60px
minimum, overflow), the single chevron and no `::before` glyph, Full
Names on flat and grouped observer cells, and the Full Names CSS.
- `tests/unit/test-xss-escape-sinks.js` now builds the real
`obsCellName` from source: escaping moved into it. Removing its
`escapeHtml` turns both observer-cell tests red.
- `test-issue-1189-composed-cell.js` and
`test-issue-2012-clear-filters-selection.js` evaluate packets.js
fragments with injected helpers and needed `obsCellName` /
`showFullNames` added.
- Existing packets E2E suites pass on this branch and on master against
a CI-prepared fixture: scope column, #1128 (layout and multi-viewport),
#1122, #1188, #1648, table sort, #1692, #1486.
- Unit: 182 of 183 suites pass locally;
`test-issue-1956-release-routing.js` could not create a temp directory
in my sandbox and is unrelated.

## Browser validation

Checked in Chromium against the fixture at 2000, 1920, 1280 (panel open
and closed), 1100 and 375px, and with the seeded grouped row expanded
and collapsed. Screenshots taken at each step; the table fills its
container exactly, grouped rows show one arrow, and the pill popover
lists full names.

## Not done, deliberately

- **No width cap on Observer under Full Names.** A cap brings back the
truncation the toggle removes, and Observer is already the first column
to give way when space is short.
- **Customizer.** The floors (60px flex minimum, 64px shrink floor, 32px
expand, 58px Rpt) are hardcoded, per AGENTS rule 8; exposing them in the
customizer can follow.
- **Wrapping paths.** Rejected because of the fixed row height above;
the pill popover is the way to read a long path.
2026-10-03 12:15:52 +02:00
n30nex 093e320c2b fix(map): keep Path Inspector candidates accessible and cover route replacement (#2082)
## Fix

Fixes #2060. Fixes #2081.

Path Inspector route drawing and candidate replacement now have
deterministic browser coverage. An idempotent SQL seed adds four
synthetic repeaters and four edges after fixture migration, without
changing packet ordering. The test uses the real API, normal clicks,
visible Leaflet geometry, and removal of every previous route object.

Restoring coverage exposed a desktop layout bug: drawing the first route
moved the map over the next candidate button. The route sidebar now
participates in the existing flex layout, including resizing and
collapse, while preserving the mobile bottom sheet.

The historical zero-candidate result did not reproduce on the current
baseline. No search thresholds or API behavior changed. No new
dependencies or customizer settings.

## Validation

- Local assertion-red history: `042db2b` requires seeded candidates;
`ca1d22a` exposes the blocked second click. Subsequent commits repair
layout and narrow-window behavior.
- All 183 standalone suites passed. Unchanged-base Go failures: #2083
readiness race (fixed in #2084) and Windows symlink privilege.
- Real Chromium: 11 Path Inspector/layout checks, 33 related map checks,
and core E2E (131 passed, 3 existing fixture skips).
- CSS-variable checks, XSS diff check, inventory, whitespace and PII
checks passed. Seed remains valid after periodic graph refresh.

Local browser validation used Chromium with a 60-second navigation
budget; the OpenClaw profile and external preflight script were
unavailable. Repository checks ran directly. [CI run
36343343569](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36343343569)
records the initial assertion-red commit. Final CI must pass before
merge.
2026-09-30 11:40:49 +02:00
n30nex ded27a8312 fix(packets): stabilize fixture tests and preserve mobile empty states (#2080)
Fixes #2054 and #2079.

Packet-row tests could outlive the fixture's default 15-minute window.
An audit of 115 wired suites identified four needing explicit fixture
windows: 180 minutes on mobile, 1440 on desktop. Mobile rejects larger
saved windows, so a desktop-only preference was insufficient.
Packet-icon checks now require a real path-bearing packet and both
detail controls.

Default-window coverage exposed #2079: responsive column hiding also hid
the spanning empty-message cell. Skip structural single-cell spanning
rows when hiding individual columns, preserving status messages and
virtual-scroll spacer height. Ordinary and partial-span data cells
retain their behavior; application defaults stay unchanged.

- Red evidence: `bae5e78` fails the aged-fixture row assertion;
`9b75cf6` fails spanning-row visibility. Guards/setup make them pass.
- Validation: 183 standalone suites; 82 targeted browser cases on a
fixture aged over 30 minutes (one pre-existing node-gesture skip).
Broader browser run: 131 passed, 3 fixture-dependent skips.
- Browser verified: fixture-backed Go server and Chromium;
`coverage/2079-mobile-empty.png`. Real mobile scrolling retains spacer
height and has no horizontal overflow.
- E2E assertions: `test-observer-iata-1188-e2e.js` default-window cases
and `test-table-fluid-e2e.js:58` shared-helper regression, both under
`tests/e2e/`.
- Three independent reviews completed; packet-icon coverage feedback
addressed in `760057d`.
- Populated mobile axe coverage remains part of #1995.
- One constant-time guard per already-visited row; no new requests or
collections.

## Preflight overrides

- External `run-all.sh` unavailable; repository checks run directly.
Local runs use UTF-8, the release harness's `GITHUB_REF_NAME`, and a 60s
navigation budget; repository timeouts are unchanged.
2026-09-30 11:40:36 +02:00
n30nex 410c82c02c fix(ui): use consistent 48px navigation and channel buttons (#2078)
Fixes #2052.

Navigation (`.nav-btn`) and channel icon (`.ch-icon-btn`) buttons now
use the shared 48px house minimum. Remove competing 44px
component/coarse-pointer rules and the legacy 32px mobile rule, and
correct comments that confused the house preference with accessibility
requirements.

At 768px, user-added channel rows already clipped Remove; the larger
Share target exposed the same constraint. Allow only those rows'
controls to wrap so both actions remain within the sidebar. Navigation
height and normal network-channel rows are unchanged.

- Red `7787d9b`: rendered target assertions fail at 375/390/768/1280px.
Green `e6ed9c8` changes CSS only.
- Validation: all 183 standalone suites; 33 computed touch-target
checks; 15 actual-page layout/click checks. Broader browser suite: 131
passed, 3 fixture-dependent skips. Three independent reviews found no
required changes.
- Browser verified: local Chromium and fixture-backed Go server;
screenshots `coverage/2052-targets-375.png` and
`coverage/2052-targets-768.png`.
- E2E assertion added: `tests/e2e/test-channel-fluid-e2e.js:112`,
covering dimensions, containment, overflow and normal Filter/Share
clicks. Existing CI selection is retained; Chromium-required runs cannot
silently skip the touch suite.
- Test-only follow-up `5144503` waits for packet readiness before
measuring navbar controls. Five fresh mobile contexts passed; size and
click assertions remain intact.
- No new settings, requests, dependencies or runtime JavaScript.

## Preflight overrides

- External `run-all.sh` unavailable; repository checks run directly.
- Local unit runner uses UTF-8 and `GITHUB_REF_NAME=local-validation`
for the existing release harness. Browser navigation uses a 60s local
budget for slow assets; repository assertions/timeouts are unchanged.
2026-09-30 11:40:22 +02:00
n30nex f1edbbef3f fix(packets): preserve observation selection in detail URLs (#2093)
Red commit: `c165087` ([CI assertion
failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36644030012)).

Fixes #2091.

Packet detail links now retain the selected observation when page
initialization or filter changes rebuild the URL. Initialization also
restores `obs` from the complete hash after the router strips its query.
Explicit ID links render the requested observation and retain it in Copy
Link state.

The shared updater reads selection from the current route, so returning
to the list or selecting another packet cannot resurrect an old
observation. Existing filter serialization and Clear Filters behavior
remain intact. No new requests, configuration, dependencies or layout
changes.

## Validation

- Real fixture browser checks use nondefault observation 502: hash/ID
load, type/observer/time-window changes, refresh, Clear, and another
refresh. Both URL and selected row are asserted.
- All 183 standalone frontend suites passed; focused filter browser
suite 11/11.
- Broader local core run reached the unrelated Live input readiness bug
#2094, reproduced on unchanged master. The complete CI browser suite
passed.
- ESLint 8: zero errors; 91 existing warnings. XSS diff, syntax and
whitespace checks passed.
- Three independent reviews found no production defects; assertions
additionally prove filter controls changed before checking selection.

E2E assertion added: `tests/e2e/test-filter-ux-e2e.js:179`.

OpenClaw profile/external preflight were unavailable; Chromium and
repository checks ran directly. Local navigation used a 60-second
budget. Final CI passed at `280ced5`, including Go, browsers, coverage
and both container architectures
([run](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36646175352)).
2026-09-30 11:40:00 +02:00
n30nex dc4db17c48 fix(nodes): remove unsupported region fields from Heard By (#2077)
Fixes #2062.

The node health API does not emit observer region data, but Heard By
rendered a Region column containing only dashes and a Regions summary
that could never appear. Remove that column, its sort control, and the
unsupported region displays from the full node page and side panel.

Keep Observer, Packets, Avg SNR and Avg RSSI, along with existing
escaping, signal placeholders, relay counts and badges. No API,
dependency or configuration changes.

- Red commit `4396492` fails on the unwanted Region header and summary;
`a360711` removes the unsupported fields.
- Validation: 183 standalone suites and 7 focused browser checks passed.
Broader browser run: 131 passed, 3 fixture-dependent skips. Independent
reviews completed; the sorting-test finding was addressed in `1452981`.
- Unit coverage evaluates the real templates. The existing CI-selected
browser suite checks four-column alignment and sorting on desktop/mobile
using the actual API field names.
- Browser verified: local Chromium against a fixture-backed Go server;
screenshots recorded in `coverage/issue-2062-heard-by-1400.png` and
`coverage/issue-2062-heard-by-390.png`.
- E2E assertion added:
`tests/e2e/test-issue-1151-orphan-separators-e2e.js:125`, the two #2062
desktop/mobile cases.
- No added requests, loops or data structures.

## Preflight overrides

- External `run-all.sh` unavailable; repository syntax, whitespace,
CSS-variable, XSS and PII checks run directly.
- Local unit execution supplies UTF-8 settings and
`GITHUB_REF_NAME=local-validation`, required by the existing
release-routing test harness.


- Local browser navigation budget raised to 60s for slow local asset
responses; repository assertions and timeout settings unchanged.
2026-09-30 11:39:51 +02:00
n30nex 31744c6da2 test(channels): isolate WS assertions from initial loading (#2087)
Fixes #2086.

The #1468 WebSocket browser checks could fail when initial channel
loading completed between their separate before/after evaluations. An
orphan message added no channel, yet unrelated loading changed the count
from 0 to 5; the positive control could also lose its sentinel when
loading replaced the list.

Each check now captures before state, processes its packet, and captures
after state in one synchronous browser evaluation. All four original
assertions and both message payloads are unchanged. This updates one
test file only, with no production, dependency or configuration changes.

## Validation

- Delayed real API loading reproduces both failures on unchanged master.
- Both corrected checks pass with loading held and with loading
completed; each uses one evaluation.
- Parent independently ran the delayed-loading harness and full browser
suite: 131 passed, three existing fixture skips.
- All 183 standalone frontend suites passed. Syntax, inventory,
whitespace and privacy checks passed.

TDD justification: test synchronization repair only. Existing assertions
demonstrate the baseline failure; no production behavior or manufactured
failing test was added.

Browser checks used Chromium against unchanged production source.
OpenClaw profile/external preflight were unavailable; repository checks
ran directly, with a 60-second local navigation budget. Final CI must
pass before merge.
2026-09-30 11:39:34 +02:00
n30nex 248d2045fd test(server): await indexes before node-path regression requests (#2084)
Fixes #2083.

Node-path regression tests could request `/paths` while background
indexes were still loading, intermittently receiving HTTP 503 instead of
exercising hop resolution or sorting. Seven fixture loads now await the
existing bounded `WaitIndexesReady` signal. The anchor-bias test uses
the same signal instead of polling.

This changes five test files only. Production readiness behavior and
every HTTP/content assertion are preserved. No dependencies,
configuration or customizer changes.

## Validation

- Unchanged baseline: 94 passed / 6 failed across 100 targeted
executions; failures were actual HTTP 503 assertions.
- Fixed setup: 200/200 targeted executions passed.
- Parent independently ran the entire server suite: exit 0, 83.4
seconds.
- All 183 standalone frontend suites and 20 real Chromium route-map
checks passed.
- Formatting, whitespace and PII checks passed.

TDD justification: test-fixture synchronization repair, with no
production logic change. The existing unchanged behavioral assertions
supplied the before/after failure proof; no fabricated failing test was
added.

Local browser checks used Chromium against the unchanged production
source; OpenClaw profile/external preflight were unavailable. The
unrelated ingestor symlink test requires a Windows privilege absent
locally; Linux CI validates that suite. Final CI must pass before merge.
2026-09-30 11:39:25 +02:00
efitenandClaude Opus 5 9eb3098867 fix(release): publish the release notes as the release body (#2076)
Three releases went out with an empty release body. Checked with `gh
release view --json body`:

| release | body | assets |
|---|---|---|
| v3.9.1 | 1206 bytes | 2 |
| v3.9.2 | 2768 bytes | 0 |
| v3.10.1 | **0** | 2 |
| v3.11.0 | **0** | 2 |
| v3.12.0 | **0** | 2 |

v3.9.1 and v3.9.2 were written by hand, which this repository then
stopped allowing because published releases are immutable. Nothing
replaced them, so the `docs/release-notes/` convention was never wired
to the release page and three releases shipped with a description of
nothing.

Reported by the fork operator, who went looking on the v3.12.0 page for
the list of fixed issues that older releases carried.

## The cause

`action-gh-release` in this workflow was given `files` and
`fail_on_unmatched_files` and nothing else. No `body`, no `body_path`,
no `generate_release_notes`.

## The change

A step resolves `docs/release-notes/${GITHUB_REF_NAME}.md` and passes it
as `body_path`, with `generate_release_notes: true` so GitHub's pull
request list lands underneath the hand-written notes. A missing notes
file is a warning rather than a failure: the release then still gets the
generated list, which is more than an empty body.

## Also in this PR

The v3.12.0 notes file gains the two sections it should have had:

- **Issues closed**, 13 of them. Collected from each pull request's
`closingIssuesReferences`, not from commit text, which is why 30 pull
requests yield 13 issues.
- **Contributors**, separating pull request authors (@efiten 22,
@liquidraver 3, @A13xB0 2, @n30nex 1, @dborup 1, @anieto 1) from commit
co-authors (@nullrouten, @anieto, SaarMesh-Bot, Openclaw) from issue
reporters (@efiten 5, @n30nex 4, @anieto 2, @liquidraver 1,
@damn-simple-scripts 1). It says outright that 22 of the 30 pull
requests are the interim maintainer's own, which is the shape of a
release cut while the owner is unreachable, not a healthy ratio.

v3.12.0's published body has already been set to exactly this content by
hand, so the release page and the file agree.

## Worth recording

`gh release edit <tag> --notes-file <f>` updates a published release
body even though `action-gh-release` cannot. The immutability that
permanently burns a tag name does not extend to the description, so a
thin release body is recoverable. That was not obvious from the existing
comments, which describe releases as immutable without qualifying which
parts.

## Verification, and its limit

`yaml.safe_load` parses the file, but that proves little: it silently
accepts duplicate mapping keys that GitHub rejects outright, which is
how a previous workflow edit passed a local check and produced no runs
at all. The real validator is this PR's own pipeline, and the change
cannot be exercised end to end until the next tag is pushed.

## Not done

No backfill for v3.10.1 or v3.11.0. Neither has a notes file, and
reconstructing them is the same retrospective work deliberately skipped
for the 3.11.0 changelog entry.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-26 18:45:50 +02:00
efitenandClaude Opus 5 6d5686dbb4 docs: release notes for v3.12.0, and the 3.11.0 entry that was never written (#2075)
Prepares the v3.12.0 release. No tag is pushed by this PR: the release
procedure requires the tagged commit to carry an `:edge` image whose
revision label matches it, so that check belongs after this merges,
against the merge commit.

## Why minor, not patch

29 commits since v3.11.0: 15 fix, 5 test, 3 perf, **3 feat**, 2 ci, 1
chore. The feats are #2047, #2067 and #2068.

## What the notes lead with

Two changes alter what a running instance does without anyone asking it
to, so they are at the top rather than in a list:

- **#2058**: the first start against a database that has never been
`ANALYZE`d builds planner statistics and stalls ingest while it does.
Measured at 3m43.9s on 9.4 GB, once per database, buffered with nothing
dropped. `db.analysisLimit` set negative skips it.
- **#2035**: `maxMemoryMB` eviction now fires where it previously did
not, because the footprint compared against the limit was undercounted.
An instance that set the limit and never saw eviction will start seeing
it.

## The 3.11.0 gap

`CHANGELOG.md` had no entry for 3.11.0 and `[Unreleased]` was empty, so
52 shipped commits were undocumented. Added as a short entry that says
outright it was written after the fact and has no notes file. The
alternative of reconstructing 52 commits for a superseded version is
error-prone work with little value, and leaving the gap silent is worse
than naming it.

## One fix beyond documentation

`deploy.yml` carried a comment that would mislead the next person
cutting a release. It still said documentation-only commits skip the
workflow "(see the `paths-ignore` above)", which is exactly how v3.10.0
lost its `:edge` image and then its tag name permanently. That filter no
longer exists: the `changes` job forces `code=true` for anything that is
not a pull request (lines 61-73), so every master commit gets an image
and a documentation commit is safe to tag. The history stays in the
comment; the false present tense does not.

## Verified rather than asserted

Both went into the notes as upgrade advice, so both were checked in the
tree:

- `node_declared_regions` is created at boot with `CREATE TABLE IF NOT
EXISTS` (`cmd/ingestor/db.go:435`), added by #2047, which found the
table was read by `region_keys.go` and `config.go` and created by
nothing. So "no manual migration step" is accurate.
- The cgo dependency attributed to #1992 in the 3.11.0 entry is the one
that stops this repo building with `CGO_ENABLED=0` today.

## Not done

- No tag, no release. Next steps, after this merges: confirm the merge
commit's pipeline is green, confirm `:edge`'s revision label is that
commit, then tag `v3.12.0` annotated, push the tag only, dispatch `CI/CD
Pipeline` on the tag ref, and verify the published digests in the
registry rather than in a green job.
- No `docs/release-notes/v3.11.0.md`, deliberately, as above.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
v3.12.0
2026-09-26 12:03:31 +02:00
efitenandClaude Opus 5 6e121ff8e6 perf(db): build planner statistics at startup when there are none (#2058) (#2074)
Follow-up to #2072, which closed #2058 but left one gap named in its own
description: the refresh ticker waits 2 minutes before its first run,
and a query arriving in that window against a database with no
statistics gets the bad plan.

Deployed to staging to measure it rather than reason about it, with
`sqlite_stat1` dropped first so the build path actually ran. That
changed two of the numbers in #2072, both in the expensive direction.

## The gap is once per database, not once per restart

`sqlite_stat1` is an ordinary table, so once `ANALYZE` has written it
the statistics stay in the file. Checked four ways:

- they survive closing the connection that wrote them
- a `mode=ro` handle reads them back, which is how `cmd/server` opens
the database
- reopening the same path through a second `OpenStore` finds them and
skips the rebuild (`TestPlannerStatsSurviveReopen_Issue2058`)
- on staging they survived a full redeploy to a different build that has
no refresh ticker at all, and that build still gets the good plan

So the window opens once, on the first start after this lands, and never
again for that database.

## The cost, corrected

#2072 said 2.0s. Observed on staging, 9.4 GB, commit `4500cfa6`:

```
13:51:26  [analyze] planner statistics refresh scheduled every 24h (analysis_limit=10000)
13:55:10  [analyze] planner statistics built in 3m43.874s (analysis_limit=10000, first run against this database)
```

**3m43.9s.** Every `ANALYZE` duration in #2072's ladder was timed warm,
run after run; cold, on a freshly started container, the same statement
takes nearly four minutes. That is the same warm/cold split #2058 work
already established for the query itself, 56.7s against 0.80s, and I
then repeated it for the `ANALYZE`. Every `2.0s` in the tree is now
marked warm and points at the cold figure.

It holds the single write connection throughout, so ingest stalls and
buffers. Per minute in `observations`:

| minute | rows |
|---|---|
| 13:49 | 220 |
| 13:50 | 106 |
| 13:51 | 0 |
| 13:52 | 0 |
| 13:53 | 0 |
| 13:54 | 0 |
| 13:55 | **1027** |
| 13:56 | 154 |

Nothing was dropped. The burst is about four minutes of traffic at the
surrounding rate, and the only ingest-buffer line in the log is the
startup one reporting `0 dropped`. The cost is a four-minute write
stall, once, not data loss.

## This cost is not introduced here

The ticker merged in #2072 pays the identical 3m43.9s two minutes later
on any database with no statistics. **Live has none, so #2072 as merged
will stall live ingest for about four minutes on its first run, with or
without this branch.** This only moves it earlier, into the startup
burst the ingest buffer is already sized for. Flagging it on #2072 as
well.

## The change

`Store.EnsurePlannerStats(analysisLimit)` checks before it builds:

- database has statistics: one `sqlite_master` query. This is every
restart after the first.
- database has none: one `ANALYZE`, and a warning first.

The warning is the part that earns its place operationally. Four minutes
of stalled ingest with no explanation in the log looks exactly like a
hang, so `EnsurePlannerStats` now says why the write path is about to
pause, what it measured on 9.4 GB, and that it happens once per
database. It stays silent on a restart, because a warning on every boot
would be worse than none.

It runs on the refresh goroutine, not the startup path, so no boot step
waits for it.

`hasPlannerStats` now gates a decision instead of only wording a log
line, so its comment says what the swallowed error costs: a query
failure reads as "no stats", which spends one unnecessary `ANALYZE`
rather than skipping a necessary one.

## Verification on staging

- before: plan drove from `idx_transmissions_payload_type`, no
`sqlite_stat1`
- after: 50 rows in `sqlite_stat1`, plan drives from
`idx_tx_channel_hash`
- dropping the table first flipped the plan back, so the causality holds
in both directions

## Tests

13 in the file. New here: builds when absent, skips when present,
disabled on a negative limit, survives close-and-reopen, warns before
building, stays quiet when statistics exist.

The reopen test is the guard on the whole design: if statistics ever
stopped living in the file, `EnsurePlannerStats` would quietly run a
four-minute `ANALYZE` on every restart and nothing else would notice.

Run locally: 13/13 on the `Issue2058` tests, `go vet` clean, `gofmt`
clean, and the rest of `cmd/ingestor` green apart from
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on
`os.Symlink` with "A required privilege is not held by the client" on
Windows without elevation, in a file this branch does not touch.

## Not done

- No query rewrite, same as #2058 and #2072.
- **No live deploy.** Live still has no statistics, so the four-minute
stall is ahead of it whenever #2072 ships there. Worth picking the
moment.
- Staging has been returned to its own fork build; the statistics it
built remain, so its next start exercises the skip path rather than the
build path.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-26 10:08:43 +02:00
efitenandClaude Opus 5 a5aa3cdc41 perf(db): refresh SQLite planner statistics with a bounded ANALYZE (#2058) (#2072)
Closes #2058.

`ANALYZE` has never run against these databases, so `sqlite_stat1` does
not exist and the planner works from built-in guesses. On the channel
queries it guesses wrong: it drives from the plain
`idx_transmissions_payload_type` instead of `idx_tx_channel_hash`, the
partial index (`WHERE payload_type = 5`) the schema already carries for
that exact filter.

@anieto's report did the diagnosis and the arithmetic. This adds the
maintenance operation that was missing, at a value measured rather than
assumed.

## The diagnosis transfers, the remedy needed measuring

Measured on our 9.4 GB staging database: 1,250,489 transmissions,
14,169,329 observations, 2.7x and 7x the reported database.
Region-filtered `GetChannels` produces the identical plan reported in
#2058, down to both temp b-trees, so the problem is the same one.

Wall time is the wrong metric here. The same query and the same plan
measure **56.7s cold and 0.80s warm** on that file, so the OS page cache
dominates. Counting page-cache misses instead:

| analysis_limit | ANALYZE | driving index | page misses |
|---|---|---|---|
| none (no statistics) | - | `idx_transmissions_payload_type` | 143,442
|
| 400 | 171 ms | `idx_transmissions_payload_type` | 143,449 |
| 1000 | 171 ms | `idx_transmissions_payload_type` | 143,450 |
| **10000** | **2.0 s** | **`idx_tx_channel_hash`** | **107,429** |
| 0 (unbounded) | 242.9 s | `idx_tx_channel_hash` | 107,429 |

400, the value SQLite's documentation offers for the bounded form,
changes nothing on this data: it samples too few rows to separate the
126,336-row partial index from the 920,700-row plain one. 10000 buys the
entire plan change for 2.0 s, and the four-minute unbounded `ANALYZE`
buys nothing beyond it.

## What it is worth, as measured

25% fewer pages read per query, 143,442 to 107,429, about 147 MB less at
a 4 KB page. Warm wall time does not move: 0.80s either way. The gain
lands on the cold path, the one that measured 56.7s, so the claim here
is fewer pages read, not a warm speedup.

This is smaller and differently shaped than the 3-4x in #2058. I cannot
reproduce that ratio on a database of this size and am not claiming it.

## The change

- `Store.RefreshPlannerStats(analysisLimit)` in `cmd/ingestor/db.go`:
`PRAGMA analysis_limit=N` then `ANALYZE`, logging the duration and
whether this was the first run.
- Wired in `cmd/ingestor/main.go` next to the existing WAL checkpoint
ticker: 24h, staggered 2 minutes past startup because it takes the write
lock.
- `db.analysisLimit` in `internal/dbconfig`, default 10000, negative
disables it.

It runs in the ingestor, not the server: `cmd/server/db.go:145` opens
`mode=ro`, and `ANALYZE` writes. This respects the read/write separation
invariant in AGENTS.md.

`analysis_limit=0` means *no* limit to SQLite rather than "use a
default", so an unset config maps to 10000 and a test covers that
specific case.

## Two faults the measurement caught in my own first commit

Both are in the history rather than hidden, because the second commit is
the one that measured:

1. **`PRAGMA optimize` was the wrong statement.** It analyzes only
tables the calling connection has itself queried during the session, and
a maintenance call has queried none. Run against staging it wrote
nothing and left `sqlite_stat1` absent; `PRAGMA optimize(0x03)` returned
no statements at all. Verified on an empty database too (SQLite 3.45.1):
`ANALYZE` creates `sqlite_stat1`, `PRAGMA optimize` does not. That
difference is what makes the behavioural test a guard instead of a
no-op.
2. **`analysis_limit=400` was the wrong value**, per the table above.

## Tests

Six cases in `cmd/ingestor/refresh_planner_stats_test.go`:

- statistics are actually written (the guard against returning to
`PRAGMA optimize`)
- the pragma reaches the connection, read back through `PRAGMA
analysis_limit`
- a negative limit leaves `sqlite_stat1` absent
- two consecutive refreshes, since a ticker calls this repeatedly
- the config default and the JSON round trip
- the default is above the range measured ineffective, so lowering it
back to 400 fails

## Not done, and one caveat

- **Correction to an earlier version of this description**, which said
the Go tests could not run locally because this box has no C compiler.
That was wrong: `CGO_ENABLED=0` and `gcc` being absent from `PATH` is
not the same as no compiler, and a mingw-w64 toolchain is installed
here. Run properly, all six tests pass locally, and so does the rest of
`cmd/ingestor` apart from
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on
`os.Symlink` with "A required privilege is not held by the client" on
Windows without elevation and lives in `stats_file_test.go`, a file this
branch does not touch. CI agrees: Go Build & Test and the ingestor race
detector are both green.
- The SQLite behaviours above were measured against 3.45.1 on the
server, not against the amalgamation `mattn/go-sqlite3` bundles.
- **No query rewrite.** #2058 explicitly left that out and so does this;
the correctness caveats it lists (per-channel most-recent-message
semantics, v2/v3 branches, `enc_` exclusion) are untouched here.
- Staging carries limit-10000 statistics, matching what this code
produces. Reversible: `DROP TABLE sqlite_stat1` was verified on a
scratch database before any of it ran.
- Whether a cold-start `ANALYZE` should also run before the 2 minute
stagger is not addressed. The first query after a restart is the
expensive one, and it can arrive first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:08:54 +02:00
efitenandClaude Opus 5 3caa847323 fix(nodes): the advert section is called "Recent Adverts", and says why (#2071)
Closes #2042, reported by @damn-simple-scripts.

## Which of the two fixes

The report offered widening the section to all packets the node
originated, or renaming it, and proposed the rename. Widening sounds
like the better fix, so I measured before agreeing. On a production
database:

| | |
|---|---|
| transmissions total | 1,251,967 |
| with `from_pubkey` populated | 191,237 |
| of those, `payload_type = 4` (ADVERT) | **191,237** |
| with `from_pubkey` and **not** an advert | **0** |

The ingestor fills that column for adverts only. Attributing a relayed
CHAN or TXT packet back to its sender is the path-resolution problem,
not a filter this section could apply — so "show all packets from this
node" is not a small change, it is a different feature resting on
attribution the data does not carry.

So the rename is correct, and the numbers say so rather than my
preference.

## The change

Both copies renamed — the full node page and the side pane. Both already
read `nodeData.recentAdverts`, so the field feeding them said "adverts"
while the heading said "packets".

The heading also gained a `title` naming `from_pubkey` as the reason it
is adverts only. Renaming without explaining invites the same report
from the next reader; the tooltip is where that explanation costs
nothing.

## Tests

`tests/unit/test-issue-2042-recent-adverts-label.js`, four cases: both
copies present and headed "Recent Adverts", no copy back to "Recent
Packets", the advert field still feeding them, and the explanation still
in place. The third matters most — it ties the label to its data source,
so pointing this section at a different field in future fails here
rather than silently making the label wrong again.

## Not done, deliberately

The reporter's follow-up idea, splitting flood adverts from zero-hop
adverts into two lists. They wrote that it can be dropped if an issue
should tackle one thing, so it is not here. It is feasible now: a
zero-hop advert is `ROUTE_TYPE_DIRECT` with an empty path, established
while fixing #2064. That deserves its own issue rather than a paragraph
in this one — say the word and I will open it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:08:28 +02:00
efitenandClaude Opus 5 2cfe9cbbc5 fix(a11y): raise the Scope Audit observed-chip contrast above 4.5:1 (#2070)
Closes #1996.

## Reproduced first

The chips tint their own background — `color-mix(in srgb,
var(--status-green) 16%, transparent)` — which darkens whatever surface
sits behind them. With `--status-green-text` (green-700, `#15803d`) on
top:

| surface | composited chip | ratio |
|---|---|---|
| `--surface-0` `#f4f5f7` | `#d2eddf` | **4.04:1** |
| `--card-bg` `#ffffff` | `#dcf6e5` | **4.38:1** |

Both under the 4.5:1 the #1719 gate requires for normal text. The first
row is the same ratio **and the same hex** @n30nex measured in the
browser, which is how I know the model in the test agrees with what a
visitor actually sees.

## Fixed on the text, not by thinning the tint

Dropping the tint to 12% reaches 4.52:1. That is two hundredths above
the line, and a margin that thin fails again the next time a surface
value moves — the same trap #2039 hit with a threshold sitting inside
the healthy band. One palette step darker gives **6.23:1** on white and
**5.74:1** on `--surface-0`, and keeps the tint that makes a chip read
as a chip rather than as plain text.

- `--palette-green-800: #166534` added. The greens ran 300–700 while the
blues already reach 900, so this fills the scale rather than inventing a
colour.
- `--sa-chip-observed-fg` defined per theme: green-800 in light, and the
bright `#22c55e` **kept** in dark, where the chip already passed at
6.48:1 / 5.73:1. Dark is deliberately untouched.
- The chip reads the variable, so the customizer still governs it and no
literal enters a component.

`.sa-chip-verified` needed no change: `scope-audit.js:106` only ever
adds it alongside `sa-chip-observed` or `sa-chip-unobserved`, and it
contributes an underline. So the fix covers both classes the issue names
— worth stating, since the title mentions both.

## Tests

`tests/unit/test-a11y-1996-scope-audit-chips.js`, 7 cases: both themes ×
both surfaces, that the chip still tints (so the suite cannot pass by
testing nothing), that its colour comes from a variable rather than a
literal, and a guard asserting green-700 **would** still fail — so a
quiet revert to `--status-green-text` turns this red instead of passing.

**Red-run confirmed:** with the old colour restored it fails at exactly
4.04:1 and 4.38:1, naming the composited `rgb(210,237,223)`.

Two deliberate choices in the test:

- **A separate suite, not a case in `test-a11y-1719`.** That suite's
`parseColor` handles hex and `rgb()` only; teaching it `color-mix()` is
a larger change than this fix. These chips are the only
contrast-critical user of the function today. If a second appears, the
two should merge, and the file says so.
- **Derived from the stylesheets, not a rendered page.** The default
Scope Audit fixture renders no chips at all, which is precisely why the
existing browser coverage missed this. A stylesheet-derived check cannot
be defeated by a fixture that shows nothing.

`check-css-vars` passes (180 definitions, 0 undefined) and
`test-test-inventory` passes with the new file classified.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:07:57 +02:00
d14317b008 fix(qa): query blacklist retention by from_pubkey (#2069)
## Summary

- query `transmissions.from_pubkey`, the column present in CoreScope's
real schema, instead of the nonexistent `from_node`
- preserve the existing SQLite parameter binding and stdin transport
- derive the unit-test table from the ingestor's committed `CREATE
TABLE` definition rather than a hand-written schema
- add positive, negative, case-normalization, injection-shaped input and
missing-column coverage
- verify the query against the committed staging-captured E2E fixture

## Why

The parameter-binding change correctly removed SQL interpolation, but
its count query and test fixture both used `from_node`. The production
schema defines `transmissions.from_pubkey`; therefore the live QA probe
could only fail with `no such column: from_node`, while the synthetic
unit fixture remained green.

The ingestor stores attributed ADVERT pubkeys as lowercase hex. The
query now uses:

```sql
SELECT COUNT(*)
FROM transmissions
WHERE from_pubkey = lower(:pubkey);
```

The value remains a bound parameter. No production schema or runtime
code changes.

## Verification

- `bash -n qa/scripts/blacklist-test.sh`
- `bash -n qa/scripts/test-blacklist-sql.sh`
- `bash qa/scripts/test-blacklist-sql.sh` — 72 passed, 0 failed
- schema fixture extracted directly from `cmd/ingestor/db.go`
- committed E2E fixture returns the same attributed-row count through
the helper query and a direct control query
- legacy fixture containing only `from_node` fails non-zero with `no
such column: from_pubkey`
- injection-shaped, empty, whitespace, multibyte and long values remain
literal bound values
- `git diff --check`

This PR intentionally contains only the schema correction and its
regression coverage.

Co-authored-by: Openclaw <openclaw@Openclaws-Mac-mini.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:32:54 +02:00
435ac25dd6 feat(nav): show the running version in the nav-drawer footer (#2068)
Takes over #1985 by @SaarMesh-Bot, as offered there on 2026-09-16 and
2026-09-17. **Their commit is the first here, unchanged and under their
authorship**; the rest clears the two review points. Closing #1985 in
favour of this so the rebase and the fixes travel together, not to
reassign the work.

## The feature, unchanged

The frontend never surfaced which build was running, though
`/api/health` has reported `{version, commit, buildTime}` all along. A
footer on the nav drawer now renders `CoreScope <version>`, linking to
the releases page, with commit and build time in the tooltip. Colours
come from existing CSS variables, the label is set with `textContent`,
and a failed health call leaves a neutral label rather than an empty
footer — all as the author wrote it.

## Review point 1: fetch on open, not on page load

The `/api/health` call sat in `buildDom()`, which runs on page load. The
drawer may never be opened, and **cannot** be opened at ≤768px, where
the module is disabled by design. So every visitor's browser was
requesting the endpoint to fill a footer most of them would never see.

Moved into `open()`, after the width gate. `fetchVersion` still caches
its promise for the page lifetime, so re-opening costs nothing.

## Review point 2: the fallback is pinned

`tests/unit/test-nav-drawer-version-footer.js`, six cases:

- a rejected fetch, a non-ok response, and a 200 without `version` each
leave the neutral `CoreScope` label — not a blank footer and not
`CoreScope undefined`, which is what an instance shows exactly when
someone is trying to read its version
- a normal response renders the version, with commit and build time in
the tooltip
- a version carrying markup lands verbatim in `textContent`, and
`innerHTML` is never touched
- the endpoint is requested once however often the footer is filled,
pinning the cache from the other side

It **slices `fetchVersion` and `fillVersion` out of the shipped
`public/nav-drawer.js`** and evaluates them rather than copying them
into the test, so it exercises what ships — same approach as
`tests/unit/test-direct-rf-heard-by.js`. Both slice markers are
asserted, so a rename fails loudly instead of quietly testing nothing.
Each case re-evaluates the slice, because `versionPromise` caches for
the page lifetime and a shared sandbox would hand the second case the
first case's answer.

Listed in `test-all.sh`, which per #2036 is the only frontend runner.

## A note on the third commit

The first version of the test used `setImmediate` to settle the promise
queue. It ran fine under node and **failed eslint**, which treats these
as browser code. I ran eslint and committed in the same command and
pushed without reading its output. Fixed in the commit after, using
`setTimeout(r, 0)`. Recording it because the PR would otherwise show a
lint failure in its history with no explanation.

## Verification

6 of 6 in the new suite, eslint clean on both changed files,
`test-test-inventory.js` passes with the new file classified.
Cherry-picked cleanly onto current master.

---------

Co-authored-by: SaarMesh-Bot <bot@saarmesh.de>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 10:58:36 +02:00
291393dcc0 feat(ingestor): log a throttled warning when the IATA whitelist drops a region (#2067)
Takes over #2008 by @nullrouten0, as offered there on 2026-09-16 and
2026-09-17. **Their commit is the first of the two here, unchanged and
under their authorship**; the second is only the fix for the one
blocker. Closing #2008 in favour of this so the rebase and the fix
travel together, not to reassign the work.

## The feature, unchanged

`observerIATAWhitelist` dropped non-whitelisted regions silently. An
allow-list fails in the dangerous direction: a legitimate but unlisted
region vanishes with nothing to show for it. One line per dropped region
now, re-logged at most every `iataWarnIntervalSec` (new optional key,
default 6h) for as long as that region keeps arriving.

The periodic re-log rather than a strict log-once is the author's call
and it is the right one: a single edge event rolls out of any scrape
window, leaving an actively-dropping region indistinguishable from a
healthy one.

## The blocker, now fixed

`ShouldWarnIATADrop` keyed its throttle map on a topic segment the
**publisher** controls and never evicted it — a remote memory sink.
Measured on the original branch: 200,000 distinct codes retained 200,000
entries and 15.1 MB of heap.

My review offered two shapes. This takes the cap rather than
shape-validation, and the reason matters: **nothing in this codebase
constrains an IATA code's shape.** It is uppercased and trimmed in
`config.go` and `db.go` and never validated. Rejecting by shape would
invent a rule operators have not agreed to, and would silently drop the
warning for anyone whose code does not fit it — the same failure mode,
one level down.

So `iataWarnMaxTracked = 512`: far above any real deployment (the
reference instance runs 43 observers across a handful of regions) and
small enough that a hostile feed gains nothing.

**Past the cap the drop is still logged**, throttled on one shared
timestamp instead of a per-code one. Swallowing it there would
reintroduce exactly the silent failure this feature exists to fix.

## Tests

The author's `iata_drop_warn_test.go` plus three:

- the map stops growing when fed 2048 distinct codes
- a new code past the cap still warns once, is then throttled, and
speaks again after the interval elapses
- an already-tracked code's throttling is unchanged, so the cap does not
alter the normal path

## Verification

`gofmt` clean, cherry-picked cleanly onto current master
(`cmd/ingestor/main.go` auto-merged). Go tests not run locally: no cgo
toolchain here since #1992, and per AGENTS.md `CGO_ENABLED=0` links a
stub that proves nothing. CI is their first run.

## Not done

The `iataWarnIntervalSec` key is undocumented outside the struct
comment. If there is a config reference that should list it, say where
and I will add it.

---------

Co-authored-by: nullrouten <nullrouten@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:58:27 +02:00
efitenandClaude Opus 5 50c4d9615b test(ingestor): wait for the boot migrations before handing over a test store (#2066)
Closes #2065. Master's `🏁 Race detector (ingestor)` job has been red
since the pushes at 2026-09-22 21:55 and 21:56.

## Correcting my own diagnosis

The issue says the fault is a goroutine outliving its test and racing a
later one, and proposes making it joinable. That was wrong, and it
matters because it changes the fix: `Close()` **already** waits on
`backfillWg` (`cmd/ingestor/db.go`), and `newTestStore` registers it as
`t.Cleanup`. The goroutines are joined before the next test starts.

The race is inside a single test.

`OpenStore` schedules two async migrations — `obs_observer_ts_idx_v1`
and `tx_last_seen_backfill_v1` — whose goroutines log while they run.
`TestHandleMessageDecodeErrorLog_PII_Issue1211` then points the standard
logger at a `bytes.Buffer` and reads it, so its **own** store's
migrations write into the buffer it reads:

```
Write by goroutine 760:  RunAsyncMigration.func1   async_migration.go:124  (log.Printf)
Read  by goroutine 757:  ...PII_Issue1211          decode_error_log_test.go:37 (buf.String)
```

## Why the helper rather than the one test

These tests capture the standard logger in **21 places across 7 files**.
Any of them that also builds a store is exposed to the same thing; the
decode-error test is just the one whose timing lost. So `newTestStore`
now waits after `OpenStore` instead of only at cleanup, and no test body
can run while a migration is in flight.

Checked before touching a shared helper: no test references either boot
migration by name, and the `pending_async` assertions in
`async_migration_test.go` use their own names with a blocking `fn`, so
they are unaffected. Cost is a few milliseconds against an empty temp
database.

## Tests

The race detector only catches this when the scheduler cooperates — it
sat latent from 2026-09-03, when those files were last touched, until it
surfaced three weeks later, and a re-run would have made it look like a
flake. So both new tests are deterministic:

- **`TestNewTestStoreWaitsForBootMigrations`** —
`tx_last_seen_backfill_v1` is scheduled unconditionally by `OpenStore`,
so on a fresh temp database it is pending at that instant and can only
read `done` if something waited. Remove the wait and this fails every
run.
- **`TestCapturedLogIsFreeOfMigrationOutput`** — asserts a captured
buffer holds no `[migration/async]` or `[async-migration]` output, which
is the failing test's own situation stated as an assertion.

## Verification

`gofmt` clean. Go tests were not run locally: no cgo toolchain on this
machine since #1992, and per AGENTS.md `CGO_ENABLED=0` links a stub that
proves nothing. CI is their first run, and the race-detector job is the
one that matters here.

## Not done

The decode-error path still logs through the standard logger, so a
future test capturing it while any other goroutine logs will race again.
An injectable logger would close that class properly. This closes the
store-boot case, which is the one that exists today.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:23:59 +02:00
6d3da77b67 perf(channels): coalesce concurrent GetChannels/GetEncryptedChannels cache misses (#2059)
Fixes #2029.

## What was wrong

`GetChannels`/`GetEncryptedChannels` (`cmd/server/db.go`) cache their
region-scoped result for 60s but had no request coalescing on a cache
miss, so every request that arrived while the cache was cold or expired
ran the region-scoped `GROUP BY` scan itself. Measured on production: 5
concurrent requests for the same never-cached region each took ~8s, no
cheaper than 5 independent runs.

`statsSF`/`regionMembershipSF` already fix the identical bug class
elsewhere in this file (#1910), so this wraps both functions'
query-build/execute/cache-populate block in a `singleflight.Group` the
same way, keyed per region, double-checking the cache inside the flight
in case a previous winner already refreshed it.

## Tests

The review on #2029 pointed out that a timing-based check ("finish
within ~2ms of each other") doesn't actually prove coalescing happened —
it would pass on a fast machine even without singleflight.
`db_channels_singleflight_test.go` uses a call counter instead, same
pattern as `TestEnsureNeighborGraph_Singleflight` (#1203 Pair A):

- `TestGetChannels_SingleflightCoalescesQueries` /
`TestGetEncryptedChannels_SingleflightCoalescesQueries`: 10 concurrent
callers against a cold cache, asserting the real query runs exactly
once. A test-only hook (`channelsQueryHook`/`encChannelsQueryHook`, nil
in production, same contract as `bgLoaderEntryHook`) increments the
counter right where the query executes, since these functions hit
`db.conn.Query` directly rather than going through an injectable builder
function.
- `TestGetChannels_SingleflightPerRegion`: two regions queried
concurrently (5 callers each) assert 2 queries, not 1 — pins that the
flight is keyed per-region and a caller for one region can't receive
another region's coalesced result.

Anti-tautology: reverting `channelsSF.Do`/`encChannelsSF.Do` back to a
bare call makes the coalescing tests observe N instead of 1.

`go build ./...`, `go vet ./...`, `gofmt -l .` clean. Full `cmd/server`
suite (race-enabled for the new concurrency tests) run in a
`golang:1.22-alpine` container, mounted repo, workdir `cmd/server` so
the sibling `internal/*` replace directives resolve:

```
=== RUN   TestGetChannels_SingleflightCoalescesQueries
--- PASS: TestGetChannels_SingleflightCoalescesQueries (0.06s)
=== RUN   TestGetChannels_SingleflightPerRegion
--- PASS: TestGetChannels_SingleflightPerRegion (0.05s)
=== RUN   TestGetEncryptedChannels_SingleflightCoalescesQueries
--- PASS: TestGetEncryptedChannels_SingleflightCoalescesQueries (0.07s)
```
Full suite: `FAIL github.com/corescope/server 98.261s`, but the only
failures are `TestHandleNodePaths_PrefixCollision_1352`,
`TestHandleNodePaths_FallbackUniquePrefix_1352`, and
`TestHandleNodePaths_FallbackUnresolvableHop_1352`, all failing on a
`503 {"error":"index loading","retryAfter":5}` — an index-build race in
this container's timing, not this change. Confirmed by running the same
three against an unmodified, freshly-cloned `master` in the same
container: they fail there too (plus
`TestHandleNodePaths_PrefixCollision_1352_FallbackBranch`, which this
run happened not to hit). Nothing in this diff touches node-path
handling.

## Not done

The deeper query-plan issue flagged in #2029 (the outer scan is driven
by `payload_type`, not region, so a cold solo request still costs
several seconds regardless of concurrency) is filed separately as #2058,
with `EXPLAIN QUERY PLAN` output and row counts against production data.
Coalescing makes one slow query serve everybody; it doesn't make the
query itself fast.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: anieto <anieto@meshtexas.org>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 10:05:59 +02:00
liquidraverandClaude Opus 5 b695b979a2 fix(nodes): keep paginating past a page that post-LIMIT filtering shortened (#2061)
## Problem

`handleNodes` runs the geo-filter, `nodeBlacklist`, `hiddenNamePrefixes`
and area
passes **after** the SQL `LIMIT/OFFSET`, and rewrites `total` to the
filtered
length. A page that loses a row is therefore short **without being the
last
page**, and neither the page length nor `total` can tell a client
whether to ask
for another page.

#1606 added the pagination loop and chose the page length as the
canonical stop.
That is correct only where nothing is ever filtered. Everywhere else the
list
truncates at the first filtered page boundary and strands every node
behind it —
the #1598 symptom reached by a different route: a node that is relaying
right
now simply stops being in the list.

The comment at `app.js:240` rejects `total` for exactly the right
reason, then
picks the signal the same code path also breaks.

## Measured on a live 2346-node deployment

Page sizes for the query the map issues:

```
offset=0     returned=500     ← full, loop continues
offset=500   returned=499     ← one row filtered AFTER the LIMIT → loop STOPS
offset=1000  returned=500     ← never requested
offset=1500  returned=500     ← never requested
offset=2000  returned=345     ← never requested
```

| stop rule | requests | nodes reached |
|---|---:|---:|
| short page (master) | 2 | **999** |
| `has_more`, else empty page | 6 | **2344** |

**1341 nodes, 57%, unreachable through the UI.**

### One hidden node truncates the whole list

The deployment this came from has no `geoFilter`
(`/api/config/geo-filter`
returns `polygon: null`) and no `nodeBlacklist`. It has a single
`hiddenNamePrefixes` entry — a deliberate operator choice — and exactly
one node
whose name starts with it:

```
public_key   d4a46ea2…1054      (64 clean hex chars)
name         🚫🔥☀️
role         repeater
last_seen    2026-09-22T10:09:38Z
```

`handleNodes` drops that row in the `IsNameHidden` pass, which runs
after the SQL
`LIMIT`. The row is counted by the `LIMIT` and by `COUNT(*)`, so the
page it lands
in comes back exactly one short — and stops every client that treats a
short page
as the end.

Isolated against SQL on the same database, seconds apart:

```
SELECT lower(public_key) FROM nodes ORDER BY last_seen DESC LIMIT 500 OFFSET 500
  -> 500 rows
GET /api/nodes?limit=500&offset=500
  -> 499 rows
comm -23 sql.txt api.txt
  -> d4a46ea2e99cab132a3286ef3d9cce9099318790af7f25671fe83de453721054
```

Deterministic — `offset=500` returned 499 on three consecutive requests.
Not a
CDN artifact either: `cf-cache-status: DYNAMIC`, origin `cache-control:
no-store`, no `age` header, and four requests with deliberately unique
cache keys
all returned 499.

So **one deliberately hidden node makes 1341 of 2344 nodes
unreachable.** The
hiding feature does exactly what it was asked to do for that one node,
and takes
57% of the network with it, silently. A single `hiddenNamePrefixes`
entry is
enough; no geo-filter, blacklist or area filter is needed to reach this
state.

### The cutoff moves, which is why this reads as intermittent

The visible set is the sum of the pages up to and including the first
short one,
so the boundary sits wherever the unreturnable row currently sorts by
`last_seen`, and jumps a whole page as ingest reorders the list. Same
deployment, same code, same config, ~2h apart:

| dropped row's rank | first short page | nodes visible |
|---|---|---:|
| inside 0–499 | page 1 | 499 |
| inside 500–999 | page 2 | 999 |

A node is visible or invisible purely by where it lands relative to that
moving
line, so affected nodes appear to vanish and return on their own. Two
operators
on this deployment reported exactly that, independently, while I was
measuring.

### A named reproduction

`HU-ZA-Lentihegy` (`5287a33f…`), reported missing from the map by an
operator
whose companion had logged its advert at 04:20 local the same morning.

Ingest was fine. The row is in `nodes` with `last_seen`
`2026-09-22T02:20:35Z` — the same advert, to the second — valid GPS,
role
`repeater`, 1033 adverts, and `/api/nodes/search?q=lentihegy` returns
it.

```
rank by last_seen : 1081
cutoff at the time:  999
```

It missed by 82 positions. Walking the same live endpoint, same moment:

| stop rule | requests | nodes reached | Lentihegy |
|---|---:|---:|---|
| short page (master) | 2 | 999 | **not reached** |
| `has_more`, else empty page | 6 | 2340 | reached |

The practical shape of this on a busy mesh: 1081 nodes had been heard
more
recently than 9.4 hours, so on that deployment **anything last heard
more than
~9 hours ago was invisible**, alive or not.

`#/nodes` compounds it — its search box filters client-side over the
truncated
set, so the server-side `?search=` never runs and an operator cannot
find the
node by searching for it either, even though the endpoint would return
it.

## Change

**Server** — `NodeListResponse` gains `has_more`, computed from the raw
SQL page
against the real `COUNT(*)` before the filter passes run, so it survives
them:

```go
hasMore := offset+len(nodes) < total
```

Always emitted (no `omitempty`) so a client can tell `false` from an old
server.
No extra request in the fixed path: `has_more` ends the loop exactly,
where the
old rule needed a probe page.

**Clients** — `app.js` `fetchAllNodes`, `nodes.js` `loadNodes` and
`area-map.html`'s inline helper stop on `has_more`, falling back to a
zero-length
page against a server that predates it. An empty page always ends the
loop, so a
`has_more` against a concurrently-shrinking table cannot spin to
`safetyCap`.

Left alone: the three loops are still three copies. Collapsing them onto
`fetchAllNodes` is a bigger change than this fix needs, and `nodes.js`
has its
own inter-page progress UI. Happy to do it separately if you want it.

## Testing

- **Unit** (`tests/unit/test-fetch-all-nodes-pagination.js`): the
fixture now
models the real handler — a row counted by the LIMIT and by `COUNT(*)`,
then
removed from the page. Three new cases. Fails on the old rule at 499 of
1199.
- **E2E** (`tests/e2e/test-map-nodes-pagination-e2e.js`, already wired
into
  `deploy.yml`): the mock drops a page-1 row and emits `has_more`.
  Mutation-checked — restoring master's stop rule fails 3 of its steps.
- **Go** (`cmd/server/nodes_pagination_has_more_test.go`): asserts
`has_more`
stays true on a page filtering shortened. Mutation-checked — recomputing
it
  after the filter block fails the test.
- Full server suite `go test -race`: ok, 41.4s. `gofmt` clean, `go vet`
passes.
- **Against a real binary**, not just mocks: fixture DB migrated with
`corescope-migrate`, `hiddenNamePrefixes: ["SKCE"]`, `limit=3`. Page 1
returns
2 of 3 with `total` rewritten to 2 and `has_more=true`. Walking the real
server
with master's rule reaches 2 nodes; with `has_more`, all 199 visible of
200,
the hidden one still hidden. The real frontend against that server loads
199
  with no JS errors.

Two existing expectations changed, both deliberate:

1. `surfaces ALL nodes past the 500 server cap` — 3 → 4 requests. That
mock emits
no `has_more`, so the 200-row final page can no longer end the loop (a
short
page is exactly what a filtered page looks like) and a zero-length probe
   follows. Against a current server `has_more` still ends it at 3.
2. `rows missing public_key are NOT collapsed into one` — its stub
returned a
constant body, which would now be paged to `safetyCap`. It serves one
page
   then empties.

Local `test-all.sh` exits 1 on two XSS-gate self-tests
(`good-2-tested.js`, `good-4-tested.js`) via a `UnicodeEncodeError`
printing an
emoji under Windows cp1252. Identical on clean `origin/master` in a
scratch
worktree, so it is pre-existing and platform-local, not this branch.

There is a second identical filter block further down `routes.go` on
another list
endpoint. Likely the same class; not touched here.

If you would rather land your own version of this, say so and I will
close mine.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 09:45:18 +02:00
efitenandClaude Opus 5 980c5c4515 fix(nodes): stop the Heard By empty state claiming the node is out of range (#2063)
Follow-up to #2057, which is correct in what it does and overclaims in
one sentence.

## The sentence

When no observer heard the node directly, the card said:

> No observer is within radio range of this node.

#2057's own rule cannot establish that:

- every **direct route** is discarded, because the firmware removes the
sender from the path before retransmitting (`Mesh.cpp`,
`removeSelfFromPath`), so the packet cannot say who transmitted it.
#2057 measured these at **38% of transmissions over 7 days**.
- **28.8% of flood observations with a path** are dropped because the
last hop resolves to more than one candidate, and the rule
under-attributes rather than guesses.

So an empty direct list is missing evidence, not evidence of missing
coverage.

## Why it matters in practice

Sampled 40 repeaters on a production instance after #2057 shipped:

| | |
|---|---|
| at least one direct observer | 24 |
| empty card, "no observer is within radio range" | **16** |
| of those 16, with a non-zero `relayObserverCount` | **16** |

Every node showing "nobody is in radio range" also showed "Seen via
relay by N observers" two lines below. An operator reading that about a
working repeater concludes they have a coverage problem they do not
have.

## The change

Wording only, in both copies of the card (full page and side pane):

> No observation proves a direct reception here, which is not the same
as being out of range.

with the reason in a `title`, so the card stays one line:

> Only flood-routed transmissions identify who was heard: a direct route
removes the sender from the path before retransmitting (firmware
`Mesh.cpp`, `removeSelfFromPath`), and an ambiguous relay hop is left
unattributed rather than guessed. So an empty list is missing evidence,
not proof of missing coverage.

No API change. #2057's rule, shape and performance work are untouched —
I verified its firmware derivation against the clone at `0679dbef`
before writing this: `Packet.h:83`, the forwarder appending with
`packet->getPathHashSize()` at `Mesh.cpp:349`, and `removeSelfFromPath`
on the direct path at `Mesh.cpp:89-105` all read as described.

## Test

`tests/unit/test-direct-rf-heard-by.js` slices this template out of
`public/nodes.js`, so it pins the shipped markup. It now asserts the old
sentence is gone and the qualifier is present.

Worth recording how that assertion was reached: my first version banned
the phrase "out of range" from the card, and it failed — on the new
line, which contains that phrase precisely in order to deny it. A word
ban was the wrong instrument. Matching the old sentence and requiring
the new qualifier is the assertion that actually distinguishes the two
states.

8 of 8 in that suite, eslint clean.

## Not in this PR

The **Regions** line and **Region** column on the same card read
`o.iata`, which `HealthObserverRow` does not emit, so both have always
been dead. #2057 named this and left it; it is now **#2062** with the
file and line references, rather than a remark inside a merged
description.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 23:56:02 +02:00
efitenandClaude Opus 5 61f565c606 fix(node-health): credit zero-hop adverts as direct reception (#2064)
Reported by @dborup on #2057. The mechanism is real; the suspected scale
is not. Both parts measured below.

## The defect

`directHeardNode` rejects every non-flood route type before it looks at
the path, so the `hop == ""` branch that credits an advert's originator
can never run for a direct route. Zero-hop adverts — the clearest
direct-RF evidence the network produces — are discarded and the node is
listed under "Seen via relay" instead.

## Firmware

Read at `0679dbef` rather than taken on trust:

- `Mesh::sendZeroHop` sets `ROUTE_TYPE_DIRECT` and `path_len = 0`,
commented there as "path_len of zero means Zero Hop". The transport
overload does the same with `ROUTE_TYPE_TRANSPORT_DIRECT`.
- `examples/simple_repeater/MyMesh.cpp` sends the periodic **local
advert** through it (the `next_local_advert` branch), as does
`sendSelfAdvertisement` when `flood` is false.
- `examples/companion_radio/MyMesh.cpp` does the same for a companion's
own advert.

An ADVERT arriving on a direct route with an empty path therefore cannot
have been forwarded: the observer received the advertiser's own
transmission, and the advert carries its pubkey in the clear.

Every other direct case keeps #2057's rule. A non-empty path on a direct
route is the **remaining** route, because the forwarder ran
`removeSelfFromPath` before retransmitting, and `advertOriginPubkey`
already returns `""` for any payload type other than ADVERT — so the
payload guard costs nothing.

## Measured, 7-day window on a production instance

| | |
|---|---|
| zero-hop advert observations currently dropped | 8,166 |
| distinct nodes they evidence | 139 |
| node-observer pairs they evidence | 185 |
| pairs **not** already credited via an empty-path flood advert | **52**
|
| pairs currently credited from flood adverts | 217 |

So the fix restores 52 node-observer pairs of direct evidence that are
invisible today, roughly a quarter more advert-based direct evidence.

## What it does not explain

The report suspected this accounts for #2057's low headline numbers
("NL-BXE-RP01 | 433 → 0", 234 of 1,860 nodes with any direct observer).
The measurement does not support that:

- Of 30 sampled nodes with zero-hop advert evidence, **29 already show
at least one direct observer**, because they also send flood adverts
which #2057 credits.
- **NL-BXE-RP01 | 433 has zero zero-hop adverts** in the window. Its
empty list is not caused by this rule.

So this mostly enriches lists that are already non-empty, and flips few
cards from empty to populated. Worth doing on correctness grounds, not
as a fix for the counts.

## Tests

Four cases added to the `TestDirectHeardNode` table, which previously
covered direct routes only with `PayloadTXT_MSG`:

- direct + ADVERT + empty path credits the advertiser
- transport-direct + ADVERT + empty path credits the advertiser
- direct + empty path + **not** an advert credits nobody
- direct + ADVERT + **non-empty** path credits nobody

The last two matter as much as the first two: they pin the exception to
exactly the shape the firmware guarantees.

`gofmt` clean. Go tests not run locally (no cgo toolchain on this
machine since #1992, and per AGENTS.md `CGO_ENABLED=0` builds a stub
that proves nothing), so CI is their first run.

## Related

#2063 fixes the empty state's wording on the same card, which asserts
the node is out of range when the data cannot establish that. The two
are independent: this one adds evidence, that one stops overclaiming
when there is none.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 23:55:55 +02:00
efitenandClaude Opus 5 25f8426d32 test: make every suite in tests/e2e run, and tell the truth when it does (#2053)
Closes #2037.

Step 1 (#2045) wired in the nine suites that already passed. This is
steps 2 and 3: the four that ran and failed, and the five that could not
run at all. After this, the count of suites in `tests/e2e` invoked by
nothing goes from 18 to 0.

## Step 2 — the four that ran and failed

Triaged against a fixture server in CI with full output kept, not the
four-line tail the first probe saved.

| suite | verdict |
|---|---|
| `test-packets-scope-column.js` | passes on master today. It failed on
2026-09-18, so something between the two fixed it. Wired in unchanged
rather than investigated. |
| `test-node-reach-e2e.js` | test wrong. It waited for the reach map
whenever any *link* had GPS; `public/node-reach.js` builds the map only
when the *node* has coordinates. Guaranteed 10s timeout on a node with
positioned neighbours and no position of its own. |
| `test-channel-modal-e2e.js` | both failures test-side. The Add
button's visible label was shortened to "+ Add" with the accessible name
moved to `aria-label` (`public/channels.js:746`); and
`.ch-section-mychannels` is conditional on the visitor having added a
channel, which has not happened at that point in the suite. |
| `test-touch-targets.js` | three test-side, two a product finding. |

The three test-side touch-target failures were the harness measuring
controls that are not shown: `.compare-btn` (the CTA was removed in
#1646, and `style.css` says so), `.ch-back-btn` (`display:none` outside
the mobile channels layout), and `.filter-toggle-btn` (`display:none` on
mobile since #1461; the control shown is the navbar mirror, which
`mobile-page-actions.js:70` builds as a `.nav-btn`, so it was already
measured). All three are dropped from the table with the reason recorded
in the file.

The remaining two are **not** a test problem: `.nav-btn` and
`.ch-icon-btn` are each declared twice in `public/style.css`, 48px in
the touch-target block and 44px in their own component rule, and the
later one wins. Rather than lower the blanket or hide the failures, the
suite now has `DEFAULT_MIN = 48` plus a `MIN_OVERRIDES` table holding
those two at their effective 44, so a third selector dropping to 44
still fails the build. The contradiction is **#2052**, with both ways
out costed; the override entries should go when it is settled.

## Step 3 — the five that could not run

None is deleted. I checked each selector and seam against the product
before deciding, and every one still targets a surface that exists and
that nothing else covers.

Four were written against `@playwright/test`, a runner the project
neither installs nor uses anywhere else. Adopting a second runner for
twelve tests costs more than porting them, and the precedent is already
set: `test-path-inspector-coverage-e2e.js` exists, as its own header
says, because `test-path-inspector-e2e.js` could not run. So they are
ported to the plain-node Chromium pattern the other 109 suites use.

- **`test-issue-1522-trace-url-sync-e2e.js`** — the trace hash in the
URL, both directions. `test-e2e-playwright.js` covers that the page
loads and searches; it never looks at the URL, which is the whole of
#1522.
- **`test-marker-outline-weight.js`** — the canvas pulse ring never
thins below 2px. There is no CSS rule to read and axe cannot see inside
a canvas, so sampling the seam is the only way. Added a guard that the
ring was actually visible, so the weight check cannot pass vacuously on
a pulse that never rendered.
- **`test-pr-1490-live-map-gpu-animations-e2e.js`** — the queue drains,
the engine sleeps again, the fading trails stay under the cap of 5, and
the canvas sits on `animationsPane` rather than under the markers.
- **`test-path-inspector-e2e.js`** — reduced to what nothing else
covers: the map side pane, the `/#/traces/<hash>` redirect, the tools
landing. Its standalone-page test duplicated the wired coverage suite
and is dropped. Its "switching candidate clears prior polyline" case
ended after the click with a comment and no assertion, which is the same
green-but-empty problem this issue is about; it now compares path
counts, and skips loudly when the fixture yields too few candidates.

The fifth, **`test-table-sort.js`**, needed `jsdom`, which was declared
nowhere. It is a unit test of `public/table-sort.js` filed under
`tests/e2e`, so: `jsdom` is a devDependency (lockfile updated, `npm ci`
stays consistent), the file moved to `tests/unit/`, and the
`domIntegration` group in `scripts/non-unit-tests.json` is gone with its
only member. It runs 22 tests. 20 passed immediately; 2 had rotted,
because #1648 M2 replaced the up/down glyphs with Phosphor sprites and
the direction moved out of `textContent` into the `<use href>`. Those
two now read the sprite ref and the `aria-sort` value, so they also
guard the accessible announcement.

## Verification

`tests/unit/test-table-sort.js` 22/22 and `test-test-inventory.js` pass
locally; the E2E suites need a fixture server, which I cannot build here
(no cgo toolchain since #1992), so CI is their first run as committed.
The triage above was measured in CI, not assumed.

## Not done

The per-assertion skips named in the second comment on #2037 are
untouched: the two flaky packet-detail cases, the fixture-data ones, and
the two `clientRxCoverage` suites that skip wholesale while reporting
success. Those need a fixture deployment with coverage enabled, which is
its own change. I have not opened it.

Option 2 from the issue, making `test-test-inventory.js` require a
`deploy.yml` line for every `tests/e2e` file, is also not here. It is
the right guard and it is now enforceable, since the list is finally at
zero, but it belongs in its own change where a red build means what it
says.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 11:45:11 +02:00
efitenandClaude Opus 5 b614badb85 fix(packets): give each observation its own wire bytes in the detail API (#2055)
Closes #1999.

## The defect, measured on a production instance

Packet `96d716f18d885e78` on a live deployment, read from the deployed
build's own API:

| | |
|---|---|
| observations in the response | 60 |
| distinct `path_json` values | 50 |
| distinct `raw_hex` values returned | **1** |
| distinct frames actually stored in SQLite | **51** |

The contradiction the issue describes, from that same response:

```
obs 38410791  path ["58C0","1403","50C7"]  ->  hex 094258C01403AF37E39624E548FB7575F195A0BF
```

Three hops in the path, two path bytes in the frame. The bytes belong to
the 2-hop observation and are served for all 60.

Across the 3000 most recent transmissions on that database: 1977 have
more than one observation and **1844 of those (93%) hold genuinely
different frames**. 32659 of 36133 observations (90%) differ from their
transmission's canonical bytes. This is the normal case, not an edge
case.

## Cause

The store deliberately does not retain `obs.RawHex`. #881 dropped it as
a memory optimisation, ~98MB measured on a 1.7M-observation store, on
the assumption that one content hash implies one frame. The firmware
hashes payload and type independently of the relay path, so that
assumption is false.

Worth adding to the issue's diagnosis: all four load and ingest paths in
`cmd/server/store.go` still `SELECT o.raw_hex` and scan it into
`obsRawHex`, then use it nowhere — LoadAll, loadChunk, `IngestNewFromDB`
and `IngestNewObservations`. The bytes are read out of SQLite and
discarded, so a cold load pays the transfer for nothing.

## The fix

Keeps the memory saving and reads the bytes back only where a human is
looking at one packet.

- **`cmd/server/db.go`** gains `ObservationRawHexForHash`: one query
returning the stored frame per observation id. Two indexed lookups
regardless of observation count — `transmissions.hash` through the
prepared `stmtTxByHash` (`idx_transmissions_hash`), then
`observations.transmission_id` (`idx_observations_transmission_id`).
Guarded by `hasObsRawHex`, because #881 made the column optional and the
query would be a SQL error without it.
- **`cmd/server/routes.go`** backfills in `handlePacketDetail`: once per
request rather than once per observation, and after the store lock is
released. Bytes already present are never overwritten, and an
observation with no stored frame still falls back to the transmission's.

Against the acceptance list: observation bytes exposed with canonical as
fallback only ✓; the store's memory optimisation untouched ✓; bounded
indexed reads with no query per observation and no work under the store
lock ✓; startup-loaded, newly ingested and DB-fallback details all
covered, because both the store path and the DB path converge on this
one backfill and both key observations by an int `id` ✓.

**No frontend change is needed.** `public/packets.js` already spreads
the selected observation over the packet (`{...pkt, ...currentObs}`) and
already reasons about per-observation bytes: the comment there says
"post-#882 per-obs raw_hex with a different path length than the
top-level packet's raw_hex still gets accurate byte highlights". The
client was built for this and has been receiving 60 copies of one frame.

## Tests

`cmd/server/obs_raw_hex_test.go`:

- the per-id mapping, with three distinct frames and a fourth
observation storing none
- the `hasObsRawHex` guard, so a schema without the column is not
queried
- the handler regression: each observation carries its own frame, the
frameless one falls back to the canonical bytes, and at least three
distinct frames come back across four observations — the last assertion
so that a regression to repeating one frame fails, rather than passing
on shape

## Verification

`gofmt` clean. **Go tests were not run locally**: no cgo toolchain on
this machine since #1992, and per AGENTS.md `CGO_ENABLED=0` builds a
stub that proves nothing. CI is their first run.

Browser validation per AGENTS.md rule 2: I verified **the defect** in a
real browser and through the deployed API, with the numbers above. I
could **not** validate the fix in a browser, because the change is
server-side Go and is not deployed anywhere yet. Saying so rather than
claiming otherwise.

## Not done

- The four scan sites that fetch `o.raw_hex` and discard it are left
alone. Removing the column from those query builders would stop
transferring roughly ten frames per transmission on every cold load, but
it touches four builders and their `scanArgs` alignment and is not
needed for this defect.
- `fetchResolvedPathForObs`, immediately next to this code in
`enrichObsWithTx`, does run one query per observation. This change
deliberately does not copy that pattern, and does not fix it either.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 11:08:17 +02:00
efitenandClaude Opus 5 d1b615fc0d fix(node-health): list only observers that heard the node on air (#2057)
Closes #2056.

## What changes

The node detail "Heard By" card now lists only observers that received
the node's **own transmission off the air**, and reports the rest as a
count.

```
HEARD BY — DIRECT (8 OBSERVERS)
OBSERVER                REGION  PACKETS  AVG SNR   AVG RSSI
BE-DUF-SiSCD-01         —        16276    7.9 dB   -108 dBm
BE-BRU-Moris  repeater  —        13775   -5.6 dB   -122 dBm
...
Seen via relay by 29 observers. Those observers heard a repeater that
forwarded this node's traffic, not this node.
```

and for a node nothing hears:

```
HEARD BY — DIRECT (0 OBSERVERS)
No observer is within radio range of this node.
Seen via relay by 2 observers. …
```

## The rule, and where it comes from

Read out of the firmware rather than assumed:

| | |
|---|---|
| `Packet.h:83` | `setPathHashSizeAndCount(sz,n) { path_len =
((sz-1)<<6) \| (n&63); }` — hash size rides in the packet's `path_len`
byte |
| `Mesh.cpp:649,678` | only `sendFlood()` sets it, so the **originator**
decides; `CommonCLI.h:69` defaults `path_hash_mode = 0`, i.e. one byte |
| `Mesh.cpp:349` | a forwarding repeater appends its hash with the
packet's size — it cannot upgrade a packet, and the **last hop is who
was heard** |
| `Mesh.cpp:89,103` | on a direct route a forwarder matches the head of
the path and calls `removeSelfFromPath` before retransmitting, so the
path is the **remaining** route and the transmitter is not in it |

So an observation credits exactly one node:

1. Route type must be `ROUTE_TYPE_FLOOD` or
`ROUTE_TYPE_TRANSPORT_FLOOD`. Direct routes never qualify (38% of
transmissions over 7 days).
2. Empty path → the originator, known only for ADVERTs.
3. Otherwise the last hop.
4. The hop must resolve to exactly one candidate. Same gate
`resolvePathForObsColdLoad` already applies: under-attribute rather than
guess. It drops 418,530 of 1,455,721 flood observations with a path over
7 days (28.8%), and it is what stops the wrong-band credits.

## Measured effect

| node | before | after |
|---|---|---|
| BE-BRU-Moris | 36 observers | 3 |
| BE-KRO-RP01 \| ON1KW | 40 | 3 |
| BE-BRE-ON8AR | 38 | 2 |
| NL-BXE-RP01 \| 433 | 35 | 0 |

Network-wide over 7 days, 234 of 1,860 nodes have at least one direct
observer (161 have exactly one, maximum 8). The direct list is therefore
empty for most nodes, with the relay count below it. That is the correct
reading: no observer is in radio range of them.

Independent corroboration on staging: for BE-WIL-3EIK-01 the eight
direct observers are exactly the top eight entries of its Neighbors
table by score and observation count.

## Perf justification

`GetNodeHealth` is fast today precisely because it never walks
observations — it uses one representative observation per transmission.
Direct-RF needs the per-observation path, and that cannot be a
per-request walk: the reference store holds **232,928 transmissions /
2,887,861 observations**, one node's `byNode` slice alone holds **55,458
transmissions / 1,450,544 observations**, and
`/api/nodes/bulk-health?limit=200` would multiply that.

So the aggregate is rebuilt by a background recomputer on the existing
`newAnalyticsRecomputer` pattern, published into an `atomic.Value`.
Reads are `O(direct observers)`, which is **cheaper than before** — the
old code built per-observer sums over every transmission in `byNode` on
every request.

Proof, `BenchmarkBuildDirectHeardIndex`:

```
BenchmarkBuildDirectHeardIndex-12    1    63067900 ns/op
```

3,000,000 observations (60,000 transmissions × 50 observations, 8-hop
paths, 64 candidate repeaters) in **63 ms**, once per recompute
interval.

Per observation the walk does one route-type check, one backward scan of
`PathJSON` for the last quoted token (no allocation, no
`json.Unmarshal`), one prefix-map lookup and one counter update.

Rebuilding wholesale also means eviction needs no bookkeeping: a pass
simply does not see evicted transmissions. The alternative — a field on
`StoreObs` updated incrementally — would have needed the call at five
construction sites (`store.go:942,1264,2854,3179`,
`chunked_load.go:609`), which is the duplication that caused #1558, plus
matching decrements at eviction.

## API

Both `GetNodeHealth` and `GetBulkHealth` carried a near-identical copy
of the observer loop; they now share one builder.

- `observers` — direct-RF only. Same field names, so no client
migration. Rows are a named `HealthObserverRow` instead of
`map[string]interface{}` (one fewer occurrence in a touched file, per
the AGENTS.md ratchet).
- `relayObserverCount` — new integer, observers that saw traffic through
the node without hearing it. `stats.totalPackets` and `stats.avgHops`
still count relayed traffic, so without this number the card would
contradict the figures printed beside it.

`docs/api-spec.md` is updated for both endpoints. It also documented an
`iata` field on these rows that the endpoint has never emitted; removed.

## Tests

- `cmd/server/direct_heard_test.go` — table test over the rule: flood
with empty path and known originator, flood whose last hop is the node,
flood whose last hop is another node, direct and transport-direct routes
(never credit), ambiguous last-hop prefix, listener-only candidate,
1-byte and 2-byte hop sizes; plus aggregation and row-building.
- `cmd/server/node_health_direct_rf_test.go` — end-to-end through the
handler: an observer that only saw relayed traffic must not appear in
`observers` but must be counted in `relayObserverCount`. Plus the
benchmark.
- `tests/unit/test-direct-rf-heard-by.js` — slices the card template out
of `public/nodes.js` and evaluates it, so it tests the shipped markup
rather than a copy: heading, empty state, relay line, singular/plural,
signal columns, listener/repeater badge tri-state.
- `cmd/server/node_health_can_relay_case_1290_test.go` — updated to seed
a genuinely direct reception, since a relay-only observer no longer
carries a badge.
- `cmd/server/analytics_recompute_after_load_test.go` — recomputer count
10 → 11.

Verified locally: `cmd/server` suite green, `sh test-all.sh` green (180
suites), `tests/e2e/test-e2e-playwright.js` 131/134 passed with 3
skipped and 0 failures against the seeded fixture, plus
`test-issue-1147-section-order-e2e.js`,
`test-issue-1151-orphan-separators-e2e.js` and
`test-issue-1281-location-row-e2e.js`, which all assert on this card.
`gofmt` clean, `vet` clean across all modules.

Browser-validated on staging: both the full detail page and the side
pane, on a node with 8 direct observers and on the 433 MHz node with
none. No console errors.

## What this does not do

`prefixMap.resolveWithContext` still guesses on ambiguous hops, so
paths, neighbor edges and analytics keep their current attribution.
Making it abstain is a much larger change and needs its own issue.

The "Regions" line and Region column on this card read `o.iata`, which
this endpoint has never emitted, so both have always been dead. Left as
found rather than widened into this change.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 07:55:37 +02:00
efitenandClaude Opus 5 5c016a210a fix(analytics): show a building state on the distance index's 202, and stop caching it (#2051)
Closes #1997.

## What was wrong

`/api/analytics/distance` answers `202 {status:"building",
retry_after_seconds:5}` with no `summary` until the lazy index (#1011)
has been built.

1. `renderDistanceTab` read `data.summary.totalHops` straight away. The
TypeError is caught by the tab's own try/catch, so it never reaches
`window.onerror`: it is painted into the tab as `Failed to load distance
analytics: Cannot read properties of undefined (reading 'totalHops')`.
2. `api()` caches any `res.ok` body, and `res.ok` is true for 202, so
the placeholder was stored for the `analyticsRF` TTL. Even a correct
retry read the cached "building" body back. That is the half that made
the broken state outlast the index build.

## What changed

- `public/app.js:173` skips the cache write when `res.status === 202`. A
200 still caches, unchanged.
- `public/analytics.js` renders a building notice and retries itself,
honouring `retry_after_seconds` clamped to [1s, 30s]. The timer is
cleared on tab switch and in `destroy()`, and a new render supersedes a
pending retry, so two renders cannot write into the same tab.

The notice uses `.text-center`/`.text-muted` rather than the `.spinner`
class used at `analytics.js:1683`, because `.spinner` has no CSS
anywhere in the repo and renders nothing.

## Tests

`tests/unit/test-issue-1997-distance-building.js` (7 assertions, wired
into `test-all.sh`) pins the two pure decisions the renderer makes and
`api()`'s refusal to cache a 202 while still caching a 200. Red-run on
the unfixed sources: 6 of 7 fail, and the "a 200 is still cached"
control stays green.

`tests/e2e/test-issue-1997-distance-building-e2e.js` (classified in
`scripts/non-unit-tests.json`, invoked from `deploy.yml` with
`CHROMIUM_REQUIRE=1`) serves both responses by route interception, so it
does not depend on whether the server under test has an index built. It
asserts: the building state appears, the tab does not paint the error
text, a retry arrives with no interaction, the retry replaces the
placeholder once the server answers 200, and no retry fires after
leaving the tab.

Check (2) deliberately asserts on the rendered text and not on
`pageerror`: the TypeError is caught, so a `pageerror` assertion would
pass on the broken build too.

## Verification

No local cgo toolchain here since #1992, so I could not build a server
to run the E2E against. Instead I ran its five steps in Playwright
against a live instance with this branch's `public/app.js` and
`public/analytics.js` injected in place of the deployed ones (both files
are byte-identical between that instance and upstream master, so the
injection is faithful):

| check | deployed build | this branch |
|---|---|---|
| (1) building state shown | fail | pass |
| (2) not rendered as data | fail | pass |
| (3) retried on its own | fail (1 request) | pass (2 requests) |
| (4) real payload after retry | fail | pass |
| (5) no retry after leaving the tab | pass | pass |

(5) passes on the broken build too: it schedules no retry at all, so it
is a control and only means anything together with (3).

The committed E2E suite has not been run as committed. CI is its first
real run.

## Not done

- The server still recomputes on every 202 poll rather than signalling
readiness.
- No other analytics tab was audited for the same assume-a-summary
pattern.
- One pre-existing unit suite (`test-preflight-xss-gate.js`) fails on
this Windows machine with a cp1252 `UnicodeEncodeError` from its Python
helper, on a clean tree as well as with this change. Unrelated, and
green on Linux CI.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 17:04:37 +02:00
efitenandClaude Opus 5 7c6b95ea53 fix(store): merge background chunks in order instead of prepending them (#2050)
Fixes #2024.

`s.packets` is declared "sorted by first_seen ASC (oldest first; newest
at tail)" (`cmd/server/store.go:177`), and retention eviction depends on
it: `evictStaleInternal` walks from the head and stops at the first
transmission inside the window. A slice out of order is therefore
**under-evicted silently** rather than failing loudly.

## What breaks it

The background chunk loader. Chunks are windowed on `last_seen` (#1690),
so a transmission first heard weeks ago and heard again recently arrives
in a *recent* chunk carrying its old `first_seen`. The chunk was then
put in front of the slice:

```go
s.packets = append(localPackets, s.packets...)
```

and never re-sorted, so the next chunk, which covers an older window,
was prepended in front of it and left that ancient row sitting behind
newer ones. `LoadChunked` re-sorts after its own load; the background
merge did not. That asymmetry is the whole bug.

It is not a corner case. On a production database, of the **236080**
transmissions in a 14 day window, **2071** have a `first_seen` more than
a day older than their `last_seen`, and **1848** more than a week.

This matters more since #2035: with the accounting fixed, `maxMemoryMB`
actually triggers, and a walk that stops early works against it.

## The fix

`mergeChunkIntoPackets` merges the two sorted runs linearly. Re-sorting
the whole slice was not an option: this runs under `s.mu` once per
chunk, so it would sort hundreds of thousands of packets while ingest
waits for the lock. The chunk already arrives sorted, since the chunk
query ends in `ORDER BY t.first_seen ASC`, so the `sort.SliceIsSorted`
guard is a contract check costing one linear pass that never sorts in
production.

## Covered

- `TestMergeChunkIntoPackets_KeepsFirstSeenOrder` pins the merge against
an interleaving, deliberately unsorted chunk.
- `BenchmarkMergeChunkIntoPackets` guards the linear cost, against a
future simplification back into a sort.

The server suite runs under `-race` in CI and is green.

## Not covered, and I would rather say it than let the PR imply
otherwise

There is **no integration test driving `loadChunk` end to end**. I wrote
one and dropped it: a faithful seed database for that path needs more of
the schema and more of the loader's preconditions than the fix itself is
worth. Two CI rounds in, the seed was still loading zero packets (the
first attempt failed at `OpenDB` on a missing `nodes` table, the second
on the window). Both attempts are in this branch's history rather than
rewritten away.

So the end-to-end claim rests on the code path quoted above and on the
production measurement, not on a test that exercises it. The unit test
covers the function where the logic now lives, which is the part that
can regress.

Also not verified locally: `cmd/server` needs cgo for the #1992 driver
and this machine has no C toolchain, so CI is the check.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 15:12:36 +02:00