mirror of
https://github.com/Kpa-clawbot/meshcore-analyzer.git
synced 2026-10-11 06:37:45 +00:00
2e7a4ebcfe7e2dc09373390ff79746ed6d4f36ee
2994
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2e7a4ebcfe |
fix(ingestor): bound the unauthenticated /neighbors report (#2122)
`handleNeighborsReport` trusted whatever an observer published on the `/neighbors` topic. The sender chose `origin_id` (whose "self" it is) and could list any pubkey as a responded neighbor; each got its `configured_scope` written with any scope string, stamped with the sender's own timestamp. The store is last-write-wins on that timestamp and `normalizeReportTS` accepted any RFC3339 time, so one report dated years ahead was written once and then blocked every genuine later report for that node until someone edited the database. The value is shown on the reach page as the confirmed scope and feeds `/api/scope-audit`. **Fix (three guards):** - pubkeys must be 64 hex chars — anything else cannot match a node anyway, so it is dropped instead of running UPDATEs that never match - the normalised scope list is capped at 256 bytes - a report stamped more than 5 minutes ahead of our clock is dropped, so a far-future timestamp can no longer lock the node **Tests:** `neighbors_guard_test.go` covers each guard, including that a genuine report still lands after a future-stamped one was rejected and that ordinary clock skew is still accepted. Full ingestor suite passes. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
a981420d21 |
fix(ingestor): cap observer-supplied string lengths (#2123)
Observer `id` and `iata` come from the MQTT topic; `origin` (name), `model`, `firmware`, `client_version` and `radio` come from the status JSON. Any publisher controls them and nothing bounded their length, so one message could store a 64 KB observer id or name. Each new id is also a new `observers` row, and that table is joined by most packet queries. **Fix:** a small `clampObserverField` helper strips control characters and truncates: ids to 128 runes, text fields to 128, IATA to 16. Applied on the status path, the packet path and in `extractObserverMeta`. Values are truncated rather than rejected, so a legitimate observer with a long name still appears. **Tests:** `observer_fields_test.go` — short values untouched, long values cut at 128 runes (not bytes, so multi-byte names are not split), control characters removed, `extractObserverMeta` caps all string fields. Full ingestor suite passes. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
548e4cd4a4 |
fix(api): cap the nodes= list on /api/packets at 50 entries (#2120)
Each entry in the comma-separated `nodes=` list on `GET /api/packets` costs one SQLite lookup (`resolveNodePubkey`) while the packet store's read lock is held. A 1 MB URL fits about 15,000 entries. On a test instance 12,000 entries took 1.3 s per request, against 0.9 ms for one entry, and the lock stalls the poller's writes for that long. A few parallel clients can keep the site busy and the live feed stale. **Fix:** lists longer than 50 entries get HTTP 400 with a clear message. No UI page sends more than a handful. **Tests:** `multi_node_cap_test.go` — 50 entries return 200, 51 return 400. Full `go test ./...` in `cmd/server` passes. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
4ce6d9f6d4 |
fix(ui): pin CDN script versions and add integrity hashes (#2121)
`leaflet.heat` and `chart.js` in `index.html`, and swagger-ui on `/api/docs`, were loaded from unpkg with no `integrity` attribute. `chart.js@4` and `swagger-ui-dist@5` also floated on a major version, so a new release would load unreviewed. A compromised CDN or package would run as our own code on every page. Leaflet itself already had a hash. **Fix:** pin `chart.js@4.5.1`, `leaflet.heat@0.2.0`, `swagger-ui-dist@5.33.1`, each with a sha384 hash and `crossorigin="anonymous"`. **How the hashes were made:** download each pinned file, `openssl dgst -sha384 -binary | base64`, then download again and check the hash matches. **Note:** bumping chart.js or swagger-ui now means updating the hash too. Vendoring them into `public/vendor/` (as markercluster already is) would remove the CDN dependency entirely; happy to do that instead if preferred. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
8b64439635 |
fix(server): keep config.json's file mode when saving the geo-filter (#2126)
`SaveGeoFilter` rewrote `config.json` through a temp file created with mode 0644, so a config an operator had made 0600 (it holds the API key and broker passwords) became world-readable after the first geo-filter save. **Fix:** stat the original and reuse its mode for the temp file. Falls back to 0644 when the stat fails. **Tests:** `config_mode_test.go` — a 0600 config stays 0600 after a save; a 0644 config stays 0644. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
523a5edfa1 |
fix(ui): escape the route string on the unknown-route page (#2124)
`app.js` wrote the URL fragment into `innerHTML` as the heading of the "Page not yet implemented" view for routes it does not know. Browsers percent-encode `<`, `>` and `"` in fragments, so this is unlikely to be exploitable today, but it is a plain `innerHTML` sink on a URL-controlled string and costs one `escapeHtml()` to close. **Tests:** one test added to `tests/unit/test-xss-escape-sinks.js` in the file's existing style. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
7309bdb5a9 |
fix(store): index a transmission once per relay key in byPathHop (#2117)
Relates to #2108 ## Problem `traffic_share_score` grows with server uptime until it is far above reality, and many relays end up clamped at 1.0. The score is the number of non-advert entries in `byPathHop[pubkey]` divided by the number of non-advert transmissions. `indexResolvedPathHops` runs once per **observation**, on the live-ingest path (`IngestNewFromDB`) and on the late-observation path (`IngestNewObservations`). `addResolvedPubkeysToPathHopIndex` only de-duplicates within one call, so every further observation through the same relay appends the transmission to that relay's bucket again. The denominator counts each transmission once. After a restart the values look right, because `buildPathHopIndex` → `retainResolvedPathHops` de-duplicates by `*StoreTx`. They then drift upwards as live observations arrive. None of the affected functions has changed on `master` since the diagnosis against `415362c`. This branch is based on `000d9ab`. ## Fix 1. **Idempotent insert, per (transmission, relay key).** `addResolvedPubkeysToPathHopIndex` keeps a side map `pathHopResolved map[*StoreTx][]string`. It holds the resolved keys each transmission is already indexed under and skips those. - Keys are compared **exactly**, so no collision can ever drop an entry. - A later observation through a **new** relay still adds that relay, once. - `StoreTx` is unchanged. 2. **Interned keys.** `pathHopKeys map[string]string` keeps one shared copy of each resolved key. The record then holds a 16-byte string header per entry, and not the string each observation's resolve allocated: a fresh `json.Unmarshal` string per persisted observation on `Load`, and a `strings.ToLower` result on the live paths. The `byPathHop` key is the same shared copy. 3. **Eviction and rebuild.** - `evictStaleInternal` drops the record of evicted transmissions, plus any interned key whose bucket it deletes. - `retainResolvedPathHops` keeps the records of live transmissions only and drops interned keys whose bucket was not carried over. When a rebuild starts from an empty index, it clears both. 4. **Defence in depth.** The single (`GetRepeaterUsefulnessScore`), batch (`GetRepeaterNodeStatsBatch`) and bulk (`computeRepeaterUsefulnessScoreMap`) scores count **distinct** non-advert transmissions per bucket, via `countDistinctNonAdvert`. - IDs are collected in a reused slice. - A bucket in ascending ID order is already distinct; only out-of-order buckets are sorted. - The bulk pass does this for full-pubkey keys only. Raw-hop buckets get one entry per transmission from `addTxToPathHopIndex`. 5. **Cache side effect.** A repeated observation through known relays no longer mutates `byPathHop`, so it no longer drops the batch relay-stats cache, as the cache contract from `1164` intends. An observation through a new relay still drops it. The other `byPathHop` consumers already de-duplicate by `tx.ID` or read only raw prefix keys: `GetNodeHopAnalytics`, `computeMultiByteCapability`, `handleNodePaths`, `computeRepeaterRelayInfoMap` and the relay info in `GetRepeaterNodeStatsBatch`. Two tests pin that they ignore duplicate entries. ## Exact keys vs. a 64-bit fingerprint The first version of this fix stored a 64-bit FNV-1a hash per key. As requested, I measured both, plus exact keys without interning, on the real `addResolvedPubkeysToPathHopIndex`. **Keys per transmission.** In the e2e fixture, the union of resolved relay keys over all observations of one transmission has a mean of 5.4 and a maximum of 17. The protocol limit is 64 path bytes (`MAX_PATH_SIZE`), i.e. 64 / hash size hops. **Memory** retained by the record, per transmission. `BenchmarkPathHopRecordMemory_2108`: 100K transmissions, keys from 4,000 relays, each transmission heard twice with freshly allocated keys. The record map and the interned keys are included. | keys per tx | exact, interned (this PR) | 64-bit fingerprint | exact, not interned | |---|---|---|---| | 2 | 88 B | 68 B | 207 B | | 5 | 136 B | 101 B | 447 B | | 17 | 344 B | 197 B | 1,423 B | - At realistic key counts, exact keys cost 20–35 B per transmission more than the fingerprint. That is about 3.5 MB per 100K transmissions, and small next to what the store already charges per transmission (`storeTxBaseBytes` alone is 384 B). - Without interning, exact keys would cost 3–4× more, and most of that would be held during every `Load`. Interning is what makes exact keys affordable. **Time.** I ran the variants interleaved, 3 rounds each, median ns/op on an Apple M4. The record step is not measurable inside the full call. The isolated record step was faster with exact keys in a separate micro-benchmark, because comparing a handful of 64-char strings is cheaper than hashing each one. | benchmark | exact, interned | fingerprint | exact, not interned | |---|---|---|---| | late observation, 20K txs | 312 | 358 | 331 | | late observation, 100K txs | 760 | 836 | 758 | | `IngestNewObservations` (SQL + resolve + index) | 531 µs | 497 µs | 527 µs | **Collision probability of the fingerprint.** For k distinct keys in one transmission it is about k(k−1)/2 / 2^64: | k | per transmission | per 10^9 transmissions | |---|---|---| | 2 | 5.4e-20 | 5.4e-11 | | 5 | 5.4e-19 | 5.4e-10 | | 17 | 7.4e-18 | 7.4e-9 | | 64 | 1.1e-16 | 1.1e-7 | For random keys this is negligible. FNV-1a is unkeyed, though, so two relay keys that collide could be found deliberately (a birthday search over about 2^32 key pairs). A collision would leave a transmission out of one relay's bucket. **Decision.** Exact keys are cheap enough once interned: the same speed within noise, and a few tens of bytes per transmission. They remove the question of collision-driven omissions entirely, so this PR uses exact keys and drops the fingerprint. The two new maps (`pathHopResolved`, `pathHopKeys`) are not added to `trackedBytes`, so the `maxMemoryMB` trigger undercounts by roughly the per-transmission figures above. Both are bounded by live state (eviction and rebuild prune them), so this is a steady proportional undercount rather than a leak. ## Tests **New: `cmd/server/pathhop_dedupe_2108_test.go`.** Behaviour tests that use only existing API. Each one runs against the real SQLite schema through `Load`, `IngestNewFromDB` and `IngestNewObservations` where noted. | test | covers | on `master` | |---|---|---| | `TestPathHopIndexOncePerTx_LateObservations_2108` | 1 + 10 late observations through the same relays: one entry per relay | **fails** (11 entries) | | `TestPathHopIndexOncePerTx_LiveIngestBatch_2108` | 11 observations in one `IngestNewFromDB` batch | **fails** (11) | | `TestPathHopIndexOncePerTx_AfterLoad_2108` | `Load` + rebuild, then one live observation | **fails** (2) | | `TestPathHopIndexAddsNewRelayFromLaterObservation_2108` | a later observation through a **new** relay adds it exactly once, and its share becomes correct | **fails** (2) | | `TestTrafficShareStableAcrossLateObservations_2108` | single, batch and bulk scores agree, equal the definition, stay put over rounds of late observations, and equal a fresh `Load` of the same data | **fails** (all four relays at 1.0, want 0.4–0.5) | | `TestPathHopIndexSizeBoundedByTransmissions_2108` | the index size does not grow with observations | **fails** | | `TestRelayStatsCacheAcrossRepeatedObservations_2108` | a repeated observation keeps the relay-stats cache, and the kept cache equals a fresh compute; a new relay drops it, and the next read sees the relay | **fails** | | `TestAddResolvedPubkeysToPathHopIndex_PerRelayIdempotent_2108` | the helper's return value and cache invalidation per (tx, relay), in any key order | **fails** | | `TestTrafficShareCountsDistinctTransmissions_2108` | all three scores count distinct transmissions in a duplicated bucket | **fails** | | `TestPathHopConsumersIgnoreDuplicateEntries_2108` | pin of the consumer audit | passes (by design) | | `TestMultiByteCapabilityIgnoresDuplicateEntries_2108` | pin of the consumer audit | passes (by design) | **New: `cmd/server/pathhop_record_2108_test.go`.** These tests use the new symbols, so they do not compile against `master`. - `TestCountDistinctNonAdvert_2108`: ascending, descending, interleaved duplicates, adverts, nils, untyped. - `TestPathHopResolvedRecordBoundedAndEvicted_2108`: the record stays at 2 keys after 11 observations. Eviction removes the record and every entry. A survivor's next observation adds nothing. Evicting everything leaves no record, no interned key and no bucket. - `TestPathHopResolvedRecordAcrossRebuild_2108`: a rebuild keeps the records of live transmissions, so the next observation adds nothing. It drops the record of a removed transmission and the interned key only that transmission used. - `TestPathHopResolvedRecordClearedWithEmptyIndex_2108`: a rebuild from an empty index clears the record and the interned keys, and the next observation puts the transmission back. - `TestPathHopResolvedRecordInternsKeys_2108`: record entries and the `byPathHop` key share one copy, even though every observation passes freshly allocated keys. **Changed: `pathhop_eviction_1908_test.go`.** It built its duplicate entries by repeating `indexResolvedPathHops`, which no longer duplicates. It now seeds the duplicates directly, so the `1908` sweep is still tested against buckets that hold one transmission several times. The expected buckets are unchanged. **Changed: `db_test.go`.** `setupTestDB` takes `testing.TB`, so the SQL benchmark can use it. **Mutants.** I applied each mutant on its own to this branch. All 14 are killed: | mutant | killed by | |---|---| | record never consulted | the OncePerTx, stable-share, size, cache and record tests | | dedupe per transmission instead of per relay | `AddsNewRelayFromLaterObservation`, `RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent` | | eviction keeps the record | `RecordBoundedAndEvicted` | | rebuild keeps records of removed transmissions | `RecordAcrossRebuild` | | rebuild from an empty index keeps the record | `RecordClearedWithEmptyIndex` | | cache dropped on every call | `RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`, existing `NoMutation_PreservesCache` | | cache kept although `byPathHop` changed | `RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`, existing `InvalidatesRelayStatsCache` | | single score counts entries | `TrafficShareCountsDistinctTransmissions` | | batch score counts entries | `TrafficShareCountsDistinctTransmissions` | | bulk score counts entries | `TrafficShareCountsDistinctTransmissions` | | distinct count trusts any bucket order | `TrafficShareCountsDistinctTransmissions`, `CountDistinctNonAdvert` | | keys not interned | `RecordInternsKeys` | | eviction keeps interned keys | `RecordBoundedAndEvicted` | | rebuild keeps interned keys of dropped buckets | `RecordAcrossRebuild` | **Commands run:** - `gofmt -l` on all tracked Go files: clean. - `go vet ./...` in all 14 modules: clean. - `cd cmd/server && go test -race ./...`: pass. - `cd cmd/ingestor && go test ./...`: pass. - `sh test-all.sh`: 184 of 186 suites pass locally. The other two, `test-issue-1956-release-routing.js` and `test-preflight-xss-gate.js`, shell out to scripts that need bash ≥ 4 (`mapfile`). They fail under macOS's bash 3.2 regardless of this change. This PR touches no frontend or script files. ## Benchmark: `master` vs. this branch I ran `master`'s sources (`000d9ab`) and this branch interleaved, 5 rounds, with the same benchmark files. Medians on an Apple M4. | benchmark | `master` | this PR | change | |---|---|---|---| | late observation through known relays, 20K txs | 446 ns, 38 B/op | 441 ns, 0 B/op | within noise | | late observation through known relays, 100K txs | 710 ns, 69 B/op | 755 ns, 0 B/op | within noise (runs overlap) | | … index size afterwards, entries/tx (20K / 100K) | 24.99 / 8.99, still growing | 4.99 / 4.99 | bounded | | `IngestNewObservations`, one observation for each of 20 txs | 532 µs | 513 µs | within noise | | bulk score pass, clean index, 20K, ingest order | 246 µs | 316 µs | +28 % | | bulk score pass, clean index, 100K, ingest order | 3.03 ms | 4.30 ms | +42 % | | bulk score pass, clean index, 20K, reversed buckets | 257 µs | 467 µs | +82 % | | bulk score pass, clean index, 100K, reversed buckets | 3.38 ms | 4.79 ms | +42 % | | bulk score pass after 11 observations per tx, 20K | 663 µs | 284 µs | −57 % | | bulk score pass after 11 observations per tx, 100K | 4.72 ms | 3.19 ms | −32 % | - On an index that is clean on both builds, the distinct count makes the bulk pass slower. That pass runs on the cache-miss path, which the background recomputer refreshes every 5 minutes by default. - On the index a live server actually holds without this fix, `master` walks every accumulated duplicate, so the real-world pass is faster after the fix. The gap grows with uptime. - The late-observation step no longer allocates. On `master` its index grows with every observation. Benchmarks: `BenchmarkLateObservationIndex_2108`, `BenchmarkIngestNewObservations_2108`, `BenchmarkTrafficShareScoreMap_2108` and `BenchmarkPathHopRecordMemory_2108`. ## Production motivation We have run this fix on two production instances. Before the fix, the summed `traffic_share_score` grew past 180 and many relays sat at the 1.0 clamp; with the fix no node reaches `≥ 0.999`. A controlled 12 h A/B run makes the drift explicit. Two servers read one identical database — one on this `master` base, one with the fix — alongside a reference server that freshly loads the same database (a fresh load is correct, because the rebuild dedupes). The unfixed server drifted to 9 relays at the 1.0 clamp and up to 0.95 absolute error per node against the reference; the fixed server stayed within 0.05 of the reference for every node, with no relay at the clamp. A smaller residual rise remains and is a separate cause (startup-vs-live hop resolution); it is deliberately left for a follow-up so its effect stays measurable. ## Out of scope That remaining slow rise comes from a separate cause: live ingest resolves some hops that the startup load does not. This PR deliberately leaves that drift alone, so its effect stays measurable. The fix will follow in a separate PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f3ae8ac25a |
feat: sync a logged-in user's settings across devices (part B) (#2130)
Part B of #2128: a logged-in user's settings follow them across devices. Log in on a phone and your own nodes, favorites, customizer and filters are there; a change on one device reaches the others within a minute or when you return to the tab. **This PR builds on #2129.** Until that one is merged, the diff here includes it. The commits for this part start at `docs(specs): settings sync for optional user management (sub-project B)`. ## The situation Everything a visitor sets up lives in one browser's `localStorage` (about 100 keys in `public/`). A second device or a cleared cache starts from zero (#895). ## What this PR adds **Storage.** `users.db` schema v2: one JSON document per user in `user_settings`, with a revision number and a generation id. A write succeeds only when the client's revision and generation match the stored ones, so two devices cannot overwrite each other silently. **Server.** `GET`, `PUT` and `DELETE /api/account/settings`, behind the same session and CSRF checks as the account routes. - The server owns the list of synced keys (61 keys, [`settings_allowlist.go`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/cmd/server/settings_allowlist.go)) and sends it to the client, so the two cannot drift. - A hard denylist, checked first, refuses `meshcore-api-key`, every `corescope_channel_*` key and `live-channel-colors` (#725). The colour map is keyed by channel hash, and for a user-added channel that hash is `user:<name>`, which would expose hashtag channel names. - Documents are capped at 256 KiB, measured like `JSON.stringify`. PUT is limited to 60 requests per hour per user. A stale revision gets 409 with the current document. **Client** ([`settings-sync.js`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/public/settings-sync.js)). Inert unless the feature is on and someone is logged in. - It wraps `localStorage.setItem` and `removeItem` for allowlisted keys only and pushes 2 seconds after the last change. - It pulls on login, page load, tab focus and every 60 seconds while the tab is visible. - **Merge:** three-way, against a per-device baseline that belongs to one user and one document generation. Lists (own nodes, favorites, saved filters) merge per item, so an item added anywhere is kept and an item removed on one device does not come back from another. Single values: the profile wins unless only this device changed it. - Remote changes are written without a push, theme and colour-blind preset are re-applied, and the current page re-renders (skipped on account pages and while the geofilter editor is open). **UI.** - Logout asks: keep my settings on this device (default), remove them from this device, or cancel. Channel keys are never removed: no copy exists anywhere else. - The account page gets a "Settings sync" section: last synced time, "Sync now", what is and is not synced, and "Delete synced settings from my account". ## Not synced Layout and device state (panel and column widths, collapsed panels, map positions, geofilter drafts), channel data (#725), the API key, and all `sessionStorage`. The full list is in the [spec](https://github.com/efiten/CoreScope/blob/feat/settings-sync/docs/specs/2026-10-06-user-settings-sync-design.md). ## Performance - One GET per page load, tab focus and minute while visible; one debounced PUT per burst of changes. - The `setItem` wrapper costs one Set lookup per write for non-synced keys. A synced write reads one small revision key, not the stored document. - The server reads or writes one row per request. ## Verification - `internal/users` and `cmd/server`: `go vet` and `go test` pass locally (22 new Go tests), including a test that every allowlisted key still occurs in `public/`, and denylist tests. - `tests/unit/test-settings-sync.js`: 79 passing (vm, real module). The cases cover the merge table, two tabs sharing one storage, stale answers after a push, delete while a push is in flight, and logout while the final push fails. - `sh test-all.sh` exits 0. - `tests/e2e/test-user-management-e2e.js` (10 steps, 4 of them new) passed locally with two browser contexts as two devices: a favorite and the packet time window travel from device 1 to device 2, a removal does not come back, and "remove from this device" clears the synced keys while a channel key stays. - Checked by hand on a staging instance with a desktop and a phone on one account. ## Not in this PR - On a shared browser where the previous user chose "keep", the next user's first login merges those settings into their own account. The user guide says to choose "remove" on shared computers. - Saved filter expressions are synced as typed, including any channel names written in them. The guide says so. - Realtime push between devices; the minute pull is the sync interval. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
d232072f85 |
feat: optional user accounts (part A: foundation) (#2129)
Part A of #2128: optional, off-by-default user accounts. With the feature off nothing changes; with it on, visitors can register and log in, and admins manage users and use the operator actions without the API key. PR #2130 (settings sync) builds on this one. The two are meant to be merged together. ## The situation - Operator actions (geofilter save and prune, backup, perf reset) need the shared `apiKey`. There is no per-person right. - Nothing in CoreScope knows who a visitor is, so the requests in #2128 that need that (#1835, #2092, #1508, #730) have nothing to build on. ## What this PR adds **Two new Go modules** - `internal/users`: a separate `users.db` (SQLite through `modernc.org/sqlite`) with users, sessions, single-use tokens, an audit log and a mail log. Passwords use argon2id. - `internal/mailer`: a `Mailer` interface with a Brevo client (send, delivery events, webhook parsing) and an in-memory fake for tests. **Server (`cmd/server`)**, active only with `userManagement.enabled` - 24 routes, all documented in OpenAPI under the `users` tag ([`auth_routes.go`](https://github.com/efiten/CoreScope/blob/feat/user-management/cmd/server/auth_routes.go)): - auth: register, activate, login, logout, me, forgot, reset; - account: profile, password, email change with confirmation, sessions, self-delete; - admin: list, detail, disable, enable, delete, role, resend activation, manual activation, mail status refresh; - a Brevo webhook, registered only when `mail.webhookSecret` is set. - `requireAdmin` replaces `requireAPIKey` at the 7 operator call sites: the API key **or** an admin session. With the feature off it is the old API-key gate (`TestRequireAdminWithoutUserManagementIsAPIKeyGate`). - `/api/config/client` gets `userManagement: {enabled: true}` only when the service started; with the feature off the response is byte-identical. **Frontend** - `auth.js` (header account control, request helper that adds the CSRF header), `account.js` (login, register, activate, forgot, reset, confirm email, my account), `admin-users.js` (`#/admin/users`, deep-linked filters), `account.css` (theme tokens only). - On phones the top-bar control is hidden, so a conditional entry goes into the bottom-nav "More" sheet and the nav drawer. - The customizer geofilter tab and the Perf "Reset stats" button use the admin session when there is one. **Config.** A `userManagement` block (`config.example.json`, [`docs/user-guide/accounts.md`](https://github.com/efiten/CoreScope/blob/feat/user-management/docs/user-guide/accounts.md)). The Brevo key can come from `CORESCOPE_BREVO_API_KEY`. The server refuses to start when the block is enabled but incomplete. ## Security choices - Session cookie `cs_session`: HttpOnly, SameSite=Lax, Secure when `publicBaseUrl` is https. Every cookie-authenticated state change needs the `X-CS-CSRF` header and a matching Origin. - Activation needs the token **and** the account password. Without the password, an attacker who keeps re-registering a known address could get the owner to activate an account that carries the attacker's password. - Register, forgot and email change answer identically for known and unknown addresses. A password reset ends all sessions, a password change ends all other sessions, and both end outstanding email-change links. - Rate limits: login 10 per 15 minutes, register and forgot 5 per hour, per IP and per address. The bucket count is capped. `trustedProxies` makes the per-IP limits see real client IPs behind a proxy. - Server logs carry `#<user id>`, never addresses, tokens or passwords; mail-provider error texts are redacted before logging. ## Performance No change to an existing hot path with the feature off. With it on: - One `users.db` lookup per authenticated request (session by token hash). - The admin user table rebuilds its `tbody` on each filter change. `users.List` caps the result at 1000 rows (`internal/users/users.go`), which bounds the rebuild. - `map[string]interface{}` in `openapi.go`: 79 before, 78 after. ## Verification - `internal/users`, `internal/mailer` and `cmd/server`: `go vet` and `go test -race` pass locally. 121 new Go tests. - `cmd/server` with `-tags e2etest`: vet and the e2e hook tests pass. - `sh test-all.sh` exits 0. `tests/unit/test-user-management-ui.js`: 67 passing (vm, real modules). - `tests/e2e/test-user-management-e2e.js` (6 steps) passed locally against an `e2etest` build with the fake mailer and against a feature-off build. CI builds the `e2etest` binary and runs the suite on a second server (`deploy.yml`). - On a staging instance with a real Brevo key: register, activation mail delivered, activate, admin table, "Refresh status" showing sent, deferred, delivered, opened and clicked. ## Not in this PR - Settings sync (#2130), the admin dashboard, approval flows and notifications (parts B to E of #2128). - A `requireReadAuth` mode (#1835). Sessions from this PR are what such a mode would accept. - Binary size and build time with `modernc.org/sqlite` linked next to `mattn/go-sqlite3` were not measured. Their driver names do not collide. #1992 discusses the driver choice. - No Brevo webhook was configured on staging; delivery status there came from "Refresh status". --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
4363403495 |
fix(nodes): match the region node filter on observer ID, not ingest-time IATA (#2115)
Follow-up to #2114, from its review. `RegionNodePubkeys` compared the IATA copied onto each observation at ingest (`StoreObs.ObserverIATA`). When an operator changes an observer's code, those copies stay as they were until a restart or until the observations age out. `/api/nodes?region=` then kept matching the old region, while the other store region filters already used the new one. Those filters resolve observers by ID through `resolveRegionObservers`, for example the packets query and `computeNodeHomeRegions`. The SQL subquery that #2114 replaced also read `observers.iata` at query time, so this restores that behaviour. ## Also fixes a regression from #2114 Found in review: `loadChunk`, the background history loader (`cmd/server/store.go`, the chunk SELECT and the `StoreObs` it builds), sets `ObserverID` on each observation but never `ObserverIATA`. With #2114 matching on that IATA, **every node whose adverts came only from background-loaded history dropped out of `/api/nodes?region=`** on current master. Matching on observer ID fixes it. `TestRegionNodePubkeysMatchesChunkLoadedObservations` pins it (an observation with the observer ID and an empty IATA) and fails on master's `region_nodes.go`. ## Change - Resolve the region to observer IDs with `resolveRegionObservers` (own mutex, 30 s cache) and match observations by `ObserverID`. Lock order: `regionNodesMu` is released before it, and `s.mu` is taken after it; none of the three is held together. - Without a database there is nothing to resolve, so `RegionNodePubkeys` reports no set and the handler keeps the SQL path. ## Tests - The region tests now seed an observers table and leave each observation's IATA at a stale value, so they can only pass through the table. - New `TestRegionNodePubkeysFollowsObserverIATAChange`: an observer that moved from SJC to SFO matches SFO and not SJC. It fails on the previous code (`got [pk_moved]` for SJC). - `TestRegionNodePubkeysMatchesChunkLoadedObservations` (second commit), see above. - `setupTestDB` takes `testing.TB` so the benchmark can use it; every existing caller passes `*testing.T` unchanged. - Full `cmd/server` suite passes locally, the region tests also with `-race`. ## Performance `BenchmarkRegionNodePubkeys` (220k adverts × 8 observations): 37 ms to 30 ms per uncached scan, a map lookup per observation instead of a string normalisation. The observer lookup is one query on the small `observers` table, cached for 30 s. ## Not done - No singleflight on a cold cache, and the 64-entry cache still resets when full. Both were non-blocking in the #2114 review. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d5e6d1b2a6 |
fix(nodes): resolve the /api/nodes region filter from the packet store (#2114)
Fixes #2101. `/api/nodes?region=` filtered nodes with a subquery that joins every advert to all of its observations and observers, with no time bound (`cmd/server/db.go`, `GetNodes`). It ran twice per request, once for `COUNT(*)` and once for the page. A client paging through nodes repeats both per page. ## Reproduced On our staging database (10.3 GB, 16.5M observations, 219,750 adverts), read-only `sqlite3`, region `BRU`: | | real | user | sys | |---|---|---|---| | one regional `COUNT(*)`, as `GetNodes` builds it | **154.8 s** | 1.9 s | 5.8 s | Almost all of it is waiting on disk. The plan walks `idx_transmissions_payload_type` for every advert and then `idx_observations_dedup` for each one's observations. Two of those per request, a few requests at once, and the reader pool is gone, which matches the report of all four database workers sitting in `GetNodes`. ## Fix - **`PacketStore.RegionNodePubkeys(region)`** (new `cmd/server/region_nodes.go`) walks the store's in-memory adverts (`byPayloadType[ADVERT]`) once. It keeps the pubkey of every node with an advert heard by an observer in the region, using the `ObserverIATA` each observation already carries. It is cached for 30 s per region, and the cache is bounded to 64 entries because its keys come from the client's parameter. Lock discipline follows `resolveAreaNodes`: the cache mutex and `s.mu` are never held together, and the ordering note in `store.go` lists it. - **`GetNodes` takes a `NodeQuery` struct** with a `RegionPubkeys` field, passed as one `json_each` parameter: `public_key IN (SELECT value FROM json_each(?))`, a primary-key lookup. An empty set matches nothing. The SQL subquery stays for a server without a store (tests, tooling). - The advert-pubkey lookup that `trackAdvertPubkey`, `untrackAdvertPubkey` and `computeNodeHomeRegions` each copied is now one helper, `advertPubkey`. The struct instead of a second `GetNodes` variant keeps the `map[string]interface{}` count unchanged in `db.go` (75) and `routes.go` (59). ## Behaviour change The region filter now covers the adverts the store holds (`packetStore.retentionHours`), which is the window the rest of the UI shows, instead of all database history. A node heard in a region only before that window no longer matches the filter. While the store is still loading after a restart, the set grows as history loads. ## Performance | | before | after | |---|---|---| | region set, uncached | 154.8 s (SQL count, staging) | 37 ms (`BenchmarkRegionNodePubkeys`: 220k adverts × 8 observations, 4,000 nodes) | | region set, within 30 s | same again | cache hit | | node count over the set | (included above) | 2 ms on staging (1,200 keys) | | 500-row page over the set | same scan again | 3 ms on staging | The scan holds `s.mu` for reading for those 37 ms, at most once per region every 30 s. ## Tests - `TestRegionNodePubkeys`: the in-region advert, an advert heard in two regions, case and whitespace in codes, a comma list, an unknown region giving an empty set, a non-advert never counting, a blank region giving no filter. - `TestRegionNodePubkeysCached`, `TestRegionNodePubkeysCacheIsBounded` (1,000 distinct regions stay within 64 entries). - `TestGetNodesRegionPubkeys`: the set combines with the role filter and counts correctly, and an empty set returns nothing even with `Region` set. - `TestHandleNodesRegionUsesStore`: an advert that only the store knows about shows up through `/api/nodes?region=`, so the handler is proven not to ask SQL. - Existing region tests (`TestGetNodesRegionFilterV2` and the `db_test.go` region cases) pass unchanged through the SQL fallback. Full `cmd/server` suite passes locally; the new tests also pass with `-race`. ## Not done - No request context on these queries, also raised in the issue. With the scan gone they take milliseconds, so I left that out of this change. - The region semantics stay "heard by an observer in the region". #1879 argues for the node's home region instead; that is a separate decision. - No frontend change. I did not check this in a browser; the nodes page and map call the same endpoint with the same parameters. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b5b230e884 |
feat(channels): show the sender's path hash size on each message (#2089)
Each channel message now shows the hash size its sender's path uses, read from bits 7-6 of the path byte that the originator writes and repeaters keep. The server sends it as path_hash_size (packetpath.HashSize); the frontend helper pathHashSize() applies the same rule and returns 0 for unknown. One rule on both pages: the packet detail Hash Size row and the hex breakdown call pathHashSize() too, so a 0-hop flood message reports the same size on Channels and on its packet page. A direct packet with no hops left reports no size, matching cmd/server/decoder.go. The cases live in test-fixtures/path-hash-size-cases.json, read by the Go and JS tests. |
||
|
|
000d9ab030 |
feat(coverage): RF noise-floor layer on the Mobile RX coverage page (#2113)
The ingestor already stores the noise floor that CoreDrive RX companions
report with each GPS fix (`client_rf_samples`, opt-in through
`clientRfSamples`, `cmd/ingestor/client_rf_sample.go:62`), but nothing
reads it back. An operator who enables it collects the data and cannot
see it. This adds the read side: `GET /api/rf-noise` and a Signal/Noise
toggle on the Mobile RX coverage page.
## What it does
- **`GET /api/rf-noise?bbox=&z=&days=`** returns a GeoJSON hex grid with
the median, quietest and noisiest noise floor per cell, in the same
shape as `/api/rx-coverage`. It is registered always and 404s unless
`clientRfSamples.enabled` is true, the same pattern as the coverage
routes.
- **Stationary samples are excluded.** A parked companion logs hundreds
of readings at one point, which would otherwise define its cell.
- **Coverage page:** a Signal/Noise toggle, rendered only when
`/api/config/client` reports `clientRfSamples: true`. The noise layer
reuses the coverage colour tokens with the axis inverted, because a
lower dBm is quieter. Tiers are at -115 and -108 dBm.
- **Deep link:** `#/rx-coverage?layer=noise` opens on the noise layer.
- **Empty and failed answers** are labelled on the map ("No RF samples
in this view yet", or a retry hint), so a blank map never reads as
"feature off".
No new configuration key: it reads the existing `clientRfSamples`
section that the ingestor already uses. Default off, so nothing changes
for an instance that has not opted in.
## Where it comes from
Ported from the efiten/CoreScope fork, where it has run on
analyzer.on8ar.eu since September (fork commits `42d09f0e`, `80581c43`,
`cde95078`). The cherry-picks conflicted with upstream's newer
`routes.go`, `types.go` and `rx-coverage.js`, so the final state was
ported by hand. The fork-only `/scopes` route that sat next to it in the
same hunk is deliberately left out.
## Performance
The query is bounded by `sampled_at` (indexed, `idx_crf_prune`) and the
bbox, aggregation is one pass plus a per-cell sort, and the response is
capped at 5000 cells. Measured on analyzer.on8ar.eu, all of Belgium
(`bbox=49.4,2.4,51.6,6.5&z=9`), from a client in Belgium, so network
time included:
| window | samples | cells | response time |
|---|---|---|---|
| 7 days | 10,835 | 241 | 0.37 s |
| 30 days | 43,303 | 429 | 0.39 s |
That table holds 45,270 rows in total. It is only read when someone
opens the noise layer, never on ingest or WebSocket paths.
## Tests
- `cmd/server/rf_noise_test.go`: 8 tests for the aggregation (median and
extremes, stationary exclusion, the cell cap, empty input) and the gate
(404 when off, even with data present). All pass, and the full
`cmd/server` suite passes locally.
- `tests/unit/test-rx-coverage-noise.js` (new, in `test-all.sh`): the
colour axis runs the right way, including both tier bounds and a string
median from the API. Slices the real function out of `rx-coverage.js`
and fails on master's copy, which has no thresholds.
- `tests/e2e/test-rx-coverage-noise-e2e.js` (new, wired in `deploy.yml`
and `scripts/non-unit-tests.json`): no toggle when the flag is off;
Noise fetches `/api/rf-noise`, draws the cells, swaps legend and
subtitle, puts `layer=noise` in the hash; a deep link opens on the noise
layer and an empty answer shows the message. 3/3 locally; against
master's `rx-coverage.js` and `roles.js`, 1/3 (only the "flag off" case
passes, as it should).
- `test-rx-coverage-viewport-e2e.js` still passes. ESLint 8 reports 0
errors on the changed files.
## Not done
- The thresholds (-115 / -108 dBm) are fitted to the fork's own data
(1,241 moving samples at the time) and are constants in
`rx-coverage.js`. Per AGENTS.md rule 8 they belong in the customizer
eventually; not in this PR.
- The E2E suite mocks the API. The real endpoint is covered by the Go
tests and by the measurements above, not by a fixture with seeded
samples.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
9c3d76e14f |
docs(release): notes and changelog for v3.13.1 (#2112)
Release notes and CHANGELOG entry for v3.13.1, a patch release that ships #2111 (node-discover replies count as coverage). Once merged, the tag goes on the merge commit so `deploy.yml` picks up `docs/release-notes/v3.13.1.md` as the release body. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>v3.13.1 |
||
|
|
00ee4f9e23 |
fix(coverage): attribute node-discover replies to the responder (#2111)
## Problem A CoreDrive RX companion that runs node-discover gets no coverage on an upstream instance. Reported today by an operator on v3.13.0: the app logged 32 discover replies from one repeater in 16 minutes (`heard 751a49f0c5dadc70 (8B, discover)`), all published, and `client_receptions` stayed at 0 rows for that companion while `client_rx_observations` held 93. Two gaps, both on the upstream side: 1. **Ingestor.** `deriveHeardKey` attributes a FLOOD `path[last]` and a 0-hop advert, and drops a 0-hop `CONTROL/DISCOVER_RESP` (firmware `CTL_TYPE_NODE_DISCOVER_RESP`). The reply carries the responder's own pubkey at offset 6 of the payload, which `decoder.go` already parses into `CtrlPubKey` (#1802), but the coverage path never used it. 2. **Server.** CoreDrive RX asks for `DISCOVER_PREFIX_ONLY`, so the firmware answers with an 8-byte prefix (`simple_repeater/MyMesh.cpp:799-805`). `coverageHeardKeyCandidates` only built the 64, 6 and 4 hex candidates, so a 16-hex `heard_key` would be stored and then matched by no per-node coverage query. ## Change - `cmd/ingestor/client_reception.go`: a third branch in `deriveHeardKey` for a discover response with no hops. Accepts exactly 8 or 32 bytes, nothing truncated, stored with `src='discover'`. - `cmd/server/rx_coverage.go`: adds the 16-hex prefix to `coverageHeardKeyCandidates`. - `docs/client-rx-coverage.md`: documents the `discover` source, the 8-byte keylen and the four-candidate lookup. The leaderboard and `/api/rx-coverage` read `client_receptions` without a key filter, so they pick the rows up without a change. Name resolution goes through `batchResolveHeardKeys`, which is a prefix lookup and handles 16 hex as is. ## Evidence from a deployment that has had this since 2026-08-19 On analyzer.on8ar.eu, `client_receptions` over the last 7 days by `src`: discover 4232, rxlog 4909, geo 548, advert 34. Discover replies are 44% of all coverage rows there (4232 of 9723); on an upstream instance those rows are not written. They cannot be backfilled afterwards either: `client_rx_observations` keeps no raw bytes, so the responder pubkey is gone. ## Tests - `TestDeriveHeardKey` and `TestBuildClientReception` gain discover cases: 8-byte and 32-byte keys accepted (32-byte uppercase input lowercased), a 3-byte and an empty key rejected, a non-discover CONTROL rejected, a discover response with hops not attributed. - `TestHandleClientPacketDiscoverRespWritesReception`: end to end, a raw 0-hop DISCOVER_RESP on the client topic writes one `client_receptions` row with `src='discover'`. - `TestCoverageHeardKeyCandidatesIncludesDiscoverPrefix`: the 16-hex prefix is among the per-node candidates. - Ran locally on Windows: `go test ./...` in `cmd/server` passes; in `cmd/ingestor` everything passes except `TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which needs the symlink privilege on Windows and fails on clean master too. Not done: no browser validation, as the change is ingestor and server only and the frontend reads the same endpoints. Not included: geographic resolution of 1-byte hops (`src='geo'`) and the RF noise layer, which are separate changes. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c2a9500060 |
docs(release): notes and changelog for v3.13.0 (#2110)
Release notes and CHANGELOG entry for v3.13.0, in the shape used for v3.12.0. Once merged, the tag goes on the merge commit so `deploy.yml` picks up `docs/release-notes/v3.13.0.md` as the release body (#2076). - 16 commits since v3.12.0: 10 fix, 4 test, 2 feat, so a minor bump. - "Read this before upgrading" covers the one-off advert route evidence backfill from #2088. Durations are measured on two instances (live 24m 1s on 16.8M observations, staging 36m 23s on 16.5M), and ingest during the backfill window is measured as unaffected on live. - Issues closed come from each PR's closing references: 14 from 16 PRs. Docs only. #2089 is deliberately not in this release. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>v3.13.0 |
||
|
|
ef12535364 |
test(map): wait for the new route sidebar before measuring it (#2109)
Step (9) of `tests/e2e/test-path-inspector-e2e.js` fails intermittently on master with `(3-9) fixture route coverage: 0 !== 700`: 3 of the 18 runs that executed it since 2026-09-30, including master runs [37120155356](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37120155356) and [37156362836](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37156362836). A red master run skips the image build, so `:edge` is not rebuilt. ## Cause The step clicks a candidate and then waits for any `.mc-rt-sidebar`. The previous route's sidebar survives the hash-only `goto` to `#/map`, and `drawPacketRoute` (`public/map.js:1080`) replaces it only after `/api/resolve-hops` answers. So the wait matched the old sidebar, and the checks that follow measured either the old sidebar or, if the replacement landed between the locator resolving and the evaluation, a detached element, whose `getBoundingClientRect()` width is 0. This is a test defect. The app keeps the previous route on screen while the next one loads, which is intended. ## Fix Remember the sidebar before the click and wait until a different one is in the DOM. Test-only, one file. ## Evidence Locally against the fixture, with `/api/resolve-hops` delayed through `page.route` in a throwaway copy of the test, measuring which sidebar the step-9 assertion sees: | Delay | Before | After | |---|---|---| | 0 ms | new sidebar, pass | new sidebar, pass | | 80 to 100 ms | old sidebar, `isConnected: false`, width 0 (3 of 5 runs) | new sidebar, pass (5 of 5) | | 1500 ms | old sidebar, pass: the step checked the wrong page | new sidebar, pass | The unmodified fixed test then passed 3 of 3 runs. ## Not done - I did not look for the same wait pattern in other suites. - The delay probe is not committed. It needs a timing window to be useful, and that window would be a fixed sleep in CI. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9e49db9725 |
feat(nodes): group adverts by observed route evidence (#2085)
Red commit: `ac60248` ([two intended label failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37137429333)). Original grouping red: `767da9b` ([CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36353796690)). Fixes #2073. #2088 is merged. Rebased onto master `415362c`; all eight patches are unchanged, including the test-first history. Recent Adverts groups the bounded sample by authoritative `advert_kind`: Flood, Mixed flood / direct (empty path), Direct (empty path), and Other / unknown. Missing evidence stays unknown. Each advert appears once, preserving reception metadata, order, packet links, counts and origin/history explanations. Wording describes observed remaining paths without inferring original send mode or RF distance. Validation at `6317baf`: - All 185 standalone suites passed; syntax, XSS and lint passed (zero errors, 92 existing warnings). - Browser verified: `http://127.0.0.1:51827` (temporary Go fixture server, now stopped), desktop pane/full view and mobile. Mixed evidence, exact links and caveats checked again after pushing. - The expanded fixture exposes the existing #2104 failure in the unmodified core harness. Supplemental integration with #2105's exact test correction passes 135 checks, with three existing skips. Both PRs remain separate. - [Fresh green CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37138366726) passed; three independent automated reviews found no must-fix issues. E2E assertion added: `tests/e2e/test-e2e-playwright.js:142`. Grouping is O(n) over the existing 20-advert sample. No new requests, settings, colors or dependencies. ## Preflight overrides External OpenClaw runner/profile unavailable; repository checks and actual local Chromium were used. |
||
|
|
99a326a6b3 |
test(packets): verify observation hex after rerenders (#2105)
Red commit: `ca553e4` ([expected assertion failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37137426691)). Fixes #2104. The observation hex test now opens the existing grouped fixture directly, selects A → B → A with stable-ID locators, and asserts the selected ID and exact raw bytes after each render. Missing fixture data fails explicitly. Distinct observation frames are seeded after schema migration; this test's fixed sleeps and silent skips are removed. Rebased onto master `415362c` after #2090 merged. Both patches are unchanged. The previous #1122 layout blocker now passes against the actual current server and assets. Validation at `5db8017`: - All 185 standalone frontend suites passed. - Core browser: 133 passed, three existing skips; #1122 layout 6/6, filter 11/11, grouped-collapse 4/4. - Syntax, YAML, fixture SQL, XSS and lint passed (zero errors, 92 existing warnings). - Fresh red CI reached the intended detached-observation assertion. [Green CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37138187496) passed; three independent automated reviews found no must-fix issues. Only test code and fixture setup change; no configuration/customizer implications. ## Preflight overrides External OpenClaw runner/profile unavailable; repository checks and actual local Chromium were used. |
||
|
|
415362c5af |
fix(packets): resolve path hops with the observer that heard them (#2097) (#2099)
Fixes #2097. Every 1-byte path prefix on this network is shared. Measured on the live deployment: 254 prefixes cover all 2043 nodes, nine of them on a single prefix, and not one node has a prefix to itself. The packets page printed one of those nine as a certainty. It could already have said otherwise. `hop-resolver.js` marks such a hop `ambiguous` and returns its full candidate list; `hop-display.js` renders a warning badge with a count and a candidate popover; `renderHop(h, observerId)` reads a per-observer cache key. None of it fired, because `HopResolver.resolve()` takes six parameters and `packets.js` passed one. Without `observerId`, `packetIata` is null, `nodeInRegion()` never runs, no candidate is flagged `regional`, `globalFallback` stays false, and `hop-display.js` computes a badge count of 0. The whole chain stayed silent while the data said it was a guess. ## What this changes **Resolution carries the observer.** `resolveHops()` takes it and writes the per-observer cache key that `renderHop()` has always read and nothing ever wrote. `resolveHopsForPackets()` groups a multi-packet call by observer so each group is filtered on its own. One site was the bug in miniature: it wrote its result under `hopNameCache[k + ':' + pkt.observer_id]` after resolving without that observer, so the key promised something the value was not. **The observer's position is the anchor.** `resolve()` has always accepted `observerLat`/`observerLon` as the anchor at the receiving end: when nothing later in the path is resolved, the observer is the next known position, which is what `pickByAffinity` needs. `observerPosition()` is the single lookup, used by both resolve paths. This matters more than the IATA route, which turns out to be dead on this network. `nodeInRegion()` looks the observer's code up in `/api/iata-coords`; **0 of 42 observers have their code in that table** (`ANR BRU GNE HEP KJK LGG MST NRW OBL OST` are all missing, while its 54 entries are `AMS`, `APC`, `ATL` and the like). So the 300 km filter has never excluded anything here. 27 of 42 observers report lat/lon directly, and that works. Filed separately as #2098, since it is a network-wide lookup that resolves nothing and hop resolution may not be its only consumer. **Each badge belongs to its pill.** The badge is a sibling after the pill, so a path read `A [8] → B [6] → C` with nothing to say which name the 8 qualified. Pill and badge now render inside one `white-space: nowrap` `.hop-group`, leaving the separator outside the pair. Only a hop that carries a badge is wrapped; wrapping every hop would change the layout of every path for nothing. That markup predates this work, but these commits are what made it visible — before them the badge essentially never appeared on this page. **The list summarises, the detail pane does not.** Every 1-byte hop getting its own badge turned a table row into a line of warning triangles. `renderPath()` takes `{ summary: true }` for the three list call sites: names with no per-hop badge, and one indicator reading "N of M hops have more than one candidate". The two detail-pane sites are unchanged and keep a badge per hop, because that is where the candidates actually get read. `HopDisplay` gains `opts.badge === false` for it, and the hop keeps its `hop-ambiguous` class either way so CSS can still mark it. ## Scope Deliberately not included: chaining the next resolved hop to narrow a candidate set to one. `pickByAffinity` already scores on graph edges and would do it; it needs neighbouring context this call does not yet supply, and that belongs in its own change to `hop-resolver.js`. What is here makes the uncertainty visible and picks a defensible candidate rather than the first in index order. ## Verification 30 unit tests in `tests/unit/test-issue-2097-hop-ambiguity-badge.js`, registered in `test-all.sh`. Full frontend suite exits 0. Four of them are structural guards over `packets.js`, and each was mutation-checked by reverting the change it protects: - no `HopResolver.resolve()` call passes the hops alone - no call passes an observer id and drops its position - one helper does the observer lookup - the detail pane calls `renderPath` without `summary` The behavioural test for the anchor runs with `iataCoords` deliberately empty, which is the live state, and lists the distant candidate first: without an anchor the resolver keeps candidate order and picks it, with one it does not, and the hop stays reported as ambiguous with all candidates listed. **Browser validated** on a staging deployment of this branch, against live data: | | list | detail pane | |---|---|---| | per-hop badges | 0 | 24 | | path indicators | 1 | 0 | | badges outside a `.hop-group` | — | 0 | On the packet that prompted the issue, the first hop now resolves to a repeater 10.6 km from the transmitter instead of one 126 km away, and the second hop resolves to the only candidate that has a recorded neighbour edge matching the next hop in the path. ## Note for reviewers The last commit exists because staging disagreed with the tests. After the anchor landed, the detail pane still picked the distant candidate: it resolves through its own `HopResolver.resolve` call, which had the observer id but a null position. The guard that now fails on exactly that shape is what would have caught it before the deploy. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8933e2c703 |
fix(channels): keep client-only state across a channel-list refresh (#2095) (#2096)
Fixes #2095. `loadChannels()` replaced the `channels` array with the server snapshot and carried nothing across, and `mergeUserChannels()` only ever ran from `init()`. So any refresh destroyed every field that exists only in the tab. **No race needed for the worst of it.** Change the region filter or toggle show-encrypted with a PSK channel open, and: - the **My Channels** section disappears — `renderChannelList()` derives it as `channels.filter(c => c.userAdded === true)`, and the server returns neither `userAdded` nor a `user:`-prefixed hash - every unread badge resets to 0 - the user's own labels vanish - `reconcileSelectionAfterChannelRefresh()` cannot find the `user:*` hash in the snapshot, so it nulls `selectedHash`, sets `messages = []`, rewrites the URL to `#/channels` and replaces the open conversation with "Choose a channel from the sidebar" ## What this does `mergeClientChannelState(fresh, prev)` mirrors `mergeWsAppendedIntoRest()`, which already does exactly this job for `messages` (#1498). Same shape: pure, takes both arrays as parameters, returns a fresh array, never aliases or mutates an input. It carries `unread`, `userAdded` and `userLabel` by hash, and keeps `lastActivityMs` / `lastSender` / `lastMessage` / `messageCount` when a WebSocket batch landed while the request was in flight and is therefore newer than the snapshot. Those four move together: a sender without its message reads as a different message. `mergeUserChannels()` now runs inside `loadChannels()`, before the render and before the reconcile, so all three call sites get it instead of `init()` alone. ## What it deliberately does not do **It does not resurrect a channel the snapshot left out.** A region-filter change legitimately narrows the list, so carrying survivors over would defeat the filter, which is a worse bug than the one being fixed. The helper only enriches rows already present in the fresh snapshot. That leaves half of finding 2 in the issue unfixed: a channel pushed by the WebSocket handler during the initial in-flight window is still dropped. That one self-heals on the channel's next packet, and the reverted preview line is overwritten by the next WS batch. Fixing it properly needs a way to tell "dropped because the snapshot is stale" from "dropped because the filter excludes it", which is a larger change than this. ## Tests `tests/unit/test-issue-2095-channels-client-state.js`, 11 cases, registered in `test-all.sh`. Part 1 exercises the helper directly. Part 2 drives the real `loadChannels()` through the existing `_channelsLoadChannelsForTest` hook with a stubbed `api()`, which is what proves the helper is wired in rather than merely defined. **Verified by mutation.** With the wiring removed from `loadChannels()` but the helper left in place, the three reproduction cases fail with exactly the reported symptoms: ``` FAIL a refresh keeps the My Channels rows My Channels lost on refresh (got ["public1"]) FAIL a refresh keeps unread badges the unread badge reset to 0 on refresh undefined !== 7 FAIL a refresh does not close an open PSK conversation the open PSK channel was deselected + null - 'user:MyPSK' ``` The fourth case, "a refresh still drops a channel the server filtered out", stays green throughout, so the fix cannot be defeating the region filter. Full frontend suite green: `sh test-all.sh` exits 0. ## Not done **No browser validation.** This is frontend JS covered by unit tests that call the production function through its own hook, but I did not run Playwright against a server, so I am not claiming a browser check I did not do. `init()` still calls `mergeUserChannels()` and `renderChannelList()` after `loadChannels()` resolves, which is now redundant and costs one extra full sidebar render per page load. Left alone deliberately: it is idempotent, and the comment there records the regression it was added to fix (`test-channel-issue-1111-e2e.js`, case 2). Worth removing separately with that e2e test watched. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3b415fb007 |
fix(live): enable node filtering after its handlers are ready (#2103)
Red commit: `8974616` ([CI assertion failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37064461392)). Fixes #2094. The Live node-filter input now starts natively disabled and becomes editable after its input, keyboard and blur handlers attach. Previously it accepted text while initial node loading was pending, then never processed that text without another keystroke. The change adds one HTML attribute and one initialization assignment. Failed initial loading still reaches the existing error handler and installs the controls. Saved filters, URL precedence, autocomplete and keyboard selection keep their existing behavior. No new settings, styling, dependencies, requests or packet-processing work. ## Validation - The existing CI-wired Live browser suite delays or aborts the real initial node request. It verifies rejected premature typing, later editability, actual suggestions, keyboard selection and URL state. Both new cases failed on the red commit; all six cases pass locally on `11fab46`. - All 183 standalone frontend suites passed, plus lint and XSS gates. Separate browser checks passed saved-filter restoration and URL override; desktop/mobile loading/ready screenshots were inspected. - A broader local core run exposed the unchanged stale observation-handle test tracked in #2104. It is outside this Live fix. Full [CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37067337557) passed at `11fab46`, including Go, browsers and both container architectures; three independent reviews are clear. E2E assertion added: `tests/e2e/test-1110-live-filter.js:69`. Preflight overrides: external script/OpenClaw profile unavailable; repository checks and local Chromium used. |
||
|
|
b9fb3d244c |
feat(analytics): split adverts using recorded route evidence (#2088)
Red commits: `dc7dcbe`, `76e3217` ([assertion failures](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37059977385)). Fixes #2041; supplies #2085's backend contract. Records at most two route bits per retained advert. Node flood counts now use the same evidence, including mixed and transport-flood adverts. Analytics failures log/count errors while core observation, path and liveness writes continue. The unrelated recent-advert limit clamp is removed. Direct adverts (empty path) describes the observed remaining path across valid hash widths. Firmware can forward an advert down to an empty path; this cannot establish original send mode or RF distance. The `zero_hop` API spelling remains compatible. Successful post-upgrade evidence writes survive observation replacement and restart. Older overwritten frames and failed writes can leave gaps. The UI explains automatic background backfill. ## Measured validation Synthetic Windows workloads; timings are not production guarantees: - 300K retained adverts, one-hour window: median 64.35→39.84 ms; 13.98 MB→145 KB allocated. Seven-day allocations unchanged; five paired samples. - 1M transmissions/11M observations: backfill 5m43s. Simple 10 Hz writer: all 3,423 observations persisted; p99 28 ms, maximum 323 ms. - Simulated pre-backfill cursor: retained evidence ready 4m21s; full catch-up 6m41s, 122 cache invalidations. Normal completed-migration startup bulk-loads current evidence. Portable opt-in scale tests and window benchmarks are included. All 183 frontend suites, full server tests, targeted races, lint and real-API desktop/mobile checks passed. Full [Linux CI](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37063210229) passed at `36def9c`; three independent reviews are clear. Local Windows testing hit the unchanged symlink-privilege test. No configuration changes. Preflight overrides: external script/OpenClaw profile unavailable; repository checks and Chromium used. |
||
|
|
a4a447acea |
fix(packets): pack short columns, one expand arrow, Full Names toggle (#2090)
The packets table wastes most of its width on short columns, draws two arrows on grouped rows, and cuts every observer and hop name short. This fixes all three and adds a way to read the full path when it does not fit. ## Why the columns were wide `makeColumnsResizable` measured each column once on load and stored the widths as **percentages** of the table. Percentages grow with the screen: on a 2000px display "17s ago" sat in a 180px column. The measurement also read the virtual-scroll spacer row — a single `colspan` cell as wide as the table — as column 0, so the expand column came out 160px wide locally and about 950px on a live instance with grouped rows. Measured on the fixture at 2000px: | Column | Before | After | |---|---|---| | Expand | 162–230px | 32px | | Time | 111px | 67px | | Size | 64px | 46px | | HB | 44px | 29px | | Scope | 80px | 48px | | Path / Details | ~180px each | ~650px each | ## The change **Column packing.** A new `fitColumnsToContent` (`app.js`) sizes every column except Path and Details to its widest cell, in px; those two split what is left. Every visible column gets an explicit width and the table gets their sum: a spacer row keeps a slot for each hidden column, and leaving any column `auto` hands those slots a share of the slack as a blank strip at the right edge. When space is short (detail panel open, narrow window) the widest fixed columns give way first, then Path and Details shrink to 60px, and only then does the wrapper scroll. The split itself is a pure function, `distributeColumnWidths`, with unit tests. It re-lays out when the container resizes (detail panel), when TableResponsive hides or reveals columns (a new `table-columns-changed` event), and after sorting or a column toggle. After each render `grow()` widens columns for the rows just inserted and never narrows them, so nothing jitters while scrolling. Drag handles store px; double-clicking one resets it. Path and Details have no handle. `makeColumnsResizable`, still used by the nodes, observers and analytics tables, now skips `colspan` rows when measuring; other rows are measured as before. **One expand arrow.** A CSS `::before` "▶" sat next to the SVG caret, which pointed up when collapsed. The glyph is gone and the caret points right, then down. **Full Names.** A toolbar toggle next to Hex Paths shows observer and hop names untruncated. It is saved per browser; a link can carry `fullNames=1`, which applies to that page without overwriting the visitor's saved preference. Paths stay on one line: the virtual scroller positions rows at one fixed height, so wrapping would make scrolling jump. **The "+N" path pill, and a full-path popover.** Hops past the column edge were meant to go behind a "+N" pill (#1124), but the pill was appended after the overflowing chips and so sat past the clipped edge: it never showed. It is now sticky at the right edge, opaque on hover, and counts hop chips only (it used to count warning icons too). Hovering or focusing it shows the full path, one hop per line and numbered; click, Enter or Space pins it until an outside click or Escape, and it closes when the page changes or its row scrolls away. Clicking the pill no longer selects the row underneath. **Two related fixes.** - Rows drawn before `/api/observers` resolves show raw 64-char pubkeys. `loadObservers` now re-renders and re-measures, instead of leaving Observer at the pubkey's width (~600px under Full Names). - Group child rows indented every cell by 20px, which misaligned them with the headers and widened every column on expand. Only the Time cell is indented now. Anyone who had dragged packet columns gets the new defaults once: the old percentage key is dropped. ## Performance `grow()` runs after every `renderVisibleRows`, so it is on the hot path. Measured in Chromium on the fixture, 300 iterations, three runs: | Case | Cost | |---|---| | Full render, 63 rows measured | +0.88–0.93ms over the forced reflow alone | | Scroll step (incremental render) | ~73 `Range` rect reads, only the inserted rows | | Large jump / re-render | ~569 rect reads | Before narrowing `grow()` to inserted rows, one scroll step with 92 rendered rows made 5,472 rect reads. The DOM row count is bounded by the virtual-scroll window, so the cost does not grow with the packet count. `distributeColumnWidths` is a binary search over at most ~10 columns. ## Tests - `tests/e2e/test-packets-compact-columns-e2e.js`, 16 cases: short columns stay under fixed px bounds at 1920px; Path and Details take the freed width; hiding both still fills the wrapper; no blank strip and no horizontal scroll, with and without the detail panel, at 1920, 1100, 1000 and 375px; responsive reveal and re-hide refit; stale percentage widths dropped; drag follows the pointer and double-click resets; expanding a group leaves widths alone; Full Names untruncates, keeps paths on one line, persists and round-trips through the URL; the Observer column shrinks back when observer names arrive late; the pill popover lists every hop vertically, closes when the pointer leaves, pins on click without selecting the row, closes on Escape and on navigation, and opens on keyboard focus. Registered in `scripts/non-unit-tests.json` and `deploy.yml`. - On `master` the suite fails where it should, and restoring each fix's old behaviour (child-row indent, the late-observer refit) turns its case red. - `tests/unit/test-packets.js`: `distributeColumnWidths` (room to spare, rounding, widest-first capping, floors, dragged columns, the 60px minimum, overflow), the single chevron and no `::before` glyph, Full Names on flat and grouped observer cells, and the Full Names CSS. - `tests/unit/test-xss-escape-sinks.js` now builds the real `obsCellName` from source: escaping moved into it. Removing its `escapeHtml` turns both observer-cell tests red. - `test-issue-1189-composed-cell.js` and `test-issue-2012-clear-filters-selection.js` evaluate packets.js fragments with injected helpers and needed `obsCellName` / `showFullNames` added. - Existing packets E2E suites pass on this branch and on master against a CI-prepared fixture: scope column, #1128 (layout and multi-viewport), #1122, #1188, #1648, table sort, #1692, #1486. - Unit: 182 of 183 suites pass locally; `test-issue-1956-release-routing.js` could not create a temp directory in my sandbox and is unrelated. ## Browser validation Checked in Chromium against the fixture at 2000, 1920, 1280 (panel open and closed), 1100 and 375px, and with the seeded grouped row expanded and collapsed. Screenshots taken at each step; the table fills its container exactly, grouped rows show one arrow, and the pill popover lists full names. ## Not done, deliberately - **No width cap on Observer under Full Names.** A cap brings back the truncation the toggle removes, and Observer is already the first column to give way when space is short. - **Customizer.** The floors (60px flex minimum, 64px shrink floor, 32px expand, 58px Rpt) are hardcoded, per AGENTS rule 8; exposing them in the customizer can follow. - **Wrapping paths.** Rejected because of the fixed row height above; the pill popover is the way to read a long path. |
||
|
|
093e320c2b |
fix(map): keep Path Inspector candidates accessible and cover route replacement (#2082)
## Fix Fixes #2060. Fixes #2081. Path Inspector route drawing and candidate replacement now have deterministic browser coverage. An idempotent SQL seed adds four synthetic repeaters and four edges after fixture migration, without changing packet ordering. The test uses the real API, normal clicks, visible Leaflet geometry, and removal of every previous route object. Restoring coverage exposed a desktop layout bug: drawing the first route moved the map over the next candidate button. The route sidebar now participates in the existing flex layout, including resizing and collapse, while preserving the mobile bottom sheet. The historical zero-candidate result did not reproduce on the current baseline. No search thresholds or API behavior changed. No new dependencies or customizer settings. ## Validation - Local assertion-red history: `042db2b` requires seeded candidates; `ca1d22a` exposes the blocked second click. Subsequent commits repair layout and narrow-window behavior. - All 183 standalone suites passed. Unchanged-base Go failures: #2083 readiness race (fixed in #2084) and Windows symlink privilege. - Real Chromium: 11 Path Inspector/layout checks, 33 related map checks, and core E2E (131 passed, 3 existing fixture skips). - CSS-variable checks, XSS diff check, inventory, whitespace and PII checks passed. Seed remains valid after periodic graph refresh. Local browser validation used Chromium with a 60-second navigation budget; the OpenClaw profile and external preflight script were unavailable. Repository checks ran directly. [CI run 36343343569](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36343343569) records the initial assertion-red commit. Final CI must pass before merge. |
||
|
|
ded27a8312 |
fix(packets): stabilize fixture tests and preserve mobile empty states (#2080)
Fixes #2054 and #2079. Packet-row tests could outlive the fixture's default 15-minute window. An audit of 115 wired suites identified four needing explicit fixture windows: 180 minutes on mobile, 1440 on desktop. Mobile rejects larger saved windows, so a desktop-only preference was insufficient. Packet-icon checks now require a real path-bearing packet and both detail controls. Default-window coverage exposed #2079: responsive column hiding also hid the spanning empty-message cell. Skip structural single-cell spanning rows when hiding individual columns, preserving status messages and virtual-scroll spacer height. Ordinary and partial-span data cells retain their behavior; application defaults stay unchanged. - Red evidence: `bae5e78` fails the aged-fixture row assertion; `9b75cf6` fails spanning-row visibility. Guards/setup make them pass. - Validation: 183 standalone suites; 82 targeted browser cases on a fixture aged over 30 minutes (one pre-existing node-gesture skip). Broader browser run: 131 passed, 3 fixture-dependent skips. - Browser verified: fixture-backed Go server and Chromium; `coverage/2079-mobile-empty.png`. Real mobile scrolling retains spacer height and has no horizontal overflow. - E2E assertions: `test-observer-iata-1188-e2e.js` default-window cases and `test-table-fluid-e2e.js:58` shared-helper regression, both under `tests/e2e/`. - Three independent reviews completed; packet-icon coverage feedback addressed in `760057d`. - Populated mobile axe coverage remains part of #1995. - One constant-time guard per already-visited row; no new requests or collections. ## Preflight overrides - External `run-all.sh` unavailable; repository checks run directly. Local runs use UTF-8, the release harness's `GITHUB_REF_NAME`, and a 60s navigation budget; repository timeouts are unchanged. |
||
|
|
410c82c02c |
fix(ui): use consistent 48px navigation and channel buttons (#2078)
Fixes #2052. Navigation (`.nav-btn`) and channel icon (`.ch-icon-btn`) buttons now use the shared 48px house minimum. Remove competing 44px component/coarse-pointer rules and the legacy 32px mobile rule, and correct comments that confused the house preference with accessibility requirements. At 768px, user-added channel rows already clipped Remove; the larger Share target exposed the same constraint. Allow only those rows' controls to wrap so both actions remain within the sidebar. Navigation height and normal network-channel rows are unchanged. - Red `7787d9b`: rendered target assertions fail at 375/390/768/1280px. Green `e6ed9c8` changes CSS only. - Validation: all 183 standalone suites; 33 computed touch-target checks; 15 actual-page layout/click checks. Broader browser suite: 131 passed, 3 fixture-dependent skips. Three independent reviews found no required changes. - Browser verified: local Chromium and fixture-backed Go server; screenshots `coverage/2052-targets-375.png` and `coverage/2052-targets-768.png`. - E2E assertion added: `tests/e2e/test-channel-fluid-e2e.js:112`, covering dimensions, containment, overflow and normal Filter/Share clicks. Existing CI selection is retained; Chromium-required runs cannot silently skip the touch suite. - Test-only follow-up `5144503` waits for packet readiness before measuring navbar controls. Five fresh mobile contexts passed; size and click assertions remain intact. - No new settings, requests, dependencies or runtime JavaScript. ## Preflight overrides - External `run-all.sh` unavailable; repository checks run directly. - Local unit runner uses UTF-8 and `GITHUB_REF_NAME=local-validation` for the existing release harness. Browser navigation uses a 60s local budget for slow assets; repository assertions/timeouts are unchanged. |
||
|
|
f1edbbef3f |
fix(packets): preserve observation selection in detail URLs (#2093)
Red commit: `c165087` ([CI assertion failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36644030012)). Fixes #2091. Packet detail links now retain the selected observation when page initialization or filter changes rebuild the URL. Initialization also restores `obs` from the complete hash after the router strips its query. Explicit ID links render the requested observation and retain it in Copy Link state. The shared updater reads selection from the current route, so returning to the list or selecting another packet cannot resurrect an old observation. Existing filter serialization and Clear Filters behavior remain intact. No new requests, configuration, dependencies or layout changes. ## Validation - Real fixture browser checks use nondefault observation 502: hash/ID load, type/observer/time-window changes, refresh, Clear, and another refresh. Both URL and selected row are asserted. - All 183 standalone frontend suites passed; focused filter browser suite 11/11. - Broader local core run reached the unrelated Live input readiness bug #2094, reproduced on unchanged master. The complete CI browser suite passed. - ESLint 8: zero errors; 91 existing warnings. XSS diff, syntax and whitespace checks passed. - Three independent reviews found no production defects; assertions additionally prove filter controls changed before checking selection. E2E assertion added: `tests/e2e/test-filter-ux-e2e.js:179`. OpenClaw profile/external preflight were unavailable; Chromium and repository checks ran directly. Local navigation used a 60-second budget. Final CI passed at `280ced5`, including Go, browsers, coverage and both container architectures ([run](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36646175352)). |
||
|
|
dc4db17c48 |
fix(nodes): remove unsupported region fields from Heard By (#2077)
Fixes #2062. The node health API does not emit observer region data, but Heard By rendered a Region column containing only dashes and a Regions summary that could never appear. Remove that column, its sort control, and the unsupported region displays from the full node page and side panel. Keep Observer, Packets, Avg SNR and Avg RSSI, along with existing escaping, signal placeholders, relay counts and badges. No API, dependency or configuration changes. - Red commit `4396492` fails on the unwanted Region header and summary; `a360711` removes the unsupported fields. - Validation: 183 standalone suites and 7 focused browser checks passed. Broader browser run: 131 passed, 3 fixture-dependent skips. Independent reviews completed; the sorting-test finding was addressed in `1452981`. - Unit coverage evaluates the real templates. The existing CI-selected browser suite checks four-column alignment and sorting on desktop/mobile using the actual API field names. - Browser verified: local Chromium against a fixture-backed Go server; screenshots recorded in `coverage/issue-2062-heard-by-1400.png` and `coverage/issue-2062-heard-by-390.png`. - E2E assertion added: `tests/e2e/test-issue-1151-orphan-separators-e2e.js:125`, the two #2062 desktop/mobile cases. - No added requests, loops or data structures. ## Preflight overrides - External `run-all.sh` unavailable; repository syntax, whitespace, CSS-variable, XSS and PII checks run directly. - Local unit execution supplies UTF-8 settings and `GITHUB_REF_NAME=local-validation`, required by the existing release-routing test harness. - Local browser navigation budget raised to 60s for slow local asset responses; repository assertions and timeout settings unchanged. |
||
|
|
31744c6da2 |
test(channels): isolate WS assertions from initial loading (#2087)
Fixes #2086. The #1468 WebSocket browser checks could fail when initial channel loading completed between their separate before/after evaluations. An orphan message added no channel, yet unrelated loading changed the count from 0 to 5; the positive control could also lose its sentinel when loading replaced the list. Each check now captures before state, processes its packet, and captures after state in one synchronous browser evaluation. All four original assertions and both message payloads are unchanged. This updates one test file only, with no production, dependency or configuration changes. ## Validation - Delayed real API loading reproduces both failures on unchanged master. - Both corrected checks pass with loading held and with loading completed; each uses one evaluation. - Parent independently ran the delayed-loading harness and full browser suite: 131 passed, three existing fixture skips. - All 183 standalone frontend suites passed. Syntax, inventory, whitespace and privacy checks passed. TDD justification: test synchronization repair only. Existing assertions demonstrate the baseline failure; no production behavior or manufactured failing test was added. Browser checks used Chromium against unchanged production source. OpenClaw profile/external preflight were unavailable; repository checks ran directly, with a 60-second local navigation budget. Final CI must pass before merge. |
||
|
|
248d2045fd |
test(server): await indexes before node-path regression requests (#2084)
Fixes #2083. Node-path regression tests could request `/paths` while background indexes were still loading, intermittently receiving HTTP 503 instead of exercising hop resolution or sorting. Seven fixture loads now await the existing bounded `WaitIndexesReady` signal. The anchor-bias test uses the same signal instead of polling. This changes five test files only. Production readiness behavior and every HTTP/content assertion are preserved. No dependencies, configuration or customizer changes. ## Validation - Unchanged baseline: 94 passed / 6 failed across 100 targeted executions; failures were actual HTTP 503 assertions. - Fixed setup: 200/200 targeted executions passed. - Parent independently ran the entire server suite: exit 0, 83.4 seconds. - All 183 standalone frontend suites and 20 real Chromium route-map checks passed. - Formatting, whitespace and PII checks passed. TDD justification: test-fixture synchronization repair, with no production logic change. The existing unchanged behavioral assertions supplied the before/after failure proof; no fabricated failing test was added. Local browser checks used Chromium against the unchanged production source; OpenClaw profile/external preflight were unavailable. The unrelated ingestor symlink test requires a Windows privilege absent locally; Linux CI validates that suite. Final CI must pass before merge. |
||
|
|
9eb3098867 |
fix(release): publish the release notes as the release body (#2076)
Three releases went out with an empty release body. Checked with `gh
release view --json body`:
| release | body | assets |
|---|---|---|
| v3.9.1 | 1206 bytes | 2 |
| v3.9.2 | 2768 bytes | 0 |
| v3.10.1 | **0** | 2 |
| v3.11.0 | **0** | 2 |
| v3.12.0 | **0** | 2 |
v3.9.1 and v3.9.2 were written by hand, which this repository then
stopped allowing because published releases are immutable. Nothing
replaced them, so the `docs/release-notes/` convention was never wired
to the release page and three releases shipped with a description of
nothing.
Reported by the fork operator, who went looking on the v3.12.0 page for
the list of fixed issues that older releases carried.
## The cause
`action-gh-release` in this workflow was given `files` and
`fail_on_unmatched_files` and nothing else. No `body`, no `body_path`,
no `generate_release_notes`.
## The change
A step resolves `docs/release-notes/${GITHUB_REF_NAME}.md` and passes it
as `body_path`, with `generate_release_notes: true` so GitHub's pull
request list lands underneath the hand-written notes. A missing notes
file is a warning rather than a failure: the release then still gets the
generated list, which is more than an empty body.
## Also in this PR
The v3.12.0 notes file gains the two sections it should have had:
- **Issues closed**, 13 of them. Collected from each pull request's
`closingIssuesReferences`, not from commit text, which is why 30 pull
requests yield 13 issues.
- **Contributors**, separating pull request authors (@efiten 22,
@liquidraver 3, @A13xB0 2, @n30nex 1, @dborup 1, @anieto 1) from commit
co-authors (@nullrouten, @anieto, SaarMesh-Bot, Openclaw) from issue
reporters (@efiten 5, @n30nex 4, @anieto 2, @liquidraver 1,
@damn-simple-scripts 1). It says outright that 22 of the 30 pull
requests are the interim maintainer's own, which is the shape of a
release cut while the owner is unreachable, not a healthy ratio.
v3.12.0's published body has already been set to exactly this content by
hand, so the release page and the file agree.
## Worth recording
`gh release edit <tag> --notes-file <f>` updates a published release
body even though `action-gh-release` cannot. The immutability that
permanently burns a tag name does not extend to the description, so a
thin release body is recoverable. That was not obvious from the existing
comments, which describe releases as immutable without qualifying which
parts.
## Verification, and its limit
`yaml.safe_load` parses the file, but that proves little: it silently
accepts duplicate mapping keys that GitHub rejects outright, which is
how a previous workflow edit passed a local check and produced no runs
at all. The real validator is this PR's own pipeline, and the change
cannot be exercised end to end until the next tag is pushed.
## Not done
No backfill for v3.10.1 or v3.11.0. Neither has a notes file, and
reconstructing them is the same retrospective work deliberately skipped
for the 3.11.0 changelog entry.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6d5686dbb4 |
docs: release notes for v3.12.0, and the 3.11.0 entry that was never written (#2075)
Prepares the v3.12.0 release. No tag is pushed by this PR: the release procedure requires the tagged commit to carry an `:edge` image whose revision label matches it, so that check belongs after this merges, against the merge commit. ## Why minor, not patch 29 commits since v3.11.0: 15 fix, 5 test, 3 perf, **3 feat**, 2 ci, 1 chore. The feats are #2047, #2067 and #2068. ## What the notes lead with Two changes alter what a running instance does without anyone asking it to, so they are at the top rather than in a list: - **#2058**: the first start against a database that has never been `ANALYZE`d builds planner statistics and stalls ingest while it does. Measured at 3m43.9s on 9.4 GB, once per database, buffered with nothing dropped. `db.analysisLimit` set negative skips it. - **#2035**: `maxMemoryMB` eviction now fires where it previously did not, because the footprint compared against the limit was undercounted. An instance that set the limit and never saw eviction will start seeing it. ## The 3.11.0 gap `CHANGELOG.md` had no entry for 3.11.0 and `[Unreleased]` was empty, so 52 shipped commits were undocumented. Added as a short entry that says outright it was written after the fact and has no notes file. The alternative of reconstructing 52 commits for a superseded version is error-prone work with little value, and leaving the gap silent is worse than naming it. ## One fix beyond documentation `deploy.yml` carried a comment that would mislead the next person cutting a release. It still said documentation-only commits skip the workflow "(see the `paths-ignore` above)", which is exactly how v3.10.0 lost its `:edge` image and then its tag name permanently. That filter no longer exists: the `changes` job forces `code=true` for anything that is not a pull request (lines 61-73), so every master commit gets an image and a documentation commit is safe to tag. The history stays in the comment; the false present tense does not. ## Verified rather than asserted Both went into the notes as upgrade advice, so both were checked in the tree: - `node_declared_regions` is created at boot with `CREATE TABLE IF NOT EXISTS` (`cmd/ingestor/db.go:435`), added by #2047, which found the table was read by `region_keys.go` and `config.go` and created by nothing. So "no manual migration step" is accurate. - The cgo dependency attributed to #1992 in the 3.11.0 entry is the one that stops this repo building with `CGO_ENABLED=0` today. ## Not done - No tag, no release. Next steps, after this merges: confirm the merge commit's pipeline is green, confirm `:edge`'s revision label is that commit, then tag `v3.12.0` annotated, push the tag only, dispatch `CI/CD Pipeline` on the tag ref, and verify the published digests in the registry rather than in a green job. - No `docs/release-notes/v3.11.0.md`, deliberately, as above. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>v3.12.0 |
||
|
|
6e121ff8e6 |
perf(db): build planner statistics at startup when there are none (#2058) (#2074)
Follow-up to #2072, which closed #2058 but left one gap named in its own description: the refresh ticker waits 2 minutes before its first run, and a query arriving in that window against a database with no statistics gets the bad plan. Deployed to staging to measure it rather than reason about it, with `sqlite_stat1` dropped first so the build path actually ran. That changed two of the numbers in #2072, both in the expensive direction. ## The gap is once per database, not once per restart `sqlite_stat1` is an ordinary table, so once `ANALYZE` has written it the statistics stay in the file. Checked four ways: - they survive closing the connection that wrote them - a `mode=ro` handle reads them back, which is how `cmd/server` opens the database - reopening the same path through a second `OpenStore` finds them and skips the rebuild (`TestPlannerStatsSurviveReopen_Issue2058`) - on staging they survived a full redeploy to a different build that has no refresh ticker at all, and that build still gets the good plan So the window opens once, on the first start after this lands, and never again for that database. ## The cost, corrected #2072 said 2.0s. Observed on staging, 9.4 GB, commit `4500cfa6`: ``` 13:51:26 [analyze] planner statistics refresh scheduled every 24h (analysis_limit=10000) 13:55:10 [analyze] planner statistics built in 3m43.874s (analysis_limit=10000, first run against this database) ``` **3m43.9s.** Every `ANALYZE` duration in #2072's ladder was timed warm, run after run; cold, on a freshly started container, the same statement takes nearly four minutes. That is the same warm/cold split #2058 work already established for the query itself, 56.7s against 0.80s, and I then repeated it for the `ANALYZE`. Every `2.0s` in the tree is now marked warm and points at the cold figure. It holds the single write connection throughout, so ingest stalls and buffers. Per minute in `observations`: | minute | rows | |---|---| | 13:49 | 220 | | 13:50 | 106 | | 13:51 | 0 | | 13:52 | 0 | | 13:53 | 0 | | 13:54 | 0 | | 13:55 | **1027** | | 13:56 | 154 | Nothing was dropped. The burst is about four minutes of traffic at the surrounding rate, and the only ingest-buffer line in the log is the startup one reporting `0 dropped`. The cost is a four-minute write stall, once, not data loss. ## This cost is not introduced here The ticker merged in #2072 pays the identical 3m43.9s two minutes later on any database with no statistics. **Live has none, so #2072 as merged will stall live ingest for about four minutes on its first run, with or without this branch.** This only moves it earlier, into the startup burst the ingest buffer is already sized for. Flagging it on #2072 as well. ## The change `Store.EnsurePlannerStats(analysisLimit)` checks before it builds: - database has statistics: one `sqlite_master` query. This is every restart after the first. - database has none: one `ANALYZE`, and a warning first. The warning is the part that earns its place operationally. Four minutes of stalled ingest with no explanation in the log looks exactly like a hang, so `EnsurePlannerStats` now says why the write path is about to pause, what it measured on 9.4 GB, and that it happens once per database. It stays silent on a restart, because a warning on every boot would be worse than none. It runs on the refresh goroutine, not the startup path, so no boot step waits for it. `hasPlannerStats` now gates a decision instead of only wording a log line, so its comment says what the swallowed error costs: a query failure reads as "no stats", which spends one unnecessary `ANALYZE` rather than skipping a necessary one. ## Verification on staging - before: plan drove from `idx_transmissions_payload_type`, no `sqlite_stat1` - after: 50 rows in `sqlite_stat1`, plan drives from `idx_tx_channel_hash` - dropping the table first flipped the plan back, so the causality holds in both directions ## Tests 13 in the file. New here: builds when absent, skips when present, disabled on a negative limit, survives close-and-reopen, warns before building, stays quiet when statistics exist. The reopen test is the guard on the whole design: if statistics ever stopped living in the file, `EnsurePlannerStats` would quietly run a four-minute `ANALYZE` on every restart and nothing else would notice. Run locally: 13/13 on the `Issue2058` tests, `go vet` clean, `gofmt` clean, and the rest of `cmd/ingestor` green apart from `TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on `os.Symlink` with "A required privilege is not held by the client" on Windows without elevation, in a file this branch does not touch. ## Not done - No query rewrite, same as #2058 and #2072. - **No live deploy.** Live still has no statistics, so the four-minute stall is ahead of it whenever #2072 ships there. Worth picking the moment. - Staging has been returned to its own fork build; the statistics it built remain, so its next start exercises the skip path rather than the build path. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5aa3cdc41 |
perf(db): refresh SQLite planner statistics with a bounded ANALYZE (#2058) (#2072)
Closes #2058. `ANALYZE` has never run against these databases, so `sqlite_stat1` does not exist and the planner works from built-in guesses. On the channel queries it guesses wrong: it drives from the plain `idx_transmissions_payload_type` instead of `idx_tx_channel_hash`, the partial index (`WHERE payload_type = 5`) the schema already carries for that exact filter. @anieto's report did the diagnosis and the arithmetic. This adds the maintenance operation that was missing, at a value measured rather than assumed. ## The diagnosis transfers, the remedy needed measuring Measured on our 9.4 GB staging database: 1,250,489 transmissions, 14,169,329 observations, 2.7x and 7x the reported database. Region-filtered `GetChannels` produces the identical plan reported in #2058, down to both temp b-trees, so the problem is the same one. Wall time is the wrong metric here. The same query and the same plan measure **56.7s cold and 0.80s warm** on that file, so the OS page cache dominates. Counting page-cache misses instead: | analysis_limit | ANALYZE | driving index | page misses | |---|---|---|---| | none (no statistics) | - | `idx_transmissions_payload_type` | 143,442 | | 400 | 171 ms | `idx_transmissions_payload_type` | 143,449 | | 1000 | 171 ms | `idx_transmissions_payload_type` | 143,450 | | **10000** | **2.0 s** | **`idx_tx_channel_hash`** | **107,429** | | 0 (unbounded) | 242.9 s | `idx_tx_channel_hash` | 107,429 | 400, the value SQLite's documentation offers for the bounded form, changes nothing on this data: it samples too few rows to separate the 126,336-row partial index from the 920,700-row plain one. 10000 buys the entire plan change for 2.0 s, and the four-minute unbounded `ANALYZE` buys nothing beyond it. ## What it is worth, as measured 25% fewer pages read per query, 143,442 to 107,429, about 147 MB less at a 4 KB page. Warm wall time does not move: 0.80s either way. The gain lands on the cold path, the one that measured 56.7s, so the claim here is fewer pages read, not a warm speedup. This is smaller and differently shaped than the 3-4x in #2058. I cannot reproduce that ratio on a database of this size and am not claiming it. ## The change - `Store.RefreshPlannerStats(analysisLimit)` in `cmd/ingestor/db.go`: `PRAGMA analysis_limit=N` then `ANALYZE`, logging the duration and whether this was the first run. - Wired in `cmd/ingestor/main.go` next to the existing WAL checkpoint ticker: 24h, staggered 2 minutes past startup because it takes the write lock. - `db.analysisLimit` in `internal/dbconfig`, default 10000, negative disables it. It runs in the ingestor, not the server: `cmd/server/db.go:145` opens `mode=ro`, and `ANALYZE` writes. This respects the read/write separation invariant in AGENTS.md. `analysis_limit=0` means *no* limit to SQLite rather than "use a default", so an unset config maps to 10000 and a test covers that specific case. ## Two faults the measurement caught in my own first commit Both are in the history rather than hidden, because the second commit is the one that measured: 1. **`PRAGMA optimize` was the wrong statement.** It analyzes only tables the calling connection has itself queried during the session, and a maintenance call has queried none. Run against staging it wrote nothing and left `sqlite_stat1` absent; `PRAGMA optimize(0x03)` returned no statements at all. Verified on an empty database too (SQLite 3.45.1): `ANALYZE` creates `sqlite_stat1`, `PRAGMA optimize` does not. That difference is what makes the behavioural test a guard instead of a no-op. 2. **`analysis_limit=400` was the wrong value**, per the table above. ## Tests Six cases in `cmd/ingestor/refresh_planner_stats_test.go`: - statistics are actually written (the guard against returning to `PRAGMA optimize`) - the pragma reaches the connection, read back through `PRAGMA analysis_limit` - a negative limit leaves `sqlite_stat1` absent - two consecutive refreshes, since a ticker calls this repeatedly - the config default and the JSON round trip - the default is above the range measured ineffective, so lowering it back to 400 fails ## Not done, and one caveat - **Correction to an earlier version of this description**, which said the Go tests could not run locally because this box has no C compiler. That was wrong: `CGO_ENABLED=0` and `gcc` being absent from `PATH` is not the same as no compiler, and a mingw-w64 toolchain is installed here. Run properly, all six tests pass locally, and so does the rest of `cmd/ingestor` apart from `TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on `os.Symlink` with "A required privilege is not held by the client" on Windows without elevation and lives in `stats_file_test.go`, a file this branch does not touch. CI agrees: Go Build & Test and the ingestor race detector are both green. - The SQLite behaviours above were measured against 3.45.1 on the server, not against the amalgamation `mattn/go-sqlite3` bundles. - **No query rewrite.** #2058 explicitly left that out and so does this; the correctness caveats it lists (per-channel most-recent-message semantics, v2/v3 branches, `enc_` exclusion) are untouched here. - Staging carries limit-10000 statistics, matching what this code produces. Reversible: `DROP TABLE sqlite_stat1` was verified on a scratch database before any of it ran. - Whether a cold-start `ANALYZE` should also run before the 2 minute stagger is not addressed. The first query after a restart is the expensive one, and it can arrive first. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3caa847323 |
fix(nodes): the advert section is called "Recent Adverts", and says why (#2071)
Closes #2042, reported by @damn-simple-scripts. ## Which of the two fixes The report offered widening the section to all packets the node originated, or renaming it, and proposed the rename. Widening sounds like the better fix, so I measured before agreeing. On a production database: | | | |---|---| | transmissions total | 1,251,967 | | with `from_pubkey` populated | 191,237 | | of those, `payload_type = 4` (ADVERT) | **191,237** | | with `from_pubkey` and **not** an advert | **0** | The ingestor fills that column for adverts only. Attributing a relayed CHAN or TXT packet back to its sender is the path-resolution problem, not a filter this section could apply — so "show all packets from this node" is not a small change, it is a different feature resting on attribution the data does not carry. So the rename is correct, and the numbers say so rather than my preference. ## The change Both copies renamed — the full node page and the side pane. Both already read `nodeData.recentAdverts`, so the field feeding them said "adverts" while the heading said "packets". The heading also gained a `title` naming `from_pubkey` as the reason it is adverts only. Renaming without explaining invites the same report from the next reader; the tooltip is where that explanation costs nothing. ## Tests `tests/unit/test-issue-2042-recent-adverts-label.js`, four cases: both copies present and headed "Recent Adverts", no copy back to "Recent Packets", the advert field still feeding them, and the explanation still in place. The third matters most — it ties the label to its data source, so pointing this section at a different field in future fails here rather than silently making the label wrong again. ## Not done, deliberately The reporter's follow-up idea, splitting flood adverts from zero-hop adverts into two lists. They wrote that it can be dropped if an issue should tackle one thing, so it is not here. It is feasible now: a zero-hop advert is `ROUTE_TYPE_DIRECT` with an empty path, established while fixing #2064. That deserves its own issue rather than a paragraph in this one — say the word and I will open it. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2cfe9cbbc5 |
fix(a11y): raise the Scope Audit observed-chip contrast above 4.5:1 (#2070)
Closes #1996. ## Reproduced first The chips tint their own background — `color-mix(in srgb, var(--status-green) 16%, transparent)` — which darkens whatever surface sits behind them. With `--status-green-text` (green-700, `#15803d`) on top: | surface | composited chip | ratio | |---|---|---| | `--surface-0` `#f4f5f7` | `#d2eddf` | **4.04:1** | | `--card-bg` `#ffffff` | `#dcf6e5` | **4.38:1** | Both under the 4.5:1 the #1719 gate requires for normal text. The first row is the same ratio **and the same hex** @n30nex measured in the browser, which is how I know the model in the test agrees with what a visitor actually sees. ## Fixed on the text, not by thinning the tint Dropping the tint to 12% reaches 4.52:1. That is two hundredths above the line, and a margin that thin fails again the next time a surface value moves — the same trap #2039 hit with a threshold sitting inside the healthy band. One palette step darker gives **6.23:1** on white and **5.74:1** on `--surface-0`, and keeps the tint that makes a chip read as a chip rather than as plain text. - `--palette-green-800: #166534` added. The greens ran 300–700 while the blues already reach 900, so this fills the scale rather than inventing a colour. - `--sa-chip-observed-fg` defined per theme: green-800 in light, and the bright `#22c55e` **kept** in dark, where the chip already passed at 6.48:1 / 5.73:1. Dark is deliberately untouched. - The chip reads the variable, so the customizer still governs it and no literal enters a component. `.sa-chip-verified` needed no change: `scope-audit.js:106` only ever adds it alongside `sa-chip-observed` or `sa-chip-unobserved`, and it contributes an underline. So the fix covers both classes the issue names — worth stating, since the title mentions both. ## Tests `tests/unit/test-a11y-1996-scope-audit-chips.js`, 7 cases: both themes × both surfaces, that the chip still tints (so the suite cannot pass by testing nothing), that its colour comes from a variable rather than a literal, and a guard asserting green-700 **would** still fail — so a quiet revert to `--status-green-text` turns this red instead of passing. **Red-run confirmed:** with the old colour restored it fails at exactly 4.04:1 and 4.38:1, naming the composited `rgb(210,237,223)`. Two deliberate choices in the test: - **A separate suite, not a case in `test-a11y-1719`.** That suite's `parseColor` handles hex and `rgb()` only; teaching it `color-mix()` is a larger change than this fix. These chips are the only contrast-critical user of the function today. If a second appears, the two should merge, and the file says so. - **Derived from the stylesheets, not a rendered page.** The default Scope Audit fixture renders no chips at all, which is precisely why the existing browser coverage missed this. A stylesheet-derived check cannot be defeated by a fixture that shows nothing. `check-css-vars` passes (180 definitions, 0 undefined) and `test-test-inventory` passes with the new file classified. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d14317b008 |
fix(qa): query blacklist retention by from_pubkey (#2069)
## Summary - query `transmissions.from_pubkey`, the column present in CoreScope's real schema, instead of the nonexistent `from_node` - preserve the existing SQLite parameter binding and stdin transport - derive the unit-test table from the ingestor's committed `CREATE TABLE` definition rather than a hand-written schema - add positive, negative, case-normalization, injection-shaped input and missing-column coverage - verify the query against the committed staging-captured E2E fixture ## Why The parameter-binding change correctly removed SQL interpolation, but its count query and test fixture both used `from_node`. The production schema defines `transmissions.from_pubkey`; therefore the live QA probe could only fail with `no such column: from_node`, while the synthetic unit fixture remained green. The ingestor stores attributed ADVERT pubkeys as lowercase hex. The query now uses: ```sql SELECT COUNT(*) FROM transmissions WHERE from_pubkey = lower(:pubkey); ``` The value remains a bound parameter. No production schema or runtime code changes. ## Verification - `bash -n qa/scripts/blacklist-test.sh` - `bash -n qa/scripts/test-blacklist-sql.sh` - `bash qa/scripts/test-blacklist-sql.sh` — 72 passed, 0 failed - schema fixture extracted directly from `cmd/ingestor/db.go` - committed E2E fixture returns the same attributed-row count through the helper query and a direct control query - legacy fixture containing only `from_node` fails non-zero with `no such column: from_pubkey` - injection-shaped, empty, whitespace, multibyte and long values remain literal bound values - `git diff --check` This PR intentionally contains only the schema correction and its regression coverage. Co-authored-by: Openclaw <openclaw@Openclaws-Mac-mini.local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
435ac25dd6 |
feat(nav): show the running version in the nav-drawer footer (#2068)
Takes over #1985 by @SaarMesh-Bot, as offered there on 2026-09-16 and 2026-09-17. **Their commit is the first here, unchanged and under their authorship**; the rest clears the two review points. Closing #1985 in favour of this so the rebase and the fixes travel together, not to reassign the work. ## The feature, unchanged The frontend never surfaced which build was running, though `/api/health` has reported `{version, commit, buildTime}` all along. A footer on the nav drawer now renders `CoreScope <version>`, linking to the releases page, with commit and build time in the tooltip. Colours come from existing CSS variables, the label is set with `textContent`, and a failed health call leaves a neutral label rather than an empty footer — all as the author wrote it. ## Review point 1: fetch on open, not on page load The `/api/health` call sat in `buildDom()`, which runs on page load. The drawer may never be opened, and **cannot** be opened at ≤768px, where the module is disabled by design. So every visitor's browser was requesting the endpoint to fill a footer most of them would never see. Moved into `open()`, after the width gate. `fetchVersion` still caches its promise for the page lifetime, so re-opening costs nothing. ## Review point 2: the fallback is pinned `tests/unit/test-nav-drawer-version-footer.js`, six cases: - a rejected fetch, a non-ok response, and a 200 without `version` each leave the neutral `CoreScope` label — not a blank footer and not `CoreScope undefined`, which is what an instance shows exactly when someone is trying to read its version - a normal response renders the version, with commit and build time in the tooltip - a version carrying markup lands verbatim in `textContent`, and `innerHTML` is never touched - the endpoint is requested once however often the footer is filled, pinning the cache from the other side It **slices `fetchVersion` and `fillVersion` out of the shipped `public/nav-drawer.js`** and evaluates them rather than copying them into the test, so it exercises what ships — same approach as `tests/unit/test-direct-rf-heard-by.js`. Both slice markers are asserted, so a rename fails loudly instead of quietly testing nothing. Each case re-evaluates the slice, because `versionPromise` caches for the page lifetime and a shared sandbox would hand the second case the first case's answer. Listed in `test-all.sh`, which per #2036 is the only frontend runner. ## A note on the third commit The first version of the test used `setImmediate` to settle the promise queue. It ran fine under node and **failed eslint**, which treats these as browser code. I ran eslint and committed in the same command and pushed without reading its output. Fixed in the commit after, using `setTimeout(r, 0)`. Recording it because the PR would otherwise show a lint failure in its history with no explanation. ## Verification 6 of 6 in the new suite, eslint clean on both changed files, `test-test-inventory.js` passes with the new file classified. Cherry-picked cleanly onto current master. --------- Co-authored-by: SaarMesh-Bot <bot@saarmesh.de> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
291393dcc0 |
feat(ingestor): log a throttled warning when the IATA whitelist drops a region (#2067)
Takes over #2008 by @nullrouten0, as offered there on 2026-09-16 and 2026-09-17. **Their commit is the first of the two here, unchanged and under their authorship**; the second is only the fix for the one blocker. Closing #2008 in favour of this so the rebase and the fix travel together, not to reassign the work. ## The feature, unchanged `observerIATAWhitelist` dropped non-whitelisted regions silently. An allow-list fails in the dangerous direction: a legitimate but unlisted region vanishes with nothing to show for it. One line per dropped region now, re-logged at most every `iataWarnIntervalSec` (new optional key, default 6h) for as long as that region keeps arriving. The periodic re-log rather than a strict log-once is the author's call and it is the right one: a single edge event rolls out of any scrape window, leaving an actively-dropping region indistinguishable from a healthy one. ## The blocker, now fixed `ShouldWarnIATADrop` keyed its throttle map on a topic segment the **publisher** controls and never evicted it — a remote memory sink. Measured on the original branch: 200,000 distinct codes retained 200,000 entries and 15.1 MB of heap. My review offered two shapes. This takes the cap rather than shape-validation, and the reason matters: **nothing in this codebase constrains an IATA code's shape.** It is uppercased and trimmed in `config.go` and `db.go` and never validated. Rejecting by shape would invent a rule operators have not agreed to, and would silently drop the warning for anyone whose code does not fit it — the same failure mode, one level down. So `iataWarnMaxTracked = 512`: far above any real deployment (the reference instance runs 43 observers across a handful of regions) and small enough that a hostile feed gains nothing. **Past the cap the drop is still logged**, throttled on one shared timestamp instead of a per-code one. Swallowing it there would reintroduce exactly the silent failure this feature exists to fix. ## Tests The author's `iata_drop_warn_test.go` plus three: - the map stops growing when fed 2048 distinct codes - a new code past the cap still warns once, is then throttled, and speaks again after the interval elapses - an already-tracked code's throttling is unchanged, so the cap does not alter the normal path ## Verification `gofmt` clean, cherry-picked cleanly onto current master (`cmd/ingestor/main.go` auto-merged). Go tests not run locally: no cgo toolchain here since #1992, and per AGENTS.md `CGO_ENABLED=0` links a stub that proves nothing. CI is their first run. ## Not done The `iataWarnIntervalSec` key is undocumented outside the struct comment. If there is a config reference that should list it, say where and I will add it. --------- Co-authored-by: nullrouten <nullrouten@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
50c4d9615b |
test(ingestor): wait for the boot migrations before handing over a test store (#2066)
Closes #2065. Master's `🏁 Race detector (ingestor)` job has been red since the pushes at 2026-09-22 21:55 and 21:56. ## Correcting my own diagnosis The issue says the fault is a goroutine outliving its test and racing a later one, and proposes making it joinable. That was wrong, and it matters because it changes the fix: `Close()` **already** waits on `backfillWg` (`cmd/ingestor/db.go`), and `newTestStore` registers it as `t.Cleanup`. The goroutines are joined before the next test starts. The race is inside a single test. `OpenStore` schedules two async migrations — `obs_observer_ts_idx_v1` and `tx_last_seen_backfill_v1` — whose goroutines log while they run. `TestHandleMessageDecodeErrorLog_PII_Issue1211` then points the standard logger at a `bytes.Buffer` and reads it, so its **own** store's migrations write into the buffer it reads: ``` Write by goroutine 760: RunAsyncMigration.func1 async_migration.go:124 (log.Printf) Read by goroutine 757: ...PII_Issue1211 decode_error_log_test.go:37 (buf.String) ``` ## Why the helper rather than the one test These tests capture the standard logger in **21 places across 7 files**. Any of them that also builds a store is exposed to the same thing; the decode-error test is just the one whose timing lost. So `newTestStore` now waits after `OpenStore` instead of only at cleanup, and no test body can run while a migration is in flight. Checked before touching a shared helper: no test references either boot migration by name, and the `pending_async` assertions in `async_migration_test.go` use their own names with a blocking `fn`, so they are unaffected. Cost is a few milliseconds against an empty temp database. ## Tests The race detector only catches this when the scheduler cooperates — it sat latent from 2026-09-03, when those files were last touched, until it surfaced three weeks later, and a re-run would have made it look like a flake. So both new tests are deterministic: - **`TestNewTestStoreWaitsForBootMigrations`** — `tx_last_seen_backfill_v1` is scheduled unconditionally by `OpenStore`, so on a fresh temp database it is pending at that instant and can only read `done` if something waited. Remove the wait and this fails every run. - **`TestCapturedLogIsFreeOfMigrationOutput`** — asserts a captured buffer holds no `[migration/async]` or `[async-migration]` output, which is the failing test's own situation stated as an assertion. ## Verification `gofmt` clean. Go tests were not run locally: no cgo toolchain on this machine since #1992, and per AGENTS.md `CGO_ENABLED=0` links a stub that proves nothing. CI is their first run, and the race-detector job is the one that matters here. ## Not done The decode-error path still logs through the standard logger, so a future test capturing it while any other goroutine logs will race again. An injectable logger would close that class properly. This closes the store-boot case, which is the one that exists today. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6d3da77b67 |
perf(channels): coalesce concurrent GetChannels/GetEncryptedChannels cache misses (#2059)
Fixes #2029. ## What was wrong `GetChannels`/`GetEncryptedChannels` (`cmd/server/db.go`) cache their region-scoped result for 60s but had no request coalescing on a cache miss, so every request that arrived while the cache was cold or expired ran the region-scoped `GROUP BY` scan itself. Measured on production: 5 concurrent requests for the same never-cached region each took ~8s, no cheaper than 5 independent runs. `statsSF`/`regionMembershipSF` already fix the identical bug class elsewhere in this file (#1910), so this wraps both functions' query-build/execute/cache-populate block in a `singleflight.Group` the same way, keyed per region, double-checking the cache inside the flight in case a previous winner already refreshed it. ## Tests The review on #2029 pointed out that a timing-based check ("finish within ~2ms of each other") doesn't actually prove coalescing happened — it would pass on a fast machine even without singleflight. `db_channels_singleflight_test.go` uses a call counter instead, same pattern as `TestEnsureNeighborGraph_Singleflight` (#1203 Pair A): - `TestGetChannels_SingleflightCoalescesQueries` / `TestGetEncryptedChannels_SingleflightCoalescesQueries`: 10 concurrent callers against a cold cache, asserting the real query runs exactly once. A test-only hook (`channelsQueryHook`/`encChannelsQueryHook`, nil in production, same contract as `bgLoaderEntryHook`) increments the counter right where the query executes, since these functions hit `db.conn.Query` directly rather than going through an injectable builder function. - `TestGetChannels_SingleflightPerRegion`: two regions queried concurrently (5 callers each) assert 2 queries, not 1 — pins that the flight is keyed per-region and a caller for one region can't receive another region's coalesced result. Anti-tautology: reverting `channelsSF.Do`/`encChannelsSF.Do` back to a bare call makes the coalescing tests observe N instead of 1. `go build ./...`, `go vet ./...`, `gofmt -l .` clean. Full `cmd/server` suite (race-enabled for the new concurrency tests) run in a `golang:1.22-alpine` container, mounted repo, workdir `cmd/server` so the sibling `internal/*` replace directives resolve: ``` === RUN TestGetChannels_SingleflightCoalescesQueries --- PASS: TestGetChannels_SingleflightCoalescesQueries (0.06s) === RUN TestGetChannels_SingleflightPerRegion --- PASS: TestGetChannels_SingleflightPerRegion (0.05s) === RUN TestGetEncryptedChannels_SingleflightCoalescesQueries --- PASS: TestGetEncryptedChannels_SingleflightCoalescesQueries (0.07s) ``` Full suite: `FAIL github.com/corescope/server 98.261s`, but the only failures are `TestHandleNodePaths_PrefixCollision_1352`, `TestHandleNodePaths_FallbackUniquePrefix_1352`, and `TestHandleNodePaths_FallbackUnresolvableHop_1352`, all failing on a `503 {"error":"index loading","retryAfter":5}` — an index-build race in this container's timing, not this change. Confirmed by running the same three against an unmodified, freshly-cloned `master` in the same container: they fail there too (plus `TestHandleNodePaths_PrefixCollision_1352_FallbackBranch`, which this run happened not to hit). Nothing in this diff touches node-path handling. ## Not done The deeper query-plan issue flagged in #2029 (the outer scan is driven by `payload_type`, not region, so a cold solo request still costs several seconds regardless of concurrency) is filed separately as #2058, with `EXPLAIN QUERY PLAN` output and row counts against production data. Coalescing makes one slow query serve everybody; it doesn't make the query itself fast. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: anieto <anieto@meshtexas.org> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b695b979a2 |
fix(nodes): keep paginating past a page that post-LIMIT filtering shortened (#2061)
## Problem `handleNodes` runs the geo-filter, `nodeBlacklist`, `hiddenNamePrefixes` and area passes **after** the SQL `LIMIT/OFFSET`, and rewrites `total` to the filtered length. A page that loses a row is therefore short **without being the last page**, and neither the page length nor `total` can tell a client whether to ask for another page. #1606 added the pagination loop and chose the page length as the canonical stop. That is correct only where nothing is ever filtered. Everywhere else the list truncates at the first filtered page boundary and strands every node behind it — the #1598 symptom reached by a different route: a node that is relaying right now simply stops being in the list. The comment at `app.js:240` rejects `total` for exactly the right reason, then picks the signal the same code path also breaks. ## Measured on a live 2346-node deployment Page sizes for the query the map issues: ``` offset=0 returned=500 ← full, loop continues offset=500 returned=499 ← one row filtered AFTER the LIMIT → loop STOPS offset=1000 returned=500 ← never requested offset=1500 returned=500 ← never requested offset=2000 returned=345 ← never requested ``` | stop rule | requests | nodes reached | |---|---:|---:| | short page (master) | 2 | **999** | | `has_more`, else empty page | 6 | **2344** | **1341 nodes, 57%, unreachable through the UI.** ### One hidden node truncates the whole list The deployment this came from has no `geoFilter` (`/api/config/geo-filter` returns `polygon: null`) and no `nodeBlacklist`. It has a single `hiddenNamePrefixes` entry — a deliberate operator choice — and exactly one node whose name starts with it: ``` public_key d4a46ea2…1054 (64 clean hex chars) name 🚫🔥☀️ role repeater last_seen 2026-09-22T10:09:38Z ``` `handleNodes` drops that row in the `IsNameHidden` pass, which runs after the SQL `LIMIT`. The row is counted by the `LIMIT` and by `COUNT(*)`, so the page it lands in comes back exactly one short — and stops every client that treats a short page as the end. Isolated against SQL on the same database, seconds apart: ``` SELECT lower(public_key) FROM nodes ORDER BY last_seen DESC LIMIT 500 OFFSET 500 -> 500 rows GET /api/nodes?limit=500&offset=500 -> 499 rows comm -23 sql.txt api.txt -> d4a46ea2e99cab132a3286ef3d9cce9099318790af7f25671fe83de453721054 ``` Deterministic — `offset=500` returned 499 on three consecutive requests. Not a CDN artifact either: `cf-cache-status: DYNAMIC`, origin `cache-control: no-store`, no `age` header, and four requests with deliberately unique cache keys all returned 499. So **one deliberately hidden node makes 1341 of 2344 nodes unreachable.** The hiding feature does exactly what it was asked to do for that one node, and takes 57% of the network with it, silently. A single `hiddenNamePrefixes` entry is enough; no geo-filter, blacklist or area filter is needed to reach this state. ### The cutoff moves, which is why this reads as intermittent The visible set is the sum of the pages up to and including the first short one, so the boundary sits wherever the unreturnable row currently sorts by `last_seen`, and jumps a whole page as ingest reorders the list. Same deployment, same code, same config, ~2h apart: | dropped row's rank | first short page | nodes visible | |---|---|---:| | inside 0–499 | page 1 | 499 | | inside 500–999 | page 2 | 999 | A node is visible or invisible purely by where it lands relative to that moving line, so affected nodes appear to vanish and return on their own. Two operators on this deployment reported exactly that, independently, while I was measuring. ### A named reproduction `HU-ZA-Lentihegy` (`5287a33f…`), reported missing from the map by an operator whose companion had logged its advert at 04:20 local the same morning. Ingest was fine. The row is in `nodes` with `last_seen` `2026-09-22T02:20:35Z` — the same advert, to the second — valid GPS, role `repeater`, 1033 adverts, and `/api/nodes/search?q=lentihegy` returns it. ``` rank by last_seen : 1081 cutoff at the time: 999 ``` It missed by 82 positions. Walking the same live endpoint, same moment: | stop rule | requests | nodes reached | Lentihegy | |---|---:|---:|---| | short page (master) | 2 | 999 | **not reached** | | `has_more`, else empty page | 6 | 2340 | reached | The practical shape of this on a busy mesh: 1081 nodes had been heard more recently than 9.4 hours, so on that deployment **anything last heard more than ~9 hours ago was invisible**, alive or not. `#/nodes` compounds it — its search box filters client-side over the truncated set, so the server-side `?search=` never runs and an operator cannot find the node by searching for it either, even though the endpoint would return it. ## Change **Server** — `NodeListResponse` gains `has_more`, computed from the raw SQL page against the real `COUNT(*)` before the filter passes run, so it survives them: ```go hasMore := offset+len(nodes) < total ``` Always emitted (no `omitempty`) so a client can tell `false` from an old server. No extra request in the fixed path: `has_more` ends the loop exactly, where the old rule needed a probe page. **Clients** — `app.js` `fetchAllNodes`, `nodes.js` `loadNodes` and `area-map.html`'s inline helper stop on `has_more`, falling back to a zero-length page against a server that predates it. An empty page always ends the loop, so a `has_more` against a concurrently-shrinking table cannot spin to `safetyCap`. Left alone: the three loops are still three copies. Collapsing them onto `fetchAllNodes` is a bigger change than this fix needs, and `nodes.js` has its own inter-page progress UI. Happy to do it separately if you want it. ## Testing - **Unit** (`tests/unit/test-fetch-all-nodes-pagination.js`): the fixture now models the real handler — a row counted by the LIMIT and by `COUNT(*)`, then removed from the page. Three new cases. Fails on the old rule at 499 of 1199. - **E2E** (`tests/e2e/test-map-nodes-pagination-e2e.js`, already wired into `deploy.yml`): the mock drops a page-1 row and emits `has_more`. Mutation-checked — restoring master's stop rule fails 3 of its steps. - **Go** (`cmd/server/nodes_pagination_has_more_test.go`): asserts `has_more` stays true on a page filtering shortened. Mutation-checked — recomputing it after the filter block fails the test. - Full server suite `go test -race`: ok, 41.4s. `gofmt` clean, `go vet` passes. - **Against a real binary**, not just mocks: fixture DB migrated with `corescope-migrate`, `hiddenNamePrefixes: ["SKCE"]`, `limit=3`. Page 1 returns 2 of 3 with `total` rewritten to 2 and `has_more=true`. Walking the real server with master's rule reaches 2 nodes; with `has_more`, all 199 visible of 200, the hidden one still hidden. The real frontend against that server loads 199 with no JS errors. Two existing expectations changed, both deliberate: 1. `surfaces ALL nodes past the 500 server cap` — 3 → 4 requests. That mock emits no `has_more`, so the 200-row final page can no longer end the loop (a short page is exactly what a filtered page looks like) and a zero-length probe follows. Against a current server `has_more` still ends it at 3. 2. `rows missing public_key are NOT collapsed into one` — its stub returned a constant body, which would now be paged to `safetyCap`. It serves one page then empties. Local `test-all.sh` exits 1 on two XSS-gate self-tests (`good-2-tested.js`, `good-4-tested.js`) via a `UnicodeEncodeError` printing an emoji under Windows cp1252. Identical on clean `origin/master` in a scratch worktree, so it is pre-existing and platform-local, not this branch. There is a second identical filter block further down `routes.go` on another list endpoint. Likely the same class; not touched here. If you would rather land your own version of this, say so and I will close mine. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
980c5c4515 |
fix(nodes): stop the Heard By empty state claiming the node is out of range (#2063)
Follow-up to #2057, which is correct in what it does and overclaims in one sentence. ## The sentence When no observer heard the node directly, the card said: > No observer is within radio range of this node. #2057's own rule cannot establish that: - every **direct route** is discarded, because the firmware removes the sender from the path before retransmitting (`Mesh.cpp`, `removeSelfFromPath`), so the packet cannot say who transmitted it. #2057 measured these at **38% of transmissions over 7 days**. - **28.8% of flood observations with a path** are dropped because the last hop resolves to more than one candidate, and the rule under-attributes rather than guesses. So an empty direct list is missing evidence, not evidence of missing coverage. ## Why it matters in practice Sampled 40 repeaters on a production instance after #2057 shipped: | | | |---|---| | at least one direct observer | 24 | | empty card, "no observer is within radio range" | **16** | | of those 16, with a non-zero `relayObserverCount` | **16** | Every node showing "nobody is in radio range" also showed "Seen via relay by N observers" two lines below. An operator reading that about a working repeater concludes they have a coverage problem they do not have. ## The change Wording only, in both copies of the card (full page and side pane): > No observation proves a direct reception here, which is not the same as being out of range. with the reason in a `title`, so the card stays one line: > Only flood-routed transmissions identify who was heard: a direct route removes the sender from the path before retransmitting (firmware `Mesh.cpp`, `removeSelfFromPath`), and an ambiguous relay hop is left unattributed rather than guessed. So an empty list is missing evidence, not proof of missing coverage. No API change. #2057's rule, shape and performance work are untouched — I verified its firmware derivation against the clone at `0679dbef` before writing this: `Packet.h:83`, the forwarder appending with `packet->getPathHashSize()` at `Mesh.cpp:349`, and `removeSelfFromPath` on the direct path at `Mesh.cpp:89-105` all read as described. ## Test `tests/unit/test-direct-rf-heard-by.js` slices this template out of `public/nodes.js`, so it pins the shipped markup. It now asserts the old sentence is gone and the qualifier is present. Worth recording how that assertion was reached: my first version banned the phrase "out of range" from the card, and it failed — on the new line, which contains that phrase precisely in order to deny it. A word ban was the wrong instrument. Matching the old sentence and requiring the new qualifier is the assertion that actually distinguishes the two states. 8 of 8 in that suite, eslint clean. ## Not in this PR The **Regions** line and **Region** column on the same card read `o.iata`, which `HealthObserverRow` does not emit, so both have always been dead. #2057 named this and left it; it is now **#2062** with the file and line references, rather than a remark inside a merged description. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
61f565c606 |
fix(node-health): credit zero-hop adverts as direct reception (#2064)
Reported by @dborup on #2057. The mechanism is real; the suspected scale is not. Both parts measured below. ## The defect `directHeardNode` rejects every non-flood route type before it looks at the path, so the `hop == ""` branch that credits an advert's originator can never run for a direct route. Zero-hop adverts — the clearest direct-RF evidence the network produces — are discarded and the node is listed under "Seen via relay" instead. ## Firmware Read at `0679dbef` rather than taken on trust: - `Mesh::sendZeroHop` sets `ROUTE_TYPE_DIRECT` and `path_len = 0`, commented there as "path_len of zero means Zero Hop". The transport overload does the same with `ROUTE_TYPE_TRANSPORT_DIRECT`. - `examples/simple_repeater/MyMesh.cpp` sends the periodic **local advert** through it (the `next_local_advert` branch), as does `sendSelfAdvertisement` when `flood` is false. - `examples/companion_radio/MyMesh.cpp` does the same for a companion's own advert. An ADVERT arriving on a direct route with an empty path therefore cannot have been forwarded: the observer received the advertiser's own transmission, and the advert carries its pubkey in the clear. Every other direct case keeps #2057's rule. A non-empty path on a direct route is the **remaining** route, because the forwarder ran `removeSelfFromPath` before retransmitting, and `advertOriginPubkey` already returns `""` for any payload type other than ADVERT — so the payload guard costs nothing. ## Measured, 7-day window on a production instance | | | |---|---| | zero-hop advert observations currently dropped | 8,166 | | distinct nodes they evidence | 139 | | node-observer pairs they evidence | 185 | | pairs **not** already credited via an empty-path flood advert | **52** | | pairs currently credited from flood adverts | 217 | So the fix restores 52 node-observer pairs of direct evidence that are invisible today, roughly a quarter more advert-based direct evidence. ## What it does not explain The report suspected this accounts for #2057's low headline numbers ("NL-BXE-RP01 | 433 → 0", 234 of 1,860 nodes with any direct observer). The measurement does not support that: - Of 30 sampled nodes with zero-hop advert evidence, **29 already show at least one direct observer**, because they also send flood adverts which #2057 credits. - **NL-BXE-RP01 | 433 has zero zero-hop adverts** in the window. Its empty list is not caused by this rule. So this mostly enriches lists that are already non-empty, and flips few cards from empty to populated. Worth doing on correctness grounds, not as a fix for the counts. ## Tests Four cases added to the `TestDirectHeardNode` table, which previously covered direct routes only with `PayloadTXT_MSG`: - direct + ADVERT + empty path credits the advertiser - transport-direct + ADVERT + empty path credits the advertiser - direct + empty path + **not** an advert credits nobody - direct + ADVERT + **non-empty** path credits nobody The last two matter as much as the first two: they pin the exception to exactly the shape the firmware guarantees. `gofmt` clean. Go tests not run locally (no cgo toolchain on this machine since #1992, and per AGENTS.md `CGO_ENABLED=0` builds a stub that proves nothing), so CI is their first run. ## Related #2063 fixes the empty state's wording on the same card, which asserts the node is out of range when the data cannot establish that. The two are independent: this one adds evidence, that one stops overclaiming when there is none. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
25f8426d32 |
test: make every suite in tests/e2e run, and tell the truth when it does (#2053)
Closes #2037. Step 1 (#2045) wired in the nine suites that already passed. This is steps 2 and 3: the four that ran and failed, and the five that could not run at all. After this, the count of suites in `tests/e2e` invoked by nothing goes from 18 to 0. ## Step 2 — the four that ran and failed Triaged against a fixture server in CI with full output kept, not the four-line tail the first probe saved. | suite | verdict | |---|---| | `test-packets-scope-column.js` | passes on master today. It failed on 2026-09-18, so something between the two fixed it. Wired in unchanged rather than investigated. | | `test-node-reach-e2e.js` | test wrong. It waited for the reach map whenever any *link* had GPS; `public/node-reach.js` builds the map only when the *node* has coordinates. Guaranteed 10s timeout on a node with positioned neighbours and no position of its own. | | `test-channel-modal-e2e.js` | both failures test-side. The Add button's visible label was shortened to "+ Add" with the accessible name moved to `aria-label` (`public/channels.js:746`); and `.ch-section-mychannels` is conditional on the visitor having added a channel, which has not happened at that point in the suite. | | `test-touch-targets.js` | three test-side, two a product finding. | The three test-side touch-target failures were the harness measuring controls that are not shown: `.compare-btn` (the CTA was removed in #1646, and `style.css` says so), `.ch-back-btn` (`display:none` outside the mobile channels layout), and `.filter-toggle-btn` (`display:none` on mobile since #1461; the control shown is the navbar mirror, which `mobile-page-actions.js:70` builds as a `.nav-btn`, so it was already measured). All three are dropped from the table with the reason recorded in the file. The remaining two are **not** a test problem: `.nav-btn` and `.ch-icon-btn` are each declared twice in `public/style.css`, 48px in the touch-target block and 44px in their own component rule, and the later one wins. Rather than lower the blanket or hide the failures, the suite now has `DEFAULT_MIN = 48` plus a `MIN_OVERRIDES` table holding those two at their effective 44, so a third selector dropping to 44 still fails the build. The contradiction is **#2052**, with both ways out costed; the override entries should go when it is settled. ## Step 3 — the five that could not run None is deleted. I checked each selector and seam against the product before deciding, and every one still targets a surface that exists and that nothing else covers. Four were written against `@playwright/test`, a runner the project neither installs nor uses anywhere else. Adopting a second runner for twelve tests costs more than porting them, and the precedent is already set: `test-path-inspector-coverage-e2e.js` exists, as its own header says, because `test-path-inspector-e2e.js` could not run. So they are ported to the plain-node Chromium pattern the other 109 suites use. - **`test-issue-1522-trace-url-sync-e2e.js`** — the trace hash in the URL, both directions. `test-e2e-playwright.js` covers that the page loads and searches; it never looks at the URL, which is the whole of #1522. - **`test-marker-outline-weight.js`** — the canvas pulse ring never thins below 2px. There is no CSS rule to read and axe cannot see inside a canvas, so sampling the seam is the only way. Added a guard that the ring was actually visible, so the weight check cannot pass vacuously on a pulse that never rendered. - **`test-pr-1490-live-map-gpu-animations-e2e.js`** — the queue drains, the engine sleeps again, the fading trails stay under the cap of 5, and the canvas sits on `animationsPane` rather than under the markers. - **`test-path-inspector-e2e.js`** — reduced to what nothing else covers: the map side pane, the `/#/traces/<hash>` redirect, the tools landing. Its standalone-page test duplicated the wired coverage suite and is dropped. Its "switching candidate clears prior polyline" case ended after the click with a comment and no assertion, which is the same green-but-empty problem this issue is about; it now compares path counts, and skips loudly when the fixture yields too few candidates. The fifth, **`test-table-sort.js`**, needed `jsdom`, which was declared nowhere. It is a unit test of `public/table-sort.js` filed under `tests/e2e`, so: `jsdom` is a devDependency (lockfile updated, `npm ci` stays consistent), the file moved to `tests/unit/`, and the `domIntegration` group in `scripts/non-unit-tests.json` is gone with its only member. It runs 22 tests. 20 passed immediately; 2 had rotted, because #1648 M2 replaced the up/down glyphs with Phosphor sprites and the direction moved out of `textContent` into the `<use href>`. Those two now read the sprite ref and the `aria-sort` value, so they also guard the accessible announcement. ## Verification `tests/unit/test-table-sort.js` 22/22 and `test-test-inventory.js` pass locally; the E2E suites need a fixture server, which I cannot build here (no cgo toolchain since #1992), so CI is their first run as committed. The triage above was measured in CI, not assumed. ## Not done The per-assertion skips named in the second comment on #2037 are untouched: the two flaky packet-detail cases, the fixture-data ones, and the two `clientRxCoverage` suites that skip wholesale while reporting success. Those need a fixture deployment with coverage enabled, which is its own change. I have not opened it. Option 2 from the issue, making `test-test-inventory.js` require a `deploy.yml` line for every `tests/e2e` file, is also not here. It is the right guard and it is now enforceable, since the list is finally at zero, but it belongs in its own change where a red build means what it says. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b614badb85 |
fix(packets): give each observation its own wire bytes in the detail API (#2055)
Closes #1999. ## The defect, measured on a production instance Packet `96d716f18d885e78` on a live deployment, read from the deployed build's own API: | | | |---|---| | observations in the response | 60 | | distinct `path_json` values | 50 | | distinct `raw_hex` values returned | **1** | | distinct frames actually stored in SQLite | **51** | The contradiction the issue describes, from that same response: ``` obs 38410791 path ["58C0","1403","50C7"] -> hex 094258C01403AF37E39624E548FB7575F195A0BF ``` Three hops in the path, two path bytes in the frame. The bytes belong to the 2-hop observation and are served for all 60. Across the 3000 most recent transmissions on that database: 1977 have more than one observation and **1844 of those (93%) hold genuinely different frames**. 32659 of 36133 observations (90%) differ from their transmission's canonical bytes. This is the normal case, not an edge case. ## Cause The store deliberately does not retain `obs.RawHex`. #881 dropped it as a memory optimisation, ~98MB measured on a 1.7M-observation store, on the assumption that one content hash implies one frame. The firmware hashes payload and type independently of the relay path, so that assumption is false. Worth adding to the issue's diagnosis: all four load and ingest paths in `cmd/server/store.go` still `SELECT o.raw_hex` and scan it into `obsRawHex`, then use it nowhere — LoadAll, loadChunk, `IngestNewFromDB` and `IngestNewObservations`. The bytes are read out of SQLite and discarded, so a cold load pays the transfer for nothing. ## The fix Keeps the memory saving and reads the bytes back only where a human is looking at one packet. - **`cmd/server/db.go`** gains `ObservationRawHexForHash`: one query returning the stored frame per observation id. Two indexed lookups regardless of observation count — `transmissions.hash` through the prepared `stmtTxByHash` (`idx_transmissions_hash`), then `observations.transmission_id` (`idx_observations_transmission_id`). Guarded by `hasObsRawHex`, because #881 made the column optional and the query would be a SQL error without it. - **`cmd/server/routes.go`** backfills in `handlePacketDetail`: once per request rather than once per observation, and after the store lock is released. Bytes already present are never overwritten, and an observation with no stored frame still falls back to the transmission's. Against the acceptance list: observation bytes exposed with canonical as fallback only ✓; the store's memory optimisation untouched ✓; bounded indexed reads with no query per observation and no work under the store lock ✓; startup-loaded, newly ingested and DB-fallback details all covered, because both the store path and the DB path converge on this one backfill and both key observations by an int `id` ✓. **No frontend change is needed.** `public/packets.js` already spreads the selected observation over the packet (`{...pkt, ...currentObs}`) and already reasons about per-observation bytes: the comment there says "post-#882 per-obs raw_hex with a different path length than the top-level packet's raw_hex still gets accurate byte highlights". The client was built for this and has been receiving 60 copies of one frame. ## Tests `cmd/server/obs_raw_hex_test.go`: - the per-id mapping, with three distinct frames and a fourth observation storing none - the `hasObsRawHex` guard, so a schema without the column is not queried - the handler regression: each observation carries its own frame, the frameless one falls back to the canonical bytes, and at least three distinct frames come back across four observations — the last assertion so that a regression to repeating one frame fails, rather than passing on shape ## Verification `gofmt` clean. **Go tests were not run locally**: no cgo toolchain on this machine since #1992, and per AGENTS.md `CGO_ENABLED=0` builds a stub that proves nothing. CI is their first run. Browser validation per AGENTS.md rule 2: I verified **the defect** in a real browser and through the deployed API, with the numbers above. I could **not** validate the fix in a browser, because the change is server-side Go and is not deployed anywhere yet. Saying so rather than claiming otherwise. ## Not done - The four scan sites that fetch `o.raw_hex` and discard it are left alone. Removing the column from those query builders would stop transferring roughly ten frames per transmission on every cold load, but it touches four builders and their `scanArgs` alignment and is not needed for this defect. - `fetchResolvedPathForObs`, immediately next to this code in `enrichObsWithTx`, does run one query per observation. This change deliberately does not copy that pattern, and does not fix it either. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d1b615fc0d |
fix(node-health): list only observers that heard the node on air (#2057)
Closes #2056. ## What changes The node detail "Heard By" card now lists only observers that received the node's **own transmission off the air**, and reports the rest as a count. ``` HEARD BY — DIRECT (8 OBSERVERS) OBSERVER REGION PACKETS AVG SNR AVG RSSI BE-DUF-SiSCD-01 — 16276 7.9 dB -108 dBm BE-BRU-Moris repeater — 13775 -5.6 dB -122 dBm ... Seen via relay by 29 observers. Those observers heard a repeater that forwarded this node's traffic, not this node. ``` and for a node nothing hears: ``` HEARD BY — DIRECT (0 OBSERVERS) No observer is within radio range of this node. Seen via relay by 2 observers. … ``` ## The rule, and where it comes from Read out of the firmware rather than assumed: | | | |---|---| | `Packet.h:83` | `setPathHashSizeAndCount(sz,n) { path_len = ((sz-1)<<6) \| (n&63); }` — hash size rides in the packet's `path_len` byte | | `Mesh.cpp:649,678` | only `sendFlood()` sets it, so the **originator** decides; `CommonCLI.h:69` defaults `path_hash_mode = 0`, i.e. one byte | | `Mesh.cpp:349` | a forwarding repeater appends its hash with the packet's size — it cannot upgrade a packet, and the **last hop is who was heard** | | `Mesh.cpp:89,103` | on a direct route a forwarder matches the head of the path and calls `removeSelfFromPath` before retransmitting, so the path is the **remaining** route and the transmitter is not in it | So an observation credits exactly one node: 1. Route type must be `ROUTE_TYPE_FLOOD` or `ROUTE_TYPE_TRANSPORT_FLOOD`. Direct routes never qualify (38% of transmissions over 7 days). 2. Empty path → the originator, known only for ADVERTs. 3. Otherwise the last hop. 4. The hop must resolve to exactly one candidate. Same gate `resolvePathForObsColdLoad` already applies: under-attribute rather than guess. It drops 418,530 of 1,455,721 flood observations with a path over 7 days (28.8%), and it is what stops the wrong-band credits. ## Measured effect | node | before | after | |---|---|---| | BE-BRU-Moris | 36 observers | 3 | | BE-KRO-RP01 \| ON1KW | 40 | 3 | | BE-BRE-ON8AR | 38 | 2 | | NL-BXE-RP01 \| 433 | 35 | 0 | Network-wide over 7 days, 234 of 1,860 nodes have at least one direct observer (161 have exactly one, maximum 8). The direct list is therefore empty for most nodes, with the relay count below it. That is the correct reading: no observer is in radio range of them. Independent corroboration on staging: for BE-WIL-3EIK-01 the eight direct observers are exactly the top eight entries of its Neighbors table by score and observation count. ## Perf justification `GetNodeHealth` is fast today precisely because it never walks observations — it uses one representative observation per transmission. Direct-RF needs the per-observation path, and that cannot be a per-request walk: the reference store holds **232,928 transmissions / 2,887,861 observations**, one node's `byNode` slice alone holds **55,458 transmissions / 1,450,544 observations**, and `/api/nodes/bulk-health?limit=200` would multiply that. So the aggregate is rebuilt by a background recomputer on the existing `newAnalyticsRecomputer` pattern, published into an `atomic.Value`. Reads are `O(direct observers)`, which is **cheaper than before** — the old code built per-observer sums over every transmission in `byNode` on every request. Proof, `BenchmarkBuildDirectHeardIndex`: ``` BenchmarkBuildDirectHeardIndex-12 1 63067900 ns/op ``` 3,000,000 observations (60,000 transmissions × 50 observations, 8-hop paths, 64 candidate repeaters) in **63 ms**, once per recompute interval. Per observation the walk does one route-type check, one backward scan of `PathJSON` for the last quoted token (no allocation, no `json.Unmarshal`), one prefix-map lookup and one counter update. Rebuilding wholesale also means eviction needs no bookkeeping: a pass simply does not see evicted transmissions. The alternative — a field on `StoreObs` updated incrementally — would have needed the call at five construction sites (`store.go:942,1264,2854,3179`, `chunked_load.go:609`), which is the duplication that caused #1558, plus matching decrements at eviction. ## API Both `GetNodeHealth` and `GetBulkHealth` carried a near-identical copy of the observer loop; they now share one builder. - `observers` — direct-RF only. Same field names, so no client migration. Rows are a named `HealthObserverRow` instead of `map[string]interface{}` (one fewer occurrence in a touched file, per the AGENTS.md ratchet). - `relayObserverCount` — new integer, observers that saw traffic through the node without hearing it. `stats.totalPackets` and `stats.avgHops` still count relayed traffic, so without this number the card would contradict the figures printed beside it. `docs/api-spec.md` is updated for both endpoints. It also documented an `iata` field on these rows that the endpoint has never emitted; removed. ## Tests - `cmd/server/direct_heard_test.go` — table test over the rule: flood with empty path and known originator, flood whose last hop is the node, flood whose last hop is another node, direct and transport-direct routes (never credit), ambiguous last-hop prefix, listener-only candidate, 1-byte and 2-byte hop sizes; plus aggregation and row-building. - `cmd/server/node_health_direct_rf_test.go` — end-to-end through the handler: an observer that only saw relayed traffic must not appear in `observers` but must be counted in `relayObserverCount`. Plus the benchmark. - `tests/unit/test-direct-rf-heard-by.js` — slices the card template out of `public/nodes.js` and evaluates it, so it tests the shipped markup rather than a copy: heading, empty state, relay line, singular/plural, signal columns, listener/repeater badge tri-state. - `cmd/server/node_health_can_relay_case_1290_test.go` — updated to seed a genuinely direct reception, since a relay-only observer no longer carries a badge. - `cmd/server/analytics_recompute_after_load_test.go` — recomputer count 10 → 11. Verified locally: `cmd/server` suite green, `sh test-all.sh` green (180 suites), `tests/e2e/test-e2e-playwright.js` 131/134 passed with 3 skipped and 0 failures against the seeded fixture, plus `test-issue-1147-section-order-e2e.js`, `test-issue-1151-orphan-separators-e2e.js` and `test-issue-1281-location-row-e2e.js`, which all assert on this card. `gofmt` clean, `vet` clean across all modules. Browser-validated on staging: both the full detail page and the side pane, on a node with 8 direct observers and on the 433 MHz node with none. No console errors. ## What this does not do `prefixMap.resolveWithContext` still guesses on ambiguous hops, so paths, neighbor edges and analytics keep their current attribution. Making it abstain is a much larger change and needs its own issue. The "Regions" line and Region column on this card read `o.iata`, which this endpoint has never emitted, so both have always been dead. Left as found rather than widened into this change. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5c016a210a |
fix(analytics): show a building state on the distance index's 202, and stop caching it (#2051)
Closes #1997. ## What was wrong `/api/analytics/distance` answers `202 {status:"building", retry_after_seconds:5}` with no `summary` until the lazy index (#1011) has been built. 1. `renderDistanceTab` read `data.summary.totalHops` straight away. The TypeError is caught by the tab's own try/catch, so it never reaches `window.onerror`: it is painted into the tab as `Failed to load distance analytics: Cannot read properties of undefined (reading 'totalHops')`. 2. `api()` caches any `res.ok` body, and `res.ok` is true for 202, so the placeholder was stored for the `analyticsRF` TTL. Even a correct retry read the cached "building" body back. That is the half that made the broken state outlast the index build. ## What changed - `public/app.js:173` skips the cache write when `res.status === 202`. A 200 still caches, unchanged. - `public/analytics.js` renders a building notice and retries itself, honouring `retry_after_seconds` clamped to [1s, 30s]. The timer is cleared on tab switch and in `destroy()`, and a new render supersedes a pending retry, so two renders cannot write into the same tab. The notice uses `.text-center`/`.text-muted` rather than the `.spinner` class used at `analytics.js:1683`, because `.spinner` has no CSS anywhere in the repo and renders nothing. ## Tests `tests/unit/test-issue-1997-distance-building.js` (7 assertions, wired into `test-all.sh`) pins the two pure decisions the renderer makes and `api()`'s refusal to cache a 202 while still caching a 200. Red-run on the unfixed sources: 6 of 7 fail, and the "a 200 is still cached" control stays green. `tests/e2e/test-issue-1997-distance-building-e2e.js` (classified in `scripts/non-unit-tests.json`, invoked from `deploy.yml` with `CHROMIUM_REQUIRE=1`) serves both responses by route interception, so it does not depend on whether the server under test has an index built. It asserts: the building state appears, the tab does not paint the error text, a retry arrives with no interaction, the retry replaces the placeholder once the server answers 200, and no retry fires after leaving the tab. Check (2) deliberately asserts on the rendered text and not on `pageerror`: the TypeError is caught, so a `pageerror` assertion would pass on the broken build too. ## Verification No local cgo toolchain here since #1992, so I could not build a server to run the E2E against. Instead I ran its five steps in Playwright against a live instance with this branch's `public/app.js` and `public/analytics.js` injected in place of the deployed ones (both files are byte-identical between that instance and upstream master, so the injection is faithful): | check | deployed build | this branch | |---|---|---| | (1) building state shown | fail | pass | | (2) not rendered as data | fail | pass | | (3) retried on its own | fail (1 request) | pass (2 requests) | | (4) real payload after retry | fail | pass | | (5) no retry after leaving the tab | pass | pass | (5) passes on the broken build too: it schedules no retry at all, so it is a control and only means anything together with (3). The committed E2E suite has not been run as committed. CI is its first real run. ## Not done - The server still recomputes on every 202 poll rather than signalling readiness. - No other analytics tab was audited for the same assume-a-summary pattern. - One pre-existing unit suite (`test-preflight-xss-gate.js`) fails on this Windows machine with a cp1252 `UnicodeEncodeError` from its Python helper, on a clean tree as well as with this change. Unrelated, and green on Linux CI. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7c6b95ea53 |
fix(store): merge background chunks in order instead of prepending them (#2050)
Fixes #2024. `s.packets` is declared "sorted by first_seen ASC (oldest first; newest at tail)" (`cmd/server/store.go:177`), and retention eviction depends on it: `evictStaleInternal` walks from the head and stops at the first transmission inside the window. A slice out of order is therefore **under-evicted silently** rather than failing loudly. ## What breaks it The background chunk loader. Chunks are windowed on `last_seen` (#1690), so a transmission first heard weeks ago and heard again recently arrives in a *recent* chunk carrying its old `first_seen`. The chunk was then put in front of the slice: ```go s.packets = append(localPackets, s.packets...) ``` and never re-sorted, so the next chunk, which covers an older window, was prepended in front of it and left that ancient row sitting behind newer ones. `LoadChunked` re-sorts after its own load; the background merge did not. That asymmetry is the whole bug. It is not a corner case. On a production database, of the **236080** transmissions in a 14 day window, **2071** have a `first_seen` more than a day older than their `last_seen`, and **1848** more than a week. This matters more since #2035: with the accounting fixed, `maxMemoryMB` actually triggers, and a walk that stops early works against it. ## The fix `mergeChunkIntoPackets` merges the two sorted runs linearly. Re-sorting the whole slice was not an option: this runs under `s.mu` once per chunk, so it would sort hundreds of thousands of packets while ingest waits for the lock. The chunk already arrives sorted, since the chunk query ends in `ORDER BY t.first_seen ASC`, so the `sort.SliceIsSorted` guard is a contract check costing one linear pass that never sorts in production. ## Covered - `TestMergeChunkIntoPackets_KeepsFirstSeenOrder` pins the merge against an interleaving, deliberately unsorted chunk. - `BenchmarkMergeChunkIntoPackets` guards the linear cost, against a future simplification back into a sort. The server suite runs under `-race` in CI and is green. ## Not covered, and I would rather say it than let the PR imply otherwise There is **no integration test driving `loadChunk` end to end**. I wrote one and dropped it: a faithful seed database for that path needs more of the schema and more of the loader's preconditions than the fix itself is worth. Two CI rounds in, the seed was still loading zero packets (the first attempt failed at `OpenDB` on a missing `nodes` table, the second on the window). Both attempts are in this branch's history rather than rewritten away. So the end-to-end claim rests on the code path quoted above and on the production measurement, not on a test that exercises it. The unit test covers the function where the logic now lives, which is the part that can regress. Also not verified locally: `cmd/server` needs cgo for the #1992 driver and this machine has no C toolchain, so CI is the check. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |