mirror of
https://github.com/Kpa-clawbot/meshcore-analyzer.git
synced 2026-10-11 19:37:43 +00:00
master
3030
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dc7ebcd6f5 |
perf(store): check observation duplicates without per-transmission maps (#2181)
## Problem Every `StoreTx` carried two maps for observation de-duplication: - `obsKeys`, keyed by `observerID + "|" + pathJSON` (a concatenated string allocated for every observation); - `observerSet`, for `UniqueObserverCount`. On a busy regional network's 7-day window (64k transmissions, 400k observations), these maps and their concatenated keys held about 60 MB of heap. Most transmissions are heard a handful of times; the busiest in that week was heard 40 times. The maps cost a few hundred bytes per transmission to de-duplicate lists that are usually under ten long. ## Change - `hasObservation` and `hasObserver` scan `tx.Observations`, comparing the observer ID and path strings the observations already hold. These are the same two values the old map key was built from, so the answers are identical. - Once a transmission passes 64 observations (`obsIndexThreshold`), it builds the maps. They're keyed by a struct that shares the observations' strings, so no key bytes are allocated, and busy transmissions keep O(1) checks (the case #543 / #355 addressed). - The five places that appended an observation now call `addObservation`, which keeps `UniqueObserverCount` and `ObservationCount` in step. Those three lines were previously copied at each site. - The memory estimate no longer charges for the per-transmission maps or the concatenated keys. ## Perf justification (AGENTS.md rule 0) - **Complexity.** - Below the threshold, a duplicate check is O(k) for a transmission with k < 64 observations, so building it costs at most 63·64/2 = 2016 string-header comparisons. - At and above 64 observations, checks are O(1) map lookups, as before. - Nothing is O(n²) in the number of transmissions or observations overall: the scan is bounded per transmission. - **Scale.** Real data peaks at 40 observations per transmission. The scan compares string headers (length first, then bytes), and 64-character observer IDs differ early. `BenchmarkObsDedup` (a miss, the common case on ingest): | observations on tx | check used | time | |---|---|---| | 4 | scan | 15 ns | | 16 | scan | 70 ns | | 40 | scan | 151 ns | | 63 | scan | 233 ns | | 64+ | map (struct key) | 15 ns | For comparison, the old map lookup was about 15 ns too, plus building and allocating the `observerID|path` key string on every check. At 40 observations that's about 3 µs extra to build a whole transmission, against roughly 2 KB of maps no longer allocated for it. **Measured on a copy of that 7-day window** (`GOMEMLIMIT=750MiB`, same queries): | | before | after | |---|---|---| | in-use heap after load | 612 MB | 568 MB | | HeapAlloc after load | 660 MB | 564 MB | | peak RSS | 1077 MB | 1009 MB | | startup to ready | 31 s | 21 s | It has also run on a staging instance (live MQTT ingest) for about nine hours alongside other store changes, with no change in observation or observer counts. ## Tests - `obs_dedup_test.go`: - duplicates (same observer and path) are rejected, while a new path or a new observer is kept; - a transmission under the threshold carries no maps; - across the threshold the maps take over with the same answers; - empty observer IDs aren't counted as observers; - `BenchmarkObsDedup`. - Red commit first (stubs, failing on assertions), then green. - `cd cmd/server && go test ./...` passes. - SonarQube: quality gate passed, no issues on added lines. Backend only: no UI, API shape or config change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
955386b63f |
test(map): settle the router's focus frame before the teardown snapshot (#2179)
## Summary Fixes an intermittent failure in `tests/e2e/test-map-nodes-pagination-e2e.js`, the "late response after map teardown" cases added for #2030 in #2048: ``` ✗ late /observers response after map teardown: old response must not change the destination page ``` It failed on #2164 ([run 38080812509](https://github.com/Kpa-clawbot/CoreScope/actions/runs/38080812509), shard 2/3), a Go-only change, and passed on the rerun. ## Cause `checkMapTeardown` compares `#app`'s HTML before and after releasing the old map's pending response. When it failed, the only difference was on the destination page's heading: ``` before: <div class="tools-landing"><h2>Tools</h2>… after: <div class="tools-landing"><h2 tabindex="-1">Tools</h2>… ``` That attribute comes from the router, not the old response. Since #630 (`app.js`, "SPA focus management") it focuses the new page's first heading in a `requestAnimationFrame` after the page renders. The before-snapshot was taken as soon as `#leaflet-map` detached, sometimes ahead of that frame. ## Fix The test now waits two frames before the before-snapshot, the same settle it already does before the after-snapshot. The router's frame is queued by then, so it always lands first. The change is test-only: no app code changes, and every assertion stays as it was. ## Tests Run locally with the e2e fixture (`tools/freshen-fixture.sh`, migrated) served by this branch's server, headless Chromium: | | runs | failed | |---|---|---| | test on `master` (`352a0ccd`) | 20 | 2 | | this branch | 50 | 0 | - **Before:** the two failures on `master` were "old response must not change the destination page" (teardown, `/config/regions` and `nodes`), plus one 30 s timeout in an "after map replacement" case in the same runs. - **After:** none of these failures or the timeout came back in 50 runs. ## Checks - **Lint:** ESLint passes with the repo config and with Node globals. SonarQube rules on the file report nothing on the added lines; its four existing minor findings are on lines this PR doesn't touch. - **Repo rules:** one logical change, explicit `git add`, no private data, no new dependencies. This is a test-only fix, so there's no red/green pair: the "red" is the intermittent failure above, measured as 2 in 20. Related: #2030, #2048 (where the test was added). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d9342bbebd |
(feat): UserManagement, Add Support for Postal Email Delivery Provider (#2154)
This PR extends email support to add another email delivery provider called Postal for user management. It is a, open source, self hosted email delivery and reputation management service. It adds 2 new additional configuration settings to the userManagement.mail object which are "postalBaseUrl" and "postalApiKey". It also adds "postal" as a supported value for the "provider" setting. All existing functionality including email delivery notification is supported by Postal. Included test cases to support the new functionality and all pipelines pass. --------- Co-authored-by: CoderNemesis <rfedor@comchan.net> |
||
|
|
a4ee3a7328 |
feat(scope-audit): observer region filter for both views (#2171)
## Summary Builds on the transport view from #2167 and #2169 (both merged); rebased onto `master` (`352a0ccd`), so the diff is this PR's six commits only. The Scope Audit was the one analysis page without the region selector. This adds the shared region filter to it, for both the declared and the transport view. `?region=SFO,SJC` means what it means on every other endpoint: count only forwarding **heard by observers in those regions**, not repeaters located there. - **API** (`GET /api/scope-audit`, both modes): - The region codes are resolved to their observers first (a handful of rows), then the audit's hop scan gets `AND o.observer_idx IN (…)` on v3, or `o.observer_id IN (…)` on v2. The 3.5M-row hop scan never joins `observers`. - `All` means no filter (shared `normalizeRegionCodes`). Codes with no observer match nothing, instead of silently widening to every observer. - Responses echo the normalised codes as `region`. The declared and transport caches and their singleflight keys now include the region, so each region set is cached separately with the existing TTLs. - Declared-region verification (`regionEvidence`) is deliberately left unfiltered: it corroborates what a repeater declares from its own traffic, which doesn't depend on where a hop was heard. - **Page:** the shared `RegionFilter` sits with the view and window controls; a change reloads the view, and the selection persists like every other page. The transport view's text box is relabelled **Scope** (it already sends `?scope=` since the rename in #2167), so the page no longer has two controls called Region. - **Docs:** `docs/api-spec.md` (parameter, both response shapes) and the OpenAPI parameters for the route (`mode`, `scope` and `region` were all missing). Partial fix for #2142 (its region filter for rollout tracking), together with #2167 and #2169. ## Performance The unfiltered path must not get slower: it's the same query with an empty clause. `BenchmarkComputeScopeTransport` (from #2167, 1,200 repeaters), the PR it stacks on vs this branch, `-count 8`: | | before | after | | |---|---|---|---| | sec/op | 220.3 ms | 218.0 ms | ~ (p=0.065) | | B/op | 43.77 MiB | 43.70 MiB | ~ | | allocs/op | 833.5k | 833.5k | ~ | **Filtered requests, measured on staging** (a copy of a production database): the first version drove `o.observer_idx IN (…)` from `idx_observations_observer_idx`, reading every observation the region's observers ever made. Fixed in `80414242` with a unary `+` (the same treatment as #2166), so the scan starts from the `first_seen` window and filters per row. Red test `67e063b1` asserts the query plan against a fixture with the real indexes. | scan, one region over 24h (13 observers) | plan | time | hop rows | |---|---|---|---| | `o.observer_idx IN (…)` | observer index | 13.44 s | 102,114 | | `+o.observer_idx IN (…)` | `first_seen` window | **0.25 s** | 102,114 | | unfiltered | `first_seen` window | 0.35 s | 225,983 | Cold end-to-end requests on staging after the fix: 24h declared, unfiltered 0.84 s · one region 0.48 s (was 16.7 s) · two regions 0.22 s · transport, one region 0.42 s · transport, six regions 0.75 s · 7d, one region 3.46 s. 24h transport, live data: 383 repeaters unfiltered, 204 heard by one country's observers. ## Tests - Red `629e4f36` (Go): extends the audit fixture with an `observers` table and `observations.observer_idx`. Covers the declared view ignoring a scope heard only by another region's observer, the transport view listing only repeaters heard in the region, `region=All`, per-region caching, a region with no observers, and the echoed codes. Five fail on assertions; `region=All` passes because the stub ignores the filter, which is the behaviour it pins. - Green `6f1b1e1c`: the implementation, plus a v2-schema case (`observer_id`), so both branches of the clause are covered. - Red `b4933d68` / green `08ecc172` (frontend): `apiPath` carries the selection in both views and adds nothing when none is selected. The region control on the page is a net-new UI surface. - `go test` (cmd/server) passes. Frontend suites: 193 of 194. The one failure, `test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally. ## Screenshot The same transport view with the shared region filter set to **SJC**: only forwarding heard by San Jose observers counts, so 7 repeaters are listed instead of 44, with their carried scopes counted from those observers alone. Taken against `test-fixtures/e2e-fixture.db` with fictional scopes.  ## Browser validation The branch's server was run locally with 14 configured regions (dropdown layout), checked in headless Chromium at 1280, 900, 700 and 390 px (touch): | check | 1280 | 900 | 700 | 390 touch | |---|---|---|---|---| | ticking SJC re-requests with `®ion=SJC` | yes | yes | yes | yes | | menu fully on screen | yes | yes | yes | yes | | page wider than viewport | no | no | no | no | | page errors | none | none | none | none | Found and fixed this way: the filter is the right-most control on desktop, and the menu (which opens rightwards) ran 66 px off the page. It now opens leftwards above 640 px. On a phone the controls wrap, the filter starts a line, and the default applies. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - **S3776 on `ScopeAuditForwardingHeardIn`:** complexity 55. The function it generalises, `ScopeAuditForwarding`, is 54 on `master`; the region clause adds one point. - **S3776 on `computeScopeAudit`:** complexity 65, unchanged from `master`. - Both pre-date this PR. Splitting them is better as a refactor PR of its own than mixed into a feature. - The page code's findings were fixed in the transport-view PR it stacks on. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
55dd28164d |
test(e2e): assert the configured-scope row after the theme-refresh re-render (#2180)
## What `test-issue-1648-m2-icons-e2e.js` now asserts the configured-scope row only after the page's `theme-refresh` re-render has fired and the detail has rendered again. ## Why On 2026-10-10 this assertion failed CI on four PRs (#2169, #2171, #2154, #2164), each time at the first live-map case: `confirmed scope row must be visible, including an empty scope`. A re-run of the failed jobs passed every time. The code path: - `customize-v2.js:814` dispatches `theme-changed` once `/api/config/theme` resolves, on every page load. - `app.js:1309-1315` turns that into `theme-refresh` 300 ms later. - `live.js:4713` re-runs `showNodeDetail(activeNodeDetailKey)` on `theme-refresh`, and `showNodeDetail` first sets `#nodeDetailContent` to "Loading…" (`live.js:2598`). This is the only caller besides the marker click (`live.js:3090`). - The test waited for `#nodeDetailContent table` and then called `row.isVisible()`, which does not wait. A refresh landing between the two fails the check. The same race could let the null-scope case pass on the "Loading…" panel, because `row.count() === 0` holds there too. The fix uses the same `addInitScript` flag pattern as `test-e2e-playwright.js:83`. ## Testing - Locally against a server on the E2E fixture: 5 of 5 runs pass, 27 of 27 checks each. - **Not reproduced locally.** The unchanged test also passed every local run: 5 plain, 3 with `/api/config/theme` delayed by 1 s, 5 with 8x CPU throttling. So the reasoning above comes from the code path and the four CI failures, not from a local red-then-green. - Test-only change, no product code touched. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4e91c5cf0b |
perf(analytics): opt-in pauseWhenIdle skips recomputes nobody reads (#2164)
## Summary The analytics recomputers (#1240) recompute eleven snapshots (topology, RF, distance, channels, hash collisions, hash sizes, roles, two clock-skew views, retransmissions, direct-heard) on a fixed interval, 300 s by default, whether or not anyone reads them. That is the right trade for a busy instance: reads always hit a warm cache. On a quiet one it means most of the server's allocation goes on results nobody looks at. Measured on one production instance (2 vCPU / 2 GB): - topology, RF and hash-sizes were each **read 23 times in 45 hours** while being recomputed every 5 minutes; - `analyticsRecomputer.runOnce` accounted for **30% of all bytes allocated** (121 GB of 391 GB in 45 h); - garbage collection was 87% of the process's CPU (30 s CPU profile: `gcDrain` 39.9 s of 45.9 s sampled), with an 805 MB live heap against a 1 GiB `GOMEMLIMIT`, so allocation rate translates directly into collector time. This PR adds an **opt-in** `analytics.pauseWhenIdle`: - a recomputer's tick is skipped when nothing has read its snapshot since the previous compute, and the recomputer is marked paused; - the next read still returns immediately with the snapshot it holds (`Load` stays non-blocking, as #1240 requires), and makes a non-blocking send that wakes the loop to recompute; reads after that get the fresh snapshot and the ticker resumes from there. **Default is off**, so behaviour is unchanged unless an operator turns it on. The trade-off for operators who do: the first read after a quiet spell can show a snapshot older than the interval (refreshed seconds later). It is documented in `config.example.json`. No existing issue covers this. ## Config example Off by default. To turn it on, add one key to the existing `analytics` block in `config.json`: ```json "analytics": { "defaultIntervalSeconds": 300, "pauseWhenIdle": true } ``` Nothing else changes: per-endpoint `recomputeIntervalSeconds` still apply while an endpoint is being read. There is no UI change, so no screenshot. ## Performance Staging A/B: the same 30 minutes of real production API traffic (5,640 requests from a day's nginx log, replayed at 2x against a copy of the production database with the live MQTT feed), on `master` and on this branch with `pauseWhenIdle: true`, each after a restart and warm-up. Allocation is the delta of `/debug/pprof/allocs` across the replay: | during the replay | master | this PR | |---|---|---| | allocated by `analyticsRecomputer.runOnce` | 459.7 MB | **53.4 MB (-88%)** | | total allocated by the process | 3,704 MB | 3,372 MB (-9%) | | GC cycles | 52 | 41 | | requests: p50 / p95 / errors | 9.9 ms / 57.0 ms / 0 | 10.4 ms / 60.1 ms / 0 | The replayed traffic included real analytics reads, so the paused recomputers were woken by them; the remaining 53 MB is that. Over 15 minutes each recomputer only ticks about three times, so on a longer quiet period the saving grows with the idle time. GC count and CPU on this shared host vary between runs on their own, so the allocation figures are the claim, not the GC count. There is no Go benchmark for this change: the behaviour is "how many computes happen in a period of no reads", which the tests below assert directly (16 computes on master vs at most 2 here over the same idle period). ## Tests - Red commit `24056c57`: config field, accessor stub and an unused recomputer flag, plus tests. Three fail on the stub: - `TestAnalyticsRecomputer_PauseWhenIdleSkipsUnreadTicks`: an unread recomputer with a 20 ms interval computes at most twice in 300 ms (16 times on the stub); - `TestConfig_AnalyticsPauseWhenIdle`: `pauseWhenIdle: true` enables it, absent or nil config leaves it off; - `TestStartAnalyticsRecomputers_AppliesPauseWhenIdle`: all eleven recomputers receive the flag. - Also added, passing on both commits on purpose: a read on a paused recomputer returns the existing snapshot and triggers a refresh; a recomputer that is being read keeps ticking; the default (off) keeps ticking unread. - Green commit `19de3636`: the behaviour, `config.example.json` entry with `_comment_pauseWhenIdle`, and the `Load` doc comment. - `go test -race` on the recomputer, warm-up (#1659) and config tests, three runs: pass. `make test`: 16 of 16 modules pass. The frontend tests that read `config.example.json` pass. ## Review notes - `roles` reads the nodes-clock-skew snapshot inside its own compute, which counts as a read, so that dependency keeps its source fresh while roles is in use. - The #1659 warm-up gate is untouched: a skipped tick never marks a first pass done. - Config documented per WORKFLOW.md section 7. Customizer (AGENTS.md rule 8): it is an operator setting, not a display value; no customizer entry proposed. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - No issues on the added lines. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
352a0ccd19 |
feat(region-filter): configurable region quick picks (#2170)
## Summary **Independent of every other open PR.** Choosing a pick saves the selection, redraws the control and notifies the page's listeners itself, exactly as a click in the control does, so it re-queries straight away whether or not #2168 (`setSelected` notifies) is merged; with both in, nothing fires twice. Region quick picks are named groups of region (IATA) codes, offered as one-tap choices inside the region filter, e.g. "Northern California" for a dozen airport codes instead of ticking each box. One deployment has been injecting these into the page through an nginx `sub_filter` and a script that wraps `window.RegionFilter`; this makes them a configured feature. ```jsonc "regionQuickPicks": [ { "name": "Northern California", "description": "Bay Area and Sacramento observers.", "regions": ["SJC", "SFO", "OAK", "SMF"] } ] ``` - **Server:** `GET /api/config/region-quick-picks` returns the groups cleaned up: names and codes trimmed, codes upper-cased and de-duplicated in their configured order, groups without a name or codes dropped, at most 20 groups of 200 codes, always a list. `/api/config/regions` is unchanged. - **Region filter:** loads the groups with the regions and offers every group with at least one observer today, inside the region control: - dropdown layout: rows of the same kind as **All** (checkbox and bold name), directly under All with a divider before the single regions; the row is ticked while that pick is the selection, ticking it selects the pick and closes the menu, unticking goes back to all; - pill layout: the picks lead the same bar, set off from the single regions by a thin divider. - A tap selects the group's codes that have an observer and re-queries straight away. - The selector then reads the group's name ("South Bay ▾") instead of "2 Regions", and the button shows as pressed. Selecting the same codes by hand does the same. - Tapping the active group again goes back to all regions. - A failed fetch, or no configured groups, leaves the filter exactly as before. - **Mobile:** the rows take the menu's 44 px row height under `(pointer: coarse)`; the pill-bar picks reuse `.region-pill` and its touch target. Inside `.filter-bar` (Packets) the existing touch rule for inputs made every region checkbox a 44 px square; the checkbox is now 20 px and the 44 px row is the tap target. - **Semantics:** like every region filter, a group means "heard by observers in these regions", not "located there"; `config.example.json` says so. No existing issue covers this. ## Screenshot Nodes page with two picks configured (example config below), after tapping **South Bay**: the pick and its codes (SJC, MRY) are selected, and the list re-queried for them. Taken against `test-fixtures/e2e-fixture.db`.  ## Config example ```json "regionQuickPicks": [ { "name": "South Bay", "description": "San Jose and Monterey observers", "regions": ["SJC", "MRY"] }, { "name": "East and North Bay", "description": "Oakland and San Francisco observers", "regions": ["OAK", "SFO"] } ] ``` Each pick needs a `name` and at least one code; `description` is optional (shown as the tooltip). Codes are trimmed, upper-cased and de-duplicated; a pick with no code that has an observer today is not shown. At most 20 picks of 200 codes. ## Browser validation The branch's server was run locally against a migrated copy of `test-fixtures/e2e-fixture.db` (observers in MRY, OAK, SFO, SJC) with three groups configured: "South Bay" (SJC, MRY, SCZ), "East and North Bay" (OAK, SFO, sts) and "Nowhere yet" (XXX). Checked in headless Chromium: | check | desktop 1280 px | phone 390 px (touch) | |---|---|---| | groups shown | South Bay, East and North Bay ("Nowhere yet" hidden) | same | | tap South Bay selects | SJC, MRY (SCZ has no observer) | same | | next `/api/packets` request | `region=SJC,MRY` | `region=SJC,MRY` | | selector label | "South Bay ▾" | "South Bay ▾" | | second tap | back to all regions | same | | button height | 24 px | 44 px | | page wider than viewport | no | no | | page errors | none | none | Checked on Packets (dropdown layout) and Nodes (pill layout). Three layout problems found this way on the phone were fixed before opening: a sideways-scrolling row clipped at the screen edge, the selector squeezed into a narrow column (the filter group is a non-wrapping flex row), and the two wrapped rows spread apart on Nodes (`align-content`). **Revision 1:** the first version put the picks in a separate row beside the selector; they moved inside the dropdown. Re-checked with a 14-region list and three picks on Packets: desktop 1280 px dark and phone 390 px light both keep the menu on screen (menu 340 px wide, page not wider than the viewport), the pick closes the menu and the trigger names the pick. **Revision 2:** picks are checkbox rows like **All** instead of chips. Re-checked on Packets at 1280 px dark and 390 px light: rows render bold under All, the divider separates the single regions, ticking a pick ticks its codes and names the trigger, no sideways scroll. Unrelated and not changed: on Nodes at 390 px the region row (existing pills included) starts flush with the screen edge. ## Tests - Red commit `81ea67d6`: the config type, a stub normaliser and endpoint returning nothing, and tests. - Go: cleaning (trim, upper-case, de-duplicate, drop incomplete groups), caps, nil config, and the endpoint's JSON. - Frontend, `tests/e2e/test-region-quick-picks.js` (moved from `tests/unit/` in `fced69a7`; it needs jsdom, so it runs in the E2E job like `test-table-sort.js`): the real `region-filter.js` in **jsdom** against faked config endpoints. Nine cases: rendering and hiding groups with no observer, selecting and re-querying, the selector label, toggling back, a hand-made selection marked active, the pill layout, no groups, a failed fetch, HTML escaping. Seven fail on the stub, all on assertions. - Green commit `0e33210e`: the implementation, styles, `config.example.json` (`regionQuickPicks` + `_comment_regionQuickPicks`), `docs/api-spec.md`, and the OpenAPI route list (`TestOpenAPICompleteness` requires it). The jsdom test also had a cross-realm comparison bug fixed in this commit (`Array.from` on arrays coming out of the jsdom window). - `make test`: 16 of 16 modules pass. Frontend suites: 195 of 196; the one failure, `test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally. - Red commit `f991060d` (revision): picks live inside the dropdown menu before the checkboxes, choosing one closes the menu, and in the pill layout they share the region bar; two fail on assertions against the old layout. Green commit `1f720e35`. Frontend 195 of 196 (same environment-only failure). - Red `134f5256` / green `da0adb50` (second revision): picks are checkbox rows between All and the single regions with a divider, no chips in the menu; ticking and unticking a pick row. The green commit also repairs `style.css`: `1f720e35` had prepended its block to the top of the file (its edit matched an earlier `.region-pill {`) and left the previous rules in place. The file now differs from `master` by 13 lines. ## Review notes - Groups are offered only when at least one of their codes appears in `/api/config/regions`, and only those codes are selected. The stored selection is pruned to known codes on load anyway, so selecting absent codes would not survive a reload. - Customizer (AGENTS.md rule 8): groups are operator configuration; a customizer editor could come later. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - **Fixed** in `b84ace4e`: - `normalizeQuickPickCodes` split out of the config normaliser (S3776, 17 > 15); - a comment on why a failed quick-pick fetch is ignored (S2486); - `for…of` in `activePick` (S4138); - in the tests: `node:` specifiers, `includes()`, and an explicit string comparator for sorted selections (S7772, S7765, S2871). - **Kept on purpose:** - S2486 is still reported on that `catch`. Swallowing the error is the intended behaviour: the filter works as before, just without picks. - S1523 on the `vm` test loader, as above. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
99efaed33d |
fix(region-filter): setSelected notifies change listeners on a real change (#2168)
## Summary `RegionFilter.setSelected()` saved the selection and redrew the control but never ran the `onChange` listeners; only a click (`toggleRegion`) did. A programmatic selection therefore updated the dropdown while the page kept showing data for the old selection until the next click. Affected today: the Packets page applies a `?region=` URL parameter through `setSelected`. It registers its listener afterwards, so its first load is fine, but any other page's listener that is still registered sees nothing. And anything that offers one-click region groups has to wrap `setSelected` to make it re-query (an nginx-injected overlay on one deployment does exactly that). This PR stands alone; #2170 (configurable quick picks) does not need it. `setSelected` now compares the new selection with the current one (order-insensitive; empty and null both mean "all") and, only when they differ, runs the listeners after saving and re-rendering, exactly as a click does. Re-applying the current selection stays silent, so a caller restoring state does not trigger a reload. No existing issue covers this. ## Tests - Red commit `75569b78`: `tests/unit/test-region-filter-setselected.js` loads the real `public/region-filter.js` in a `vm` sandbox. Two cases fail on `master` (a programmatic selection notifies once with the new codes; a clear notifies once with `null`). Two pass on both commits on purpose: re-applying the current selection in another order, and clearing an empty selection, must not notify. - Green commit `8a7ddfe1`: the fix. - Listed in `test-all.sh` in sorted order; `test-test-inventory.js` passes. `test-1108-region-hide-nodes.js` (the other region-filter suite) passes. Frontend suites: 194 of 195; the one failure, `test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally. ## Review notes - `packets.js` "clear all filters" calls `setSelected([])` and then reloads itself. When a region was selected, that path now also triggers the page's region listener, so the packets list is requested twice. It is a cosmetic double fetch on an explicit user action; left as is to keep this PR to the component. - No performance claim; this is a correctness fix. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - **Fixed** in `9e732706`: `node:` specifiers (S7772). - **Kept on purpose:** S1523 on `vm.runInContext`, which loads the repo's own `public/region-filter.js` as the other frontend unit tests do. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
deb55dc82e |
feat(scope-audit): transport view on the Scope Audit page (#2142, page) (#2169)
> **Requires #2167 to be merged first.** This page calls the transport view API that #2167 adds (`/api/scope-audit?mode=transport`), so it cannot be split from it; until #2167 merges this PR also shows its commits. ## Summary Page half of #2142. **Stacked on #2167** (the API); review that first. #2171 adds the observer region filter on top. Partial fix for #2142 until all three are in. Adds a **Declared / Transport** toggle to `#/scope-audit`. The transport view calls `?mode=transport` and lists every repeater seen forwarding in the window: - **Seen carrying:** one chip per region scope with its packet count (first/last seen in the tooltip); - **Declared:** the repeater's declared regions coloured by whether they were seen carried (the declared view's own chips), or *not asked* when no declared-regions answer exists. Never asked is worded differently from declaring nothing; - **Other traffic:** unscoped, unmatched and ambiguous-hop counts. A **Region** box (suggestions from the scopes seen) adds a *Carries X* column, lists the repeaters not yet carrying it first, and summarises carriers against active non-carriers: the rollout report the issue asks for. Window, view and region are deep-linked (`#/scope-audit?window=7d&mode=transport&scope=be`) and survive a reload. The page says plainly that *transported* means seen carrying, not configured, and that over a short window a quiet region looks the same as a missing one. The intro text changes with the view. ## Screenshot The transport view with scope `be` typed in: every repeater seen forwarding in the last 24 h, whether it carries `be` (non-carriers first), the scopes it was seen carrying with counts, and its declared regions or *not asked*. Taken against `test-fixtures/e2e-fixture.db` with fictional `be` / `be-van` / `nl` scopes.  ## Browser validation (staging, real data) Headless Chromium against a staging copy of a production database with its live feed: - `#/scope-audit?window=24h`: the declared view, 16 rows, as before. - Clicking **Transport**: 385 rows; the hash gains `mode=transport`. - Typing a scope as `#XX-YY` in capitals: normalised to `xx-yy`; the hash gains it (now `scope=xx-yy`); summary "385 repeaters seen forwarding in the last 24h · 16 with a declared-regions answer · 51 carried xx-yy · 334 active but not seen carrying it"; the first rows are the non-carriers. - Reloading that hash restores view and region. Switching back to Declared hides the Region box and restores the declared intro. No page errors. - Checked at 1280 px and 390 px wide. Two layout bugs found this way were fixed before opening: the region box inherited `width: 100%` from `.nodes-search`, and `display: flex` overrode the `hidden` attribute on the region bar. ## Tests `tests/unit/test-frontend-helpers.js` gains 10 cases on rendered markup and state, using the page's exported internals like the existing scope-audit tests: hash and API paths (and that a region never leaks into the declared view's link), region normalisation (`#`, capitals, empty), *not asked* vs declared chips (with the `*` wildcard), carried-scope chips with counts and the filtered region marked, the carries column only with a region, HTML escaping of names, the summary counts, region suggestions, and the per-view intro. 768 passing in that file. Net-new UI surface, so per TDD.md the tests land in this PR rather than as a preceding red commit. Frontend suites on this branch: 193 of 194; the one failure, `test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and fails identically on `master` locally (#2149 and #2160 address that). ## Review notes - Colours only through existing CSS variables (no hex literals); the filtered region is marked with the link colour, not a new status colour. - The window note's wording is in the same spirit as the declared view's `windowHonestyNote`. - No API calls per row: one request per view/window/region, cached client-side for 30 s like the declared view. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - **Fixed** in `df53e770`: - one `modeQuery` helper for the page hash and the API path (S3358; the two copies were identical); - the carries-scope cell built before the row (S3358); - `localeCompare` for the region suggestions (S2871), `for…of` and `.dataset` (S4138, S7761), optional chaining (S6582); - object spread in the helper tests (S6661). - Optional chaining and spread are already used throughout `public/`. No behaviour change. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
696906315a |
feat(scope-audit): transport view of every repeater seen forwarding (#2142, API) (#2167)
## Summary Backend half of #2142 (the page is the stacked PR #2169). `GET /api/scope-audit` lists only repeaters that have answered a declared-regions request: 17 of almost 200 nodes on the instance the request came from. This adds a transport view, `GET /api/scope-audit?mode=transport&window=…[&scope=X]`, with one row per repeater **seen forwarding** in the window: - `transported`: the named region scopes it was seen carrying, each with packets and first/last seen (most packets first), plus unscoped, unmatched and ambiguous-hop counts; - `asked`, `declaredRegions`, `notObserved`, `configState`, `declaredAt` where a declared answer exists; `asked: false` with null declared fields where it does not ("not asked" is not "declares nothing"); - with `scope=X` (spelled `be`, `#be` or `BE`): `carriesScope` on every row and `carrying` / `notCarrying` counts, which is the rollout report the issue asks for. `mode=declared` (or no `mode`) is unchanged. An unknown `mode` is a 400. Partial fix for #2142: this PR provides the API; the page is #2169 and the observer region filter #2171. The issue should only be considered done once all three are in. ## How it follows the issue's notes - **The scan does not change.** It is the same single `ScopeAuditForwarding` pass, with every repeater and room in `nodes` (plus any declared target) as a target instead of only the declared ones. Same attribution rules: FLOOD-family route types only, `minForwarderHopHexLen` (no 1-byte hops), and a hop whose prefix matches more than one repeater is credited to none of them. - **Identity lookup with 1,000+ pubkeys:** `scopeAuditNodeIdentities` runs one `IN` query; exercised at 1,200 keys in the benchmark below. - **Row count and response size:** measured below. - **Cache:** the transport view is cached per window with the declared view's TTLs (30 s; 5 min for 7d), under the same singleflight group with a `transport|` key. The scope filter is applied to a copy of the cached result, so one scan serves every scope. (`scope=`, not `region=`: across the API `region=` is the observer IATA filter.) - **Wording:** "transported" is documented as *seen carrying*, not configured for; the declared side stays separate. ## Performance `BenchmarkComputeScopeTransport`: 1,200 repeaters (#1975 quotes an instance with 1,179), 30k flood transmissions in the last 24h on 4-hop paths of 3-byte prefixes, a third region-scoped: ``` BenchmarkComputeScopeTransport-12 5 219253727 ns/op 425.4 response-KB 1200 rows ``` 219 ms per compute, served from cache for 30 s (24h) or 5 min (7d); a 425 KB uncompressed response for 1,200 rows. This is new code with no "before"; the declared view's numbers are untouched (same scan, same cache). ### On real data (a staging copy of a production database and its live feed) | request | rows | result | time | |---|---|---|---| | declared view, 24h | 16 | (unchanged) | | | `mode=transport`, 24h | **385** repeaters seen forwarding, 16 of them with a declared answer | | 1.5 s cold | | `mode=transport&scope=<scope A>`, 7d | 664 | **51 carrying it, 613 not** | 14.4 s cold | | `mode=transport&scope=%23<SCOPE B>` (`#` and capitals), 7d, same window now cached | 664 | 237 carrying, 427 not | 0.35 s | The 7d cold compute is the same full-window hop scan the declared view already pays for 7d (its own TTL comment measures 16.7 s on a live-shaped database); like the declared view it is then cached for 5 minutes, and every scope filter is served from that one scan. - Commit `ceaeb56b` renamed the filter from `region=` to `scope=` after opening; the real-data figures above were taken before the rename and are otherwise the same code. ## Tests - Red commit `2a9b1286`: response types and a `mode=transport` stub that returns no rows, plus tests. Four fail on the stub: - repeaters that never answered are listed with `asked: false` and null declared fields (and the answered one with its declared list and `notObserved`); - `scope=fr` / `%23fr` / `FR` marks carriers and non-carriers and counts them; no scope means no `carriesScope` or counts; - an unknown mode is a 400; - the transport and declared views are cached separately (requesting one does not serve the other). - Also added, passing on the stub by construction: a repeater with no traffic in the window is not listed; the node blacklist applies. - Green commit `60c669f4`: the implementation, `docs/api-spec.md` (new `mode` and `scope` parameters, a Transport view section with the response shape), and the benchmark. - `make test`: 16 of 16 modules pass, including every existing scope-audit test. ## Review notes - Hidden-name and blacklist filtering use the same rules as the declared view. - `newTestStoreWithDB` now takes `testing.TB` so the benchmark can use it (callers unchanged). - No new `map[string]interface{}`; the response is a named struct. - No config added. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - **Fixed** in `060668b6`: S3776 (cognitive complexity 41 > 15) in the new `computeScopeTransport`. It is split into target, visibility, row and sort helpers; the response is the same and the transport tests are unchanged. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8a3b923e81 |
perf(scope-stats): keep the per-region count on the first_seen index (#2166)
## Summary `/api/scope-stats` averaged **2.6 s on one production instance (max 17.7 s)**, whatever the window. Four of its five queries take 10-100 ms on that database. The fifth, the per-region count, took **2.5-2.9 s for every window, 1h included**: SQLite chose `idx_tx_scope_name` (`scope_name > ''`) and walked every scoped transmission ever stored (45 days of them), filtering by `first_seen` afterwards. ``` SEARCH transmissions USING INDEX idx_tx_scope_name (scope_name>?) | USE TEMP B-TREE FOR ORDER BY ``` A unary `+` on every `scope_name` reference in that query (`WHERE … +scope_name …`, `GROUP BY +scope_name`) stops the planner using that index. The query then searches the window on `idx_transmissions_first_seen` and groups in a temp b-tree, the same idiom the `advertsByRole` query in the same function already uses for `payload_type`. Putting `+` in the `WHERE` alone is not enough: the planner then walks `idx_tx_scope_name` to avoid sorting for the `GROUP BY`, which is just as slow; the `GROUP BY +scope_name` is what fixes it. ``` SEARCH transmissions USING INDEX idx_transmissions_first_seen (first_seen>?) | USE TEMP B-TREE FOR GROUP BY | USE TEMP B-TREE FOR ORDER BY ``` Related to #2146, which lists `/api/scope-stats` (2.3 s average on one instance) and records the 24 h summary as 7 ms warm there. On the instance measured here the per-region query planned onto `idx_tx_scope_name` and took 2.5-2.9 s for every window; this PR fixes that plan. Not marked as a fix: #2146 is about pool and lock waits. ## Results are identical On a copy of the production database, the old and new query run inside **one read transaction** (same snapshot) return exactly the same rows for 1h, 24h, 7d and 45d (5, 10, 12 and 18 regions; 73, 2,024, 13,822 and 76,083 transmissions). `TestScopeStatsByRegion_CountsMatchWindow` also checks the per-region counts against a direct count of the window. ## Performance ### The query alone, on a copy of the production database 311k transmissions, 76k of them region-scoped over 45 days; best of three: | window | before | after | |---|---|---| | 1h | 2,530 ms | 0.1 ms | | 24h | 2,570 ms | 11.6 ms | | 7d | 2,876 ms | 124 ms | ### The endpoint, on staging Cold `GET /api/scope-stats` (the response cache is 30 s, so each sample is a full compute) on a staging host with a copy of a production database and its live feed; two samples each, 32 s apart: | window | master | this PR | |---|---|---| | 1h | 3.08 s, 2.83 s | **16 ms, 12 ms** | | 24h | 2.69 s, 2.79 s | **131 ms, 101 ms** | | 7d | 3.35 s, 4.53 s | **626 ms, 663 ms** | ### Go benchmark `BenchmarkGetScopeStats24h`: all five scope-stats queries for 24h, on 300k transmissions spread over 45 days (a quarter region-scoped), with production's tables and both competing indexes and `ANALYZE` run as production does (#2072). 10 runs: ``` cpu: AMD Ryzen 5 3600XT 6-Core Processor │ master │ this PR │ GetScopeStats24h-12 33.186m ± 3% 8.058m ± 6% -75.72% (p=0.000 n=10) ``` The fixture is smaller and better cached than production, which is why the absolute times are lower than above; the plan change is the same. ## Tests - Red commit `a7dcf98f`: moves the query into `scopeStatsByRegionQuery` (unchanged), adds the fixture, the benchmark, `TestScopeStatsByRegion_UsesFirstSeenIndex` (`EXPLAIN QUERY PLAN` must use `idx_transmissions_first_seen` and not `idx_tx_scope_name`; fails on `master` with exactly the production plan shown above) and `TestScopeStatsByRegion_CountsMatchWindow` (passes on both commits; guards correctness). - Green commit `f9ba986b`: the `+` hints and a comment explaining them. - All existing scope-stats tests pass; `make test`: 16 of 16 modules pass. ## Review notes - Only the per-region query changes; the summary, non-transport, time-series and adverts-by-role queries already search by `first_seen`. - If `idx_transmissions_first_seen` were ever missing, the query would still run correctly (full scan); `+` only removes `idx_tx_scope_name` from consideration, unlike `INDEXED BY`, which would make the query fail. - `/api/scope-stats` still has no singleflight around its 30 s cache (unlike stats, channels and the scope audit). With the query now well under a second it matters much less; left for a separate change. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - No issues on the added lines. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
95672ac2be |
perf(server): share /api/nodes responses for identical queries for 15 s (#2165)
## Summary
`/api/nodes` scans every node row into a map, enriches it (hash size,
multi-byte capability, relay info, usefulness scores, declared scope)
and encodes the page on **every** request. On one production instance (2
vCPU / 2 GB) that was **95 GB of allocation in 45 hours, about 6.7 MB
per request and a quarter of all bytes allocated**, on a process where
garbage collection was 87% of CPU.
The queries repeat heavily. In that instance's nginx log for one day,
three pages of one query (`limit=500&offset=0|500|1000&lastHeard=30d`)
were 10,000 of the 14,053 `/api/nodes` requests.
This PR keeps the encoded response for 15 s per distinct query:
- `handleNodes` serves the cached bytes when it can; otherwise it builds
through `buildNodesResponse` (the old handler body, unchanged apart from
returning instead of writing) inside a singleflight, so concurrent
identical misses share one build.
- **Cache key:** blacklist generation | hidden-prefix generation |
sorted query string. Parameter order does not split the cache, and any
change to either filter list misses it at once. This is the pattern
`/api/nodes/{pubkey}/reach` already uses (#1629).
- `setGeoFilter` drops the cache, so a geo-filter change through the
admin endpoint shows immediately.
- **Bounded twice:** by `maxCacheEntries` (256) like the other keyed
caches, and by bytes, since a full-list response is several MB. Expired
entries are dropped on every insert, the total is capped at 32 MB, and a
single response over 8 MB is not cached.
- **15 s** matches the relay/usefulness caches the response is built
from (#1257); the client already keeps its own node list for 90 s.
Related to #2146: `/api/nodes` is in its endpoint table (144 s max on
one instance). This PR reduces how often the handler runs and allocates,
not the SQL wait the issue is about, so it is not a fix for it.
## Performance
### Go benchmarks
1,100 nodes, half repeaters, 10 runs each:
```
cpu: AMD Ryzen 5 3600XT 6-Core Processor
│ master │ this PR │
│ sec/op │ sec/op vs base │
HandleNodesRepeated-12 5991.14µ ± 3% 41.21µ ± 7% -99.31% (p=0.000 n=10)
HandleNodesUnique-12 6.010m ± 2% 5.967m ± 0% ~ (p=0.393 n=10)
│ B/op │ B/op vs base │
HandleNodesRepeated-12 1611.0Ki ± 1% 240.9Ki ± 0% -85.05% (p=0.000 n=10)
HandleNodesUnique-12 1.566Mi ± 1% 1.753Mi ± 1% +11.96% (p=0.000 n=10)
```
`Repeated` is one query asked over and over (the production pattern).
`Unique` makes every request a different query, so nothing is served
from cache: the time is unchanged, and it allocates 12% more because the
body is now built in memory so it can be kept.
### Expected hit rate
Replaying that day's 14,053 `/api/nodes` requests through a cache of
each lifetime:
| TTL | served from cache |
|---|---|
| 5 s | 39% |
| 10 s | 47% |
| **15 s** | **52%** |
| 30 s | 58% |
| 60 s | 69% |
### Staging A/B
The same 30 minutes of real production API traffic (5,640 requests, 279
of them `/api/nodes`), replayed at 2x on a staging host against a copy
of a production database, on `master` and on this branch:
| `/api/nodes` | master | this PR |
|---|---|---|
| p50 | 47.1 ms | **31.0 ms** |
| p95 | 165.4 ms | **114.4 ms** |
| max | 759.9 ms | 321.7 ms |
| allocated by `handleNodes` during the replay | 1,000 MB | **594 MB
(-41%)** |
| total allocated by the process | 3,704 MB | 3,349 MB (-10%) |
The -41% matches the expectation from the hit rate and the miss-path
cost: about 0.48 × 1.12 + 0.52 × 0.15 ≈ 0.61 of the old allocation.
## Tests
- Red commit `ab33bc29`: a test-only build hook, the tests and both
benchmarks. Three fail on `master`: identical queries are built once,
reordered parameters share an entry, ten concurrent misses share one
build. Three more guard against over-caching and pass on both commits:
different queries are built separately, an expired entry is rebuilt, a
geo-filter change drops entries.
- Green commit `ea0d1cf2`: the cache, plus
`TestHandleNodes_FilterListChangesMissCache` (hidden-prefix and
blacklist changes miss the cache) and `TestNodesCache_BoundedByBytes`
(stays under 32 MB; an oversized response is not cached).
- One existing test changed: `TestHandleNodesRegionUsesStore` already
clears the store's own 30 s region cache by hand before its second
request; it now clears this cache on the same line.
`TestHiddenNamePrefix_1181_NodesList`, which changes the hidden-prefix
list between requests, passes unchanged because the generation is in the
key.
- `go test -race` on the new tests passes. `make test`: 16 of 16 modules
pass.
## Review notes
- The response depends only on the query string and server state, not on
who is asking (no auth or header-dependent output in the handler), so a
shared cache is safe.
- `writeJSONBytes` writes the cached bytes with the same `Content-Type`
and status as `writeJSON`; the body has the trailing newline
`json.Encoder` adds, so output is byte-identical.
- No new `map[string]interface{}` added (the existing node maps are
unchanged; converting them to a struct is a separate, larger job).
- The TTL and byte budget are constants. They could be operator settings
later (AGENTS.md rule 8); not added here.
## SonarLint
SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.
- **S3776 on `buildNodesResponse`:** complexity 91. That is upstream's
existing handler body moved unchanged (it is also 91 on `master`); this
PR makes the handler itself simpler.
- **Fixed** in `25d40950`: the test hook moved to the cache-miss path,
so the builder stays identical to the old body rather than one point
above it. Splitting that function is a refactor of its own.
## Checked against the repo rules
- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
323b8040d4 |
perf(frontend): throttle nav stats to one refresh per 15 s, none in hidden tabs (#2162)
## Summary The nav bar's packet / node / observer counts were refreshed three ways at once: a 15 s timer, every batch of live WebSocket packets (`debouncedOnWS`, 250 ms), and the `/stats` client cache being cleared every 5 s while packets arrive. None of them checked whether the tab was visible. So every open tab, including hidden ones, asked `/api/stats` every few seconds for as long as it stayed open. On one production instance's nginx log for a single day, `/api/stats` was **135,654 of 187,411** CoreScope API requests (72%), about 20,000 a day for each browser that kept a tab open. Each response is cheap thanks to the 10 s server-side cache (#1910), so this is about request volume, bandwidth and log noise, not server CPU. This PR adds `createVisibleThrottle()` to `app.js` and routes all three triggers through one throttled refresh: - at most one `/api/stats` request per 15 s from a visible tab (the old timer's own period); - none from a hidden tab; - one immediate refresh when a hidden tab becomes visible again. No existing issue covers this. ## Performance Measured with headless Chromium, counting `/api/stats` requests from a real tab: 60 s visible, then 60 s with the page hidden (`document.hidden` and a `visibilitychange` event), then back to visible. | `/api/stats` requests | production (current release) | staging, `master` | staging, this PR | |---|---|---|---| | visible tab, per minute | 9 | 4 | 4 (at +10, +25, +40, +55 s) | | hidden tab, per minute | 6 | 5 | **0** | | on return to visible | n/a | n/a | 1, immediately | The same check confirmed the nav bar still renders its counts (`… pkts · … nodes · … obs`). ## Tests - Red commit `2c81a3e4`: `tests/unit/test-nav-stats-throttle.js` loads the real `public/app.js` in a `vm` context. `createVisibleThrottle` lands as a pass-through stub, so the test fails on its first assertion (40 refreshes where 1 is expected). - Green commit `f0ba92ab`: the throttle, and the nav stats timer, WS hook and a new `visibilitychange` listener calling one throttled trigger. All 6 cases pass: one refresh per burst, again after the interval, at most 241 an hour, none while hidden, one on return, and a source check that `updateNavStats` is wrapped and not called by the timer directly. - Added to `test-all.sh` in sorted order; `test-test-inventory.js` passes. - Frontend suites: 194 of 195 pass. The one failure, `test-issue-1956-release-routing.js`, needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally. ## Review notes - `home.js` keeps its own one-shot `/stats` load on the home page; only the nav bar's polling changed. - A tab opened in the background loads its counts when it is first shown. - The 15 s interval is a constant at the call site. Could be exposed in the customizer later (AGENTS.md rule 8); not added here. - No Playwright E2E added: the change is timing behaviour, covered by the unit test against the real file and by the browser measurement above. Happy to add an E2E if wanted. ## SonarLint SonarQube Community rules (the SonarLint engine) were run on the files this PR touches, and only issues on lines it adds are listed. - **Fixed** in `9eb94ea1`: `node:` specifiers (S7772). - **Kept on purpose:** - S7773 wants `Number.isFinite`/`Number.parseInt` for the `isFinite` and `parseInt` handed to the `vm` sandbox, but `app.js` calls the global functions, so those are the ones it needs. - S1523 flags `vm.runInContext`, which is how the frontend unit tests load the repo's own `public/app.js`. ## Checked against the repo rules - Rebased onto `master` (`673442cc`). Each red commit was re-run and still fails on its assertions; each green commit passes (TDD.md). - `make test`: all 16 modules pass. Frontend suites: all pass except `test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME` and fails the same way on `master` locally (#2160 addresses it). - No private data: test fixtures and examples use the codes the existing tests use (SJC/SFO/LAX, `#be`-style scopes). - No new `map[string]interface{}`, no hard-coded colours. `git add` was explicit, and there is one logical change per commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
be6aeffe45 |
perf(store): drop the duplicate subpath count map (#2163)
## Summary
The subpath index keeps every contiguous run of 2-8 hops from every
stored transmission's path, in two maps:
- `spIndex`: key → count
- `spTxIndex`: key → transmissions containing that subpath, one entry
per occurrence
Both were always updated together (the only production callers were
`addTxToSubpathIndexFull` / `removeTxFromSubpathIndexFull`, which touch
both), so every count equalled `len(spTxIndex[key])`. The existing
`TestSubpathTxIndexPopulated` asserted exactly that. The index therefore
held its 1M+ keys twice.
This PR removes `spIndex`. The two analytics consumers read `len(txs)`,
and the add/remove helpers take only the transaction map. No API
response changes.
No existing upstream issue covers this; it came out of profiling a
production instance where garbage collection was 87% of CPU (30 s CPU
profile: `gcDrain` 39.9 s of 45.9 s sampled) with an 805 MB live heap
against a 1 GiB `GOMEMLIMIT`. That heap profile put 186 MB in this
index; on that instance's 7-day window there are 1.2M distinct subpath
keys (79% seen once) and 3.07M entries.
## Performance
### Go benchmark
`BenchmarkBuildSubpathIndex` builds the index for a synthetic,
production-shaped window: 41k transmissions with path lengths drawn from
the production hop-count histogram, and hops from random walks over a
400-node graph with 3-byte prefixes. That gives 1.48M keys, 85% seen
once (production: 1.2M, 79%). No real data is used. 10 runs each:
```
cpu: AMD Ryzen 5 3600XT 6-Core Processor
│ master │ this PR │
│ sec/op │ sec/op vs base │
BuildSubpathIndex-12 2.157 ± 1% 1.610 ± 5% -25.33% (p=0.000 n=10)
│ retained-B │ retained-B vs base │
BuildSubpathIndex-12 238.5Mi ± 0% 187.7Mi ± 0% -21.31% (p=0.000 n=10)
│ B/op │ B/op vs base │
BuildSubpathIndex-12 572.0Mi ± 0% 465.6Mi ± 0% -18.61% (p=0.000 n=10)
```
(`retained-B` is the live heap the index keeps after a forced GC: 250.1
MB → 196.8 MB.)
### Staging heap
Live heap after a forced GC (`/debug/pprof/heap?gc=1`), on a staging
host running a copy of a production database and its live MQTT feed,
attributed to the subpath index builder:
| | master | this PR |
|---|---|---|
| subpath index (`buildSubpathIndex`, cum) | 172.4 MB | 116.1 MB |
| packets in memory at the time | 27,521 | 24,086 |
| per packet | 6.27 KB | 4.82 KB (**-23%**) |
The two snapshots were taken minutes apart with different numbers of
packets loaded, hence the per-packet figure; it agrees with the
benchmark's -21%.
The same A/B replay used for the other PRs (30 minutes of real
production API traffic) showed no latency change (p50 10.3 ms vs 10.7
ms, p95 59.5 ms vs 52.9 ms, 0 errors). GC count and CPU were lower in
this run (44 → 28 collections, 13.0% → 9.6% CPU), but runs on this
shared host vary by that much on their own, so no claim is made from
them.
## Tests
- Red commit `ded359ff`: the fixture, `TestSubpathFixtureShape` (keeps
the fixture production-like: 60-95% singleton keys),
`BenchmarkBuildSubpathIndex`, and `TestSubpathIndexRetainedHeap`, which
allows at most 225 MB for the 41k fixture. It fails on `master` with
250.1 MB.
- Green commit `e399c561`: the change. The heap test passes at about 197
MB.
- `make test`: 16 of 16 modules pass; `go test -race` over the subpath,
eviction, path-hop and resolved-index tests passes.
Test changes in the green commit are mechanical, as required for a
structural change:
- struct literals stop setting `spIndex`; where a literal did not set
`spTxIndex` it now does, because appending to a nil map panics where the
old code skipped the transaction map when passed `nil`;
- count assertions read `len(spTxIndex[key])`;
- the old "spIndex and spTxIndex agree" check becomes "no key is kept
with an empty list", which is what keeps `len()` a correct count.
## Review notes
- Removal still takes out one occurrence per subpath occurrence and
deletes a key whose list becomes empty, so `len()` stays equal to the
old count through eviction.
- The store's memory estimate (`perSubpathEntryBytes`) is unchanged
here; accounting accuracy is a separate item.
- No new `map[string]interface{}`.
## SonarLint
SonarQube Community rules (the SonarLint engine) were run on the files
this PR touches, and only issues on lines it adds are listed.
- **Fixed** in `8af9582f`: S3776 (cognitive complexity 16 > 15) in the
test fixture. Path generation is split into graph and hop-count helpers;
the random draws stay in the same order, so the fixture is unchanged.
## Checked against the repo rules
- Rebased onto `master` (`673442cc`). Each red commit was re-run and
still fails on its assertions; each green commit passes (TDD.md).
- `make test`: all 16 modules pass. Frontend suites: all pass except
`test-issue-1956-release-routing.js`, which needs CI's `GITHUB_REF_NAME`
and fails the same way on `master` locally (#2160 addresses it).
- No private data: test fixtures and examples use the codes the existing
tests use (SJC/SFO/LAX, `#be`-style scopes).
- No new `map[string]interface{}`, no hard-coded colours. `git add` was
explicit, and there is one logical change per commit.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01KaR4xu8gVykQBkH8yG8m6H
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
7f2c999266 |
fix(ingestor): record cancelled background migrations separately (#2161)
Refs #1739. Depends on the small local-validation fix in #2160; merge that first. This branch contains its prerequisite commit plus one migration-status commit. An interrupted background migration currently records `failed`, making an operator shutdown look like a migration fault. Errors matching `context.Canceled`, including wrapped errors, now record `cancelled` and retain the error and completion time without logging `FAILED`. Ordinary errors, deadline expiry and panics remain failures even when the context is cancelled; a successful callback still records `done`. Cancelled work follows the existing retry path on the next run. Regression tests cover classification and clearing stale completion/error fields before retrying, then successful completion. The migration runbook documents the distinction. This covers the existing ingestor bookkeeping portion of #1739. The issue's server status mapper and banner depend on closed, unmerged #1735 and are absent from current master; this PR does not add that API/UI and leaves #1739 open for the remaining scope. No new configuration, customizer setting or schema migration is required. Validation: direct and wrapped cancellation failed the new regression before the change; focused tests and the race check pass afterward. Native full server/ingestor suites and all 194 standalone frontend suites pass with #2160. The browser smoke suite passed 148 checks, with three fixture-dependent skips and two pre-existing disabled cases. Windows-only filesystem capability skips are documented in #2160; Linux CI retains the security assertions. Performance impact: one error-chain check when a background migration finishes, using the existing single status update. No packet-ingestion loop, batching, polling or retention behavior changes, and no speedup is claimed. |
||
|
|
a5b60ca594 |
test: make local validation independent of runner environment (#2160)
Fixes #2159. Related to #1995. Standalone release tests failed when `GITHUB_REF_NAME` was absent and could use an unrelated runner tag when it was present. Supply the simulated tag explicitly; exercise the real release-notes step with both missing and unrelated host values, including present and absent notes files. Two filesystem fixtures now report unsupported Windows capabilities explicitly: POSIX permission bits, and symlink creation denied with error 1314. Unix mode checks and all symlink safety assertions remain intact; other symlink setup errors still fail. Production code and workflows are unchanged. This is 29 added lines across three test files. Validation: reproduced all three failures before fixing them; focused regressions, the full standalone frontend suite, and native Go server/ingestor suites pass. The POSIX permission and unprivileged Windows symlink fixtures are explicitly skipped on this host and remain enabled in Linux CI. The browser smoke suite passed 148 checks, with three fixture-dependent skips and two pre-existing disabled cases. No configuration or customizer changes. |
||
|
|
0737f10a5d |
feat(notify): optional external node alerts feed and the node.external event (#2177) (#2178)
Closes #2177. ## What Node notifications can mail alerts computed outside CoreScope. With `userManagement.notifications.externalAlerts` configured, users can choose a new event, `node.external`, which is off by default and only offered when configured: ```json "notifications": { "enabled": true, "externalAlerts": { "url": "http://coverage-job/alerts.json", "maxAgeHours": 48, "label": "Coverage" } } ``` The feed: ```json { "generatedAt": "2026-10-11T01:35:12Z", "alerts": [ { "pubkey": "<64 hex>", "key": "SE", "text": "never heard to the south-east", "url": "https://example.org/coverage" } ] } ``` Mail lines: `<name>: Coverage: never heard to the south-east, <time>` when an alert appears, `<name>: Coverage resolved (SE), <time>` when it is gone from a fresh feed. ## How - `cmd/server/notify_external.go`: feed types, `parseExternalFeed` (RFC 3339 `generatedAt`, age check, per-alert validation), `fetchExternalFeed` (200 only, 1 MB cap, 10 s timeout). - `notify_eval.go`: `evalExternal`. Subjects are `<pubkey>/<key>`; a `<pubkey>/*` row is the per-node baseline, so the first fresh feed after watching stores the current alerts without a mail (like `foreign.new`), and later alerts mail at once. `External == nil` (not configured, nobody chose it, or stale) leaves every `node.external` state as it is. - `notify_service.go`: the feed is fetched only when configured and at least one user chose the event. Each new reason the feed is unusable is logged once, and so is the recovery. Mail text and link: the alert's `url`, otherwise the node page. - `notify_handlers.go`: `availableEvents` and the new `externalLabel` include the event only when configured; PUT answers 400 for `node.external` otherwise. - `internal/users/notify.go`: event constant, `OptionalNodeNotifyEvents` (canonical order: node, optional, admin), `RemoveWatch` also deletes the node's `<pubkey>/...` rows. - Account page label, OpenAPI descriptions, `config.example.json` comment, `docs/user-guide/accounts.md` section. Decisions on the three open questions in #2177: the generic name `node.external`; the resolved mail names the key instead of repeating the text, so there is **no `users.db` schema change** (a bump would also make older binaries refuse the file); one feed. ## Invariants - Read/write separation (#1283): the server only reads the feed over HTTP. State rows go to `users.db` through `internal/users`, like the other events. - A dead, garbled or old feed never mails "resolved": `TestEvaluateExternalStaleFeedChangesNothing` and `TestNotifierMailsExternalAlertsAndSkipsAStaleFeed` (HTTP 502 and a 72 h old feed in a row, then a fresh empty feed resolves). ## Performance The evaluator adds one pass over the stored states (already read every tick) plus, per user with the event on and per watched node, the node's current alerts. Bounded by `maxWatchesPerUser` (50) times the alerts per node. The feed is fetched once per tick (default 5 minutes) only when someone chose the event, capped at 1 MB. No packet-store lock is taken for it. Nothing changes for instances without `externalAlerts`. ## Tests - `cmd/server`: `go test` green locally (Windows; `TestSaveGeoFilterPreservesFileMode` skipped as it needs Unix mode bits). New: `notify_external_test.go` (feed parsing and limits, HTTP fetch 200/404/oversized, config resolution, evaluator baseline/new/resolved/returning/stale/not chosen/notifications off, notifier end to end with a fake feed, account routes). - `internal/users`: green, 2 new tests (optional event and canonical order, `RemoveWatch` deletes external subjects only for that node). - Frontend: `test-all.sh` green, 1 new test in `test-notifications-ui.js` (label from `externalLabel`, escaped). - `scripts/check-xss-sinks.sh --file public/notifications.js`: clean. ## Not done - No E2E test: the account page change is one label, covered by the unit test. - Not run against a live instance yet. The first feed in use is the coverage job of on8ar.eu (`alerts.json`, 29 alerts at the moment). Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c2a590b741 |
feat(companions): linked-only ingest, account Companions and My coverage (#2176)
Part 3 of 3 for #2173 (companion linking): the optional linked-only ingest filter, the account page and the "My coverage" toggle, plus the docs. Builds on #2174 and #2175 (both merged); `upstream/master` is merged in, so the diff shows only this part. ## Changes in this PR **Ingestor: linked-only ingest** (`cmd/ingestor/linked_companions.go`) - With `clientRxCoverage.requireLinkedCompanion` on (default off) and user management enabled, every `meshcore/client/<pubkey>/...` message whose pubkey is not in `companion_links` is dropped before any client handler runs, and counted (`client_unlinked_dropped` in the stats file). - The ingestor opens `users.db` read-only (`mode=ro`, raw SQL, no `internal/users` import), like the approved-channels reader, so the #1283 write invariant holds. `TestLinkedCompanionSetIsReadOnly` pins it. - The set refreshes every 60 s. A miss re-reads at most once per 5 s, so a companion linked a moment ago is accepted within seconds while a flood of unknown pubkeys costs one read per 5 s. - It passes everything until its first successful read, and when `users.db` has no `companion_links` table yet (a server older than this). After the first read it keeps the last set on a read error. It is a filter, not a security boundary: anyone with the shared broker credential can still publish under a linked pubkey. **Frontend** - Account page: device tokens show under Devices as "CoreDrive RX" with their label; a new Companions section lists linked companions (name, short key, linked since, last seen) with an Unlink button. - Coverage page (`#/rx-coverage`): a "My coverage" toggle for logged-in users (`mine=1`). Only the newest coverage reply is drawn, and a refused `mine=1` turns the toggle off. Under "My coverage" the page does not ask for the "nothing received" cells from #2172, and their legend entry is hidden. **Docs**: user guide (accounts: devices, companions), `docs/client-rx-coverage.md` (linking, linked-only ingest, the CORS exception), `config.example.json`, and the spec status. ## Performance The filter runs on every client message. Benchmarks with 1000 linked companions (12-core dev box, `go test -bench`, 3 runs each): | Path | Time | Allocations | |---|---|---| | linked pubkey (lock-free map lookup) | 54 ns/op | 0 | | linked pubkey, parallel | 11-14 ns/op | 0 | | unlinked pubkey inside the 5 s miss gap | 85-95 ns/op | 0 | | one `users.db` re-read | 0.38 ms | 3030 | The unlinked benchmark asserts that no extra `users.db` read happens. A re-read happens at most once per 5 s on a miss and once a minute from the refresh loop. ## Verification - `cmd/ingestor`: 15 new tests (filter off changes nothing, linked passes, unlinked dropped before every handler, miss refresh accepts a new link, the 5 s cap holds under a flood, unlink honoured after the periodic refresh, missing table and missing file, fail-closed only after the first read, read-only open, startup, stats counter) and 4 benchmarks. Full ingestor suite green. - Frontend unit tests for the account page (device tokens, companions list escaping, empty state, Unlink confirm and cancel, load errors) and `tests/unit/test-rx-coverage-mine.js`. - E2E `tests/e2e/test-user-management-e2e.js` gains a companion step (device token under Devices, a companion linked with a real Ed25519 signature under Companions, Unlink). Run locally against the e2etest build: 28/28. The coverage suites (`test-rx-coverage-gaps-e2e.js`, `-noise-`, `-viewport-`) still pass. - Frontend unit suite (`test-all.sh`) green. ## Not in this PR - Narrowing the "nothing received" cells to the caller's companions under "My coverage" (the track query takes one companion today). - Per-user broker credentials or ACLs, which would turn the filter into an access control (non-goal in the spec). --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
b2cda21a27 |
feat(server): companion linking routes, device tokens and bearer auth (#2175)
Part 2 of 3 for #2173 (companion linking): the server routes, bearer auth and the link handshake. Builds on #2174 (merged); `upstream/master` is merged in, so the diff shows only this part. Everything here is behind `userManagement.enabled` and answers 404 with it off. ## Changes in this PR **Device tokens and bearer auth** - `POST /api/auth/device-token`: email + password returns a bearer token, stored as a `kind = 'device'` session with a label. Same failure answers and the same rate limit as login. - Server-enforced scope for a bearer token: `/api/auth/me`, `/api/auth/logout`, `/api/account/companions*`, `/api/account/settings`. Everything else answers 403. - A bearer header only counts when there is no session cookie, so an auth proxy that adds its own `Authorization` header does not hide a logged-in browser. Cookies never accept a device token. - The sessions list on the account API carries `kind` and `label`; logout with a bearer token revokes that device. **Linking a companion** - `POST /api/account/companions/challenge` `{pubkey}` returns `{challenge, expiresAt, host}`: 32 random bytes, single use, 5 minutes, bound to the caller and the pubkey. - `POST /api/account/companions` `{pubkey, challenge, signature, name}`: the signature is the companion's Ed25519 signature over `corescope-link:<host>:<challenge>` (`internal/sigvalidate.VerifyMessage`). The challenge is consumed in every case. - A companion linked to another account moves to the caller (the newest proof wins). Both users get a `companion.transfer` audit row and the previous owner a mail. That mail is a security notice: it is sent whatever the node notification settings are. - `GET /api/account/companions` (with `lastSeenAt` from `client_receptions`) and `DELETE /api/account/companions/{pubkey}`. - Linking does not touch the synced `meshcore-my-nodes`: a companion is not a node to monitor. - Admin user detail lists linked companions; the account data export includes companions and device sessions. **Coverage and config** - `GET /api/rx-coverage?mine=1`: only coverage of the caller's linked companions (401 when logged out, 403 for a device token). With `gaps=1` (#2172) the reply carries no `gaps` member under `mine=1`, because the track query takes one companion and cannot narrow to a set yet (`TestRxCoverageMineOmitsGaps`). - `clientRxCoverage.requireLinkedCompanion` (default off) is read and reported in `/api/config/client`, which also announces `userManagement.companionLinking`, so CoreDrive RX shows its login only where linking exists. The ingestor side is part 3. - CORS: for origins in `corsAllowedOrigins`, the device-token routes alone also allow POST/PUT/DELETE and the `Authorization` and `Content-Type` headers. No credentialed CORS. Documented in `config.example.json`. - OpenAPI and `docs/api-spec.md` cover the new routes. ## Verification - 26 new tests in `cmd/server` and `internal/sigvalidate`, among them: `TestCompanionLinkFlow`, `TestCompanionLinkRejections` (signature by another key, malformed signature, challenge bound to another pubkey, unknown challenge; reuse and expiry are covered by the store tests in #2174), `TestCompanionTransfer`, `TestCompanionTransferMailIgnoresNotificationSettings`, `TestBearerScopedRoutes`, `TestBearerIgnoredWithCookie`, `TestDeviceTokenSharesLoginRateLimit`, `TestCORSBearerRoutes`, `TestCompanionRoutesAbsentWhenOff`, `TestRxCoverageMine`, `TestRxCoverageMineOmitsGaps` (fails without the `mine == nil` guard), `TestVerifyMessage`. - `cmd/server`, `internal/sigvalidate` and `internal/users`: `go vet` and `go test ./...` green, `gofmt` clean. - The linking flow runs on analyzer.on8ar.eu since 9 October 2026 with CoreDrive RX v1.19.0 (a real companion linked through the signed challenge). ## Not in this PR - The ingestor's linked-only filter, the account page (Devices, Companions), the "My coverage" toggle and the user docs: part 3. - Per-user broker credentials and per-user API keys (non-goals in the spec). --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
9897d52cf6 |
feat(users): companion linking store (schema v6, device sessions, link challenges) (#2174)
Part 1 of 3 for #2173 (companion linking). This PR adds only the `users.db` side: the schema and store methods the server (part 2) and ingestor (part 3) build on. No route, no UI and no behaviour change for a running server. ## Changes in this PR - **Design spec** `docs/specs/2026-10-08-companion-linking-design.md`, including the optional linked-only ingest. - **Schema v6** (`internal/users/schema.go`): - `sessions` gains `kind` (`'web'` or `'device'`), `label` and `scopes` (`''` = full, for web sessions). - `companion_links (pubkey PRIMARY KEY, user_id, name, linked_at)`, cascading on user delete. - `link_challenges`: single-use challenges bound to a user and a pubkey, with an expiry. - **Device sessions** (`sessions.go`): create a device session with a label and scopes, list sessions with kind and label. - **Challenges** (`challenges.go`): issue and consume a challenge exactly once; consumed, expired, or bound to another user or pubkey all fail. - **Companion links** (`companions.go`): link, list, unlink. Linking a pubkey that belongs to another user moves it and reports the previous owner (transfer detection), so the server can audit and mail. - **Janitor**: the existing janitor also prunes expired challenges (`cmd/server/auth_service.go`, 3 lines). ## Upgrade note The store migrates `users.db` to v6 on start. A binary that only knows v5 refuses to open a v6 file (`users: database schema version 6 is newer than this binary supports (5)`), so rolling back after this lands needs the pre-upgrade `users.db` backup. ## Verification - `internal/users`: new tests for the v6 migration from v5, device sessions, challenges (single use, expiry, binding) and companion links (transfer, cascade on user delete). `go vet` and `go test ./...` green. - `cmd/server` full suite green against the v6 store. ## Not in this PR - Routes, bearer auth and the link handshake: part 2. - Ingestor filter, account page and "My coverage": part 3. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
f3004d977e |
feat(rx-coverage): show where a companion drove and received nothing (#2172)
The Mobile RX coverage page (`#/rx-coverage`) shows only cells where a companion received something. Where a companion drove and received nothing is just as useful for operators, and the data is already there: the RF sample track (`client_rf_samples`, #1905/#1906) and the receptions (`client_receptions`) use the same hex grid (`hexCellAt`, `zoomToHexRes`). This PR adds those cells as a hatched layer under the reception cells. ## Changes in this PR **Server** - `GET /api/rx-coverage?...&gaps=1` adds a `gaps` member to the FeatureCollection: cells with moving (non-stationary) track samples and no reception in the same `days` window and `bbox`, each with `samples`. `gaps_truncated` is set when more than 5000 cells qualify (the most-sampled are kept, same pattern as `coverageFeatureCap` and `rfNoiseFeatureCap`). - Gaps are only returned when `clientRfSamples` is on and there is no `node` filter: with `?node=` the question would be "this node was not heard", which these cells do not answer. `?rx=` limits both the track and the receptions to that companion. - Without `gaps=1` the payload is unchanged: the `gaps` member is omitted. - New code lives in `cmd/server/rx_coverage_gaps.go` (pure `aggregateGaps` plus `queryTrackPoints`); the handler change in `rx_dashboard.go` is 14 lines. - `/api/rx-coverage` is now documented in `openapi.go` and removed from `openapi_known_gaps.json` (#1670). **Frontend** (`public/rx-coverage.js`, `public/node-reach-coverage.css`) - A "Nothing received" toggle on the signal layer, on by default, only rendered with `MC_CLIENT_RF_SAMPLES`. Hidden on the noise layer. - Gap cells use a 45-degree hatch (an SVG pattern in a dedicated renderer, so they sit under the reception cells). The hatch is denser for cells driven through more often (1 / 2-3 / 4+ samples). Tooltip: sample count. Legend entry: "driven, nothing received". - Off state is deep-linkable as `gaps=0`. - New token `--nq-cov-gap` for light and dark. Plain grey (`--nq-cov-grey`) already means "received, no SNR", so the hatch keeps the two apart. ## Threshold: one sample Measured on a live network (analyzer.on8ar.eu, Belgium, September to October 2026), with cells of about 250 m: a cell with track samples and no reception in one half of a month, driven again in the other half. How often did it receive something the second time? | Minimum samples | Cells "nothing received" | Driven again | Received something later | |---|---|---|---| | 1 | 3971 | 1263 | 29% | | 2 | 1565 | 592 | 32% | | 3 | 619 | 243 | 35% | | 5 | 142 | 63 | 40% | A higher threshold hides cells without making the rest more reliable, so the threshold is 1 and the wording is "nothing received in this period", never "no coverage". Using every decoded packet as reception evidence (including 1-byte path hashes) gave 25% to 32%, so the capture rule for 1-byte hashes is not what drives this. The median cell has 2 samples. ## Verification - Go: 7 new tests in `rx_coverage_gaps_test.go` (gap/no gap, stationary excluded, single sample, cap keeps the most-sampled, endpoint, `?rx=` filter, no `gaps` member when not asked / with `?node=` / with RF samples off). The endpoint tests fail without the handler change. Full `cmd/server` suite green locally, `go vet` and `gofmt` clean. - E2E: `tests/e2e/test-rx-coverage-gaps-e2e.js` (4 steps: no toggle with RF samples off, on by default with hatched cells under the reception cells, toggle off with `gaps=0` and deep link, noise layer hides it). 2 steps fail without the frontend change. Existing `test-rx-coverage-noise-e2e.js` and `test-rx-coverage-viewport-e2e.js` still pass. Listed in `scripts/non-unit-tests.json`. - Frontend unit suite (`test-all.sh`) green. - Real data on a staging instance (7 days, bbox 50.8,4.9,51.3,5.8, z=12): 478 reception cells and 379 gap cells, no cell in both, reception features identical with and without `gaps=1`, 0.48 s to 0.52 s per request both ways. - Browser: checked on staging with real data (light theme), and light, dark and 390 px wide with mocked data. ## Not in this PR - A per-node variant on the Reach page ("this repeater transmitted while a companion was nearby, and was not heard"). That needs the transmit moments from `observations` and is a heavier query. - `deploy.yml` does not get a line for the new E2E file, like the other entries that are only in `non-unit-tests.json` (#2037). Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
673442ccb9 |
perf(ingestor): refresh planner statistics 24h after the last refresh, not after every start (#2158)
Refs #2146. ## Situation The daily planner statistics refresh (#2058) runs its first `ANALYZE` 2 minutes after every ingestor start. On a cold page cache that bounded `ANALYZE` (analysis_limit 10000) takes minutes on a 10-11 GB database, and it holds the ingestor's single write connection the whole time: | Where | Logged after a restart | |---|---| | production instance (11.5 GB) | `component=analyze duration=469855.2ms`, plus `[neighbor-build] SLOW tick: 7m49.98s` behind it | | staging instance, 12 restarts logged on 2026-10-09 | 4m11s to 9m11s each; 7 of the 12 took over 7 minutes | `EnsurePlannerStats` already documents what that costs: ingest stalls and buffers for those minutes. The server's reads compete with the scan for the disk during that window. Every restart pays this for statistics that are at most a day old. ## My change in this PR - `RefreshPlannerStats` records when it succeeded, in a file next to the database (`<db>.planner-stats-refreshed`, RFC 3339). - The refresh loop waits until that time is 24 hours old: at least the existing 2-minute stagger, at most 24 hours (`nextPlannerStatsRefresh`). A restart within a day of the last refresh runs no `ANALYZE`. - No recorded time (the first start after upgrading, or a new database) refreshes after 2 minutes as before, and records it. `EnsurePlannerStats` records its build too, so a new database no longer runs a second `ANALYZE` 2 minutes after the first. - A failed refresh keeps the daily rhythm instead of retrying every 2 minutes, the same as the old ticker. ## Verification On staging (`/var/lib/docker/volumes/.../meshcore.db`, 168 h window): ``` first start: [analyze] next planner statistics refresh in 2m0s [analyze] planner statistics refreshed in 7m26.086s (analysis_limit=10000) [analyze] next planner statistics refresh in 24h0m0s meshcore.db.planner-stats-refreshed: 2026-10-09T21:08:34Z docker restart, then: [analyze] next planner statistics refresh in 24h0m0s (no ANALYZE and no db-slow-writer line in the following 6 minutes) ``` Tests: - `TestNextPlannerStatsRefresh`: never refreshed, 1 hour ago, 23h59m ago, two days ago, a recorded time in the future. - `TestRefreshPlannerStatsRecordsWhen`: a refresh records about now, and a reopened store reads it back as a wait of about 24 hours. - `TestRefreshPlannerStatsDisabledRecordsNothing`. - The existing #2058 planner-statistics tests pass unchanged. `cmd/ingestor`: `go test ./...` green. ## Not in this PR - The time lives in a file, not in the database: the schema has no metadata table, and a new table would need a migration for one timestamp. If the file is lost, the only cost is one refresh 2 minutes after the next start, as today. - The cold `ANALYZE` itself is unchanged; it now runs once a day instead of after every restart. Its time of day is wherever the last refresh landed. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a2b5a4931e |
perf(server): battery history via the primary key, cached /api/perf row counts (#2156)
Refs #2146 (items: `/api/nodes/:pubkey/battery`, `/api/perf`). ## Situation Two endpoints ran SQL that reads more than it needs, on every request. Timings are read-only on a production database (18.3M observations, 283k `observer_metrics` rows), first run / second run: | Query | Now | With this PR | |---|---|---| | battery history, `WHERE LOWER(observer_id) = ?` | 174 / 29 ms (plan: `idx_observer_metrics_timestamp`) | 5.9 / 5.5 ms (plan: primary key `(observer_id, timestamp)`) | | `/api/perf`: `SELECT COUNT(*) FROM observations` | 118 / 116 ms on every call | once per 60 s | ## My changes in this PR 1. **Battery history** (`5bdd3c7f`). `LOWER(observer_id) = ?` keeps SQLite off the primary-key index, so it walked the timestamp index over every observer's metrics in the window. The query now matches `observer_id IN (?, ?)` with the pubkey lowercased and uppercased, the two casings observer IDs are stored in. 2. **`/api/perf` row counts** (`e24c24cf`). The four `COUNT(*)` (transmissions, observations, nodes, observers) are cached for 60 s. File, WAL and freelist sizes are still read on every call. ## Tests - `TestNodeBatteryQueryUsesObserverIndex` loads the `sqlite_stat1` values of that database and checks the plan uses `sqlite_autoindex_observer_metrics_1`. With the `LOWER()` form it fails with `SEARCH observer_metrics USING INDEX idx_observer_metrics_timestamp (timestamp>?)`. - `TestGetNodeBatteryHistory_LowercaseStoredID`: a lowercase-stored observer ID is still found. - `TestGetDBSizeStatsTypedCachesRowCounts`: the cached count is served within the TTL and refreshed after it. - `cmd/server`: `go test ./...` green. ## Not in this PR Other per-request queries listed in #2146 measured too small to change on the same database: `COUNT(*) FROM dropped_packets` 0.0 ms, `SELECT DISTINCT iata FROM observers` 0.2 ms, the observers packet-count query 1.3 ms, the 7-day active node count 0.1 ms. An observer ID stored in mixed case would no longer match the battery query; in that database all 41 distinct `observer_metrics.observer_id` values are uppercase. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
be9f382526 |
perf(server): /paths reads canonical paths in batches and skips redundant collision checks (#2157)
Refs #2146 (item: `/api/nodes/:pubkey/paths`). ## Situation Per-step timing in a staging build (3 most-connected nodes, 2 calls each) showed where a `/paths` request spends its time: | Step | Time | |---|---| | build candidates under the read lock | 73-115 ms | | collision check: one `INSTR(LOWER(resolved_path), ?)` query per index hit, 13,367-16,790 per call | 1.3-2.7 s (first call after start: 13.5 s) | | canonical resolved_path per survivor (`bestResolvedPath`, one or two queries each) | 1.7-2.9 s | | aggregation under the read lock | 82-120 ms | So 3.2-4.1 s of a 3.4-5.1 s request was SQL. On a production database (busiest node, 27k transmissions, read-only): the collision check costs 1.85 s one by one and still 1.69 s batched, because SQLite reads and lowercases every stored path either way. Reading the canonical path (the longest observation's stored path) costs 0.18 s one by one and 0.09 s batched by id. ## My changes in this PR - **The collision check runs only where it decides something.** The aggregation keeps a transmission with a canonical path only if that path names the node (`if !containsTarget { continue }`). If no stored path names the node, the canonical one does not either. So for index hits with a canonical path the check changes nothing, and it now runs only for index hits without one, as before. - **Canonical paths are read in batches.** `loadCanonicalResolvedPaths` reads the stored path of each candidate's longest observation, one query per 500 transmissions. When that path is missing, empty or unparseable it calls `bestResolvedPath` for that transmission, so the result is the same as before. - The snapshot of observation IDs and `path_json` strings is taken in the existing candidate read lock (string copies, no parsing). The PR has two commits: a first version that batched both queries and did not get faster on staging, and the change described above, whose message has the measurements. Squash-merging is fine. ## Measurements Staging, after the ingestor's startup `ANALYZE`, same probe for both builds (ms): | | without this PR | with this PR | |---|---|---| | node 1 (degree 55), call 1 / call 2 | 22,237 / 3,364 | 1,793 / 1,686 | | node 2 (degree 47) | 4,427 / 3,615 | 1,828 / 1,809 | | node 3 (degree 46) | 5,171 / 4,035 | 3,050 / 2,408 | | 4 concurrent clients over 40 nodes, 120 s: calls / p50 / p95 | 73 / 6,717 / 15,160 | 125 / 2,861 / 10,630 | Both staging builds also contain #2147 and #2155. ## Tests - `TestLoadCanonicalResolvedPathsMatchesBestResolvedPath` compares the batched result with `bestResolvedPath` on seven cases: the #810 case (longest observation without a stored path), equal lengths, an empty stored path, case, an observation missing from the snapshot, longest wins, nothing stored. - `TestLoadCanonicalResolvedPathsBatchesQueries`: 1,001 transmissions take 3 batched queries and no per-transmission lookups. - `TestHandleNodePaths_CanonicalPathDecidesIndexHits`: the longest observation names A, a shorter one names B; both are index hits. A gets the transmission, B does not, with no per-transmission query. Run against master, the A/B results are the same and only the query-count check fails (1 there, 0 here). - The existing `/paths` tests pass unchanged (#810, #1278, #1352 collision and fallback cases, anchor bias). - `cmd/server`: `go test ./...` and `go test -race` green. ## Not in this PR - `nodeInResolvedPathViaIndex` still calls the per-transmission check; it has no production caller, only tests. - The concurrent p95 is still 10.6 s on staging. What is left there is the 4-connection read pool shared with other endpoints; a staging A/B with 8 connections gave no gain (`/paths` throughput dropped from 133-146 to 81-103 calls per 120 s), so the pool stays at 4. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3e21b179ec |
fix(server): drive the reach scan from the observations time window (#2155)
Refs #2146 (item: `/api/nodes/:pubkey/reach`). ## Situation The reach scan joins `observations` to `transmissions` and filters on `o.timestamp >= ?` plus one `o.path_json LIKE '%"XX"%'` per token. With the planner statistics `ANALYZE` produces on a large database, SQLite runs it as `SCAN t`: every transmission, with an index probe into `observations` (`idx_observations_tx_ts`) per row. The cost is then the size of `transmissions`, whatever node is asked for and whatever the window. ## My change in this PR `scanReachRows` uses `CROSS JOIN` for the transmissions join. In SQLite that fixes the join order: `observations` is the outer loop, read through `idx_observations_timestamp` for the window, and each row looks up its transmission by primary key. Results are the same rows; it is still an inner join. The query text moves into `reachScanSQL(nTokens)`, so a test can inspect the plan. ## Measurements Read-only on a database with 18.3M observations and 1.5M transmissions (Python `sqlite3` 3.45.1 on the host), 7-day window: | node (rows returned) | `JOIN`, first run | `JOIN`, warm | `CROSS JOIN`, both runs | |---|---|---|---| | busiest relay (54,387) | 10.29 s | 3.27 s | 1.95 s | | quartile (3,913) | 3.16 s | 3.15 s | 1.81 s | | median (1,310) | 3.17 s | 3.18 s | 1.82 s | | a repeater reported as timing out in the browser (34,209) | 14.11 s | 3.25 s | 1.93 s | On that instance `/api/nodes/:pubkey/reach` reports p95 188.6 s and max 192.6 s over 34 calls in `/api/perf`; a single uncached call from the host took 3.4 to 3.6 s, so most of that is waiting for the 2 build slots and the read pool. On a staging instance (smaller database), with unique cache keys and `days` 6 or 7, one cycle per build after the ingestor's startup `ANALYZE` had finished: | | without this PR | with this PR | |---|---|---| | sequential, 12 calls, p50 | 3402-3801 ms | 2078-2467 ms | | 2 concurrent clients (= `reachMaxConcurrentBuilds`), p50 / p95 | 3575-4087 / 4049-5291 ms | 2271-2462 / 2814-2928 ms | | calls completed in 120 s | 58-68 | 96-101 | The staging builds also contain #2147. ## Verification - `TestReachScanStartsFromObservationsTimeIndex` creates the two production indexes, loads the `sqlite_stat1` values from the 18.3M-observation database, and checks that the plan starts with `SEARCH o USING INDEX idx_observations_timestamp`. With the old `JOIN` it fails with the production plan (`SCAN t`, then `SEARCH o USING INDEX idx_observations_tx_ts`). - The existing reach tests (`TestScanReachRows_DecodesRows`, `TestScanReachRows_CapTruncates`, attribution tests) pass unchanged. - `cmd/server`: `go test ./...` green. ## Not in this PR - The scan still reads every observation in the window and applies the `LIKE` per row, so its cost grows with the window and with traffic. Removing the scan needs an index on hop tokens (for example a table the ingestor maintains). I have not measured what that would cost in storage. - The build-slot limit (2) and the 5-minute cache are unchanged. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f50312767f |
fix(server): no SQL while holding PacketStore.mu (#2147)
Refs #2146 (first item: "Stop issuing SQL while holding `store.mu`").
## Situation
Several store functions ran SQLite queries while holding
`PacketStore.mu`. The server's read pool has 4 connections
(`db.go:153`), so a store-lock holder could wait for a pool connection.
`mu` is a `sync.RWMutex`: once the ingest poller asks for the write
lock, every new reader queues behind it, including `/api/healthz`. A
test that scans the package (below) finds 16 such call sites on `master`
at
|
||
|
|
9dbc287579 |
docs: release notes for v3.14.0 (#2143)
Release notes and the CHANGELOG entry for v3.14.0. The release workflow reads `docs/release-notes/v3.14.0.md` from the tagged commit as the release body, so this lands before the tag. 23 commits since v3.13.1: 13 fix, 8 feat, 1 ci, and this docs commit. The headline is optional user accounts (#2128). "Read this before upgrading" covers the new limits on unauthenticated endpoints, the `traffic_share_score` drop (#2117), the region filter window (#2114) and the config file mode (#2126). After merge: tag this commit as `v3.14.0` once its `:edge` image is built, then dispatch the CI/CD pipeline on the tag. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>v3.14.0 |
||
|
|
57395e2b28 |
fix(ui): escape observer iata, name and id, and node names, in eight render sinks (#2118)
Anyone who can publish to an MQTT broker that CoreScope reads controls three strings: the observer `iata` (a topic segment), the observer `name` (the JSON `origin` field) and the observer `id` (a topic segment). Anyone with a radio controls the node `name` in an ADVERT. Eight places wrote these into the page with `innerHTML` without escaping, so a crafted value runs as script in every visitor's browser: | file | what | |---|---| | `packets.js` | Region badge in the group, child and single packet rows | | `observers.js` | Region badge in the observers table and in the slide-over | | `live.js` | "Heard By — Regions" header in the node detail | | `nodes.js` | observer name in the clock-skew evidence panel | | `analytics.js` | `observer_id` inside `data-observer=` and `id=` attributes | | `app.js` | node name in the favorites dropdown | | `customize-v2.js` | node name in the geofilter prune preview (admin page) | **Fix:** each value now goes through the existing `escapeHtml()` / `esc()` helper. No behaviour change for normal values. **Tests:** one test per sink added to `tests/unit/test-xss-escape-sinks.js`, following the file's existing pattern (pull the template out of the source, render it with a hostile value, check the output is escaped). Reverting any single fix makes its test fail. 41/41 pass. **Notes for reviewers:** - The ingestor uppercases `iata` before storing it. That is not a defence: tag and attribute names are case-insensitive, and script can be written with numeric character references. - An operator who has set an `observerIATAWhitelist` is protected from the `iata` sinks but not from the name/id sinks. - Follow-ups worth a separate PR: validate `iata` and observer `id` to a safe character set at ingest, cap the `origin` length, and add a `Content-Security-Policy` header. This PR is escaping only so it is easy to review and safe to ship. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
29095658e9 |
fix(nodes): explain packet and observation counts (#2132)
Red commit: `d8a6d42` ([CI assertion](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37695889461)). Review-regression red: `f73643b` ([CI assertion](https://github.com/Kpa-clawbot/CoreScope/actions/runs/37703995494)). Fixes #2131. Partial fix for #1154: U2 only; other audit items remain open. Both node views explain distinct transmissions versus recorded observations. The shared help is reachable by keyboard and touch, stays readable at either scroll edge, and remains open when the pointer enters its text. Escape dismisses help while preserving node context and focus; a second Escape keeps the existing navigation behavior. Count expressions and API requests are preserved. One controller per nodes page handles one active help element and one queued animation frame. Scroll/resize work is coalesced; listeners are removed on page teardown. Work is constant relative to packet/node counts, with no dataset scans or new settings/dependencies. E2E assertion added: `tests/e2e/test-e2e-playwright.js:114`. Thirteen cases cover exact counts, accessible description, keyboard/touch, hover transfer, Escape, and viewport/scroll bounds. Existing metric assertions are unchanged. Browser verified: http://localhost:63508 using Chromium, controlled fixtures and actual API data. Desktop/mobile screenshots retained locally; additional checks exercised resizing and repeated page cleanup. Local validation: 189 standalone suites; core browser 148 passed, 3 fixture-dependent skips (existing manual skips unchanged); location/metric suite 6 passed; syntax, theme variables and lint passed (0 errors, 92 existing warnings). ## Preflight overrides - OpenClaw preflight/profile unavailable; repository checks and local Chromium used. - Windows ingestor symlink fixture lacks permission. Backend files are unchanged; full Linux Go CI is required. |
||
|
|
2456762179 |
feat: account data export and users.db backup (#2141)
A small follow-up to parts A to E of #2128: users can download all data the instance holds about them, and the server backs up `users.db`. Stacked on #2140. Review the commits after that branch. ## The situation - An instance with accounts holds personal data, but a user has no way to get a copy (GDPR articles 15 and 20). Self-delete already exists. - `users.db` is the only file the server writes, and nothing backs it up: `GET /api/backup` snapshots the analyzer database only. ## What this PR adds **Data export.** `GET /api/account/export` and a "Download my data" button on the account page give one JSON file with: - the profile, including activation and a pending email change; - sessions (times and user agent); - synced settings, own proposals, notification prefs and watches; - audit rows where the user is actor or target, with other accounts as an id only; - the mail log. Password hash, session and link tokens and the unsubscribe token are left out: they are credentials. Each export is audited as `user.export`. **users.db backup** - Daily `VACUUM INTO` snapshots through `internal/users`, on by default when user management is on. They go to `backups/` next to `users.db`, the 7 newest are kept, the directory is 0700 and the files 0600. Rotation only touches `users-<timestamp>.db` files and never the snapshot it just wrote. Config: `userManagement.backup {enabled, dir, keep}`. - `GET /api/admin/users-backup`: an admin downloads a fresh snapshot to keep off the server; audited as `user.backup`. - `docs/user-guide/accounts.md` gains a restore procedure. It covers what a restore brings back (accounts deleted after the snapshot, old passwords, revoked sessions) and how to clean that up afterwards. Spec: [`docs/specs/2026-10-08-account-export-and-users-backup-design.md`](https://github.com/efiten/CoreScope/blob/feat/account-export-backup/docs/specs/2026-10-08-account-export-and-users-backup-design.md). ## Verification - `internal/users` with `-race` and the `cmd/server` suite pass locally, 30 new Go tests. The file-mode tests skip on Windows. - `sh test-all.sh` exits 0; the XSS gate in diff mode passes. - User-management E2E: 27 of 27 steps locally. - On our production instance since 8 October 2026; the first snapshot there was 143 kB. ## Not in this PR - Importing an export into another instance. - Off-site upload of the snapshots. The admin download covers that by hand. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
de364a8bb6 |
feat: mail notifications for watched nodes (user management part E) (#2140)
Part E of #2128: a logged-in user watches nodes and gets one mail when a watched node goes offline, comes back, or reports a low battery. Admins can add instance events: a new foreign node, and an observer going offline or back. This covers the per-user mail part of #775. Stacked on #2139 (part D, which builds on #2138). Review the commits after that branch. ## The situation A sysop learns that a repeater went silent, or that its battery is running down, only by opening CoreScope. #775 asks for notifications. Accounts (part A) now give a verified address to mail and a mailer with delivery status. ## What this PR adds **Events.** They use the instance thresholds, so a mail does not disagree with the node page: - `node.offline`: silent for the role's `healthThresholds` window. Last heard comes from the packet store, and repeaters and rooms count relayed traffic (#1598). - `node.battery`: advert telemetry below `batteryThresholds.lowMv`, recovered at `lowMv + 100`. - Admins only: `foreign.new` (once per node) and `observer.offline`. **Store.** Schema v5 with `notification_prefs`, `notification_watches` and `notification_state`. **Server** (only with `userManagement.notifications.enabled`) - A notifier next to the janitor, every `intervalMinutes` (5). The evaluator is a pure function of watches, prefs, stored states and node snapshots. - No burst after a restart or a new watch: a subject's first evaluation stores its state without mailing. The loop waits for the startup load. While ingest is stale (newest packet older than 30 minutes) the offline checks pause, and after recovery they wait one silent window. - One mail per user per check. Limits: 20 per user and 100 per instance per rolling 24 hours, 50 watches per user. A change over a limit is recorded and never mailed later. - Every mail has a one-click unsubscribe (`List-Unsubscribe` and `List-Unsubscribe-Post`). The GET only redirects to a confirm page, so link scanners cannot unsubscribe anyone. Node names are cleaned of control and bidi characters before they go into a mail. - Routes: `GET`/`PUT /api/account/notifications`, `PUT`/`DELETE /api/account/notifications/watches/{pubkey}`, `POST /api/account/notifications/watch-my-nodes` (copies the synced "my nodes" list), `GET`/`POST /api/notifications/unsubscribe`. `mailer.Message` gains `Headers`. **Frontend.** A "Notify me" toggle on the node pages, a Notifications section on the account page (`#/account?section=notifications`), an unsubscribe page, and the notification mail count on the admin overview. Spec: [`docs/specs/2026-10-07-node-notifications-design.md`](https://github.com/efiten/CoreScope/blob/feat/notifications/docs/specs/2026-10-07-node-notifications-design.md). ## Performance Each check reads watches, prefs and states from `users.db` (one query each), the watched nodes from the analyzer DB in chunks of 500, and last-heard times from the packet store. | Measurement | Result | |---|---| | One check on our staging instance, 130,000 packets in memory, 1 watcher | 6.9 ms, of which 0.93 ms holding the store read lock | | Evaluator benchmark, 100 users with 50 watches each, 2,000 nodes | 1.75 ms | No work is added to ingest, broadcast or any request path. ## Verification - `internal/users` and `internal/mailer` with `-race`; the `cmd/server` suite plus `-race` on the notifier tests; 74 new Go tests. The same Windows-only failure as noted in #2138 applies. - `sh test-all.sh` exits 0; the XSS gate in diff mode passes. - User-management E2E: 26 of 26 steps locally. With notifications off the new steps fail, so they do test the flag. - On our staging and production instance since 7 October 2026. ## Not in this PR - Other channels (Discord, Telegram, webhooks). Detection is separate from delivery, so one can be added. - The topology, RF and anomaly alerts of #775. - If the whole server starts after a feed outage that already ended, the grace window is unknown and a watcher can get one wrong offline mail. The user guide says so. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
abbe6846ff |
feat: propose and approve hashtag channels (user management part D) (#2139)
Part D of #2128: logged-in users propose a hashtag channel, an admin approves, rejects or later revokes it, and the ingestor decrypts approved channels without a restart. It is the first consumer of a generic proposal table, and answers the request in #2092. Stacked on #2138 (part C). Review the commits after that branch; GitHub shows C's commits here until it is merged. ## The situation - An instance decrypts only the hashtag channels listed in `hashChannels`. Users who want a city or club channel shown have to ask the operator, who edits `config.json` and restarts the ingestor. - #2092 asked for a suggest-and-approve flow; a downstream fork (dborup/CoreScope#99) runs one with anonymous suggestions and a file queue. ## What this PR adds **Name rules (`internal/channel`).** `ValidateHashtagName`, mirrored in the browser: - at most 31 UTF-8 bytes including `#` (firmware `ChannelDetails.name[32]`), case preserved, `Public` refused in any case; - control, bidi, separator and format characters refused except ZWJ, plus invisible fillers (`Other_Default_Ignorable_Code_Point`, U+2800, non-ASCII spaces). Go and JS give the same verdict for every code point. **Store.** Schema v4, one generic `proposals` table (kind, subject, status pending/approved/rejected/revoked, proposer, reviewer, note). One row per (kind, subject); limits are checked in the same transaction as the change. **Server** (only with `userManagement.channelProposals.enabled`) - `POST /api/proposals`, `GET /api/account/proposals`, `GET /api/admin/proposals`, `POST /api/admin/proposals/{id}/approve|reject|revoke`. - 400 for an invalid name. 409 for a duplicate, a rejected name, or a name already in `hashChannels`/`channelKeys` ("this channel is already decrypted on this instance"). 429 over a limit (`maxPending` 100, `maxApproved` 128, `perUserPerDay` 5). - `GET /api/channels` gains `approvedChannels`, so approved channels are listed before they have traffic. With the feature off the response is byte-identical. **Ingestor.** Opens `users.db` with `mode=ro` every 60 seconds, re-validates the names, and swaps its key map through an `atomic.Pointer`. Configured channels always win, and a failed read keeps the last good set. The ingestor never writes `users.db`; AGENTS.md now says so. **Frontend.** "Propose for everyone" in the add-channel dialog, an "approved" marker in the channel list, "My proposals" on the account page, and an admin "Proposals" tab with a warning on approve. The packet detail pane now escapes the channel name: until now only the operator chose those names. Spec: [`docs/specs/2026-10-07-channel-proposals-design.md`](https://github.com/efiten/CoreScope/blob/feat/channel-proposals/docs/specs/2026-10-07-channel-proposals-design.md). ## Performance Every key is tried on each GRP_TXT that no key opens (`decodeGrpTxt`). Benchmark of that worst case: | Keys | Time per packet | |---|---| | 320 configured | 204 to 216 µs | | 320 configured + 128 approved | 272 to 320 µs | That is about +33%, linear in the key count and capped by `maxApproved`. With the feature off the decoder gets the identical map and no ticker runs. ## Verification - `internal/channel`, `internal/users` (with `-race`), `cmd/server` and `cmd/ingestor` pass locally, 53 new Go tests. The ingestor's `TestWriteStatsAtomic_SymlinkAtDestIsReplaced` needs symlink rights on Windows and fails on `master` too. - `sh test-all.sh` exits 0; the XSS gate in diff mode passes. - User-management E2E: 21 of 21 steps locally (propose, approve, the channel appears, revoke removes it). - On our staging and production instance since 7 October 2026. ## Not in this PR - Private (PSK) channels: their keys stay in the browser (#725). - Anonymous proposals. - Names that need ZWNJ (some Persian and Urdu spellings) and subdivision flags are refused, because ZWNJ and tag characters are format characters. Allowing them later only widens the rule, so no approved name gets stranded. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
18a634c115 |
feat: admin dashboard with overview and audit log (user management part C) (#2138)
Part C of #2128: an admin area with an overview, the user table and a global audit log. Off by default like the rest of user management: with `userManagement` off nothing changes, and non-admins get no new UI. This is the first of 4 stacked PRs (C, D, E, then a small export/backup follow-up). Each later one contains this branch; review them in order. ## The situation - An admin could manage users one by one in `#/admin/users`, but nothing showed whether the instance needs attention: accounts stuck in activation, bouncing mail, someone guessing a password, an MQTT source that dropped. - The audit log existed (`internal/users/audit.go`) but could only be read per user, and logins were not recorded. ## What this PR adds **Store (`internal/users`)** - Schema v3: an index on `audit_log(at)`. - `AuditList` with filters (action or group prefix like `user.login.*`, user as actor or target, period) and keyset pagination; `PruneAudit`; `Stats` for the user figures. **Server** - `GET /api/admin/stats` (typed struct): accounts by status, admins, registrations and active users over 7 and 30 days, logins and failed logins in 24 hours, mail by final status, and the attention items computed server-side. - `GET /api/admin/audit`: filtered, newest first, `next` cursor. - `GET /api/admin/users` gains `bouncing=1`. - Logins are audited as `user.login` and `user.login.failed` (reason `wrong_password`, `pending` or `disabled`). An unknown address writes no row. The writes are asynchronous, so the login response does not wait on `users.db`. Login rows are pruned after 90 days. **Frontend** - `#/admin?tab=overview|users|audit`: `admin.js` (tab shell), `admin-overview.js` ("Needs attention", Users card, System card from the existing health, MQTT and observer endpoints), `admin-audit.js` (filters in the URL, "Load more"). `admin-users.js` becomes the Users tab; the old `#/admin/users` link rewrites to it. - `/api/healthz` is read on open and on Refresh only, not on the 60-second timer, because it walks every packet under a read lock. Spec: [`docs/specs/2026-10-07-admin-dashboard-design.md`](https://github.com/efiten/CoreScope/blob/feat/admin-dashboard/docs/specs/2026-10-07-admin-dashboard-design.md). ## Verification - `internal/users` and `cmd/server`: `go vet` and `go test` pass locally, 26 new Go tests. One upstream test, `TestSaveGeoFilterPreservesFileMode`, also fails on Windows on plain `master` (file modes) and is unrelated. - `sh test-all.sh` exits 0; the XSS gate in diff mode passes. - User-management E2E with the `e2etest` build: 13 of 13 steps locally. - Running on our staging and production instance since 7 October 2026. ## Not in this PR - Restricting existing pages (perf, MQTT status) to logged-in users or admins. The admin area adds no access mechanism of its own, so that stays a later router rule plus endpoint check. - Server-enforced customizer tab restrictions (#1508). - The 24-hour, 5-attempt and 10-minute attention thresholds are constants for now (AGENTS.md rule 8: customizer later). --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a4445b4985 |
ci: run E2E, Go tests and the image build side by side (#2135)
Closes #2134. ## What changes The test jobs and the image build now run side by side. Only the GHCR push waits for all of them. ``` changes ─┬─ go-test ─────────────────────────┐ ├─ race-test (ingestor changes only) │ ├─ e2e-shard ×3 ── e2e-test (gate) ───┼─ build-and-publish → deploy, publish └─ image-check (two-arch build) ──────┘ ``` - **`e2e-shard` (new, matrix of 3).** It needs only `changes`. The suite list stays in the workflow: each line now reads `suite <shard> <file> [VAR=value ...]`, with the same environment variables as before. - Shards are balanced on the measured suite times, about 6 minutes each. - A shard number outside 1..3 fails the step, so a typo cannot drop a suite silently. - `test-issue-1648-m4-icons-e2e.js` stays in the same shard as, and after, the a11y suite. Its distance-tab check only counts once the lazy distance index (#1011) has been built, and the a11y suite is what first requests it. - **`e2e-test`** is now a gate under the old check name. It runs only if all three shards passed, checks that three results arrived, and builds the same `e2e-badges` artifact as before. - **`image-check` (new).** It runs the two-arch build (no push) and the arm64 QEMU smoke on every event, beside the tests. It also computes the build metadata once, which the push reuses. - **`build-and-publish`** now needs `go-test`, `e2e-test` and `image-check`. On a PR every step is skipped. On push and tag refs it only pushes, with the same build args as `image-check`, so every layer should be a cache hit. - **The unused `docker compose build` of the staging image is gone.** `docker compose config --quiet` still validates the file. - `AXE_SCREENSHOT_DIR` now actually reaches the a11y suite. - `tests/unit/test-issue-1956-release-routing.js` follows the new jobs, and checks more than before: - a failure in `go-test`, `e2e-shard`, `e2e-test` or `image-check` skips `build-and-publish`; - the only non-publishing build lives in `image-check` and runs on every event; - the push takes its build args from `image-check`. ## Measurements The old numbers are from upstream; the new ones are from PR runs on my fork (a branch on top of this commit, plus one temporary commit that also triggers the workflow for PRs into a test base branch). | | Before (PR run 37766553643) | After (fork runs 37812660655 / 37814554422) | |---|---|---| | Go Build & Test | 4.9 min | 7.2 / 7.5 min, in parallel | | E2E | 18.4 min | 3 shards of 7.1–9.0 min, in parallel | | Image build | 6.9 min, after E2E | 5.9 / 5.4 min, in parallel | | **Total** | **30.5 min** | **9.6 / 10.6 min** | **Same tests:** I compared the second fork run with the E2E log of the upstream run, suite by suite. All 118 suites ran, and each one passed the same number of checks (1161 in total). **Cost:** more runner minutes, because each shard repeats about 2.5 minutes of setup. Standard runners are free for a public repository. ## Not verified here - **A master push.** On a fork the GHCR push cannot run. The expectation is that `Build and push to GHCR` gets cache hits from `image-check` and finishes in a minute or two instead of 4–4.5. The first master run after merge will show whether it does. - **Branch protection.** If it requires other job names than `🎭 Playwright E2E Tests` and `🏗️ Build & Publish Docker Image`, those may need adding. I cannot read the protection settings, so I could not check. ## Tests run locally - `sh test-all.sh` passes (on Windows with `PYTHONUTF8=1`). - `test-issue-1956-release-routing.js` passes, and fails as expected when `go-test` is removed from the `needs` of `build-and-publish`. - actionlint 1.7.7 reports only the existing self-hosted runner label. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
caa66356bd |
fix(api): cap concurrent cold reach scans at 2, answer 429 beyond that (#2127)
`GET /api/nodes/{pubkey}/reach` is unauthenticated and a cold-cache
request runs a full scan. `singleflight` only collapses identical keys,
so distinct `(pubkey, days)` requests each start a scan, and they share
the 4-connection SQLite pool with every other handler. A handful of
requests for different nodes can hold every reader and stall the whole
API.
**Fix:** a small semaphore allows two cold scans at once. A cold request
that finds both slots busy gets `429` with `Retry-After: 5` instead of
queueing. Cached answers are unaffected, and the two in-flight scans
still complete normally.
**Tests:** `node_reach_concurrency_test.go` — with both slots held a
cold request returns 429 with `Retry-After`; after release the same
request runs normally. Full server suite passes.
**Note:** our instance runs a fork that already bounds reach work
differently (an async job queue), so this exact change is not what we
run in production. The test suite is the verification here.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com>
|
||
|
|
0a6870c508 |
fix(api): cap request bodies on /api/decode and /api/packets/observations (#2119)
`POST /api/decode` and `POST /api/packets/observations` are unauthenticated and decoded the request body with no size limit. `json.Decoder` buffers the whole value in memory, and `/api/decode` hex-decodes the string before the 184-byte payload check runs. On a test instance one 100 MB request raised RSS by about 180 MB. Many requests in parallel can push a container past its memory limit, and a restart reloads the dataset for several minutes. **Fix:** `http.MaxBytesReader` on both handlers, the same way the geo-filter PUT and `/api/paths/inspect` already do it. | endpoint | cap | why | |---|---|---| | `/api/decode` | 4 KB | a packet is at most ~512 hex chars | | `/api/packets/observations` | 64 KB | 200 hashes of 64 chars is ~14 KB | Oversized bodies get HTTP 413 instead of a generic 400. **Tests:** `body_limits_test.go` — normal bodies still return 200, oversized bodies return 413, for both endpoints. Full `go test ./...` in `cmd/server` passes. **Note:** deployments behind a reverse proxy with its own body limit were already safe. The shipped Caddy config has no body limit, so default installs were not. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
2e7a4ebcfe |
fix(ingestor): bound the unauthenticated /neighbors report (#2122)
`handleNeighborsReport` trusted whatever an observer published on the `/neighbors` topic. The sender chose `origin_id` (whose "self" it is) and could list any pubkey as a responded neighbor; each got its `configured_scope` written with any scope string, stamped with the sender's own timestamp. The store is last-write-wins on that timestamp and `normalizeReportTS` accepted any RFC3339 time, so one report dated years ahead was written once and then blocked every genuine later report for that node until someone edited the database. The value is shown on the reach page as the confirmed scope and feeds `/api/scope-audit`. **Fix (three guards):** - pubkeys must be 64 hex chars — anything else cannot match a node anyway, so it is dropped instead of running UPDATEs that never match - the normalised scope list is capped at 256 bytes - a report stamped more than 5 minutes ahead of our clock is dropped, so a far-future timestamp can no longer lock the node **Tests:** `neighbors_guard_test.go` covers each guard, including that a genuine report still lands after a future-stamped one was rejected and that ordinary clock skew is still accepted. Full ingestor suite passes. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
a981420d21 |
fix(ingestor): cap observer-supplied string lengths (#2123)
Observer `id` and `iata` come from the MQTT topic; `origin` (name), `model`, `firmware`, `client_version` and `radio` come from the status JSON. Any publisher controls them and nothing bounded their length, so one message could store a 64 KB observer id or name. Each new id is also a new `observers` row, and that table is joined by most packet queries. **Fix:** a small `clampObserverField` helper strips control characters and truncates: ids to 128 runes, text fields to 128, IATA to 16. Applied on the status path, the packet path and in `extractObserverMeta`. Values are truncated rather than rejected, so a legitimate observer with a long name still appears. **Tests:** `observer_fields_test.go` — short values untouched, long values cut at 128 runes (not bytes, so multi-byte names are not split), control characters removed, `extractObserverMeta` caps all string fields. Full ingestor suite passes. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
548e4cd4a4 |
fix(api): cap the nodes= list on /api/packets at 50 entries (#2120)
Each entry in the comma-separated `nodes=` list on `GET /api/packets` costs one SQLite lookup (`resolveNodePubkey`) while the packet store's read lock is held. A 1 MB URL fits about 15,000 entries. On a test instance 12,000 entries took 1.3 s per request, against 0.9 ms for one entry, and the lock stalls the poller's writes for that long. A few parallel clients can keep the site busy and the live feed stale. **Fix:** lists longer than 50 entries get HTTP 400 with a clear message. No UI page sends more than a handful. **Tests:** `multi_node_cap_test.go` — 50 entries return 200, 51 return 400. Full `go test ./...` in `cmd/server` passes. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
4ce6d9f6d4 |
fix(ui): pin CDN script versions and add integrity hashes (#2121)
`leaflet.heat` and `chart.js` in `index.html`, and swagger-ui on `/api/docs`, were loaded from unpkg with no `integrity` attribute. `chart.js@4` and `swagger-ui-dist@5` also floated on a major version, so a new release would load unreviewed. A compromised CDN or package would run as our own code on every page. Leaflet itself already had a hash. **Fix:** pin `chart.js@4.5.1`, `leaflet.heat@0.2.0`, `swagger-ui-dist@5.33.1`, each with a sha384 hash and `crossorigin="anonymous"`. **How the hashes were made:** download each pinned file, `openssl dgst -sha384 -binary | base64`, then download again and check the hash matches. **Note:** bumping chart.js or swagger-ui now means updating the hash too. Vendoring them into `public/vendor/` (as markercluster already is) would remove the CDN dependency entirely; happy to do that instead if preferred. Running in production on our instance since 2026-10-07. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
8b64439635 |
fix(server): keep config.json's file mode when saving the geo-filter (#2126)
`SaveGeoFilter` rewrote `config.json` through a temp file created with mode 0644, so a config an operator had made 0600 (it holds the API key and broker passwords) became world-readable after the first geo-filter save. **Fix:** stat the original and reuse its mode for the temp file. Falls back to 0644 when the stat fails. **Tests:** `config_mode_test.go` — a 0600 config stays 0600 after a save; a 0644 config stays 0644. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
523a5edfa1 |
fix(ui): escape the route string on the unknown-route page (#2124)
`app.js` wrote the URL fragment into `innerHTML` as the heading of the "Page not yet implemented" view for routes it does not know. Browsers percent-encode `<`, `>` and `"` in fragments, so this is unlikely to be exploitable today, but it is a plain `innerHTML` sink on a URL-controlled string and costs one `escapeHtml()` to close. **Tests:** one test added to `tests/unit/test-xss-escape-sinks.js` in the file's existing style. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Mythos 5.1 <noreply@anthropic.com> |
||
|
|
7309bdb5a9 |
fix(store): index a transmission once per relay key in byPathHop (#2117)
Relates to #2108 ## Problem `traffic_share_score` grows with server uptime until it is far above reality, and many relays end up clamped at 1.0. The score is the number of non-advert entries in `byPathHop[pubkey]` divided by the number of non-advert transmissions. `indexResolvedPathHops` runs once per **observation**, on the live-ingest path (`IngestNewFromDB`) and on the late-observation path (`IngestNewObservations`). `addResolvedPubkeysToPathHopIndex` only de-duplicates within one call, so every further observation through the same relay appends the transmission to that relay's bucket again. The denominator counts each transmission once. After a restart the values look right, because `buildPathHopIndex` → `retainResolvedPathHops` de-duplicates by `*StoreTx`. They then drift upwards as live observations arrive. None of the affected functions has changed on `master` since the diagnosis against `415362c`. This branch is based on `000d9ab`. ## Fix 1. **Idempotent insert, per (transmission, relay key).** `addResolvedPubkeysToPathHopIndex` keeps a side map `pathHopResolved map[*StoreTx][]string`. It holds the resolved keys each transmission is already indexed under and skips those. - Keys are compared **exactly**, so no collision can ever drop an entry. - A later observation through a **new** relay still adds that relay, once. - `StoreTx` is unchanged. 2. **Interned keys.** `pathHopKeys map[string]string` keeps one shared copy of each resolved key. The record then holds a 16-byte string header per entry, and not the string each observation's resolve allocated: a fresh `json.Unmarshal` string per persisted observation on `Load`, and a `strings.ToLower` result on the live paths. The `byPathHop` key is the same shared copy. 3. **Eviction and rebuild.** - `evictStaleInternal` drops the record of evicted transmissions, plus any interned key whose bucket it deletes. - `retainResolvedPathHops` keeps the records of live transmissions only and drops interned keys whose bucket was not carried over. When a rebuild starts from an empty index, it clears both. 4. **Defence in depth.** The single (`GetRepeaterUsefulnessScore`), batch (`GetRepeaterNodeStatsBatch`) and bulk (`computeRepeaterUsefulnessScoreMap`) scores count **distinct** non-advert transmissions per bucket, via `countDistinctNonAdvert`. - IDs are collected in a reused slice. - A bucket in ascending ID order is already distinct; only out-of-order buckets are sorted. - The bulk pass does this for full-pubkey keys only. Raw-hop buckets get one entry per transmission from `addTxToPathHopIndex`. 5. **Cache side effect.** A repeated observation through known relays no longer mutates `byPathHop`, so it no longer drops the batch relay-stats cache, as the cache contract from `1164` intends. An observation through a new relay still drops it. The other `byPathHop` consumers already de-duplicate by `tx.ID` or read only raw prefix keys: `GetNodeHopAnalytics`, `computeMultiByteCapability`, `handleNodePaths`, `computeRepeaterRelayInfoMap` and the relay info in `GetRepeaterNodeStatsBatch`. Two tests pin that they ignore duplicate entries. ## Exact keys vs. a 64-bit fingerprint The first version of this fix stored a 64-bit FNV-1a hash per key. As requested, I measured both, plus exact keys without interning, on the real `addResolvedPubkeysToPathHopIndex`. **Keys per transmission.** In the e2e fixture, the union of resolved relay keys over all observations of one transmission has a mean of 5.4 and a maximum of 17. The protocol limit is 64 path bytes (`MAX_PATH_SIZE`), i.e. 64 / hash size hops. **Memory** retained by the record, per transmission. `BenchmarkPathHopRecordMemory_2108`: 100K transmissions, keys from 4,000 relays, each transmission heard twice with freshly allocated keys. The record map and the interned keys are included. | keys per tx | exact, interned (this PR) | 64-bit fingerprint | exact, not interned | |---|---|---|---| | 2 | 88 B | 68 B | 207 B | | 5 | 136 B | 101 B | 447 B | | 17 | 344 B | 197 B | 1,423 B | - At realistic key counts, exact keys cost 20–35 B per transmission more than the fingerprint. That is about 3.5 MB per 100K transmissions, and small next to what the store already charges per transmission (`storeTxBaseBytes` alone is 384 B). - Without interning, exact keys would cost 3–4× more, and most of that would be held during every `Load`. Interning is what makes exact keys affordable. **Time.** I ran the variants interleaved, 3 rounds each, median ns/op on an Apple M4. The record step is not measurable inside the full call. The isolated record step was faster with exact keys in a separate micro-benchmark, because comparing a handful of 64-char strings is cheaper than hashing each one. | benchmark | exact, interned | fingerprint | exact, not interned | |---|---|---|---| | late observation, 20K txs | 312 | 358 | 331 | | late observation, 100K txs | 760 | 836 | 758 | | `IngestNewObservations` (SQL + resolve + index) | 531 µs | 497 µs | 527 µs | **Collision probability of the fingerprint.** For k distinct keys in one transmission it is about k(k−1)/2 / 2^64: | k | per transmission | per 10^9 transmissions | |---|---|---| | 2 | 5.4e-20 | 5.4e-11 | | 5 | 5.4e-19 | 5.4e-10 | | 17 | 7.4e-18 | 7.4e-9 | | 64 | 1.1e-16 | 1.1e-7 | For random keys this is negligible. FNV-1a is unkeyed, though, so two relay keys that collide could be found deliberately (a birthday search over about 2^32 key pairs). A collision would leave a transmission out of one relay's bucket. **Decision.** Exact keys are cheap enough once interned: the same speed within noise, and a few tens of bytes per transmission. They remove the question of collision-driven omissions entirely, so this PR uses exact keys and drops the fingerprint. The two new maps (`pathHopResolved`, `pathHopKeys`) are not added to `trackedBytes`, so the `maxMemoryMB` trigger undercounts by roughly the per-transmission figures above. Both are bounded by live state (eviction and rebuild prune them), so this is a steady proportional undercount rather than a leak. ## Tests **New: `cmd/server/pathhop_dedupe_2108_test.go`.** Behaviour tests that use only existing API. Each one runs against the real SQLite schema through `Load`, `IngestNewFromDB` and `IngestNewObservations` where noted. | test | covers | on `master` | |---|---|---| | `TestPathHopIndexOncePerTx_LateObservations_2108` | 1 + 10 late observations through the same relays: one entry per relay | **fails** (11 entries) | | `TestPathHopIndexOncePerTx_LiveIngestBatch_2108` | 11 observations in one `IngestNewFromDB` batch | **fails** (11) | | `TestPathHopIndexOncePerTx_AfterLoad_2108` | `Load` + rebuild, then one live observation | **fails** (2) | | `TestPathHopIndexAddsNewRelayFromLaterObservation_2108` | a later observation through a **new** relay adds it exactly once, and its share becomes correct | **fails** (2) | | `TestTrafficShareStableAcrossLateObservations_2108` | single, batch and bulk scores agree, equal the definition, stay put over rounds of late observations, and equal a fresh `Load` of the same data | **fails** (all four relays at 1.0, want 0.4–0.5) | | `TestPathHopIndexSizeBoundedByTransmissions_2108` | the index size does not grow with observations | **fails** | | `TestRelayStatsCacheAcrossRepeatedObservations_2108` | a repeated observation keeps the relay-stats cache, and the kept cache equals a fresh compute; a new relay drops it, and the next read sees the relay | **fails** | | `TestAddResolvedPubkeysToPathHopIndex_PerRelayIdempotent_2108` | the helper's return value and cache invalidation per (tx, relay), in any key order | **fails** | | `TestTrafficShareCountsDistinctTransmissions_2108` | all three scores count distinct transmissions in a duplicated bucket | **fails** | | `TestPathHopConsumersIgnoreDuplicateEntries_2108` | pin of the consumer audit | passes (by design) | | `TestMultiByteCapabilityIgnoresDuplicateEntries_2108` | pin of the consumer audit | passes (by design) | **New: `cmd/server/pathhop_record_2108_test.go`.** These tests use the new symbols, so they do not compile against `master`. - `TestCountDistinctNonAdvert_2108`: ascending, descending, interleaved duplicates, adverts, nils, untyped. - `TestPathHopResolvedRecordBoundedAndEvicted_2108`: the record stays at 2 keys after 11 observations. Eviction removes the record and every entry. A survivor's next observation adds nothing. Evicting everything leaves no record, no interned key and no bucket. - `TestPathHopResolvedRecordAcrossRebuild_2108`: a rebuild keeps the records of live transmissions, so the next observation adds nothing. It drops the record of a removed transmission and the interned key only that transmission used. - `TestPathHopResolvedRecordClearedWithEmptyIndex_2108`: a rebuild from an empty index clears the record and the interned keys, and the next observation puts the transmission back. - `TestPathHopResolvedRecordInternsKeys_2108`: record entries and the `byPathHop` key share one copy, even though every observation passes freshly allocated keys. **Changed: `pathhop_eviction_1908_test.go`.** It built its duplicate entries by repeating `indexResolvedPathHops`, which no longer duplicates. It now seeds the duplicates directly, so the `1908` sweep is still tested against buckets that hold one transmission several times. The expected buckets are unchanged. **Changed: `db_test.go`.** `setupTestDB` takes `testing.TB`, so the SQL benchmark can use it. **Mutants.** I applied each mutant on its own to this branch. All 14 are killed: | mutant | killed by | |---|---| | record never consulted | the OncePerTx, stable-share, size, cache and record tests | | dedupe per transmission instead of per relay | `AddsNewRelayFromLaterObservation`, `RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent` | | eviction keeps the record | `RecordBoundedAndEvicted` | | rebuild keeps records of removed transmissions | `RecordAcrossRebuild` | | rebuild from an empty index keeps the record | `RecordClearedWithEmptyIndex` | | cache dropped on every call | `RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`, existing `NoMutation_PreservesCache` | | cache kept although `byPathHop` changed | `RelayStatsCacheAcrossRepeatedObservations`, `PerRelayIdempotent`, existing `InvalidatesRelayStatsCache` | | single score counts entries | `TrafficShareCountsDistinctTransmissions` | | batch score counts entries | `TrafficShareCountsDistinctTransmissions` | | bulk score counts entries | `TrafficShareCountsDistinctTransmissions` | | distinct count trusts any bucket order | `TrafficShareCountsDistinctTransmissions`, `CountDistinctNonAdvert` | | keys not interned | `RecordInternsKeys` | | eviction keeps interned keys | `RecordBoundedAndEvicted` | | rebuild keeps interned keys of dropped buckets | `RecordAcrossRebuild` | **Commands run:** - `gofmt -l` on all tracked Go files: clean. - `go vet ./...` in all 14 modules: clean. - `cd cmd/server && go test -race ./...`: pass. - `cd cmd/ingestor && go test ./...`: pass. - `sh test-all.sh`: 184 of 186 suites pass locally. The other two, `test-issue-1956-release-routing.js` and `test-preflight-xss-gate.js`, shell out to scripts that need bash ≥ 4 (`mapfile`). They fail under macOS's bash 3.2 regardless of this change. This PR touches no frontend or script files. ## Benchmark: `master` vs. this branch I ran `master`'s sources (`000d9ab`) and this branch interleaved, 5 rounds, with the same benchmark files. Medians on an Apple M4. | benchmark | `master` | this PR | change | |---|---|---|---| | late observation through known relays, 20K txs | 446 ns, 38 B/op | 441 ns, 0 B/op | within noise | | late observation through known relays, 100K txs | 710 ns, 69 B/op | 755 ns, 0 B/op | within noise (runs overlap) | | … index size afterwards, entries/tx (20K / 100K) | 24.99 / 8.99, still growing | 4.99 / 4.99 | bounded | | `IngestNewObservations`, one observation for each of 20 txs | 532 µs | 513 µs | within noise | | bulk score pass, clean index, 20K, ingest order | 246 µs | 316 µs | +28 % | | bulk score pass, clean index, 100K, ingest order | 3.03 ms | 4.30 ms | +42 % | | bulk score pass, clean index, 20K, reversed buckets | 257 µs | 467 µs | +82 % | | bulk score pass, clean index, 100K, reversed buckets | 3.38 ms | 4.79 ms | +42 % | | bulk score pass after 11 observations per tx, 20K | 663 µs | 284 µs | −57 % | | bulk score pass after 11 observations per tx, 100K | 4.72 ms | 3.19 ms | −32 % | - On an index that is clean on both builds, the distinct count makes the bulk pass slower. That pass runs on the cache-miss path, which the background recomputer refreshes every 5 minutes by default. - On the index a live server actually holds without this fix, `master` walks every accumulated duplicate, so the real-world pass is faster after the fix. The gap grows with uptime. - The late-observation step no longer allocates. On `master` its index grows with every observation. Benchmarks: `BenchmarkLateObservationIndex_2108`, `BenchmarkIngestNewObservations_2108`, `BenchmarkTrafficShareScoreMap_2108` and `BenchmarkPathHopRecordMemory_2108`. ## Production motivation We have run this fix on two production instances. Before the fix, the summed `traffic_share_score` grew past 180 and many relays sat at the 1.0 clamp; with the fix no node reaches `≥ 0.999`. A controlled 12 h A/B run makes the drift explicit. Two servers read one identical database — one on this `master` base, one with the fix — alongside a reference server that freshly loads the same database (a fresh load is correct, because the rebuild dedupes). The unfixed server drifted to 9 relays at the 1.0 clamp and up to 0.95 absolute error per node against the reference; the fixed server stayed within 0.05 of the reference for every node, with no relay at the clamp. A smaller residual rise remains and is a separate cause (startup-vs-live hop resolution); it is deliberately left for a follow-up so its effect stays measurable. ## Out of scope That remaining slow rise comes from a separate cause: live ingest resolves some hops that the startup load does not. This PR deliberately leaves that drift alone, so its effect stays measurable. The fix will follow in a separate PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f3ae8ac25a |
feat: sync a logged-in user's settings across devices (part B) (#2130)
Part B of #2128: a logged-in user's settings follow them across devices. Log in on a phone and your own nodes, favorites, customizer and filters are there; a change on one device reaches the others within a minute or when you return to the tab. **This PR builds on #2129.** Until that one is merged, the diff here includes it. The commits for this part start at `docs(specs): settings sync for optional user management (sub-project B)`. ## The situation Everything a visitor sets up lives in one browser's `localStorage` (about 100 keys in `public/`). A second device or a cleared cache starts from zero (#895). ## What this PR adds **Storage.** `users.db` schema v2: one JSON document per user in `user_settings`, with a revision number and a generation id. A write succeeds only when the client's revision and generation match the stored ones, so two devices cannot overwrite each other silently. **Server.** `GET`, `PUT` and `DELETE /api/account/settings`, behind the same session and CSRF checks as the account routes. - The server owns the list of synced keys (61 keys, [`settings_allowlist.go`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/cmd/server/settings_allowlist.go)) and sends it to the client, so the two cannot drift. - A hard denylist, checked first, refuses `meshcore-api-key`, every `corescope_channel_*` key and `live-channel-colors` (#725). The colour map is keyed by channel hash, and for a user-added channel that hash is `user:<name>`, which would expose hashtag channel names. - Documents are capped at 256 KiB, measured like `JSON.stringify`. PUT is limited to 60 requests per hour per user. A stale revision gets 409 with the current document. **Client** ([`settings-sync.js`](https://github.com/efiten/CoreScope/blob/feat/settings-sync/public/settings-sync.js)). Inert unless the feature is on and someone is logged in. - It wraps `localStorage.setItem` and `removeItem` for allowlisted keys only and pushes 2 seconds after the last change. - It pulls on login, page load, tab focus and every 60 seconds while the tab is visible. - **Merge:** three-way, against a per-device baseline that belongs to one user and one document generation. Lists (own nodes, favorites, saved filters) merge per item, so an item added anywhere is kept and an item removed on one device does not come back from another. Single values: the profile wins unless only this device changed it. - Remote changes are written without a push, theme and colour-blind preset are re-applied, and the current page re-renders (skipped on account pages and while the geofilter editor is open). **UI.** - Logout asks: keep my settings on this device (default), remove them from this device, or cancel. Channel keys are never removed: no copy exists anywhere else. - The account page gets a "Settings sync" section: last synced time, "Sync now", what is and is not synced, and "Delete synced settings from my account". ## Not synced Layout and device state (panel and column widths, collapsed panels, map positions, geofilter drafts), channel data (#725), the API key, and all `sessionStorage`. The full list is in the [spec](https://github.com/efiten/CoreScope/blob/feat/settings-sync/docs/specs/2026-10-06-user-settings-sync-design.md). ## Performance - One GET per page load, tab focus and minute while visible; one debounced PUT per burst of changes. - The `setItem` wrapper costs one Set lookup per write for non-synced keys. A synced write reads one small revision key, not the stored document. - The server reads or writes one row per request. ## Verification - `internal/users` and `cmd/server`: `go vet` and `go test` pass locally (22 new Go tests), including a test that every allowlisted key still occurs in `public/`, and denylist tests. - `tests/unit/test-settings-sync.js`: 79 passing (vm, real module). The cases cover the merge table, two tabs sharing one storage, stale answers after a push, delete while a push is in flight, and logout while the final push fails. - `sh test-all.sh` exits 0. - `tests/e2e/test-user-management-e2e.js` (10 steps, 4 of them new) passed locally with two browser contexts as two devices: a favorite and the packet time window travel from device 1 to device 2, a removal does not come back, and "remove from this device" clears the synced keys while a channel key stays. - Checked by hand on a staging instance with a desktop and a phone on one account. ## Not in this PR - On a shared browser where the previous user chose "keep", the next user's first login merges those settings into their own account. The user guide says to choose "remove" on shared computers. - Saved filter expressions are synced as typed, including any channel names written in them. The guide says so. - Realtime push between devices; the minute pull is the sync interval. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
d232072f85 |
feat: optional user accounts (part A: foundation) (#2129)
Part A of #2128: optional, off-by-default user accounts. With the feature off nothing changes; with it on, visitors can register and log in, and admins manage users and use the operator actions without the API key. PR #2130 (settings sync) builds on this one. The two are meant to be merged together. ## The situation - Operator actions (geofilter save and prune, backup, perf reset) need the shared `apiKey`. There is no per-person right. - Nothing in CoreScope knows who a visitor is, so the requests in #2128 that need that (#1835, #2092, #1508, #730) have nothing to build on. ## What this PR adds **Two new Go modules** - `internal/users`: a separate `users.db` (SQLite through `modernc.org/sqlite`) with users, sessions, single-use tokens, an audit log and a mail log. Passwords use argon2id. - `internal/mailer`: a `Mailer` interface with a Brevo client (send, delivery events, webhook parsing) and an in-memory fake for tests. **Server (`cmd/server`)**, active only with `userManagement.enabled` - 24 routes, all documented in OpenAPI under the `users` tag ([`auth_routes.go`](https://github.com/efiten/CoreScope/blob/feat/user-management/cmd/server/auth_routes.go)): - auth: register, activate, login, logout, me, forgot, reset; - account: profile, password, email change with confirmation, sessions, self-delete; - admin: list, detail, disable, enable, delete, role, resend activation, manual activation, mail status refresh; - a Brevo webhook, registered only when `mail.webhookSecret` is set. - `requireAdmin` replaces `requireAPIKey` at the 7 operator call sites: the API key **or** an admin session. With the feature off it is the old API-key gate (`TestRequireAdminWithoutUserManagementIsAPIKeyGate`). - `/api/config/client` gets `userManagement: {enabled: true}` only when the service started; with the feature off the response is byte-identical. **Frontend** - `auth.js` (header account control, request helper that adds the CSRF header), `account.js` (login, register, activate, forgot, reset, confirm email, my account), `admin-users.js` (`#/admin/users`, deep-linked filters), `account.css` (theme tokens only). - On phones the top-bar control is hidden, so a conditional entry goes into the bottom-nav "More" sheet and the nav drawer. - The customizer geofilter tab and the Perf "Reset stats" button use the admin session when there is one. **Config.** A `userManagement` block (`config.example.json`, [`docs/user-guide/accounts.md`](https://github.com/efiten/CoreScope/blob/feat/user-management/docs/user-guide/accounts.md)). The Brevo key can come from `CORESCOPE_BREVO_API_KEY`. The server refuses to start when the block is enabled but incomplete. ## Security choices - Session cookie `cs_session`: HttpOnly, SameSite=Lax, Secure when `publicBaseUrl` is https. Every cookie-authenticated state change needs the `X-CS-CSRF` header and a matching Origin. - Activation needs the token **and** the account password. Without the password, an attacker who keeps re-registering a known address could get the owner to activate an account that carries the attacker's password. - Register, forgot and email change answer identically for known and unknown addresses. A password reset ends all sessions, a password change ends all other sessions, and both end outstanding email-change links. - Rate limits: login 10 per 15 minutes, register and forgot 5 per hour, per IP and per address. The bucket count is capped. `trustedProxies` makes the per-IP limits see real client IPs behind a proxy. - Server logs carry `#<user id>`, never addresses, tokens or passwords; mail-provider error texts are redacted before logging. ## Performance No change to an existing hot path with the feature off. With it on: - One `users.db` lookup per authenticated request (session by token hash). - The admin user table rebuilds its `tbody` on each filter change. `users.List` caps the result at 1000 rows (`internal/users/users.go`), which bounds the rebuild. - `map[string]interface{}` in `openapi.go`: 79 before, 78 after. ## Verification - `internal/users`, `internal/mailer` and `cmd/server`: `go vet` and `go test -race` pass locally. 121 new Go tests. - `cmd/server` with `-tags e2etest`: vet and the e2e hook tests pass. - `sh test-all.sh` exits 0. `tests/unit/test-user-management-ui.js`: 67 passing (vm, real modules). - `tests/e2e/test-user-management-e2e.js` (6 steps) passed locally against an `e2etest` build with the fake mailer and against a feature-off build. CI builds the `e2etest` binary and runs the suite on a second server (`deploy.yml`). - On a staging instance with a real Brevo key: register, activation mail delivered, activate, admin table, "Refresh status" showing sent, deferred, delivered, opened and clicked. ## Not in this PR - Settings sync (#2130), the admin dashboard, approval flows and notifications (parts B to E of #2128). - A `requireReadAuth` mode (#1835). Sessions from this PR are what such a mode would accept. - Binary size and build time with `modernc.org/sqlite` linked next to `mattn/go-sqlite3` were not measured. Their driver names do not collide. #1992 discusses the driver choice. - No Brevo webhook was configured on staging; delivery status there came from "Refresh status". --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
4363403495 |
fix(nodes): match the region node filter on observer ID, not ingest-time IATA (#2115)
Follow-up to #2114, from its review. `RegionNodePubkeys` compared the IATA copied onto each observation at ingest (`StoreObs.ObserverIATA`). When an operator changes an observer's code, those copies stay as they were until a restart or until the observations age out. `/api/nodes?region=` then kept matching the old region, while the other store region filters already used the new one. Those filters resolve observers by ID through `resolveRegionObservers`, for example the packets query and `computeNodeHomeRegions`. The SQL subquery that #2114 replaced also read `observers.iata` at query time, so this restores that behaviour. ## Also fixes a regression from #2114 Found in review: `loadChunk`, the background history loader (`cmd/server/store.go`, the chunk SELECT and the `StoreObs` it builds), sets `ObserverID` on each observation but never `ObserverIATA`. With #2114 matching on that IATA, **every node whose adverts came only from background-loaded history dropped out of `/api/nodes?region=`** on current master. Matching on observer ID fixes it. `TestRegionNodePubkeysMatchesChunkLoadedObservations` pins it (an observation with the observer ID and an empty IATA) and fails on master's `region_nodes.go`. ## Change - Resolve the region to observer IDs with `resolveRegionObservers` (own mutex, 30 s cache) and match observations by `ObserverID`. Lock order: `regionNodesMu` is released before it, and `s.mu` is taken after it; none of the three is held together. - Without a database there is nothing to resolve, so `RegionNodePubkeys` reports no set and the handler keeps the SQL path. ## Tests - The region tests now seed an observers table and leave each observation's IATA at a stale value, so they can only pass through the table. - New `TestRegionNodePubkeysFollowsObserverIATAChange`: an observer that moved from SJC to SFO matches SFO and not SJC. It fails on the previous code (`got [pk_moved]` for SJC). - `TestRegionNodePubkeysMatchesChunkLoadedObservations` (second commit), see above. - `setupTestDB` takes `testing.TB` so the benchmark can use it; every existing caller passes `*testing.T` unchanged. - Full `cmd/server` suite passes locally, the region tests also with `-race`. ## Performance `BenchmarkRegionNodePubkeys` (220k adverts × 8 observations): 37 ms to 30 ms per uncached scan, a map lookup per observation instead of a string normalisation. The observer lookup is one query on the small `observers` table, cached for 30 s. ## Not done - No singleflight on a cold cache, and the 64-entry cache still resets when full. Both were non-blocking in the #2114 review. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d5e6d1b2a6 |
fix(nodes): resolve the /api/nodes region filter from the packet store (#2114)
Fixes #2101. `/api/nodes?region=` filtered nodes with a subquery that joins every advert to all of its observations and observers, with no time bound (`cmd/server/db.go`, `GetNodes`). It ran twice per request, once for `COUNT(*)` and once for the page. A client paging through nodes repeats both per page. ## Reproduced On our staging database (10.3 GB, 16.5M observations, 219,750 adverts), read-only `sqlite3`, region `BRU`: | | real | user | sys | |---|---|---|---| | one regional `COUNT(*)`, as `GetNodes` builds it | **154.8 s** | 1.9 s | 5.8 s | Almost all of it is waiting on disk. The plan walks `idx_transmissions_payload_type` for every advert and then `idx_observations_dedup` for each one's observations. Two of those per request, a few requests at once, and the reader pool is gone, which matches the report of all four database workers sitting in `GetNodes`. ## Fix - **`PacketStore.RegionNodePubkeys(region)`** (new `cmd/server/region_nodes.go`) walks the store's in-memory adverts (`byPayloadType[ADVERT]`) once. It keeps the pubkey of every node with an advert heard by an observer in the region, using the `ObserverIATA` each observation already carries. It is cached for 30 s per region, and the cache is bounded to 64 entries because its keys come from the client's parameter. Lock discipline follows `resolveAreaNodes`: the cache mutex and `s.mu` are never held together, and the ordering note in `store.go` lists it. - **`GetNodes` takes a `NodeQuery` struct** with a `RegionPubkeys` field, passed as one `json_each` parameter: `public_key IN (SELECT value FROM json_each(?))`, a primary-key lookup. An empty set matches nothing. The SQL subquery stays for a server without a store (tests, tooling). - The advert-pubkey lookup that `trackAdvertPubkey`, `untrackAdvertPubkey` and `computeNodeHomeRegions` each copied is now one helper, `advertPubkey`. The struct instead of a second `GetNodes` variant keeps the `map[string]interface{}` count unchanged in `db.go` (75) and `routes.go` (59). ## Behaviour change The region filter now covers the adverts the store holds (`packetStore.retentionHours`), which is the window the rest of the UI shows, instead of all database history. A node heard in a region only before that window no longer matches the filter. While the store is still loading after a restart, the set grows as history loads. ## Performance | | before | after | |---|---|---| | region set, uncached | 154.8 s (SQL count, staging) | 37 ms (`BenchmarkRegionNodePubkeys`: 220k adverts × 8 observations, 4,000 nodes) | | region set, within 30 s | same again | cache hit | | node count over the set | (included above) | 2 ms on staging (1,200 keys) | | 500-row page over the set | same scan again | 3 ms on staging | The scan holds `s.mu` for reading for those 37 ms, at most once per region every 30 s. ## Tests - `TestRegionNodePubkeys`: the in-region advert, an advert heard in two regions, case and whitespace in codes, a comma list, an unknown region giving an empty set, a non-advert never counting, a blank region giving no filter. - `TestRegionNodePubkeysCached`, `TestRegionNodePubkeysCacheIsBounded` (1,000 distinct regions stay within 64 entries). - `TestGetNodesRegionPubkeys`: the set combines with the role filter and counts correctly, and an empty set returns nothing even with `Region` set. - `TestHandleNodesRegionUsesStore`: an advert that only the store knows about shows up through `/api/nodes?region=`, so the handler is proven not to ask SQL. - Existing region tests (`TestGetNodesRegionFilterV2` and the `db_test.go` region cases) pass unchanged through the SQL fallback. Full `cmd/server` suite passes locally; the new tests also pass with `-race`. ## Not done - No request context on these queries, also raised in the issue. With the scan gone they take milliseconds, so I left that out of this change. - The region semantics stay "heard by an observer in the region". #1879 argues for the node's home region instead; that is a separate decision. - No frontend change. I did not check this in a browser; the nodes page and map call the same endpoint with the same parameters. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b5b230e884 |
feat(channels): show the sender's path hash size on each message (#2089)
Each channel message now shows the hash size its sender's path uses, read from bits 7-6 of the path byte that the originator writes and repeaters keep. The server sends it as path_hash_size (packetpath.HashSize); the frontend helper pathHashSize() applies the same rule and returns 0 for unknown. One rule on both pages: the packet detail Hash Size row and the hex breakdown call pathHashSize() too, so a 0-hop flood message reports the same size on Channels and on its packet page. A direct packet with no hops left reports no size, matching cmd/server/decoder.go. The cases live in test-fixtures/path-hash-size-cases.json, read by the Go and JS tests. |
||
|
|
000d9ab030 |
feat(coverage): RF noise-floor layer on the Mobile RX coverage page (#2113)
The ingestor already stores the noise floor that CoreDrive RX companions
report with each GPS fix (`client_rf_samples`, opt-in through
`clientRfSamples`, `cmd/ingestor/client_rf_sample.go:62`), but nothing
reads it back. An operator who enables it collects the data and cannot
see it. This adds the read side: `GET /api/rf-noise` and a Signal/Noise
toggle on the Mobile RX coverage page.
## What it does
- **`GET /api/rf-noise?bbox=&z=&days=`** returns a GeoJSON hex grid with
the median, quietest and noisiest noise floor per cell, in the same
shape as `/api/rx-coverage`. It is registered always and 404s unless
`clientRfSamples.enabled` is true, the same pattern as the coverage
routes.
- **Stationary samples are excluded.** A parked companion logs hundreds
of readings at one point, which would otherwise define its cell.
- **Coverage page:** a Signal/Noise toggle, rendered only when
`/api/config/client` reports `clientRfSamples: true`. The noise layer
reuses the coverage colour tokens with the axis inverted, because a
lower dBm is quieter. Tiers are at -115 and -108 dBm.
- **Deep link:** `#/rx-coverage?layer=noise` opens on the noise layer.
- **Empty and failed answers** are labelled on the map ("No RF samples
in this view yet", or a retry hint), so a blank map never reads as
"feature off".
No new configuration key: it reads the existing `clientRfSamples`
section that the ingestor already uses. Default off, so nothing changes
for an instance that has not opted in.
## Where it comes from
Ported from the efiten/CoreScope fork, where it has run on
analyzer.on8ar.eu since September (fork commits `42d09f0e`, `80581c43`,
`cde95078`). The cherry-picks conflicted with upstream's newer
`routes.go`, `types.go` and `rx-coverage.js`, so the final state was
ported by hand. The fork-only `/scopes` route that sat next to it in the
same hunk is deliberately left out.
## Performance
The query is bounded by `sampled_at` (indexed, `idx_crf_prune`) and the
bbox, aggregation is one pass plus a per-cell sort, and the response is
capped at 5000 cells. Measured on analyzer.on8ar.eu, all of Belgium
(`bbox=49.4,2.4,51.6,6.5&z=9`), from a client in Belgium, so network
time included:
| window | samples | cells | response time |
|---|---|---|---|
| 7 days | 10,835 | 241 | 0.37 s |
| 30 days | 43,303 | 429 | 0.39 s |
That table holds 45,270 rows in total. It is only read when someone
opens the noise layer, never on ingest or WebSocket paths.
## Tests
- `cmd/server/rf_noise_test.go`: 8 tests for the aggregation (median and
extremes, stationary exclusion, the cell cap, empty input) and the gate
(404 when off, even with data present). All pass, and the full
`cmd/server` suite passes locally.
- `tests/unit/test-rx-coverage-noise.js` (new, in `test-all.sh`): the
colour axis runs the right way, including both tier bounds and a string
median from the API. Slices the real function out of `rx-coverage.js`
and fails on master's copy, which has no thresholds.
- `tests/e2e/test-rx-coverage-noise-e2e.js` (new, wired in `deploy.yml`
and `scripts/non-unit-tests.json`): no toggle when the flag is off;
Noise fetches `/api/rf-noise`, draws the cells, swaps legend and
subtitle, puts `layer=noise` in the hash; a deep link opens on the noise
layer and an empty answer shows the message. 3/3 locally; against
master's `rx-coverage.js` and `roles.js`, 1/3 (only the "flag off" case
passes, as it should).
- `test-rx-coverage-viewport-e2e.js` still passes. ESLint 8 reports 0
errors on the changed files.
## Not done
- The thresholds (-115 / -108 dBm) are fitted to the fork's own data
(1,241 moving samples at the time) and are constants in
`rx-coverage.js`. Per AGENTS.md rule 8 they belong in the customizer
eventually; not in this PR.
- The E2E suite mocks the API. The real endpoint is covered by the Go
tests and by the measurements above, not by a fixture with seeded
samples.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
9c3d76e14f |
docs(release): notes and changelog for v3.13.1 (#2112)
Release notes and CHANGELOG entry for v3.13.1, a patch release that ships #2111 (node-discover replies count as coverage). Once merged, the tag goes on the merge commit so `deploy.yml` picks up `docs/release-notes/v3.13.1.md` as the release body. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>v3.13.1 |