Commit Graph
12 Commits
Author SHA1 Message Date
n30nex f1edbbef3f fix(packets): preserve observation selection in detail URLs (#2093)
Red commit: `c165087` ([CI assertion
failure](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36644030012)).

Fixes #2091.

Packet detail links now retain the selected observation when page
initialization or filter changes rebuild the URL. Initialization also
restores `obs` from the complete hash after the router strips its query.
Explicit ID links render the requested observation and retain it in Copy
Link state.

The shared updater reads selection from the current route, so returning
to the list or selecting another packet cannot resurrect an old
observation. Existing filter serialization and Clear Filters behavior
remain intact. No new requests, configuration, dependencies or layout
changes.

## Validation

- Real fixture browser checks use nondefault observation 502: hash/ID
load, type/observer/time-window changes, refresh, Clear, and another
refresh. Both URL and selected row are asserted.
- All 183 standalone frontend suites passed; focused filter browser
suite 11/11.
- Broader local core run reached the unrelated Live input readiness bug
#2094, reproduced on unchanged master. The complete CI browser suite
passed.
- ESLint 8: zero errors; 91 existing warnings. XSS diff, syntax and
whitespace checks passed.
- Three independent reviews found no production defects; assertions
additionally prove filter controls changed before checking selection.

E2E assertion added: `tests/e2e/test-filter-ux-e2e.js:179`.

OpenClaw profile/external preflight were unavailable; Chromium and
repository checks ran directly. Local navigation used a 60-second
budget. Final CI passed at `280ced5`, including Go, browsers, coverage
and both container architectures
([run](https://github.com/Kpa-clawbot/CoreScope/actions/runs/36646175352)).
2026-09-30 11:40:00 +02:00
n30nex dc4db17c48 fix(nodes): remove unsupported region fields from Heard By (#2077)
Fixes #2062.

The node health API does not emit observer region data, but Heard By
rendered a Region column containing only dashes and a Regions summary
that could never appear. Remove that column, its sort control, and the
unsupported region displays from the full node page and side panel.

Keep Observer, Packets, Avg SNR and Avg RSSI, along with existing
escaping, signal placeholders, relay counts and badges. No API,
dependency or configuration changes.

- Red commit `4396492` fails on the unwanted Region header and summary;
`a360711` removes the unsupported fields.
- Validation: 183 standalone suites and 7 focused browser checks passed.
Broader browser run: 131 passed, 3 fixture-dependent skips. Independent
reviews completed; the sorting-test finding was addressed in `1452981`.
- Unit coverage evaluates the real templates. The existing CI-selected
browser suite checks four-column alignment and sorting on desktop/mobile
using the actual API field names.
- Browser verified: local Chromium against a fixture-backed Go server;
screenshots recorded in `coverage/issue-2062-heard-by-1400.png` and
`coverage/issue-2062-heard-by-390.png`.
- E2E assertion added:
`tests/e2e/test-issue-1151-orphan-separators-e2e.js:125`, the two #2062
desktop/mobile cases.
- No added requests, loops or data structures.

## Preflight overrides

- External `run-all.sh` unavailable; repository syntax, whitespace,
CSS-variable, XSS and PII checks run directly.
- Local unit execution supplies UTF-8 settings and
`GITHUB_REF_NAME=local-validation`, required by the existing
release-routing test harness.


- Local browser navigation budget raised to 60s for slow local asset
responses; repository assertions and timeout settings unchanged.
2026-09-30 11:39:51 +02:00
efitenandClaude Opus 5 3caa847323 fix(nodes): the advert section is called "Recent Adverts", and says why (#2071)
Closes #2042, reported by @damn-simple-scripts.

## Which of the two fixes

The report offered widening the section to all packets the node
originated, or renaming it, and proposed the rename. Widening sounds
like the better fix, so I measured before agreeing. On a production
database:

| | |
|---|---|
| transmissions total | 1,251,967 |
| with `from_pubkey` populated | 191,237 |
| of those, `payload_type = 4` (ADVERT) | **191,237** |
| with `from_pubkey` and **not** an advert | **0** |

The ingestor fills that column for adverts only. Attributing a relayed
CHAN or TXT packet back to its sender is the path-resolution problem,
not a filter this section could apply — so "show all packets from this
node" is not a small change, it is a different feature resting on
attribution the data does not carry.

So the rename is correct, and the numbers say so rather than my
preference.

## The change

Both copies renamed — the full node page and the side pane. Both already
read `nodeData.recentAdverts`, so the field feeding them said "adverts"
while the heading said "packets".

The heading also gained a `title` naming `from_pubkey` as the reason it
is adverts only. Renaming without explaining invites the same report
from the next reader; the tooltip is where that explanation costs
nothing.

## Tests

`tests/unit/test-issue-2042-recent-adverts-label.js`, four cases: both
copies present and headed "Recent Adverts", no copy back to "Recent
Packets", the advert field still feeding them, and the explanation still
in place. The third matters most — it ties the label to its data source,
so pointing this section at a different field in future fails here
rather than silently making the label wrong again.

## Not done, deliberately

The reporter's follow-up idea, splitting flood adverts from zero-hop
adverts into two lists. They wrote that it can be dropped if an issue
should tackle one thing, so it is not here. It is feasible now: a
zero-hop advert is `ROUTE_TYPE_DIRECT` with an empty path, established
while fixing #2064. That deserves its own issue rather than a paragraph
in this one — say the word and I will open it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:08:28 +02:00
efitenandClaude Opus 5 2cfe9cbbc5 fix(a11y): raise the Scope Audit observed-chip contrast above 4.5:1 (#2070)
Closes #1996.

## Reproduced first

The chips tint their own background — `color-mix(in srgb,
var(--status-green) 16%, transparent)` — which darkens whatever surface
sits behind them. With `--status-green-text` (green-700, `#15803d`) on
top:

| surface | composited chip | ratio |
|---|---|---|
| `--surface-0` `#f4f5f7` | `#d2eddf` | **4.04:1** |
| `--card-bg` `#ffffff` | `#dcf6e5` | **4.38:1** |

Both under the 4.5:1 the #1719 gate requires for normal text. The first
row is the same ratio **and the same hex** @n30nex measured in the
browser, which is how I know the model in the test agrees with what a
visitor actually sees.

## Fixed on the text, not by thinning the tint

Dropping the tint to 12% reaches 4.52:1. That is two hundredths above
the line, and a margin that thin fails again the next time a surface
value moves — the same trap #2039 hit with a threshold sitting inside
the healthy band. One palette step darker gives **6.23:1** on white and
**5.74:1** on `--surface-0`, and keeps the tint that makes a chip read
as a chip rather than as plain text.

- `--palette-green-800: #166534` added. The greens ran 300–700 while the
blues already reach 900, so this fills the scale rather than inventing a
colour.
- `--sa-chip-observed-fg` defined per theme: green-800 in light, and the
bright `#22c55e` **kept** in dark, where the chip already passed at
6.48:1 / 5.73:1. Dark is deliberately untouched.
- The chip reads the variable, so the customizer still governs it and no
literal enters a component.

`.sa-chip-verified` needed no change: `scope-audit.js:106` only ever
adds it alongside `sa-chip-observed` or `sa-chip-unobserved`, and it
contributes an underline. So the fix covers both classes the issue names
— worth stating, since the title mentions both.

## Tests

`tests/unit/test-a11y-1996-scope-audit-chips.js`, 7 cases: both themes ×
both surfaces, that the chip still tints (so the suite cannot pass by
testing nothing), that its colour comes from a variable rather than a
literal, and a guard asserting green-700 **would** still fail — so a
quiet revert to `--status-green-text` turns this red instead of passing.

**Red-run confirmed:** with the old colour restored it fails at exactly
4.04:1 and 4.38:1, naming the composited `rgb(210,237,223)`.

Two deliberate choices in the test:

- **A separate suite, not a case in `test-a11y-1719`.** That suite's
`parseColor` handles hex and `rgb()` only; teaching it `color-mix()` is
a larger change than this fix. These chips are the only
contrast-critical user of the function today. If a second appears, the
two should merge, and the file says so.
- **Derived from the stylesheets, not a rendered page.** The default
Scope Audit fixture renders no chips at all, which is precisely why the
existing browser coverage missed this. A stylesheet-derived check cannot
be defeated by a fixture that shows nothing.

`check-css-vars` passes (180 definitions, 0 undefined) and
`test-test-inventory` passes with the new file classified.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 00:07:57 +02:00
435ac25dd6 feat(nav): show the running version in the nav-drawer footer (#2068)
Takes over #1985 by @SaarMesh-Bot, as offered there on 2026-09-16 and
2026-09-17. **Their commit is the first here, unchanged and under their
authorship**; the rest clears the two review points. Closing #1985 in
favour of this so the rebase and the fixes travel together, not to
reassign the work.

## The feature, unchanged

The frontend never surfaced which build was running, though
`/api/health` has reported `{version, commit, buildTime}` all along. A
footer on the nav drawer now renders `CoreScope <version>`, linking to
the releases page, with commit and build time in the tooltip. Colours
come from existing CSS variables, the label is set with `textContent`,
and a failed health call leaves a neutral label rather than an empty
footer — all as the author wrote it.

## Review point 1: fetch on open, not on page load

The `/api/health` call sat in `buildDom()`, which runs on page load. The
drawer may never be opened, and **cannot** be opened at ≤768px, where
the module is disabled by design. So every visitor's browser was
requesting the endpoint to fill a footer most of them would never see.

Moved into `open()`, after the width gate. `fetchVersion` still caches
its promise for the page lifetime, so re-opening costs nothing.

## Review point 2: the fallback is pinned

`tests/unit/test-nav-drawer-version-footer.js`, six cases:

- a rejected fetch, a non-ok response, and a 200 without `version` each
leave the neutral `CoreScope` label — not a blank footer and not
`CoreScope undefined`, which is what an instance shows exactly when
someone is trying to read its version
- a normal response renders the version, with commit and build time in
the tooltip
- a version carrying markup lands verbatim in `textContent`, and
`innerHTML` is never touched
- the endpoint is requested once however often the footer is filled,
pinning the cache from the other side

It **slices `fetchVersion` and `fillVersion` out of the shipped
`public/nav-drawer.js`** and evaluates them rather than copying them
into the test, so it exercises what ships — same approach as
`tests/unit/test-direct-rf-heard-by.js`. Both slice markers are
asserted, so a rename fails loudly instead of quietly testing nothing.
Each case re-evaluates the slice, because `versionPromise` caches for
the page lifetime and a shared sandbox would hand the second case the
first case's answer.

Listed in `test-all.sh`, which per #2036 is the only frontend runner.

## A note on the third commit

The first version of the test used `setImmediate` to settle the promise
queue. It ran fine under node and **failed eslint**, which treats these
as browser code. I ran eslint and committed in the same command and
pushed without reading its output. Fixed in the commit after, using
`setTimeout(r, 0)`. Recording it because the PR would otherwise show a
lint failure in its history with no explanation.

## Verification

6 of 6 in the new suite, eslint clean on both changed files,
`test-test-inventory.js` passes with the new file classified.
Cherry-picked cleanly onto current master.

---------

Co-authored-by: SaarMesh-Bot <bot@saarmesh.de>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 10:58:36 +02:00
liquidraverandClaude Opus 5 b695b979a2 fix(nodes): keep paginating past a page that post-LIMIT filtering shortened (#2061)
## Problem

`handleNodes` runs the geo-filter, `nodeBlacklist`, `hiddenNamePrefixes`
and area
passes **after** the SQL `LIMIT/OFFSET`, and rewrites `total` to the
filtered
length. A page that loses a row is therefore short **without being the
last
page**, and neither the page length nor `total` can tell a client
whether to ask
for another page.

#1606 added the pagination loop and chose the page length as the
canonical stop.
That is correct only where nothing is ever filtered. Everywhere else the
list
truncates at the first filtered page boundary and strands every node
behind it —
the #1598 symptom reached by a different route: a node that is relaying
right
now simply stops being in the list.

The comment at `app.js:240` rejects `total` for exactly the right
reason, then
picks the signal the same code path also breaks.

## Measured on a live 2346-node deployment

Page sizes for the query the map issues:

```
offset=0     returned=500     ← full, loop continues
offset=500   returned=499     ← one row filtered AFTER the LIMIT → loop STOPS
offset=1000  returned=500     ← never requested
offset=1500  returned=500     ← never requested
offset=2000  returned=345     ← never requested
```

| stop rule | requests | nodes reached |
|---|---:|---:|
| short page (master) | 2 | **999** |
| `has_more`, else empty page | 6 | **2344** |

**1341 nodes, 57%, unreachable through the UI.**

### One hidden node truncates the whole list

The deployment this came from has no `geoFilter`
(`/api/config/geo-filter`
returns `polygon: null`) and no `nodeBlacklist`. It has a single
`hiddenNamePrefixes` entry — a deliberate operator choice — and exactly
one node
whose name starts with it:

```
public_key   d4a46ea2…1054      (64 clean hex chars)
name         🚫🔥☀️
role         repeater
last_seen    2026-09-22T10:09:38Z
```

`handleNodes` drops that row in the `IsNameHidden` pass, which runs
after the SQL
`LIMIT`. The row is counted by the `LIMIT` and by `COUNT(*)`, so the
page it lands
in comes back exactly one short — and stops every client that treats a
short page
as the end.

Isolated against SQL on the same database, seconds apart:

```
SELECT lower(public_key) FROM nodes ORDER BY last_seen DESC LIMIT 500 OFFSET 500
  -> 500 rows
GET /api/nodes?limit=500&offset=500
  -> 499 rows
comm -23 sql.txt api.txt
  -> d4a46ea2e99cab132a3286ef3d9cce9099318790af7f25671fe83de453721054
```

Deterministic — `offset=500` returned 499 on three consecutive requests.
Not a
CDN artifact either: `cf-cache-status: DYNAMIC`, origin `cache-control:
no-store`, no `age` header, and four requests with deliberately unique
cache keys
all returned 499.

So **one deliberately hidden node makes 1341 of 2344 nodes
unreachable.** The
hiding feature does exactly what it was asked to do for that one node,
and takes
57% of the network with it, silently. A single `hiddenNamePrefixes`
entry is
enough; no geo-filter, blacklist or area filter is needed to reach this
state.

### The cutoff moves, which is why this reads as intermittent

The visible set is the sum of the pages up to and including the first
short one,
so the boundary sits wherever the unreturnable row currently sorts by
`last_seen`, and jumps a whole page as ingest reorders the list. Same
deployment, same code, same config, ~2h apart:

| dropped row's rank | first short page | nodes visible |
|---|---|---:|
| inside 0–499 | page 1 | 499 |
| inside 500–999 | page 2 | 999 |

A node is visible or invisible purely by where it lands relative to that
moving
line, so affected nodes appear to vanish and return on their own. Two
operators
on this deployment reported exactly that, independently, while I was
measuring.

### A named reproduction

`HU-ZA-Lentihegy` (`5287a33f…`), reported missing from the map by an
operator
whose companion had logged its advert at 04:20 local the same morning.

Ingest was fine. The row is in `nodes` with `last_seen`
`2026-09-22T02:20:35Z` — the same advert, to the second — valid GPS,
role
`repeater`, 1033 adverts, and `/api/nodes/search?q=lentihegy` returns
it.

```
rank by last_seen : 1081
cutoff at the time:  999
```

It missed by 82 positions. Walking the same live endpoint, same moment:

| stop rule | requests | nodes reached | Lentihegy |
|---|---:|---:|---|
| short page (master) | 2 | 999 | **not reached** |
| `has_more`, else empty page | 6 | 2340 | reached |

The practical shape of this on a busy mesh: 1081 nodes had been heard
more
recently than 9.4 hours, so on that deployment **anything last heard
more than
~9 hours ago was invisible**, alive or not.

`#/nodes` compounds it — its search box filters client-side over the
truncated
set, so the server-side `?search=` never runs and an operator cannot
find the
node by searching for it either, even though the endpoint would return
it.

## Change

**Server** — `NodeListResponse` gains `has_more`, computed from the raw
SQL page
against the real `COUNT(*)` before the filter passes run, so it survives
them:

```go
hasMore := offset+len(nodes) < total
```

Always emitted (no `omitempty`) so a client can tell `false` from an old
server.
No extra request in the fixed path: `has_more` ends the loop exactly,
where the
old rule needed a probe page.

**Clients** — `app.js` `fetchAllNodes`, `nodes.js` `loadNodes` and
`area-map.html`'s inline helper stop on `has_more`, falling back to a
zero-length
page against a server that predates it. An empty page always ends the
loop, so a
`has_more` against a concurrently-shrinking table cannot spin to
`safetyCap`.

Left alone: the three loops are still three copies. Collapsing them onto
`fetchAllNodes` is a bigger change than this fix needs, and `nodes.js`
has its
own inter-page progress UI. Happy to do it separately if you want it.

## Testing

- **Unit** (`tests/unit/test-fetch-all-nodes-pagination.js`): the
fixture now
models the real handler — a row counted by the LIMIT and by `COUNT(*)`,
then
removed from the page. Three new cases. Fails on the old rule at 499 of
1199.
- **E2E** (`tests/e2e/test-map-nodes-pagination-e2e.js`, already wired
into
  `deploy.yml`): the mock drops a page-1 row and emits `has_more`.
  Mutation-checked — restoring master's stop rule fails 3 of its steps.
- **Go** (`cmd/server/nodes_pagination_has_more_test.go`): asserts
`has_more`
stays true on a page filtering shortened. Mutation-checked — recomputing
it
  after the filter block fails the test.
- Full server suite `go test -race`: ok, 41.4s. `gofmt` clean, `go vet`
passes.
- **Against a real binary**, not just mocks: fixture DB migrated with
`corescope-migrate`, `hiddenNamePrefixes: ["SKCE"]`, `limit=3`. Page 1
returns
2 of 3 with `total` rewritten to 2 and `has_more=true`. Walking the real
server
with master's rule reaches 2 nodes; with `has_more`, all 199 visible of
200,
the hidden one still hidden. The real frontend against that server loads
199
  with no JS errors.

Two existing expectations changed, both deliberate:

1. `surfaces ALL nodes past the 500 server cap` — 3 → 4 requests. That
mock emits
no `has_more`, so the 200-row final page can no longer end the loop (a
short
page is exactly what a filtered page looks like) and a zero-length probe
   follows. Against a current server `has_more` still ends it at 3.
2. `rows missing public_key are NOT collapsed into one` — its stub
returned a
constant body, which would now be paged to `safetyCap`. It serves one
page
   then empties.

Local `test-all.sh` exits 1 on two XSS-gate self-tests
(`good-2-tested.js`, `good-4-tested.js`) via a `UnicodeEncodeError`
printing an
emoji under Windows cp1252. Identical on clean `origin/master` in a
scratch
worktree, so it is pre-existing and platform-local, not this branch.

There is a second identical filter block further down `routes.go` on
another list
endpoint. Likely the same class; not touched here.

If you would rather land your own version of this, say so and I will
close mine.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 09:45:18 +02:00
efitenandClaude Opus 5 980c5c4515 fix(nodes): stop the Heard By empty state claiming the node is out of range (#2063)
Follow-up to #2057, which is correct in what it does and overclaims in
one sentence.

## The sentence

When no observer heard the node directly, the card said:

> No observer is within radio range of this node.

#2057's own rule cannot establish that:

- every **direct route** is discarded, because the firmware removes the
sender from the path before retransmitting (`Mesh.cpp`,
`removeSelfFromPath`), so the packet cannot say who transmitted it.
#2057 measured these at **38% of transmissions over 7 days**.
- **28.8% of flood observations with a path** are dropped because the
last hop resolves to more than one candidate, and the rule
under-attributes rather than guesses.

So an empty direct list is missing evidence, not evidence of missing
coverage.

## Why it matters in practice

Sampled 40 repeaters on a production instance after #2057 shipped:

| | |
|---|---|
| at least one direct observer | 24 |
| empty card, "no observer is within radio range" | **16** |
| of those 16, with a non-zero `relayObserverCount` | **16** |

Every node showing "nobody is in radio range" also showed "Seen via
relay by N observers" two lines below. An operator reading that about a
working repeater concludes they have a coverage problem they do not
have.

## The change

Wording only, in both copies of the card (full page and side pane):

> No observation proves a direct reception here, which is not the same
as being out of range.

with the reason in a `title`, so the card stays one line:

> Only flood-routed transmissions identify who was heard: a direct route
removes the sender from the path before retransmitting (firmware
`Mesh.cpp`, `removeSelfFromPath`), and an ambiguous relay hop is left
unattributed rather than guessed. So an empty list is missing evidence,
not proof of missing coverage.

No API change. #2057's rule, shape and performance work are untouched —
I verified its firmware derivation against the clone at `0679dbef`
before writing this: `Packet.h:83`, the forwarder appending with
`packet->getPathHashSize()` at `Mesh.cpp:349`, and `removeSelfFromPath`
on the direct path at `Mesh.cpp:89-105` all read as described.

## Test

`tests/unit/test-direct-rf-heard-by.js` slices this template out of
`public/nodes.js`, so it pins the shipped markup. It now asserts the old
sentence is gone and the qualifier is present.

Worth recording how that assertion was reached: my first version banned
the phrase "out of range" from the card, and it failed — on the new
line, which contains that phrase precisely in order to deny it. A word
ban was the wrong instrument. Matching the old sentence and requiring
the new qualifier is the assertion that actually distinguishes the two
states.

8 of 8 in that suite, eslint clean.

## Not in this PR

The **Regions** line and **Region** column on the same card read
`o.iata`, which `HealthObserverRow` does not emit, so both have always
been dead. #2057 named this and left it; it is now **#2062** with the
file and line references, rather than a remark inside a merged
description.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 23:56:02 +02:00
efitenandClaude Opus 5 d1b615fc0d fix(node-health): list only observers that heard the node on air (#2057)
Closes #2056.

## What changes

The node detail "Heard By" card now lists only observers that received
the node's **own transmission off the air**, and reports the rest as a
count.

```
HEARD BY — DIRECT (8 OBSERVERS)
OBSERVER                REGION  PACKETS  AVG SNR   AVG RSSI
BE-DUF-SiSCD-01         —        16276    7.9 dB   -108 dBm
BE-BRU-Moris  repeater  —        13775   -5.6 dB   -122 dBm
...
Seen via relay by 29 observers. Those observers heard a repeater that
forwarded this node's traffic, not this node.
```

and for a node nothing hears:

```
HEARD BY — DIRECT (0 OBSERVERS)
No observer is within radio range of this node.
Seen via relay by 2 observers. …
```

## The rule, and where it comes from

Read out of the firmware rather than assumed:

| | |
|---|---|
| `Packet.h:83` | `setPathHashSizeAndCount(sz,n) { path_len =
((sz-1)<<6) \| (n&63); }` — hash size rides in the packet's `path_len`
byte |
| `Mesh.cpp:649,678` | only `sendFlood()` sets it, so the **originator**
decides; `CommonCLI.h:69` defaults `path_hash_mode = 0`, i.e. one byte |
| `Mesh.cpp:349` | a forwarding repeater appends its hash with the
packet's size — it cannot upgrade a packet, and the **last hop is who
was heard** |
| `Mesh.cpp:89,103` | on a direct route a forwarder matches the head of
the path and calls `removeSelfFromPath` before retransmitting, so the
path is the **remaining** route and the transmitter is not in it |

So an observation credits exactly one node:

1. Route type must be `ROUTE_TYPE_FLOOD` or
`ROUTE_TYPE_TRANSPORT_FLOOD`. Direct routes never qualify (38% of
transmissions over 7 days).
2. Empty path → the originator, known only for ADVERTs.
3. Otherwise the last hop.
4. The hop must resolve to exactly one candidate. Same gate
`resolvePathForObsColdLoad` already applies: under-attribute rather than
guess. It drops 418,530 of 1,455,721 flood observations with a path over
7 days (28.8%), and it is what stops the wrong-band credits.

## Measured effect

| node | before | after |
|---|---|---|
| BE-BRU-Moris | 36 observers | 3 |
| BE-KRO-RP01 \| ON1KW | 40 | 3 |
| BE-BRE-ON8AR | 38 | 2 |
| NL-BXE-RP01 \| 433 | 35 | 0 |

Network-wide over 7 days, 234 of 1,860 nodes have at least one direct
observer (161 have exactly one, maximum 8). The direct list is therefore
empty for most nodes, with the relay count below it. That is the correct
reading: no observer is in radio range of them.

Independent corroboration on staging: for BE-WIL-3EIK-01 the eight
direct observers are exactly the top eight entries of its Neighbors
table by score and observation count.

## Perf justification

`GetNodeHealth` is fast today precisely because it never walks
observations — it uses one representative observation per transmission.
Direct-RF needs the per-observation path, and that cannot be a
per-request walk: the reference store holds **232,928 transmissions /
2,887,861 observations**, one node's `byNode` slice alone holds **55,458
transmissions / 1,450,544 observations**, and
`/api/nodes/bulk-health?limit=200` would multiply that.

So the aggregate is rebuilt by a background recomputer on the existing
`newAnalyticsRecomputer` pattern, published into an `atomic.Value`.
Reads are `O(direct observers)`, which is **cheaper than before** — the
old code built per-observer sums over every transmission in `byNode` on
every request.

Proof, `BenchmarkBuildDirectHeardIndex`:

```
BenchmarkBuildDirectHeardIndex-12    1    63067900 ns/op
```

3,000,000 observations (60,000 transmissions × 50 observations, 8-hop
paths, 64 candidate repeaters) in **63 ms**, once per recompute
interval.

Per observation the walk does one route-type check, one backward scan of
`PathJSON` for the last quoted token (no allocation, no
`json.Unmarshal`), one prefix-map lookup and one counter update.

Rebuilding wholesale also means eviction needs no bookkeeping: a pass
simply does not see evicted transmissions. The alternative — a field on
`StoreObs` updated incrementally — would have needed the call at five
construction sites (`store.go:942,1264,2854,3179`,
`chunked_load.go:609`), which is the duplication that caused #1558, plus
matching decrements at eviction.

## API

Both `GetNodeHealth` and `GetBulkHealth` carried a near-identical copy
of the observer loop; they now share one builder.

- `observers` — direct-RF only. Same field names, so no client
migration. Rows are a named `HealthObserverRow` instead of
`map[string]interface{}` (one fewer occurrence in a touched file, per
the AGENTS.md ratchet).
- `relayObserverCount` — new integer, observers that saw traffic through
the node without hearing it. `stats.totalPackets` and `stats.avgHops`
still count relayed traffic, so without this number the card would
contradict the figures printed beside it.

`docs/api-spec.md` is updated for both endpoints. It also documented an
`iata` field on these rows that the endpoint has never emitted; removed.

## Tests

- `cmd/server/direct_heard_test.go` — table test over the rule: flood
with empty path and known originator, flood whose last hop is the node,
flood whose last hop is another node, direct and transport-direct routes
(never credit), ambiguous last-hop prefix, listener-only candidate,
1-byte and 2-byte hop sizes; plus aggregation and row-building.
- `cmd/server/node_health_direct_rf_test.go` — end-to-end through the
handler: an observer that only saw relayed traffic must not appear in
`observers` but must be counted in `relayObserverCount`. Plus the
benchmark.
- `tests/unit/test-direct-rf-heard-by.js` — slices the card template out
of `public/nodes.js` and evaluates it, so it tests the shipped markup
rather than a copy: heading, empty state, relay line, singular/plural,
signal columns, listener/repeater badge tri-state.
- `cmd/server/node_health_can_relay_case_1290_test.go` — updated to seed
a genuinely direct reception, since a relay-only observer no longer
carries a badge.
- `cmd/server/analytics_recompute_after_load_test.go` — recomputer count
10 → 11.

Verified locally: `cmd/server` suite green, `sh test-all.sh` green (180
suites), `tests/e2e/test-e2e-playwright.js` 131/134 passed with 3
skipped and 0 failures against the seeded fixture, plus
`test-issue-1147-section-order-e2e.js`,
`test-issue-1151-orphan-separators-e2e.js` and
`test-issue-1281-location-row-e2e.js`, which all assert on this card.
`gofmt` clean, `vet` clean across all modules.

Browser-validated on staging: both the full detail page and the side
pane, on a node with 8 direct observers and on the 433 MHz node with
none. No console errors.

## What this does not do

`prefixMap.resolveWithContext` still guesses on ambiguous hops, so
paths, neighbor edges and analytics keep their current attribution.
Making it abstain is a much larger change and needs its own issue.

The "Regions" line and Region column on this card read `o.iata`, which
this endpoint has never emitted, so both have always been dead. Left as
found rather than widened into this change.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 07:55:37 +02:00
efitenandClaude Opus 5 5c016a210a fix(analytics): show a building state on the distance index's 202, and stop caching it (#2051)
Closes #1997.

## What was wrong

`/api/analytics/distance` answers `202 {status:"building",
retry_after_seconds:5}` with no `summary` until the lazy index (#1011)
has been built.

1. `renderDistanceTab` read `data.summary.totalHops` straight away. The
TypeError is caught by the tab's own try/catch, so it never reaches
`window.onerror`: it is painted into the tab as `Failed to load distance
analytics: Cannot read properties of undefined (reading 'totalHops')`.
2. `api()` caches any `res.ok` body, and `res.ok` is true for 202, so
the placeholder was stored for the `analyticsRF` TTL. Even a correct
retry read the cached "building" body back. That is the half that made
the broken state outlast the index build.

## What changed

- `public/app.js:173` skips the cache write when `res.status === 202`. A
200 still caches, unchanged.
- `public/analytics.js` renders a building notice and retries itself,
honouring `retry_after_seconds` clamped to [1s, 30s]. The timer is
cleared on tab switch and in `destroy()`, and a new render supersedes a
pending retry, so two renders cannot write into the same tab.

The notice uses `.text-center`/`.text-muted` rather than the `.spinner`
class used at `analytics.js:1683`, because `.spinner` has no CSS
anywhere in the repo and renders nothing.

## Tests

`tests/unit/test-issue-1997-distance-building.js` (7 assertions, wired
into `test-all.sh`) pins the two pure decisions the renderer makes and
`api()`'s refusal to cache a 202 while still caching a 200. Red-run on
the unfixed sources: 6 of 7 fail, and the "a 200 is still cached"
control stays green.

`tests/e2e/test-issue-1997-distance-building-e2e.js` (classified in
`scripts/non-unit-tests.json`, invoked from `deploy.yml` with
`CHROMIUM_REQUIRE=1`) serves both responses by route interception, so it
does not depend on whether the server under test has an index built. It
asserts: the building state appears, the tab does not paint the error
text, a retry arrives with no interaction, the retry replaces the
placeholder once the server answers 200, and no retry fires after
leaving the tab.

Check (2) deliberately asserts on the rendered text and not on
`pageerror`: the TypeError is caught, so a `pageerror` assertion would
pass on the broken build too.

## Verification

No local cgo toolchain here since #1992, so I could not build a server
to run the E2E against. Instead I ran its five steps in Playwright
against a live instance with this branch's `public/app.js` and
`public/analytics.js` injected in place of the deployed ones (both files
are byte-identical between that instance and upstream master, so the
injection is faithful):

| check | deployed build | this branch |
|---|---|---|
| (1) building state shown | fail | pass |
| (2) not rendered as data | fail | pass |
| (3) retried on its own | fail (1 request) | pass (2 requests) |
| (4) real payload after retry | fail | pass |
| (5) no retry after leaving the tab | pass | pass |

(5) passes on the broken build too: it schedules no retry at all, so it
is a control and only means anything together with (3).

The committed E2E suite has not been run as committed. CI is its first
real run.

## Not done

- The server still recomputes on every 202 poll rather than signalling
readiness.
- No other analytics tab was audited for the same assume-a-summary
pattern.
- One pre-existing unit suite (`test-preflight-xss-gate.js`) fails on
this Windows machine with a cp1252 `UnicodeEncodeError` from its Python
helper, on a clean tree as well as with this change. Unrelated, and
green on Linux CI.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 17:04:37 +02:00
liquidraver 84eba41592 fix(live): anchor the legend toggle to .live-page so the VCR bar stops eating its clicks (#2049)
The PACKET TYPES legend could not be dismissed on Live: clicking the palette toggle did nothing, while document.querySelector('#legendToggleBtn').click() from the console worked. That asymmetry was the diagnosis, since a synthesised click skips hit testing.

Residual half of #1833. That fix corrected the button's offset to calc(var(--vcr-bar-height) + 10px) but left the anchor: .legend-toggle-btn was position: fixed, so the offset resolved against the viewport, while the .vcr-bar it clears is absolute inside .live-page, whose height subtracts --bottom-nav-reserve (56px + safe-area at <=768). The two anchors disagreed by exactly the reserve, dropping the button into the bar's band, and .vcr-bar at z-index 1000 against the button's 500 took every real click. The window is 641-768px: below 641 both buttons are display:none, above 768 the reserve is 0, which is why neither the desktop nor the phone layout test saw it.

Both buttons are now position: absolute, resolving against the same containing block as the bar. .feed-show-btn carried the identical defect but missed the bar by 16px because --legend-toggle-stack parks it a row higher, so that change prevents a break rather than repairs one.

Three review rounds, each verified rather than argued:
- The new E2E was classified in non-unit-tests.json but had no deploy.yml line, so it would never have run (#2037). Now wired in, and it covers both buttons.
- The '.feed-show-btn is display:none at <=768' claim was wrong (it is <=640), which the author measured and corrected.
- Its first real execution failed on its own setup: getComputedStyle().getPropertyValue() on an unregistered custom property returns the unevaluated declaration ('calc(56px + 0px)'), so the non-zero check could not work. It now asserts the measured gap instead, and deliberately does not pin 56 because the gap measures 58 on CI.

Final run: 14 passed, 0 failed. The discriminating assertion is (f), each button's offset net of the bar being identical at 720 and 1440 (11px and 123px slack); under the old anchor those differ by the reserve. (g) pins the desktop layout unmoved.

Merged by the interim maintainer without a second human reviewer: CI and the review above are the independent checks.

Fixes #1833
2026-09-20 14:00:45 +02:00
liquidraver bbf54cfe1b fix(rx-coverage): open at the configured map default, with its own saved viewport (#2033)
The coverage page opened at a hardcoded [51.0, 4.8] zoom 8 regardless of deployment, ignoring /api/config/map (#2032). It now follows the same precedence as the main map (URL hash, then saved position, then /api/config/map, then [37.6, -122.1] zoom 9) and persists its own position across visits, syncing lat/lon/zoom into the hash so a view is shareable.

The saved position lives under its own key, rx-coverage-view, and the page never reads or writes the main map's map-view: sharing the configured default was the bug, sharing the session position was not. Both suites assert map-view stays untouched after a pan, so reintroducing a shared write fails instead of passing quietly.

Also fixed here: selectedRx is now percent-encoded into the hash, and a generation counter stops a late /api/config/map response or a stale 150ms layout timer from building a map for a page that was already left.

Reviewed twice. Verified by mutation rather than by reading: writing map-view too, ignoring the saved coverage position, and dropping the /api/config/map fetch each make the unit suite exit 1, so it covers the feature, the fix and the rejected alternative. The deploy.yml invocation was confirmed to land inside Run Playwright E2E tests (fail-fast) by parsing the workflow, and CI run 35261689694 is the E2E suite's first real execution: Go prints 'RX coverage viewport regressions OK' and Playwright prints 'RX coverage viewport browser regressions OK'.

Worth recording for the next reviewer: that E2E asserted localStorage.getItem('map-view'), so wiring it into deploy.yml without updating the assertion would have turned the job red on its first ever run. It was registered in scripts/non-unit-tests.json but invoked by nothing, the gap tracked as #2037.

Merged by the interim maintainer without a second human reviewer: CI and the mutation checks above are the independent checks.

Fixes #2032
2026-09-17 22:02:36 +02:00
Alex B 893773338e chore(tests): move root test-*.js into tests/unit and tests/e2e (#2036)
Moves 290 root test-*.js into tests/unit (177, listed in test-all.sh) and tests/e2e (113, classified in scripts/non-unit-tests.json), per #1981 and PR-D of #1385. Root goes from 348 entries to 48. test-all.sh and test-fixtures/ stay put. The inventory guard now fails if a test reappears in the root or sits in the wrong folder.

Verified independently of the diff: the invoked sets are unchanged (test-all.sh 177 before and after, deploy.yml 96 before and after, both identical as sets), and a full local run of test-all.sh on master and on the branch produced 4702 output lines each whose only differences are absolute paths, stack-trace line numbers shifted by the REPO_ROOT line, the inventory wording and two perf ratios. The guard was mutation-checked: a test back in the root, a unit suite in tests/e2e, and a suite dropped from test-all.sh each make it exit 1. CI run 35246304316 ran 97 suites from tests/e2e and is green.

Follow-up 9335c51d finished the instruction files: no bare root test command is left in AGENTS.md, the squad charters, .github or docs, and every tests/ path they name resolves.

Merged by the interim maintainer without a second human reviewer: CI and the local runs above are the independent checks.

Known and deliberately out of scope: 18 of the 113 files in tests/e2e are invoked by no runner at all, and one of them cannot run anywhere because it requires jsdom, which is not a declared dependency. Tracked separately.
2026-09-17 19:06:07 +02:00