Files
meshcore-analyzer/docs/client-rx-coverage.md
T
efitenandClaude Opus 5 00ee4f9e23 fix(coverage): attribute node-discover replies to the responder (#2111)
## Problem

A CoreDrive RX companion that runs node-discover gets no coverage on an
upstream instance. Reported today by an operator on v3.13.0: the app
logged 32 discover replies from one repeater in 16 minutes (`heard
751a49f0c5dadc70 (8B, discover)`), all published, and
`client_receptions` stayed at 0 rows for that companion while
`client_rx_observations` held 93.

Two gaps, both on the upstream side:

1. **Ingestor.** `deriveHeardKey` attributes a FLOOD `path[last]` and a
0-hop advert, and drops a 0-hop `CONTROL/DISCOVER_RESP` (firmware
`CTL_TYPE_NODE_DISCOVER_RESP`). The reply carries the responder's own
pubkey at offset 6 of the payload, which `decoder.go` already parses
into `CtrlPubKey` (#1802), but the coverage path never used it.
2. **Server.** CoreDrive RX asks for `DISCOVER_PREFIX_ONLY`, so the
firmware answers with an 8-byte prefix
(`simple_repeater/MyMesh.cpp:799-805`). `coverageHeardKeyCandidates`
only built the 64, 6 and 4 hex candidates, so a 16-hex `heard_key` would
be stored and then matched by no per-node coverage query.

## Change

- `cmd/ingestor/client_reception.go`: a third branch in `deriveHeardKey`
for a discover response with no hops. Accepts exactly 8 or 32 bytes,
nothing truncated, stored with `src='discover'`.
- `cmd/server/rx_coverage.go`: adds the 16-hex prefix to
`coverageHeardKeyCandidates`.
- `docs/client-rx-coverage.md`: documents the `discover` source, the
8-byte keylen and the four-candidate lookup.

The leaderboard and `/api/rx-coverage` read `client_receptions` without
a key filter, so they pick the rows up without a change. Name resolution
goes through `batchResolveHeardKeys`, which is a prefix lookup and
handles 16 hex as is.

## Evidence from a deployment that has had this since 2026-08-19

On analyzer.on8ar.eu, `client_receptions` over the last 7 days by `src`:
discover 4232, rxlog 4909, geo 548, advert 34. Discover replies are 44%
of all coverage rows there (4232 of 9723); on an upstream instance those
rows are not written. They cannot be backfilled afterwards either:
`client_rx_observations` keeps no raw bytes, so the responder pubkey is
gone.

## Tests

- `TestDeriveHeardKey` and `TestBuildClientReception` gain discover
cases: 8-byte and 32-byte keys accepted (32-byte uppercase input
lowercased), a 3-byte and an empty key rejected, a non-discover CONTROL
rejected, a discover response with hops not attributed.
- `TestHandleClientPacketDiscoverRespWritesReception`: end to end, a raw
0-hop DISCOVER_RESP on the client topic writes one `client_receptions`
row with `src='discover'`.
- `TestCoverageHeardKeyCandidatesIncludesDiscoverPrefix`: the 16-hex
prefix is among the per-node candidates.
- Ran locally on Windows: `go test ./...` in `cmd/server` passes; in
`cmd/ingestor` everything passes except
`TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which needs the symlink
privilege on Windows and fails on clean master too.

Not done: no browser validation, as the change is ingestor and server
only and the frontend reads the same endpoints. Not included: geographic
resolution of 1-byte hops (`src='geo'`) and the RF noise layer, which
are separate changes.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:15:50 +02:00

309 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Client RX Coverage
Crowdsourced RF coverage from mobile clients: a phone connects over BLE to a MeshCore
*companion* radio, captures which nodes the companion hears (with SNR/RSSI), tags each reception
with the phone's GPS position, and publishes it to MQTT. CoreScope ingests these into
`client_receptions` and renders per-node H3-style hex coverage on the Reach page.
## Companion app — where to get it
The mobile capture side is **[corescope-rx](https://github.com/efiten/corescope-rx)** — an
open-source (GPL-3.0) Android PWA. Operators who enable coverage point their users at it: it connects
over BLE to a MeshCore companion radio, captures directly-heard nodes + the phone's GPS, and publishes
the payload defined below. It's self-hostable and generic — a runtime `config.json` aims it at your
own MQTT broker + CoreScope instance (see its README).
## Enabling coverage (operators)
Coverage is **off by default**. To turn it on:
1. In CoreScope's `config.json`, set `"clientRxCoverage": { "enabled": true }` and restart the server
and ingestor. This is a **single flag read by both processes** — the ingestor and server each parse
the same `config.json`, so you set `clientRxCoverage.enabled` once and it gates both the ingest write
path and the read endpoints. There is no separate per-process flag.
2. **Required: an ACL-capable broker.** Bind `meshcore/client/{PUBLIC_KEY}/packets` so each client may
publish **only** under its own pubkey (e.g. an EMQX ACL keyed on the connected client's identity).
This is the trust boundary, not an optimization — see [Trust](#trust). The ingestor already
subscribes under `meshcore/#`.
3. Optionally set `retention.clientRxDays` to bound the coverage tables (see
[Storage](#storage--client_receptions-ingestor-owned)).
4. Point your users at [corescope-rx](https://github.com/efiten/corescope-rx) and they start
contributing. Results show on each node's Reach page (coverage toggle) and the `#/rx-coverage`
dashboard. **Warn them first that their contribution is world-readable and a per-observer view can
reconstruct their movements — see [Privacy](#privacy--contributor-location-is-public).**
The rest of this document is the MQTT payload contract the companion app implements.
## Companion BLE source (verified against firmware)
The mobile app's RX data comes from the companion's **`PUSH_CODE_LOG_RX_DATA` (0x88)** BLE frame:
`[0x88][snr×4 int8][rssi int8][raw packet bytes]`. This is emitted for **every** received
packet (promiscuous, incl. overheard flood traffic), not just messages addressed to the device:
- `src/Dispatcher.cpp:198` calls `logRxRaw(getLastSNR(), getLastRSSI(), raw, len)` in `checkRecv()`
**unconditionally** — NOT behind `#if MESH_PACKET_LOGGING`. So it works on stock firmware.
- `examples/companion_radio/MyMesh.cpp:283` overrides it to write the 0x88 frame whenever the app
is connected over BLE (`_serial->isConnected()`).
So per received packet the app gets SNR + RSSI + the raw bytes. It decodes the raw packet (standard
MeshCore format) to derive the directly-heard node (`path[last]`, a 0-hop advert pubkey, or a 0-hop node-discover response pubkey) and pairs it
with the phone's GPS. The bare advert push (`PUSH_CODE_ADVERT` 0x80) carries only a pubkey (no SNR/
RSSI/path) and is NOT used — 0x88 already covers adverts (the raw advert is in its payload).
Caveats: 0x88 is only sent while the app is BLE-connected; packets larger than `MAX_FRAME_SIZE` are
skipped; the firmware doc labels 0x88 "can be ignored" (messaging-app view) — for coverage it is the
primary frame. GPS is always the phone's, never the companion's.
## MQTT topic & payload
Topic: `meshcore/client/{PUBLIC_KEY}/packets` — `{PUBLIC_KEY}` is the companion's pubkey. The
broker (EMQX) should ACL-restrict each client to publish only under its own pubkey, which is how
"a connected companion may only inject under the keys that apply" is enforced.
Payload — meshcoretomqtt-compatible packet, plus a `gps` object:
```json
{
"origin": "<companion name>",
"origin_id": "<companion pubkey hex>",
"timestamp": "2026-06-09T12:00:00Z",
"type": "PACKET",
"direction": "rx",
"raw": "<packet hex>",
"SNR": -7,
"RSSI": -92,
"gps": { "lat": 51.05, "lon": 3.72, "acc_m": 8 }
}
```
- The discriminator is the `gps` object. A packet without `gps` is dropped (coverage needs a position).
- `raw` is decoded server-side to derive the directly-heard node and the path; `hash`/`path` fields
are not required.
- Subscription: the ingestor's default subscription (`meshcore/#`) already covers this topic. Sources
configured with an explicit topic list must add `meshcore/client/+/packets`.
### Region answers — `meshcore/client/{PUBLIC_KEY}/regions`
CoreDrive RX can also ask a repeater it hears directly which regions it is configured to flood, and
publishes the answer here. It is accepted under the same `clientRxCoverage.enabled` switch and lands
in `node_declared_regions`, which the Scope Audit page and the `autoRegionKeys` tier read alongside
the observer `/neighbors` source (`nodes.configured_scope`). Newest answer per repeater wins across
both, by the answer's own `timestamp`, so a buffered upload arriving late cannot overwrite a fresher
one.
```json
{
"origin": "<companion name>",
"origin_id": "<companion pubkey hex>",
"timestamp": "2026-09-18T09:09:32.120Z",
"type": "REGIONS",
"target": "<repeater pubkey, 64 hex>",
"regions": ["*", "hu"],
"truncated": false
}
```
- `regions: []` is stored: the repeater answered that it floods nothing. A missing or non-array
`regions` is dropped, never read as an empty answer.
- `target` must be a full 64-hex pubkey. Names containing a comma, longer than 64 characters, or past
the 64th entry are dropped and the answer is flagged truncated.
- The broker ACL must allow the client to publish to `/regions` as well as `/packets` and `/rf`.
Explicit topic lists need `meshcore/client/+/regions` (or `meshcore/client/#`).
## Capture HARD RULE — only what was heard directly
The app and ingestor record **only the node the companion physically received**, never upstream
relayers:
- **FLOOD** packet **with a path** (≥1 hop) → record `path[len-1]` (the last forwarder = the
immediate RF transmitter). Confirmed against firmware `Mesh.cpp` (`routeRecvPacket` appends the
forwarder's hash to the END of the path) and CoreScope's `neighbor_builder.go:226-228`.
- **DIRECT** packet **with a path** → **NOT attributable, discarded.** Direct forwarders consume the
next hop from the FRONT (`Mesh.cpp removeSelfFromPath`), so `path[len-1]` is the route's
destination-side end, NOT the node we heard. Attributing it credits the SNR to the wrong (often
far-away) node. Only FLOOD routes (0,1) are recorded from a path.
- Packet **with no path** (0 hops) **and** an advert → record the advertiser's full pubkey.
- `direction` must be `rx`. 1-byte (2 hex char) prefixes are excluded (collision-prone, like Reach).
- The RSSI/SNR belong to the directly-received transmission, so they attach to the recorded node.
- The rest of the path is discarded for coverage.
## Storage — `client_receptions` (ingestor-owned)
A roaming companion is a mobile observer with a moving position, so it gets its own table (not
`observations`, which assumes a fixed observer location). Per the #1283 read/write invariant, the
table and all writes live in `cmd/ingestor/`.
```
client_receptions(
id, rx_pubkey, heard_key, heard_keylen, rssi, snr,
lat, lon, pos_acc_m, rx_at, ingested_at, src,
UNIQUE(rx_pubkey, heard_key, rx_at)) -- idempotent re-ingest
```
`heard_keylen` is 32 for a full pubkey (0-hop advert, or a node-discover response that carried the
full key) or 2, 3 or 8 for a multibyte prefix. `src` is `advert`, `rxlog` or `discover`.
`discover` rows come from a 0-hop CONTROL/DISCOVER_RESP, which carries the responder's OWN pubkey
in its payload — a first-hand identification, not a path inference. corescope-rx asks for
DISCOVER_PREFIX_ONLY to save airtime, so those keys are normally 8 bytes; per-node coverage queries
must therefore include the 16-hex prefix among their heard_key candidates. No hex cell is stored — binning is computed server-side from lat/lon.
Indexes: a composite `(heard_key, heard_keylen, lat, lon)` and a `(lat, lon)` index back the coverage
queries; the per-node query matches a sargable `heard_key IN (pubkey, prefix16, prefix6, prefix4)` list so the
composite is used instead of a table scan (see the benchmark in `cmd/ingestor`).
Retention: the table grows on every submission, so set `retention.clientRxDays` (ingestor) to delete
rows older than N days (and stale `client_observers`); `0` disables it. Without it the table is
unbounded.
## Diagnostic observations — `client_rx_observations` (ingestor-owned)
The client topic may also carry packets the companion could not attribute to a directly-heard
node — a DIRECT-route packet with a path, for instance (see the capture HARD RULE above). Those
packets are still decodable, and are optionally recorded as a diagnostic RF observation,
independent of whether they produced a coverage row.
**Not literally every decodable packet, though.** A packet still needs a `gps` fix and
`direction: "rx"` to reach the decoder/observation write at all — `handleClientPacket` returns
early (before `DecodePacket` even runs) when `gps` is missing or its `lat`/`lon` don't parse, and
`direction: "tx"` (a companion's own outgoing transmission) is decoded but explicitly excluded
from the observation write, the same as it already was from coverage. A diagnostic table silently
requiring a GPS fix is a bit surprising, so: no `gps` → no observation row either, same constraint
as coverage.
- Written to `client_rx_observations` only, **never** to `client_receptions` — the coverage
invariant (only directly-heard nodes) is unchanged and unaffected by this feature.
- Gated by its own flag, `"clientRxObservations": { "enabled": true }` — a **top-level** `Config`
field, not nested inside `clientRxCoverage` in the JSON. It IS gated behind `clientRxCoverage` in
the *control flow*: `handleClientPacket` (where the observation write lives) is only reached when
`clientRxCoverage.enabled` is true, so observations require coverage to be enabled even though the
two keys are siblings on disk:
```json
{
"clientRxCoverage": { "enabled": true },
"clientRxObservations": { "enabled": true }
}
```
Config loading is plain `json.Unmarshal` with no `DisallowUnknownFields`, so nesting
`clientRxObservations` under `clientRxCoverage` as written above is silently ignored — the key is
never read, the feature stays off, and nothing logs or errors. An ingestor without
`clientRxObservations.enabled` simply drops these packets (no table writes, no error).
- Enabling the companion app's `fullRfLog` flag while `clientRxObservations.enabled` is `false` here
is pure waste: the phone spends mobile data uploading packets this ingestor decodes and discards,
with no row written and no warning anywhere. `fullRfLog` multiplies normal upload volume — see the
corescope-rx README.
- The JSON payload shape from the companion app is **unchanged** either way — this is purely an
ingestor-side decision based on what `raw` decodes to, not a new field the app must send.
- Captures routing detail the coverage path discards: `route_type`, `payload_type`,
`code1`/`code2` transport codes (route types 0/3 only), `scope_name` (matched against configured
region keys), `hash_size`, `hop_count`, the full forwarder path (`path_json`), and — for FLOOD
routes only — the immediate `forwarder`.
- `rx_at` is stored at millisecond precision (unlike `client_receptions.rx_at`), because
`UNIQUE(rx_pubkey, pkt_hash, rx_at)` deliberately allows multiple rows per `pkt_hash`: each row is
one forwarder's copy of the same flood, and that multiplicity is the flood-amplification signal
this table exists to capture. Retention is `retention.clientRxObsDays` (separate from, and
typically shorter than, `retention.clientRxDays` — this table is diagnostic, not archival).
- `pkt_hash` (`ComputeContentHash`) deliberately excludes both the transport-code bytes and the
path bytes, so distinctness inside `UNIQUE(rx_pubkey, pkt_hash, rx_at)` rests entirely on `rx_at`.
On the happy path that's fine — real receive times come from the envelope timestamp at
millisecond resolution, and same-millisecond collisions aren't physical on a half-duplex LoRa
radio. But on any fallback path (missing/unparseable/implausible timestamp — see
`resolveRxTimeCore`), every packet in a buffered upload batch is stamped with the same ingest-time
`rx_at`, and distinct forwarder copies of one flood inside that batch collapse into a single row
via `ON CONFLICT DO NOTHING`. Not a correctness bug — the constraint is doing exactly what it's
told — but it means a buffered/late upload with a bad envelope timestamp under-reports flood
amplification for that batch. This isn't limited to the server-side fallback path either: the
companion app stamps `rx_at` at BLE-frame *processing* time, not true RF receive time (`app.js`),
so two forwarder copies processed in the same millisecond collapse just as effectively even when
the envelope timestamp itself is fine.
- A 0-hop advert gets `forwarder = NULL` in `client_rx_observations` even though the transmitter is
known — it's the advert's own pubkey, which the coverage path records separately with
`src='advert'` (see `client_receptions` above). Don't mistake this `NULL` for "unknown".
## Read API — coverage GeoJSON
`GET /api/nodes/{pubkey}/rx-coverage?bbox={minLat,minLon,maxLat,maxLon}&z={zoom}`
Returns a GeoJSON `FeatureCollection` of hexagons covering where clients heard the node, aggregated
server-side (read-only). Each feature:
```json
{ "type": "Feature",
"geometry": { "type": "Polygon", "coordinates": [[[lon,lat], ...]] },
"properties": { "cell": "9:123:-45", "count": 7, "best_snr": -6, "has_sig": true,
"nodes": [{ "prefix": "aabbcc", "name": "Alice", "snr": -6, "count": 3 }],
"nodes_truncated": false } }
```
- Hex binning is a pure-Go pointy-top grid over Web Mercator (`cmd/server/hexgrid.go`). We do **not**
use `uber/h3-go`, which would add a C dependency for no benefit (the SQLite driver
is cgo since the mattn/go-sqlite3 move, but that is a library we need). Latitude is only
defined within ±85.05° (Web Mercator limit) and is clamped to that range.
- `z` (Leaflet zoom) selects the hex resolution (zoom-adaptive). Raw points never leave the server
(privacy: contributors' tracks are not exposed).
- `best_snr` / `has_sig` drive the colour: green→orange by best SNR, grey when no signal metric.
- Features are sorted by `cell` for a deterministic (cacheable) payload.
- **Bounds:** the per-cell `nodes` list is capped (with `nodes_truncated`), and the collection is
capped at a fixed feature count — when exceeded, the densest cells are kept and the top-level
`truncated` flag is set. The per-node endpoint also returns `mobile_receptions` and `mobile_clients`
totals (node-wide, independent of the bbox).
## Frontend
Shown only in the Reach view (`#/nodes/{pubkey}/reach`), as a toggleable hex layer drawn on the
existing Leaflet map (`public/node-reach-coverage.js`), deep-linked via `?coverage=1`. No new
frontend dependencies. Colours come from CSS variables in `public/node-reach.css`
(`--nq-cov-strong|mid|weak|grey`).
## Trust
Identity = the companion pubkey (`rx_pubkey`), taken from the `{PUBLIC_KEY}` topic segment.
**The feature requires an ACL-capable broker.** The reported GPS position is the contributor's own
claim, so the only thing anchoring a reception to a real identity is the broker ACL binding
`meshcore/client/{PUBLIC_KEY}/packets` to the client that holds that key. **Without such an ACL, the
topic — and therefore the GPS and the heard-node attribution — is spoofable:** anyone who can publish
to the broker could inject coverage under any pubkey. Do not enable this feature on an open/no-ACL
broker if you trust the resulting map.
Server/ingestor-side defense-in-depth (these reduce blast radius but do **not** replace the ACL):
- The ingestor rejects any topic pubkey that is not lowercase hex before writing, and never falls back
to a payload-supplied id (`cmd/ingestor/client_reception.go`, #2/#10).
- A blacklisted operator cannot contribute via the client topic (the blacklist is enforced before the
coverage write, #1).
- The frontend HTML-escapes the pubkey it renders, so a junk pubkey can't inject markup (#14).
- `/api/nodes/resolve` and coverage tooltips never reveal blacklisted or hidden-prefix node identities
(#15).
## Privacy — contributor location is public
⚠️ **Enabling coverage publishes contributors' GPS-tagged receptions, and the per-observer view can
reconstruct a contributor's movements.** The hex map is read without authentication. The leaderboard
exposes each companion's pubkey, and clicking one filters the map to that single companion
(`/api/rx-coverage?rx=<pubkey>`); at high zoom over the retention window this is effectively a public
movement trail (home / work / commute) of whoever carries that companion. **A pseudonymous companion
name does not mitigate this** — the *locations themselves* are identifying (overnight clustering = home),
and all of one contributor's points are linked by the pubkey.
This is an accepted tradeoff of the feature, not a bug: fine resolution is what makes the aggregate
coverage map useful, the feature is opt-in and OFF by default, and contributors choose to run the
companion. But the consent must be **informed**:
- **Operators:** tell your users, before they contribute, that their coverage (including a per-observer
view of their own track) is world-readable for as long as `retention.clientRxDays` keeps it.
- **Contributors:** do not contribute from a device you carry on your person if a public record of where
you have been is a concern. Use a dedicated/stationary node, or accept that the trail is public.
Operators who want to harden this further can lower `retention.clientRxDays`, run the dashboard behind
their own auth/proxy, or (future hardening) coarsen stored coordinates / apply a k-anonymity threshold
to the per-observer view.
Optional future hardening: have the companion sign a broker-issued token (the firmware exposes
on-device signing) — not required for the MVP, tracked as a follow-up.
## Configurable values (future customizer)
Hardcoded initially, tracked for the customizer per AGENTS.md rule 8: hex resolution per zoom
(`zoomToHexRes`), colour SNR thresholds (`coverageColorVar`), and any `rx_at` max-age validation.