Files
meshcore-analyzer/docs/client-rx-coverage.md
T
Sylvain Rabot 6c86c3141f perf(sqlite): swap modernc.org/sqlite for mattn/go-sqlite3, cross-built with zig
Scans and joins are ~2x faster, prepared lookups ~1.3x. The cost is that the
build is cgo now, paid with `zig cc` so cross-compilation still works from one
builder.

`modernc.org/sqlite` is a pure-Go transpilation of SQLite; this repo is
read-heavy (`cmd/server` chunk-loads a graph at startup and fans out
neighbour/topology/analytics queries per request) and pays for that on
exactly those paths. Head-to-head on the same 120k-transmission /
240k-observation database, running our own hot-path SQL under both drivers
(medians of 5):

| workload                                            | modernc | mattn |
|-----------------------------------------------------|--------:|------:|
| chunk load (`chunked_load.go` v3 join, 20k tx)      |  449ms  | 196ms |
| aggregate scan (240k-row join + GROUP BY)           |  276ms  | 137ms |
| 1500 prepared-statement lookups                     |  512ms  | 403ms |

Allocations fall with it: 1.12M vs 1.64M allocs and 21MB vs 30MB on the
chunk load. Bundled SQLite goes 3.46.0 -> 3.53.4.

Verified on production data (358,266 transmissions / 2,931,500 observations):
full store load in 33s with no errors, and a two-arch image built and run
under QEMU.

## The build

`CGO_ENABLED=0` still *builds*, which is the trap: mattn links a stub and the
binary dies on its first query with "go-sqlite3 requires cgo to work. This is a
stub". `GOOS=linux go build` genuinely cannot cross-compile any more. A new root
`Makefile` is the entry point; its
`crossbuild` target uses `zig cc -target {x86_64,aarch64}-linux-musl` and
links static, so the artifact stays a single self-contained file and the
`alpine:3.20` runtime no longer depends on the base image's libc at all.
`-Wl,-s` is load-bearing: Go's own `-s -w` does not reach the musl objects
zig links in, and without it the server binary is 19.8MB instead of 12.1MB.

`netgo,osusergo` keep the pure-Go resolver and user lookup the binaries had
under `CGO_ENABLED=0`, so enabling cgo for SQLite does not quietly move DNS
resolution to the C resolver.

The Dockerfile keeps its single `$BUILDPLATFORM` builder and gains a
checksum-pinned zig plus BuildKit cache mounts; without those, an image
build recompiles the amalgamation from cold and takes over half an hour.

## Five behavioural differences, not one

**Statement preparation is eager.** modernc's `newStmt` stored the SQL and
compiled lazily; mattn calls `sqlite3_prepare_v2` inside `Prepare`, so SQL
referencing a missing table fails at *open*. 59 server tests failed on this
alone, all fixtures with partial schemas. `OpenDB` keeps failing loudly
(#1901, and `main.go` gates on `dbschema.AssertReady` anyway); the fixtures
now declare what they are prepared against via `ensurePreparable`, and
`TestEnsurePreparableMatchesPrepareStatements` fails if a new prepared
statement outgrows it. It also exposed nine `nodes(pubkey ...)`
declarations across seven files, when production has only ever had `public_key`.

**It surfaced a real bug.** `stmtInsertObservation` resolves its ON CONFLICT
against `idx_observations_dedup`, which `cmd/ingestor/db.go` only ever
created inside the branch that creates `observations` for the first time.
Databases predating that branch never had one, so the UPSERT had no conflict
target: modernc failed on first insert, mattn fails at `OpenStore`. Same bug,
found earlier. `internal/dbschema` now creates it unconditionally, collapsing
pre-existing duplicates first, because `test-fixtures/e2e-fixture.db` held one.

Replaying that UPSERT faithfully is subtler than it looks, and a first cut of
this got it wrong in two ways. `COALESCE(excluded.x, x)` means the *incoming*
value wins, so down a group in id order the survivor keeps the LAST non-NULL
value, not the first; and the UPSERT names exactly five columns, so every other
column must keep the surviving row's own value rather than be merged. Taking the
first non-NULL, across all columns, silently discarded newer readings. Merge,
delete and index creation now also share one transaction: split apart, a writer
inserting a duplicate in the gap fails the index creation while leaving the
deletions committed.

**`synchronous` silently dropped FULL -> NORMAL.** mattn defaults it to
NORMAL and executes the pragma unconditionally, where SQLite's own default
(what modernc left alone) is FULL. In WAL mode that weakens durability under
power loss. Pinned with `_synchronous=FULL` in `dbschema.WriterDSN`, which both
writers now share: `cmd/migrate` kept a bare path at first and so quietly wrote
at NORMAL, which is what a second copy of a DSN buys you. `TestOpenStorePragmas`
reads every pragma back through the store's own connection and
`TestWriterDSNPragmas` covers the DSN itself, because a separate connection or
the startup log line would prove nothing.

**The DSN dialects are mutually invisible.** modernc understood only
`_pragma=name(value)`, mattn only `_`-prefixed parameters, and neither errors
on the other's form — a driver-only rename would have dropped every pragma in
silence. All five rewritten. `_journal_mode=WAL` is gone from the server's
read handle: modernc ignored it, mattn honours it, and setting journal_mode on
a read-only connection is a write. Dropping `_busy_timeout` with it costs
nothing since mattn defaults to 5000ms, which means the read handle finally
gets the busy timeout it had silently lacked.

**`mode=ro` survives, but for a non-obvious reason.** mattn always passes
`READWRITE|CREATE` and its amalgamation has `SQLITE_USE_URI=0`; what makes the
URI work is its C wrapper ORing `SQLITE_OPEN_URI` in. So the #1283/#1289
invariant holds with no build flags, but it depends on the `file:` prefix.
`TestOpenDBRefusesMissingDatabase` fails if that stops holding. `cmd/decrypt`
had been building its DSN without that prefix, so its `mode=ro` had never
applied and a missing path was created read-write — fixed in passing, never a
migration regression.

## What this does not change

No modernc-specific API was in use: no `RegisterFunction`, no `*sqlite.Conn`,
no `sqlite/lib` error constants, no `sql.Register`. No `time.Time` is ever
bound as a query argument, so driver time handling is not in play. Both
drivers convert declared DATE/DATETIME/TIMESTAMP columns to `time.Time`, so
`/api/dropped-packets` keeps emitting `dropped_at` as RFC3339 — an earlier
draft of this change "fixed" that with a CAST and would have been the
regression.

## Memory

GOMEMLIMIT covers what the Go runtime manages - heap, stacks, runtime
structures - so it is not an RSS ceiling now that SQLite allocates in C. What is
bounded is the page cache specifically: both DSNs pin `_cache_size=-2000`, ~2MiB
per connection, ~8MiB across the server's `SetMaxOpenConns(4)`, comfortably
inside the 1.5x headroom `applyMemoryLimit` derives. That caps the page cache,
not every native allocation. Note `processRSSMB - goSysMB` is NOT the C share:
goSysMB is reserved address space, not resident memory.

## CI

`go-test` gains test execution for `cmd/migrate` and `internal/dbschema`,
which had none and both open the database. A PR-time two-arch build plus an
arm64 QEMU smoke gate is new: the GHCR push is push/tag-only, so without it
nothing on a PR would exercise zig, static musl linking, or arm64, and the
first signal would arrive on master. `cache-dependency-path` widens from 2 of
the 5 tracked `go.sum` files to all of them.

Constraint: cmd/server must stay read-only (#1283/#1289) — mode=ro is load-bearing, not decoration
Constraint: binaries must stay single self-contained files on alpine/scratch
Constraint: durability contract must not change silently — hence _synchronous=FULL
Rejected: CGO_CFLAGS=-DSQLITE_USE_URI=1 | unnecessary; mattn's C wrapper already ORs SQLITE_OPEN_URI, and it would have needed propagating to every build and test path
Rejected: tolerant prepareStatements leaving failed statements nil | would have avoided the fixture work but weakened #1901's fail-loud contract
Rejected: CAST(dropped_at AS TEXT) | modernc already returned time.Time, so the CAST was the regression, not the fix
Rejected: QEMU-native container builds instead of zig | compiles the amalgamation under emulation on every cache miss
Rejected: refusing to start on duplicate observations | without the index the ingestor cannot prepare its UPSERT, so refusing is not the safer option
Directive: never re-add CGO_ENABLED=0 — it no longer builds this repo
Directive: keep Makefile and Dockerfile build tags in step; netgo/osusergo are why DNS still resolves in pure Go
Directive: cross-build recipes set CC=zig; native build/vet/-race must not, or -race stops testing the host toolchain
Reviewed-by: Codex (gpt-6-astra), which caught the dedup merge ordering, the
  non-atomic index creation, and cmd/migrate's durability regression
Confidence: high
Scope-risk: broad
Not-tested: the 2-2.3x figures come from a standalone head-to-head harness, not from this load under the old driver
Not-tested: SQLite 3.46.0 -> 3.53.4 query-planner differences on queries with no total ORDER BY
Not-tested: sustained live ingest through the new writer DSN (verified against a static snapshot only)
Not-tested: the duplicate collapse at production scale against a real ingestor writing concurrently; measured at 4.1s on 2.4M synthetic rows holding 5 duplicates, holding the write lock throughout
2026-09-09 22:47:53 +02:00

17 KiB
Raw Blame History

Client RX Coverage

Crowdsourced RF coverage from mobile clients: a phone connects over BLE to a MeshCore companion radio, captures which nodes the companion hears (with SNR/RSSI), tags each reception with the phone's GPS position, and publishes it to MQTT. CoreScope ingests these into client_receptions and renders per-node H3-style hex coverage on the Reach page.

Companion app — where to get it

The mobile capture side is corescope-rx — an open-source (GPL-3.0) Android PWA. Operators who enable coverage point their users at it: it connects over BLE to a MeshCore companion radio, captures directly-heard nodes + the phone's GPS, and publishes the payload defined below. It's self-hostable and generic — a runtime config.json aims it at your own MQTT broker + CoreScope instance (see its README).

Enabling coverage (operators)

Coverage is off by default. To turn it on:

  1. In CoreScope's config.json, set "clientRxCoverage": { "enabled": true } and restart the server and ingestor. This is a single flag read by both processes — the ingestor and server each parse the same config.json, so you set clientRxCoverage.enabled once and it gates both the ingest write path and the read endpoints. There is no separate per-process flag.
  2. Required: an ACL-capable broker. Bind meshcore/client/{PUBLIC_KEY}/packets so each client may publish only under its own pubkey (e.g. an EMQX ACL keyed on the connected client's identity). This is the trust boundary, not an optimization — see Trust. The ingestor already subscribes under meshcore/#.
  3. Optionally set retention.clientRxDays to bound the coverage tables (see Storage).
  4. Point your users at corescope-rx and they start contributing. Results show on each node's Reach page (coverage toggle) and the #/rx-coverage dashboard. Warn them first that their contribution is world-readable and a per-observer view can reconstruct their movements — see Privacy.

The rest of this document is the MQTT payload contract the companion app implements.

Companion BLE source (verified against firmware)

The mobile app's RX data comes from the companion's PUSH_CODE_LOG_RX_DATA (0x88) BLE frame: [0x88][snr×4 int8][rssi int8][raw packet bytes]. This is emitted for every received packet (promiscuous, incl. overheard flood traffic), not just messages addressed to the device:

  • src/Dispatcher.cpp:198 calls logRxRaw(getLastSNR(), getLastRSSI(), raw, len) in checkRecv() unconditionally — NOT behind #if MESH_PACKET_LOGGING. So it works on stock firmware.
  • examples/companion_radio/MyMesh.cpp:283 overrides it to write the 0x88 frame whenever the app is connected over BLE (_serial->isConnected()).

So per received packet the app gets SNR + RSSI + the raw bytes. It decodes the raw packet (standard MeshCore format) to derive the directly-heard node (path[last] or 0-hop advert pubkey) and pairs it with the phone's GPS. The bare advert push (PUSH_CODE_ADVERT 0x80) carries only a pubkey (no SNR/ RSSI/path) and is NOT used — 0x88 already covers adverts (the raw advert is in its payload).

Caveats: 0x88 is only sent while the app is BLE-connected; packets larger than MAX_FRAME_SIZE are skipped; the firmware doc labels 0x88 "can be ignored" (messaging-app view) — for coverage it is the primary frame. GPS is always the phone's, never the companion's.

MQTT topic & payload

Topic: meshcore/client/{PUBLIC_KEY}/packets — {PUBLIC_KEY} is the companion's pubkey. The broker (EMQX) should ACL-restrict each client to publish only under its own pubkey, which is how "a connected companion may only inject under the keys that apply" is enforced.

Payload — meshcoretomqtt-compatible packet, plus a gps object:

{
  "origin": "<companion name>",
  "origin_id": "<companion pubkey hex>",
  "timestamp": "2026-06-09T12:00:00Z",
  "type": "PACKET",
  "direction": "rx",
  "raw": "<packet hex>",
  "SNR": -7,
  "RSSI": -92,
  "gps": { "lat": 51.05, "lon": 3.72, "acc_m": 8 }
}
  • The discriminator is the gps object. A packet without gps is dropped (coverage needs a position).
  • raw is decoded server-side to derive the directly-heard node and the path; hash/path fields are not required.
  • Subscription: the ingestor's default subscription (meshcore/#) already covers this topic. Sources configured with an explicit topic list must add meshcore/client/+/packets.

Capture HARD RULE — only what was heard directly

The app and ingestor record only the node the companion physically received, never upstream relayers:

  • FLOOD packet with a path (≥1 hop) → record path[len-1] (the last forwarder = the immediate RF transmitter). Confirmed against firmware Mesh.cpp (routeRecvPacket appends the forwarder's hash to the END of the path) and CoreScope's neighbor_builder.go:226-228.
  • DIRECT packet with a path → NOT attributable, discarded. Direct forwarders consume the next hop from the FRONT (Mesh.cpp removeSelfFromPath), so path[len-1] is the route's destination-side end, NOT the node we heard. Attributing it credits the SNR to the wrong (often far-away) node. Only FLOOD routes (0,1) are recorded from a path.
  • Packet with no path (0 hops) and an advert → record the advertiser's full pubkey.
  • direction must be rx. 1-byte (2 hex char) prefixes are excluded (collision-prone, like Reach).
  • The RSSI/SNR belong to the directly-received transmission, so they attach to the recorded node.
  • The rest of the path is discarded for coverage.

Storage — client_receptions (ingestor-owned)

A roaming companion is a mobile observer with a moving position, so it gets its own table (not observations, which assumes a fixed observer location). Per the #1283 read/write invariant, the table and all writes live in cmd/ingestor/.

client_receptions(
  id, rx_pubkey, heard_key, heard_keylen, rssi, snr,
  lat, lon, pos_acc_m, rx_at, ingested_at, src,
  UNIQUE(rx_pubkey, heard_key, rx_at))   -- idempotent re-ingest

heard_keylen is 32 for a full pubkey (0-hop advert) or 2/3 for a multibyte prefix. src is advert or rxlog. No hex cell is stored — binning is computed server-side from lat/lon.

Indexes: a composite (heard_key, heard_keylen, lat, lon) and a (lat, lon) index back the coverage queries; the per-node query matches a sargable heard_key IN (pubkey, prefix6, prefix4) list so the composite is used instead of a table scan (see the benchmark in cmd/ingestor).

Retention: the table grows on every submission, so set retention.clientRxDays (ingestor) to delete rows older than N days (and stale client_observers); 0 disables it. Without it the table is unbounded.

Diagnostic observations — client_rx_observations (ingestor-owned)

The client topic may also carry packets the companion could not attribute to a directly-heard node — a DIRECT-route packet with a path, for instance (see the capture HARD RULE above). Those packets are still decodable, and are optionally recorded as a diagnostic RF observation, independent of whether they produced a coverage row.

Not literally every decodable packet, though. A packet still needs a gps fix and direction: "rx" to reach the decoder/observation write at all — handleClientPacket returns early (before DecodePacket even runs) when gps is missing or its lat/lon don't parse, and direction: "tx" (a companion's own outgoing transmission) is decoded but explicitly excluded from the observation write, the same as it already was from coverage. A diagnostic table silently requiring a GPS fix is a bit surprising, so: no gps → no observation row either, same constraint as coverage.

  • Written to client_rx_observations only, never to client_receptions — the coverage invariant (only directly-heard nodes) is unchanged and unaffected by this feature.
  • Gated by its own flag, "clientRxObservations": { "enabled": true } — a top-level Config field, not nested inside clientRxCoverage in the JSON. It IS gated behind clientRxCoverage in the control flow: handleClientPacket (where the observation write lives) is only reached when clientRxCoverage.enabled is true, so observations require coverage to be enabled even though the two keys are siblings on disk:
    {
      "clientRxCoverage": { "enabled": true },
      "clientRxObservations": { "enabled": true }
    }
    
    Config loading is plain json.Unmarshal with no DisallowUnknownFields, so nesting clientRxObservations under clientRxCoverage as written above is silently ignored — the key is never read, the feature stays off, and nothing logs or errors. An ingestor without clientRxObservations.enabled simply drops these packets (no table writes, no error).
  • Enabling the companion app's fullRfLog flag while clientRxObservations.enabled is false here is pure waste: the phone spends mobile data uploading packets this ingestor decodes and discards, with no row written and no warning anywhere. fullRfLog multiplies normal upload volume — see the corescope-rx README.
  • The JSON payload shape from the companion app is unchanged either way — this is purely an ingestor-side decision based on what raw decodes to, not a new field the app must send.
  • Captures routing detail the coverage path discards: route_type, payload_type, code1/code2 transport codes (route types 0/3 only), scope_name (matched against configured region keys), hash_size, hop_count, the full forwarder path (path_json), and — for FLOOD routes only — the immediate forwarder.
  • rx_at is stored at millisecond precision (unlike client_receptions.rx_at), because UNIQUE(rx_pubkey, pkt_hash, rx_at) deliberately allows multiple rows per pkt_hash: each row is one forwarder's copy of the same flood, and that multiplicity is the flood-amplification signal this table exists to capture. Retention is retention.clientRxObsDays (separate from, and typically shorter than, retention.clientRxDays — this table is diagnostic, not archival).
  • pkt_hash (ComputeContentHash) deliberately excludes both the transport-code bytes and the path bytes, so distinctness inside UNIQUE(rx_pubkey, pkt_hash, rx_at) rests entirely on rx_at. On the happy path that's fine — real receive times come from the envelope timestamp at millisecond resolution, and same-millisecond collisions aren't physical on a half-duplex LoRa radio. But on any fallback path (missing/unparseable/implausible timestamp — see resolveRxTimeCore), every packet in a buffered upload batch is stamped with the same ingest-time rx_at, and distinct forwarder copies of one flood inside that batch collapse into a single row via ON CONFLICT DO NOTHING. Not a correctness bug — the constraint is doing exactly what it's told — but it means a buffered/late upload with a bad envelope timestamp under-reports flood amplification for that batch. This isn't limited to the server-side fallback path either: the companion app stamps rx_at at BLE-frame processing time, not true RF receive time (app.js), so two forwarder copies processed in the same millisecond collapse just as effectively even when the envelope timestamp itself is fine.
  • A 0-hop advert gets forwarder = NULL in client_rx_observations even though the transmitter is known — it's the advert's own pubkey, which the coverage path records separately with src='advert' (see client_receptions above). Don't mistake this NULL for "unknown".

Read API — coverage GeoJSON

GET /api/nodes/{pubkey}/rx-coverage?bbox={minLat,minLon,maxLat,maxLon}&z={zoom}

Returns a GeoJSON FeatureCollection of hexagons covering where clients heard the node, aggregated server-side (read-only). Each feature:

{ "type": "Feature",
  "geometry": { "type": "Polygon", "coordinates": [[[lon,lat], ...]] },
  "properties": { "cell": "9:123:-45", "count": 7, "best_snr": -6, "has_sig": true,
                  "nodes": [{ "prefix": "aabbcc", "name": "Alice", "snr": -6, "count": 3 }],
                  "nodes_truncated": false } }
  • Hex binning is a pure-Go pointy-top grid over Web Mercator (cmd/server/hexgrid.go). We do not use uber/h3-go, which would add a C dependency for no benefit (the SQLite driver is cgo since the mattn/go-sqlite3 move, but that is a library we need). Latitude is only defined within ±85.05° (Web Mercator limit) and is clamped to that range.
  • z (Leaflet zoom) selects the hex resolution (zoom-adaptive). Raw points never leave the server (privacy: contributors' tracks are not exposed).
  • best_snr / has_sig drive the colour: green→orange by best SNR, grey when no signal metric.
  • Features are sorted by cell for a deterministic (cacheable) payload.
  • Bounds: the per-cell nodes list is capped (with nodes_truncated), and the collection is capped at a fixed feature count — when exceeded, the densest cells are kept and the top-level truncated flag is set. The per-node endpoint also returns mobile_receptions and mobile_clients totals (node-wide, independent of the bbox).

Frontend

Shown only in the Reach view (#/nodes/{pubkey}/reach), as a toggleable hex layer drawn on the existing Leaflet map (public/node-reach-coverage.js), deep-linked via ?coverage=1. No new frontend dependencies. Colours come from CSS variables in public/node-reach.css (--nq-cov-strong|mid|weak|grey).

Trust

Identity = the companion pubkey (rx_pubkey), taken from the {PUBLIC_KEY} topic segment.

The feature requires an ACL-capable broker. The reported GPS position is the contributor's own claim, so the only thing anchoring a reception to a real identity is the broker ACL binding meshcore/client/{PUBLIC_KEY}/packets to the client that holds that key. Without such an ACL, the topic — and therefore the GPS and the heard-node attribution — is spoofable: anyone who can publish to the broker could inject coverage under any pubkey. Do not enable this feature on an open/no-ACL broker if you trust the resulting map.

Server/ingestor-side defense-in-depth (these reduce blast radius but do not replace the ACL):

  • The ingestor rejects any topic pubkey that is not lowercase hex before writing, and never falls back to a payload-supplied id (cmd/ingestor/client_reception.go, #2/#10).
  • A blacklisted operator cannot contribute via the client topic (the blacklist is enforced before the coverage write, #1).
  • The frontend HTML-escapes the pubkey it renders, so a junk pubkey can't inject markup (#14).
  • /api/nodes/resolve and coverage tooltips never reveal blacklisted or hidden-prefix node identities (#15).

Privacy — contributor location is public

⚠️ Enabling coverage publishes contributors' GPS-tagged receptions, and the per-observer view can reconstruct a contributor's movements. The hex map is read without authentication. The leaderboard exposes each companion's pubkey, and clicking one filters the map to that single companion (/api/rx-coverage?rx=<pubkey>); at high zoom over the retention window this is effectively a public movement trail (home / work / commute) of whoever carries that companion. A pseudonymous companion name does not mitigate this — the locations themselves are identifying (overnight clustering = home), and all of one contributor's points are linked by the pubkey.

This is an accepted tradeoff of the feature, not a bug: fine resolution is what makes the aggregate coverage map useful, the feature is opt-in and OFF by default, and contributors choose to run the companion. But the consent must be informed:

  • Operators: tell your users, before they contribute, that their coverage (including a per-observer view of their own track) is world-readable for as long as retention.clientRxDays keeps it.
  • Contributors: do not contribute from a device you carry on your person if a public record of where you have been is a concern. Use a dedicated/stationary node, or accept that the trail is public.

Operators who want to harden this further can lower retention.clientRxDays, run the dashboard behind their own auth/proxy, or (future hardening) coarsen stored coordinates / apply a k-anonymity threshold to the per-observer view.

Optional future hardening: have the companion sign a broker-issued token (the firmware exposes on-device signing) — not required for the MVP, tracked as a follow-up.

Configurable values (future customizer)

Hardcoded initially, tracked for the customizer per AGENTS.md rule 8: hex resolution per zoom (zoomToHexRes), colour SNR thresholds (coverageColorVar), and any rx_at max-age validation.