The compose split (2026-08-09) renamed the data plane to meshcore-infra-*
containers; the hardcoded meshcore-analytics-mosquitto-1/-timescaledb-1
references broke user creation. Resolve both container names from docker
ps at script start (infra names preferred, generic suffix fallback).
Contact page (site footer -> Contact) with Discord sudogadget and
ukmesh@proton.me, plus an observer-onboarding blurb. Wired into
publicRoutes (sitemap: true), lazy-loaded ContactPage chunk, and the
shared SiteLayout footer (app + website surfaces). Re-baselined
website-live-manifest.json for the new website image (e5db6790).
The cold WS initial-state fetch (fetchInitialState -> getRecentPackets +
getRecentMessages) called meshcore_canonical_node_id() ~100k times per
fetch (per-row subquery against node_identity_aliases). 6.97s cold on the
2-vCPU box; under load it exceeded the 8s WS_INITIAL_STATE_DB_TIMEOUT_MS
abort, so the map never received an initial snapshot (empty map +
synthetic_check_failed). Replace the per-row function calls with
node_identity_aliases LEFT JOINs (COALESCE(alias.canonical_node_id,
upper(btrim(node_id))) - exact function semantics, table ~20 rows).
Measured: full getRecentMessages 6968ms -> 1772ms (4x); getRecentPackets
same pattern. Also raise the WS_INITIAL_STATE_DB_TIMEOUT_MS env cap
9s -> 60s (default stays 8s; deploy sets 30s) so slow-but-working fetches
complete and warm the cache instead of aborting forever.
281/281 backend tests pass; tsc clean. Live: initial_state 502ms warm,
all synthetic checks green.
Two live-E2E bugs:
1. mosquitto only re-reads password_file on SIGHUP; the reconciler reloads
every ~60s but a device may connect before then and gets CONNACK 135.
The script now sends kill -HUP 1 after mosquitto_passwd -b and -D so
credentials are valid/revoked immediately.
2. The flock was held through the entire discovery watch (up to hours),
blocking every other newuser run. Release before the passive log-poll
and re-acquire before the mutation phase.
Root cause: node_identity_nodes was estimated at one scoped row, so the combined two-endpoint filter formed a nested-loop Cartesian search and issued roughly 174 million composite link-index probes. The 14-row radio-report table was not the bottleneck.
Split the scoped-node and viable-link reads, intersect endpoints in memory, and preserve the existing all-time neighbor SUM/MAX fields. Thread an 8s AbortSignal through every initial-state database read so a pathological cold fetch releases its gate slot before the 10s synthetic deadline.
Evidence:
- before EXPLAIN (plan only): Nested Loop with node_identity_links_pkey and net_nodes rows=1
- fixed EXPLAIN ANALYZE (ARRAY['ukmesh']): 258.325ms + 61.944ms = 320.269ms server time; 0.52s psql wall
- production getViableLinks('ukmesh'): 20,448 rows in 312.1ms
- semantics check: 20,448 links; neighbor rows=11, reports=15, best SNR=12 dB
- npm test: 280 passed, 0 failed
- npx tsc --noEmit: PASS
The hourly rollup rebuild was filtering source rows with
PUBLIC_PACKET_PRIVACY_SQL, so a backfill reproduced the public-only
rollup (~26k/24h) even after the incremental writer (packetBatch.ts)
was changed to count all rows (~101k/24h). The tool's hourly apply
slice + dry-run + catchUpCurrentHour now have NO privacy filter,
matching the runtime writer. The daily max-hop and observer-region
phases KEEP the privacy filter - their runtime writers still filter.
The hourly/daily chart series come from packet_hourly_stats, whose
incremental writer (packetBatch.ts) filtered WHERE path_is_valid AND
NOT is_private - so the charts only ever showed public-valid rows
(~26k/24h) while the totals now count all observations (~101k/24h).
Per Ben's spec (every packet received by every observer counts), the
rollup must include private rows too.
Changes:
- packetBatch.ts: rollup INSERT drops the path_is_valid/is_private
filter (keeps the meshcore-test topic exclusion) - all rows counted
- statsRepository.ts: aggregate raw-slice CTEs (fresh hour + rollup
gap) use totalFilters; legacy/observer chart series (perHour,
perDay, types, hops, routes, transports) use totalFilters; region
summary counts all rows
- rollup history will be rebuilt via stats:backfill-rollups (the
backfill tool already counts all rows)
Verified: totals ~101k/24h; charts will sum to the same after backfill.
Per Ben's spec the "Observed packets" totals must count EVERY packet
received by EVERY observer - duplicate receptions of the same packet by
multiple stations each count. Reverts the COUNT(DISTINCT packet_hash)
semantics from 8dff7d7 and additionally drops the privacy visibility
filter from the totals so they match the chart rollup path (which has
never been privacy-filtered - aggregateScope filters by network only).
Changes:
- packetsDay, totalPackets24h, totalPackets7d: COUNT(*) over all rows
(network/test-topic scoped, rx_node_id present), no visibility filter
- channelTraffic: COUNT(*) all rows (percentages stay consistent with
the all-rows total denominator)
- networkFilters(): new includePrivacy option (default true) so total
queries can opt out of visibility conditions
- legacy/observer-scoped chart series back to COUNT(*) (row semantics
everywhere; transportCodes alias fix from 40be809 retained)
- tests: includePrivacy coverage added
Verified live: 24h total ~99k rows (vs 25k privacy-filtered, vs 5.8k
distinct) - consistent with the hourly chart series which already
counted all rows.
The packets table stores one row per observer per packet, so COUNT(*)
overcounted by the average observation multiplicity (~4.3x public, ~9.5x
overall). A packet heard by 5 stations counted 5 times.
Converted to COUNT(DISTINCT packet_hash) on every surface that presents a
"packets" number:
- /api/stats summary packetsDay (app + home dashboard card)
- charts summary totalPackets24h / totalPackets7d (StatsPage cards)
- legacy/observer-scoped chart series: packetsPerHour, packetsPerDay,
packetTypes, hopDistribution, routeTypes, transportCodes
- channelTraffic channel counts (keeps allPct consistent with the
distinct denominator)
The canonical rollup path (packet_hourly_stats, incrementally maintained)
remains row-based - it cannot maintain distinct counts incrementally and
stays as the trend-shape source for the public charts.
Verified: public distinct 24h = 5,835 vs 25,236 public rows vs 99,079
total rows (74% of traffic is from privacy-marked nodes).
df7cfbe migrated listNodeLinks to node_identity_nodes but kept the
un-aliased filters.nodes fragment, whose EXISTS subquery references
nodes.node_id with no nodes table in scope, so every GET /nodes/:id/links
500'd (missing FROM-clause entry for table nodes) and the repeater page
showed no simulated neighbours for any node. Use filters.nodesAlias like
the sibling identity queries do.
The overlay still carried pre-migration 172.18.30.x addresses; the VPS
meshcore-analytics_default bridge is 172.30.0.0/24 (containers already run
on 172.30.0.x). Stale IPs made compose unable to (re)create app-ukmesh
(no subnet contains 172.18.30.10).
- topic.ts/brokerLog.ts/client.ts: accept the neighbours suffix (official
MeshCore firmware publishes the UK spelling; our MQTT fork uses neighbors)
- aclManager.ts: render + verify BOTH spellings in owner ACL grants (ukmesh
and test scopes); inventoryOwnerAuthorization regex updated to match
- statsService.ts: chart bucket labels are now machine-readable ISO instead
of server-timezone HH:MM/en-GB strings (backend container runs UTC;
viewers saw UTC hours)
- StatsPage.tsx: axis ticks + tooltips + summaries formatted via
statsTimeFormat.ts (Intl local timezone, raw fallback for legacy cached
snapshots) + unit tests