Follow-up to #2072, which closed #2058 but left one gap named in its own description: the refresh ticker waits 2 minutes before its first run, and a query arriving in that window against a database with no statistics gets the bad plan. Deployed to staging to measure it rather than reason about it, with `sqlite_stat1` dropped first so the build path actually ran. That changed two of the numbers in #2072, both in the expensive direction. ## The gap is once per database, not once per restart `sqlite_stat1` is an ordinary table, so once `ANALYZE` has written it the statistics stay in the file. Checked four ways: - they survive closing the connection that wrote them - a `mode=ro` handle reads them back, which is how `cmd/server` opens the database - reopening the same path through a second `OpenStore` finds them and skips the rebuild (`TestPlannerStatsSurviveReopen_Issue2058`) - on staging they survived a full redeploy to a different build that has no refresh ticker at all, and that build still gets the good plan So the window opens once, on the first start after this lands, and never again for that database. ## The cost, corrected #2072 said 2.0s. Observed on staging, 9.4 GB, commit `4500cfa6`: ``` 13:51:26 [analyze] planner statistics refresh scheduled every 24h (analysis_limit=10000) 13:55:10 [analyze] planner statistics built in 3m43.874s (analysis_limit=10000, first run against this database) ``` **3m43.9s.** Every `ANALYZE` duration in #2072's ladder was timed warm, run after run; cold, on a freshly started container, the same statement takes nearly four minutes. That is the same warm/cold split #2058 work already established for the query itself, 56.7s against 0.80s, and I then repeated it for the `ANALYZE`. Every `2.0s` in the tree is now marked warm and points at the cold figure. It holds the single write connection throughout, so ingest stalls and buffers. Per minute in `observations`: | minute | rows | |---|---| | 13:49 | 220 | | 13:50 | 106 | | 13:51 | 0 | | 13:52 | 0 | | 13:53 | 0 | | 13:54 | 0 | | 13:55 | **1027** | | 13:56 | 154 | Nothing was dropped. The burst is about four minutes of traffic at the surrounding rate, and the only ingest-buffer line in the log is the startup one reporting `0 dropped`. The cost is a four-minute write stall, once, not data loss. ## This cost is not introduced here The ticker merged in #2072 pays the identical 3m43.9s two minutes later on any database with no statistics. **Live has none, so #2072 as merged will stall live ingest for about four minutes on its first run, with or without this branch.** This only moves it earlier, into the startup burst the ingest buffer is already sized for. Flagging it on #2072 as well. ## The change `Store.EnsurePlannerStats(analysisLimit)` checks before it builds: - database has statistics: one `sqlite_master` query. This is every restart after the first. - database has none: one `ANALYZE`, and a warning first. The warning is the part that earns its place operationally. Four minutes of stalled ingest with no explanation in the log looks exactly like a hang, so `EnsurePlannerStats` now says why the write path is about to pause, what it measured on 9.4 GB, and that it happens once per database. It stays silent on a restart, because a warning on every boot would be worse than none. It runs on the refresh goroutine, not the startup path, so no boot step waits for it. `hasPlannerStats` now gates a decision instead of only wording a log line, so its comment says what the swallowed error costs: a query failure reads as "no stats", which spends one unnecessary `ANALYZE` rather than skipping a necessary one. ## Verification on staging - before: plan drove from `idx_transmissions_payload_type`, no `sqlite_stat1` - after: 50 rows in `sqlite_stat1`, plan drives from `idx_tx_channel_hash` - dropping the table first flipped the plan back, so the causality holds in both directions ## Tests 13 in the file. New here: builds when absent, skips when present, disabled on a negative limit, survives close-and-reopen, warns before building, stays quiet when statistics exist. The reopen test is the guard on the whole design: if statistics ever stopped living in the file, `EnsurePlannerStats` would quietly run a four-minute `ANALYZE` on every restart and nothing else would notice. Run locally: 13/13 on the `Issue2058` tests, `go vet` clean, `gofmt` clean, and the rest of `cmd/ingestor` green apart from `TestWriteStatsAtomic_SymlinkAtDestIsReplaced`, which fails on `os.Symlink` with "A required privilege is not held by the client" on Windows without elevation, in a file this branch does not touch. ## Not done - No query rewrite, same as #2058 and #2072. - **No live deploy.** Live still has no statistics, so the four-minute stall is ahead of it whenever #2072 ships there. Worth picking the moment. - Staging has been returned to its own fork build; the statistics it built remain, so its next start exercises the skip path rather than the build path. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_013YAR8fdNTzqjtsggq4xCX6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
MeshCore MQTT Ingestor (Go)
Standalone MQTT ingestion service for CoreScope. Connects to MQTT brokers, decodes raw MeshCore packets, and writes to the same SQLite database used by the Node.js web server.
This is the first step of a larger Go rewrite — separating MQTT ingestion from the web server.
Architecture
MQTT Broker(s) → Go Ingestor → SQLite DB ← Node.js Web Server
(this binary) (shared)
- Single static binary — no runtime dependencies, no CGO
- SQLite via
github.com/mattn/go-sqlite3(cgo; cross-compiled withzig cc, see the rootMakefile) - MQTT via
github.com/eclipse/paho.mqtt.golang - Runs alongside the Node.js server — they share the DB file
- Does NOT serve HTTP/WebSocket — that stays in Node.js
Build
Requires Go 1.22+.
cd cmd/ingestor
go build -o corescope-ingestor .
Cross-compile for Linux (e.g., for the production VM):
GOOS=linux GOARCH=amd64 go build -o corescope-ingestor .
Run
./corescope-ingestor -config /path/to/config.json
The config file uses the same format as the Node.js config.json. The ingestor reads the mqttSources array (or legacy mqtt object) and dbPath fields.
Environment Variables
| Variable | Description | Default |
|---|---|---|
DB_PATH |
SQLite database path | data/meshcore.db |
MQTT_BROKER |
Single MQTT broker URL (overrides config) | — |
MQTT_TOPIC |
MQTT topic (used with MQTT_BROKER) |
meshcore/# |
CORESCOPE_INGESTOR_STATS |
Path to the per-second stats JSON file consumed by the server's /api/perf/io and /api/perf/write-sources endpoints (#1120) |
/tmp/corescope-ingestor-stats.json |
Stats file (CORESCOPE_INGESTOR_STATS)
Every second the ingestor publishes a JSON snapshot of its counters
(tx_inserted, obs_inserted, walCommits, backfillUpdates.*, etc.) plus
a procIO block sampled from /proc/self/io (read/write/cancelled bytes per
second + syscall counts). The server reads this file and surfaces the data on
the Perf page so operators can self-diagnose write-volume anomalies.
The writer uses O_NOFOLLOW | O_CREAT | O_TRUNC mode 0o600, so a
pre-planted symlink at the path cannot be used to clobber an arbitrary file.
Security note: the default lives in /tmp, which is world-writable on
most hosts (sticky bit only protects deletion, not creation). On
shared/multi-tenant hosts, override CORESCOPE_INGESTOR_STATS to point at a
private directory (e.g. /var/lib/corescope/ingestor-stats.json) that only
the corescope user can write to.
Minimal Config
{
"dbPath": "data/meshcore.db",
"mqttSources": [
{
"name": "local",
"broker": "mqtt://localhost:1883",
"topics": ["meshcore/#"]
}
]
}
Full Config (same as Node.js)
The ingestor reads these fields from the existing config.json:
mqttSources[]— array of MQTT broker connectionsname— display name for loggingbroker— MQTT URL (mqtt://,mqtts://)username/password— auth credentialstopics— array of topic patterns to subscribeiataFilter— optional regional filterclientId: optional MQTT ClientID. Defaultcorescope-<name>-<random>, new suffix per ingestor start. Two running ingestors must not share a value on one broker
mqtt— legacy single-broker config (auto-converted tomqttSources)dbPath— SQLite DB path (default:data/meshcore.db)
Test
cd cmd/ingestor
go test -v ./...
What It Does
- Connects to configured MQTT brokers with auto-reconnect
- Subscribes to mesh packet topics (e.g.,
meshcore/+/+/packets) - Receives raw hex packets via JSON messages (
{ "raw": "...", "SNR": ..., "RSSI": ... }) - Decodes MeshCore packet headers, paths, and payloads (ported from
decoder.js) - Computes content hashes (path-independent, SHA-256-based)
- Writes to SQLite:
transmissions+observationstables - Upserts
nodesfrom decoded ADVERT packets (with validation) - Upserts
observersfrom MQTT topic metadata
Schema Compatibility
The Go ingestor creates the same v3 schema as the Node.js server:
transmissions— deduplicated by content hashobservations— per-observer sightings withobserver_idx(rowid reference)nodes— mesh nodes discovered from advertsobservers— MQTT feed sources
Both processes can write to the same DB concurrently (SQLite WAL mode).
What's Not Ported (Yet)
- Companion bridge format (Format 2 —
meshcore/advertisement, channel messages, etc.) - Channel key decryption (GRP_TXT encrypted payload decryption)
- WebSocket broadcast to browsers
- In-memory packet store
- Cache invalidation
These stay in the Node.js server for now.
Files
cmd/ingestor/
main.go — entry point, MQTT connect, message handler
decoder.go — MeshCore packet decoder (ported from decoder.js)
decoder_test.go — decoder tests (25 tests, golden fixtures)
db.go — SQLite writer (schema-compatible with db.js)
db_test.go — DB tests (schema validation, insert/upsert, E2E)
config.go — config struct + loader
util.go — shared utilities
go.mod / go.sum — Go module definition