mirror of
https://github.com/Kpa-clawbot/meshcore-analyzer.git
synced 2026-10-07 00:37:29 +00:00
Swaps the SQLite driver from `modernc.org/sqlite` (pure Go, SQLite
3.46.0) to `github.com/mattn/go-sqlite3` (cgo, bundled SQLite 3.53.4),
and pays the resulting cross-compilation cost with `zig cc`.
Draft because the riskiest part of this deletes rows — see [Please
review this part first](#please-review-this-part-first) — and because
three things remain unverified at the bottom.
`modernc.org/sqlite` is a transpilation of the C amalgamation. This repo
is read-heavy: `cmd/server` chunk-loads a graph at startup and fans out
neighbour/topology/analytics queries per request, and it pays for that
transpilation on exactly those paths. Head-to-head on the same
120k-transmission / 240k-observation database, running our own hot-path
SQL under both drivers (Apple M4, `-count=5`, medians):
| workload | modernc | mattn | |
|---|---:|---:|---|
| chunk load (`chunked_load.go` v3 join, 20k tx) | 449ms | 196ms |
**2.3×** |
| aggregate scan (240k-row join + `GROUP BY`) | 276ms | 137ms | **2.0×**
|
| 1500 prepared-statement lookups | 512ms | 403ms | **1.3×** |
Allocations fall with it: 1.12M vs 1.64M allocs and 21MB vs 30MB on the
chunk load.
**Superseded by a production run.** @efiten measured both drivers on a
real instance — 11,077,038 observations, 9.7GB database, 4-core arm64 —
as server-only containers against the same live volume, one at a time,
with round 2 reversing the order so the page cache favours the old
driver:
| | audit 7d | audit 24h | background fill (13 chunks) | start →
/api/health |
|---|---:|---:|---:|---:|
| modernc, round 1 | 16.67s | 2.27s | 130.2s | 16.6s |
| mattn, round 1 | 7.87s | 1.34s | 93.8s | 13.5s |
| mattn, round 2 | 8.15s | 1.35s | 96.4s | 13.0s |
| modernc, round 2 | 13.46s | 2.29s | 137.8s | 15.5s |
Warm, the old driver improves to 13.46s on the 7d audit and still loses
by ~1.8×. Chunk load is ~1.4×. `/api/nodes?limit=500` is 0.039s against
0.037s — nothing.
**So the real gain is ~1.4–1.8× on the paths that matter, not 2–2.3×.**
The shape the harness predicted holds — scans and joins gain, small
lookups do not — which is more reassuring than the magnitude would have
been. Quote these numbers.
**The counterweight**, cold and native on that machine: a build goes
from **52s to 163s**. An instance that builds its own image pays that
per deploy.
## The build is cgo now, and one thing about that is a trap
**`CGO_ENABLED=0` still builds.** mattn links a stub, and the binary
dies on its first query with `go-sqlite3 requires cgo to work. This is a
stub`. A green build is not evidence of anything here, which is why
`AGENTS.md` now says so explicitly. `GOOS=linux go build` genuinely
cannot cross-compile any more.
A new root `Makefile` is the entry point. `make crossbuild` uses `zig cc
-target {x86_64,aarch64}-linux-musl` and links static, so each artifact
stays a single self-contained file and the `alpine:3.20` runtime no
longer depends on the base image's libc at all.
`-Wl,-s` is load-bearing: Go's own `-s -w` does not reach the musl
objects zig links in, and without it the server binary is 19.8MB instead
of 12.1MB.
The Dockerfile keeps its single `$BUILDPLATFORM` builder — still no QEMU
for compilation — and gains a checksum-pinned zig plus BuildKit cache
mounts. The mounts are not a nicety: without them an image build
recompiles the amalgamation from cold and takes over half an hour.
## Please review this part first
`internal/dbschema/dedup_index.go` **deletes observation rows**. It is
the one part of this change that can lose data, and it exists because
the migration exposed a real bug rather than causing one.
`stmtInsertObservation` resolves its `ON CONFLICT` against
`idx_observations_dedup`, which `cmd/ingestor/db.go` only ever created
inside the branch that creates the `observations` table for the first
time. Any database whose table predates that branch never got one, so
the UPSERT had no conflict target. modernc failed on the first insert;
mattn fails at `OpenStore`. Same bug, found earlier.
Creating the index unconditionally repairs it — but the index is what
was supposed to prevent duplicates, so a database that never had it can
already hold rows violating it. **`test-fixtures/e2e-fixture.db` in this
repo holds one.** So duplicates are collapsed first. Refusing is not the
safer option: without the index the ingestor cannot prepare its UPSERT,
so it cannot start at all.
Replaying that UPSERT faithfully is subtler than it looks, and a first
cut of this got it wrong twice:
- `COALESCE(excluded.x, x)` means the **incoming** value wins, so down a
group in id order the survivor keeps the **last** non-NULL value. Taking
the first silently discarded newer readings.
- The UPSERT names exactly five columns (`snr`, `rssi`, `score`,
`raw_hex`, `resolved_path`). Every other column must keep the surviving
row's own value; merging those too invents history the ingestor would
never have written.
Merge, delete and `CREATE UNIQUE INDEX` now share one transaction. Split
apart, a writer inserting a duplicate in the gap fails the index
creation while leaving the deletions committed — rows destroyed and no
index to show for it.
Cost, measured on 2.4M synthetic rows holding 5 duplicates: **4.1s**,
holding the write lock throughout, once, at ingestor startup before MQTT
subscribe. Materialising the duplicate-group scan once rather than per
column took that from 9.7s; the pathological case (400k of 600k rows
duplicated) is 5.7s, slightly worse than the 4.2s it was before that
change.
## Four more behavioural differences
Full detail in `docs/sqlite-driver-migration.md`. Briefly:
**Statement preparation is eager.** modernc's `newStmt` stored the SQL
and compiled lazily; mattn calls `sqlite3_prepare_v2` inside `Prepare`,
so SQL naming a missing table fails at *open*. 59 server tests failed on
this alone, all fixtures with partial schemas. `OpenDB` keeps failing
loudly (#1901; `main.go` gates on `dbschema.AssertReady` anyway) and the
fixtures now declare what they are prepared against via
`ensurePreparable`. This also exposed nine `nodes(pubkey …)`
declarations across seven files, where production has only ever had
`public_key` — lazy compilation had hidden the mismatch for as long as
it existed.
**`synchronous` silently dropped FULL → NORMAL.** mattn defaults it to
NORMAL and executes the pragma unconditionally, where SQLite's own
default (what modernc left alone) is FULL. In WAL mode that weakens
durability under power loss. Pinned in `dbschema.WriterDSN`, which both
writers now share — `cmd/migrate` kept a bare path at first and so
quietly wrote at NORMAL, which is what a second copy of a DSN buys you.
**The DSN dialects are mutually invisible.** modernc understood only
`_pragma=name(value)`, mattn only `_`-prefixed parameters, and neither
errors on the other's form — a driver-only rename would have dropped
every pragma in silence. `_journal_mode=WAL` is also gone from the
server's read handle: modernc ignored it, mattn honours it, and setting
`journal_mode` on a read-only connection is a write. Dropping
`_busy_timeout` with it costs nothing, since mattn already defaults to
5000ms — which means the read handle finally *gets* the busy timeout it
had silently lacked.
**`mode=ro` survives for a non-obvious reason.** mattn always passes
`READWRITE|CREATE` and its amalgamation has `SQLITE_USE_URI=0`; what
makes the URI work is its C wrapper ORing `SQLITE_OPEN_URI` in. So the
#1283/#1289 invariant holds with no build flags — but it depends on the
`file:` prefix. `cmd/decrypt` had been building its DSN without one, so
its `mode=ro` had never applied and a missing path was created
read-write. Fixed in passing; never a migration regression.
## What did not change
No modernc-specific API was in use: no `RegisterFunction`, no
`*sqlite.Conn`, no `sqlite/lib` error constants, no `sql.Register`. No
`time.Time` is ever bound as a query argument, so driver time handling
is not in play. Both drivers convert declared
`DATE`/`DATETIME`/`TIMESTAMP` columns to `time.Time`, so
`/api/dropped-packets` keeps emitting `dropped_at` as RFC3339 — an
earlier draft "fixed" that with a `CAST` and would have been the
regression.
## Tests and CI
New regression tests, each written because something got through without
it:
- `TestEnsureObservationsDedupIndexKeepsLatestValues` — the merge
ordering. The original test used complementary NULLs, which passes
whichever direction you pick, which is why the bug survived it.
- `TestCollapseDuplicatesAndIndexIsAtomic` — a failed index creation
must roll the deletions back.
- `TestOpenStorePragmas` / `TestWriterDSNPragmas` — every writer pragma,
read back through the store's own connection. A separate `sqlite3`
session or the startup log line would prove nothing.
- `TestOpenDBRefusesMissingDatabase` — the read-only invariant, which
now rests on a detail of the driver's C wrapper.
- `TestEnsurePreparableMatchesPrepareStatements` — fails when a new
prepared statement outgrows the fixture helper.
CI gains test execution for `cmd/migrate` and `internal/dbschema`, which
had none and both open the database. A PR-time two-arch build plus an
arm64 QEMU smoke gate is new: the GHCR push is push/tag-only, so without
it nothing on a PR would exercise zig, static musl linking or arm64, and
the first signal would arrive on master. `cache-dependency-path` widens
from 2 of the 5 tracked `go.sum` files to all of them.
`make test` passes across all 14 modules, `cmd/server` also under `-race
-count=2` with no failures and no races. `gofmt` and `go vet` clean.
Release-routing and Dockerfile COPY-invariant gates pass.
## Verified by running
- All 8 cross-builds static and correct-architecture; both arches of the
container image built, exported and run under QEMU, serving
`/api/health` and `/api/nodes` against a 2.9M-observation production
snapshot.
- The `migrate` binary repairing that snapshot's duplicate on bare
Alpine.
- `CGO_ENABLED=0` producing a binary that builds and then fails on first
query.
## Not verified
- ~~The 2–2.3× figures come from a standalone harness, not this load
under the old driver.~~ **Closed** by @efiten's production run above,
which also corrected the multiplier.
- SQLite 3.46.0 → 3.53.4 query-planner differences on queries with no
total `ORDER BY`.
- Sustained live ingest through the new writer DSN, and the duplicate
collapse against a database an ingestor is actively writing to. Verified
against a static snapshot only, and the collapse is measured at 4.1s on
2.4M synthetic rows with 5 duplicates — well short of an 11M-row
instance. @efiten has offered a staging instance taking real MQTT
traffic; **this is the item to close before the PR leaves draft.**
An earlier revision of this branch shipped the dedup merge in the wrong
direction with a green test suite, and review then found three more
things in the same file: the repair gated on an error string, a
non-atomic TEMP table drop aimed at the wrong connection, and a deletion
whose only record was a row count. All fixed in ac7e8d38. Passing tests
did not establish safety here, which is why the deletion path wanted a
second pair of eyes rather than a rubber stamp.
616 lines
18 KiB
Go
616 lines
18 KiB
Go
package main
|
||
|
||
import (
|
||
"database/sql"
|
||
"encoding/json"
|
||
"fmt"
|
||
"net/http"
|
||
"net/http/httptest"
|
||
"path/filepath"
|
||
"sync"
|
||
"sync/atomic"
|
||
"testing"
|
||
"time"
|
||
|
||
"github.com/gorilla/mux"
|
||
_ "github.com/mattn/go-sqlite3"
|
||
)
|
||
|
||
// createTestDBMultiDay creates a test DB with packets spread across numDays days.
|
||
// txPerDay transmissions are inserted per day, oldest day first.
|
||
// Packets within each day are spaced 1 minute apart.
|
||
func createTestDBMultiDay(t *testing.T, numDays, txPerDay int) string {
|
||
t.Helper()
|
||
dir := t.TempDir()
|
||
dbPath := filepath.Join(dir, "test.db")
|
||
|
||
conn, err := sql.Open("sqlite3", dbPath+"?_journal_mode=WAL")
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer conn.Close()
|
||
|
||
execOrFail := func(s string) {
|
||
if _, err := conn.Exec(s); err != nil {
|
||
t.Fatalf("createTestDBMultiDay setup: %v", err)
|
||
}
|
||
}
|
||
execOrFail(`CREATE TABLE transmissions (id INTEGER PRIMARY KEY, raw_hex TEXT, hash TEXT, first_seen TEXT, route_type INTEGER, payload_type INTEGER, payload_version INTEGER, decoded_json TEXT)`)
|
||
execOrFail(`CREATE TABLE observations (id INTEGER PRIMARY KEY, transmission_id INTEGER, observer_id TEXT, observer_name TEXT, direction TEXT, snr REAL, rssi REAL, score INTEGER, path_json TEXT, timestamp TEXT, raw_hex TEXT)`)
|
||
execOrFail(`CREATE TABLE observers (rowid INTEGER PRIMARY KEY, id TEXT, name TEXT, iata TEXT, inactive INTEGER)`)
|
||
execOrFail(`CREATE TABLE nodes (public_key TEXT PRIMARY KEY, name TEXT, role TEXT, lat REAL, lon REAL, last_seen TEXT, first_seen TEXT, frequency REAL)`)
|
||
execOrFail(`CREATE TABLE schema_version (version INTEGER)`)
|
||
execOrFail(`INSERT INTO schema_version (version) VALUES (1)`)
|
||
execOrFail(`CREATE INDEX idx_tx_first_seen ON transmissions(first_seen)`)
|
||
|
||
id := 1
|
||
now := time.Now().UTC()
|
||
for day := numDays; day >= 1; day-- {
|
||
// Offset by +30 minutes so day boundaries don't coincide exactly with
|
||
// hotStartupHours/retentionHours cutoffs, preventing timing-boundary flakiness.
|
||
// E.g. for numDays=3: day3 starts at now-71.5h, day2 at now-47.5h, day1 at now-23.5h.
|
||
base := now.Add(-time.Duration(day)*24*time.Hour + 30*time.Minute)
|
||
for i := 0; i < txPerDay; i++ {
|
||
ts := base.Add(time.Duration(i) * time.Minute).Format(time.RFC3339)
|
||
hash := fmt.Sprintf("hash%06d", id)
|
||
if _, err := conn.Exec("INSERT INTO transmissions VALUES (?,?,?,?,0,4,1,?)", id, "aa", hash, ts, `{}`); err != nil {
|
||
t.Fatalf("createTestDBMultiDay insert tx: %v", err)
|
||
}
|
||
if _, err := conn.Exec("INSERT INTO observations VALUES (?,?,?,?,?,?,?,?,?,?,?)", id, id, "obs1", "Obs1", "RX", -10.0, -80.0, 5, `[]`, ts, ""); err != nil {
|
||
t.Fatalf("createTestDBMultiDay insert obs: %v", err)
|
||
}
|
||
id++
|
||
}
|
||
}
|
||
return dbPath
|
||
}
|
||
|
||
// waitForBackgroundLoad polls backgroundLoadDone until true or timeout.
|
||
func waitForBackgroundLoad(t *testing.T, store *PacketStore, timeout time.Duration) {
|
||
t.Helper()
|
||
deadline := time.Now().Add(timeout)
|
||
for time.Now().Before(deadline) {
|
||
if store.backgroundLoadDone.Load() {
|
||
return
|
||
}
|
||
time.Sleep(50 * time.Millisecond)
|
||
}
|
||
t.Fatalf("background load did not complete within %v", timeout)
|
||
}
|
||
|
||
func TestHotStartupConfig_Clamp(t *testing.T) {
|
||
dbPath := createTestDB(t, 10)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
// hotStartupHours > retentionHours → must be clamped
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 24,
|
||
HotStartupHours: 48,
|
||
})
|
||
if store.hotStartupHours != 24 {
|
||
t.Errorf("expected hotStartupHours clamped to retentionHours=24, got %f", store.hotStartupHours)
|
||
}
|
||
}
|
||
|
||
func TestHotStartupConfig_ZeroIsDisabled(t *testing.T) {
|
||
dbPath := createTestDB(t, 10)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 24,
|
||
HotStartupHours: 0,
|
||
})
|
||
if store.hotStartupHours != 0 {
|
||
t.Errorf("expected hotStartupHours=0, got %f", store.hotStartupHours)
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_LoadsOnlyHotWindow(t *testing.T) {
|
||
// 50 old packets (48h ago), 10 recent (30min ago)
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1, // load only last 1 hour
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
// Only the 10 recent packets should be in memory
|
||
if len(store.packets) != 10 {
|
||
t.Errorf("expected 10 recent packets in hot window, got %d", len(store.packets))
|
||
}
|
||
// oldestLoaded should be ~1h ago
|
||
if store.oldestLoaded == "" {
|
||
t.Fatal("oldestLoaded must be set after Load()")
|
||
}
|
||
oldest, _ := time.Parse(time.RFC3339, store.oldestLoaded)
|
||
diff := time.Since(oldest)
|
||
if diff < 30*time.Minute || diff > 90*time.Minute {
|
||
t.Errorf("oldestLoaded %s should be ~1h ago, got diff=%v", store.oldestLoaded, diff)
|
||
}
|
||
// backgroundLoadDone must not be set by Load() itself
|
||
if store.backgroundLoadDone.Load() {
|
||
t.Error("backgroundLoadDone must not be true after Load()")
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_DisabledWhenZero(t *testing.T) {
|
||
// 50 old (48h ago), 10 recent (30min ago) — all within 72h retention
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 0, // disabled → load all retentionHours as before
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
// All 60 packets should be loaded (both old and recent within 72h)
|
||
if len(store.packets) != 60 {
|
||
t.Errorf("expected 60 packets with hotStartupHours=0, got %d", len(store.packets))
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_loadChunk_AddsOlderData(t *testing.T) {
|
||
// 50 old packets (48h ago), 10 recent (30min ago)
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
if len(store.packets) != 10 {
|
||
t.Fatalf("setup: expected 10 packets after hot Load, got %d", len(store.packets))
|
||
}
|
||
|
||
// Load the old chunk (covers the 50 old packets at ~48h ago)
|
||
chunkEnd := time.Now().UTC().Add(-1 * time.Hour)
|
||
chunkStart := time.Now().UTC().Add(-72 * time.Hour)
|
||
if err := store.loadChunk(chunkStart, chunkEnd); err != nil {
|
||
t.Fatalf("loadChunk failed: %v", err)
|
||
}
|
||
|
||
// Should have 10 recent + 50 old
|
||
if len(store.packets) != 60 {
|
||
t.Errorf("expected 60 packets after loadChunk, got %d", len(store.packets))
|
||
}
|
||
// Packets must remain sorted ASC by first_seen
|
||
for i := 1; i < len(store.packets); i++ {
|
||
if store.packets[i].FirstSeen < store.packets[i-1].FirstSeen {
|
||
t.Fatalf("packets not in ASC order at index %d: %s < %s",
|
||
i, store.packets[i].FirstSeen, store.packets[i-1].FirstSeen)
|
||
}
|
||
}
|
||
// byHash must include the old packets
|
||
if len(store.byHash) != 60 {
|
||
t.Errorf("expected byHash len=60, got %d", len(store.byHash))
|
||
}
|
||
// byObserver must reflect all 60 observations for obs1
|
||
if len(store.byObserver["obs1"]) != 60 {
|
||
t.Errorf("expected byObserver[obs1] len=60, got %d", len(store.byObserver["obs1"]))
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_BackgroundFillsToRetention(t *testing.T) {
|
||
// 3 days × 50 tx/day = 150 total
|
||
dbPath := createTestDBMultiDay(t, 3, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 24,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
// After hot Load: only ~50 packets (day 1 = last 24h)
|
||
afterHot := len(store.packets)
|
||
if afterHot < 1 || afterHot > 60 {
|
||
t.Errorf("expected ~50 packets after hot Load, got %d", afterHot)
|
||
}
|
||
|
||
// Start background fill
|
||
go store.loadBackgroundChunks()
|
||
waitForBackgroundLoad(t, store, 15*time.Second)
|
||
|
||
// After background fill: all 150 packets should be loaded
|
||
store.mu.RLock()
|
||
total := len(store.packets)
|
||
store.mu.RUnlock()
|
||
|
||
if total != 150 {
|
||
t.Errorf("expected 150 packets after background load, got %d", total)
|
||
}
|
||
if !store.backgroundLoadDone.Load() {
|
||
t.Error("backgroundLoadDone must be true after loadBackgroundChunks returns")
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_ChunkErrorRecovery(t *testing.T) {
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
// intentional: closed early to simulate chunk-load failures; no defer
|
||
db.conn.Close()
|
||
|
||
done := make(chan struct{})
|
||
go func() {
|
||
store.loadBackgroundChunks()
|
||
close(done)
|
||
}()
|
||
|
||
select {
|
||
case <-done:
|
||
// Good — completed without hanging.
|
||
case <-time.After(10 * time.Second):
|
||
t.Fatal("loadBackgroundChunks hung after DB close")
|
||
}
|
||
|
||
// #1690: backgroundLoadFailed must be true (chunk errors AND coverage
|
||
// fell short); backgroundLoadDone stays false because the in-memory
|
||
// store does NOT reflect the on-disk DB. Pre-#1690 the test asserted
|
||
// Done=true on errors — that was the very lie the issue documents.
|
||
if !store.backgroundLoadFailed.Load() {
|
||
t.Error("backgroundLoadFailed must be true after all chunks fail (#1690)")
|
||
}
|
||
if store.backgroundLoadDone.Load() {
|
||
t.Error("backgroundLoadDone must remain false when the store does not reflect the DB (#1690)")
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_SQLFallback_TriggeredForOldDate(t *testing.T) {
|
||
// 50 old packets (48h ago), 10 recent (30min ago)
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
// Hot load: only last 1h → 10 recent packets in memory
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
if len(store.packets) != 10 {
|
||
t.Fatalf("setup: expected 10 in-memory packets, got %d", len(store.packets))
|
||
}
|
||
|
||
// Query with Since = 49h ago (before oldestLoaded ~1h ago) → SQL fallback
|
||
since49h := time.Now().UTC().Add(-49 * time.Hour).Format(time.RFC3339)
|
||
result := store.QueryPackets(PacketQuery{Since: since49h, Limit: 100, Order: "ASC"})
|
||
|
||
// SQL fallback returns all packets newer than Since: 50 old (48h ago) + 10 recent (30min ago) = 60
|
||
if result.Total != 60 {
|
||
t.Errorf("expected SQL fallback to return 60 packets for Since=49h ago, got %d", result.Total)
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_PerfStats(t *testing.T) {
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
stats := store.GetPerfStoreStats()
|
||
|
||
if v, ok := stats["hotStartupHours"]; !ok || v.(float64) != 1 {
|
||
t.Errorf("expected hotStartupHours=1 in stats, got %v", v)
|
||
}
|
||
if v, ok := stats["backgroundLoadComplete"]; !ok || v.(bool) != false {
|
||
t.Errorf("expected backgroundLoadComplete=false in stats, got %v", v)
|
||
}
|
||
if _, ok := stats["backgroundLoadProgress"]; !ok {
|
||
t.Error("expected backgroundLoadProgress in stats")
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_SQLFallback_NotTriggeredForRecentDate(t *testing.T) {
|
||
// 50 old packets (48h ago), 10 recent (30min ago)
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
// Hot load: last 1h → 10 recent packets in memory
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
// Query with Since = 45min ago (after oldestLoaded ~1h ago) → in-memory path
|
||
since45m := time.Now().UTC().Add(-45 * time.Minute).Format(time.RFC3339)
|
||
result := store.QueryPackets(PacketQuery{Since: since45m, Limit: 100, Order: "ASC"})
|
||
|
||
// In-memory path: returns only the 10 recent packets (all within last 30min)
|
||
if result.Total != 10 {
|
||
t.Errorf("expected 10 in-memory packets for recent Since query, got %d", result.Total)
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_SQLFallback_Until(t *testing.T) {
|
||
// 50 old packets (48h ago), 10 recent (30min ago)
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
// Hot load: only last 1h → 10 recent in memory, oldestLoaded ~1h ago
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
if len(store.packets) != 10 {
|
||
t.Fatalf("setup: expected 10 in-memory packets, got %d", len(store.packets))
|
||
}
|
||
|
||
// Until = 2h ago (before oldestLoaded ~1h ago) → SQL fallback
|
||
until2h := time.Now().UTC().Add(-2 * time.Hour).Format(time.RFC3339)
|
||
result := store.QueryPackets(PacketQuery{Until: until2h, Limit: 100, Order: "ASC"})
|
||
|
||
// SQL fallback returns the 50 old packets (stored at ~48h ago, all before Until)
|
||
if result.Total != 50 {
|
||
t.Errorf("expected SQL fallback to return 50 old packets for Until before oldestLoaded, got %d", result.Total)
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_PerfStoreHTTP(t *testing.T) {
|
||
dbPath := createTestDBWithAgedPackets(t, 10, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 1,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
srv := NewServer(db, &Config{Port: 3000}, NewHub())
|
||
srv.store = store
|
||
router := mux.NewRouter()
|
||
srv.RegisterRoutes(router)
|
||
|
||
req := httptest.NewRequest("GET", "/api/perf", nil)
|
||
w := httptest.NewRecorder()
|
||
router.ServeHTTP(w, req)
|
||
|
||
if w.Code != http.StatusOK {
|
||
t.Fatalf("expected 200, got %d", w.Code)
|
||
}
|
||
var body map[string]interface{}
|
||
if err := json.Unmarshal(w.Body.Bytes(), &body); err != nil {
|
||
t.Fatalf("invalid JSON: %v", err)
|
||
}
|
||
ps, ok := body["packetStore"].(map[string]interface{})
|
||
if !ok {
|
||
t.Fatalf("missing packetStore in /api/perf response")
|
||
}
|
||
for _, field := range []string{"hotStartupHours", "backgroundLoadComplete", "backgroundLoadProgress"} {
|
||
if _, ok := ps[field]; !ok {
|
||
t.Errorf("missing field %q in packetStore", field)
|
||
}
|
||
}
|
||
if v, ok := ps["hotStartupHours"].(float64); !ok || v != 1 {
|
||
t.Errorf("expected hotStartupHours=1, got %v", ps["hotStartupHours"])
|
||
}
|
||
}
|
||
|
||
func TestHotStartup_ConcurrentQueryDuringBackgroundLoad(t *testing.T) {
|
||
// 5 days × 200 tx/day = 1000 total — small enough to run in CI fast,
|
||
// large enough to give pollers >=1 query during the background fill.
|
||
dbPath := createTestDBMultiDay(t, 5, 200)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
defer db.conn.Close()
|
||
|
||
// Hot load: only last 24h → ~200 packets in memory
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 120,
|
||
HotStartupHours: 24,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
preLen := len(store.packets)
|
||
|
||
// Real invariant (Munger r2 #5): while background fill is running,
|
||
// the result set for a fixed [since, until] window must be monotonic
|
||
// in TIME — rows only appear, never disappear. The query window must
|
||
// straddle the moving oldestLoaded boundary so we exercise both the
|
||
// SQL fallback (since < oldestLoaded) and the in-memory path
|
||
// (oldestLoaded shrinks below since as chunks merge).
|
||
//
|
||
// since=200h ago covers everything; as oldestLoaded retreats from
|
||
// 24h ago to 120h ago, the answer source switches from SQL fallback
|
||
// to in-memory; Total must never decrease across that switch.
|
||
since := time.Now().UTC().Add(-200 * time.Hour).Format(time.RFC3339)
|
||
q := PacketQuery{Since: since, Limit: 5000, Order: "ASC"}
|
||
|
||
// Start background fill.
|
||
go store.loadBackgroundChunks()
|
||
|
||
// Pollers: each goroutine keeps querying until the loader is done,
|
||
// asserting that within its own series Total only grows or stays equal.
|
||
// A shrink — even by one row — is a real-invariant violation that
|
||
// the trivial Total>=0 / postLen>=preLen tests could not catch.
|
||
var wg sync.WaitGroup
|
||
pollers := 8
|
||
totalSamples := atomicSamples{}
|
||
for i := 0; i < pollers; i++ {
|
||
wg.Add(1)
|
||
go func(i int) {
|
||
defer wg.Done()
|
||
lastTotal := -1
|
||
for !store.backgroundLoadDone.Load() {
|
||
r := store.QueryPackets(q)
|
||
if r == nil {
|
||
continue
|
||
}
|
||
if lastTotal >= 0 && r.Total < lastTotal {
|
||
t.Errorf("poller %d: result set shrank (%d → %d) — non-monotonic across moving oldestLoaded boundary",
|
||
i, lastTotal, r.Total)
|
||
}
|
||
lastTotal = r.Total
|
||
totalSamples.inc()
|
||
}
|
||
r := store.QueryPackets(q)
|
||
if r != nil {
|
||
if lastTotal >= 0 && r.Total < lastTotal {
|
||
t.Errorf("poller %d: post-load result set shrank (%d → %d)", i, lastTotal, r.Total)
|
||
}
|
||
totalSamples.inc()
|
||
}
|
||
}(i)
|
||
}
|
||
wg.Wait()
|
||
|
||
waitForBackgroundLoad(t, store, 60*time.Second)
|
||
|
||
store.mu.RLock()
|
||
postLen := len(store.packets)
|
||
store.mu.RUnlock()
|
||
|
||
if postLen < preLen {
|
||
t.Errorf("expected packet count after background load (%d) >= pre-background (%d)", postLen, preLen)
|
||
}
|
||
if totalSamples.get() == 0 {
|
||
t.Error("pollers observed zero samples — test did not actually exercise the invariant")
|
||
}
|
||
}
|
||
|
||
type atomicSamples struct {
|
||
n int64
|
||
}
|
||
|
||
func (a *atomicSamples) inc() { atomic.AddInt64(&a.n, 1) }
|
||
func (a *atomicSamples) get() int64 {
|
||
return atomic.LoadInt64(&a.n)
|
||
}
|
||
|
||
// TestHotStartup_BackgroundLoadFailureSurfacesInPerf asserts that when every
|
||
// background chunk errors, the store does NOT report backgroundLoadComplete=true
|
||
// — instead it surfaces backgroundLoadFailed=true via GetPerfStoreStats so
|
||
// operators see a visible failure rather than silent data loss. Munger r2 #3.
|
||
func TestHotStartup_BackgroundLoadFailureSurfacesInPerf(t *testing.T) {
|
||
dbPath := createTestDBMultiDay(t, 3, 50)
|
||
|
||
db, err := OpenDB(dbPath)
|
||
if err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
store := NewPacketStore(db, &PacketStoreConfig{
|
||
RetentionHours: 72,
|
||
HotStartupHours: 24,
|
||
})
|
||
if err := store.Load(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
// Force every loadChunk call to fail by closing the read connection.
|
||
// loadBackgroundChunks must then NOT report "complete" — it must report failed.
|
||
if err := db.conn.Close(); err != nil {
|
||
t.Fatal(err)
|
||
}
|
||
|
||
store.loadBackgroundChunks()
|
||
|
||
perf := store.GetPerfStoreStats()
|
||
failed, hasFailedKey := perf["backgroundLoadFailed"].(bool)
|
||
|
||
if !hasFailedKey {
|
||
t.Fatalf("expected backgroundLoadFailed key in /api/perf payload, got keys: %v", perf)
|
||
}
|
||
if !failed {
|
||
t.Errorf("expected backgroundLoadFailed=true after every chunk errored, got false (observability lying)")
|
||
}
|
||
}
|