mirror of
https://github.com/Kpa-clawbot/meshcore-analyzer.git
synced 2026-09-27 14:28:07 +00:00
Refs #2023, #1659, #1724 ### Problem `main.go:258` waits only for the first load chunk, then `main.go:402` starts the analytics recomputers. `Start()` computes immediately on that chunk (`analytics_recomputer.go:86` on master) and the next compute waits a full interval (`:93`, 5 min default). The chunk loader walks by ascending id, so that chunk holds the oldest transmissions. - RF, topology, channels: the #1659 gate checked `LoadComplete()` after the compute (`analytics_warmup_1659.go:122`). `LoadComplete` flips at the end of the hot window (`chunked_load.go:489`), before the background fill (`store.go:1455`), so the gate could open on a snapshot without the background fill, and the 60 s force timeout (`:73`, `:162`) opened it on the first-chunk snapshot. In an end-to-end test on master, `/api/analytics/rf` returned 200 with 8 of 100 packets before the background fill ran. - Distance, hash-collisions, hash-sizes, roles, observers-clock-skew, nodes-clock-skew: no gate, partial snapshot served from the start. - Distance additionally served a snapshot from the previous index for up to one interval after each lazy index build. On a staging instance, a default analytics request returned 5,911 packets with hours-old last buckets until the next recompute (about 74k). ### Change - `StartupLoadDone()` (`chunked_load.go:108`): closed when `RunStartupLoad` returns, on every path (`chunked_load.go:202`). Closing it drops the hash-size info cache (15 s TTL) and the clock-skew engine throttle (30 s, `clock_skew.go:225`), both read by the post-load computes. - `recomputeWhenLoaded` (`analytics_recomputer.go:172`): on that signal, recompute each recomputer once, sequentially, via `RecomputeNow` (`:154`), which runs on the recomputer's own loop and restarts its ticker (`:106`). Order (`:255`): rf, topology, channels, distance, hash-collisions, hash-sizes, observers-clock-skew, nodes-clock-skew, roles (roles reads the nodes-clock-skew snapshot). Logs one line with per-recomputer durations. - Warm-up gate: now the same signal (`:343-354`), sampled before the compute starts (`:129`), so a pass that began on partial data never opens it. 503 + `Retry-After: 5` and the force timeout are unchanged; a forced-open snapshot is replaced by the post-load recompute. - Ungated endpoints: no new 503s (their API has none); snapshot replaced right after the load. - Distance: the lazy index build refreshes the distance recomputer before reporting built (`store.go:4476`). - Recompute intervals and config unchanged. ### Performance One extra compute per recomputer per process start, run sequentially so they do not all hold the store read lock at once. Ticker phases afterwards are offset by the cumulative post-load compute durations instead of all starting within the first-chunk compute window (relevant to #1724; the effect on lock waves is not measured). ### Tests `analytics_recompute_after_load_test.go`: signal open during background fill, closed after success and failure; cache drops; immediate and ordered post-load recompute; gate not opened by a pass started before the load; forced-open snapshot replaced on load; ticker restart; distance refresh before 202 ends; end to end with recomputers started before the background fill (RF 503 until load, then `totalTransmissions` equals the full store; six ungated endpoints 200 during load; all nine recomputed after load). 9 of these failed on master with stubs; 8 single-line mutations each caught. `go test ./...` in `cmd/server` passes. ### Staging validation Deployed together with the review follow-ups of #2015-#2023 (build `c646310f`), container restart: ``` 16:35:20 [store] first chunk ready (chunkSize=10000) 16:35:25 [store] LoadChunked complete ... starting background fill loader 16:36:58 [store] background load complete: 121120/121282 packets in memory (coverage=99.9%) 16:37:03 [analytics-recompute] startup load done: recomputed 10 snapshots in 5.155s (rf=955ms topology=1.684s channels=43ms distance=49ms hash-collisions=30ms hash-sizes=338ms observers-clock-skew=369ms nodes-clock-skew=692ms roles=2ms retransmissions=994ms) ``` Right after that line, `/api/analytics/rf` reported `totalTransmissions` 121,121 against 120,700 packets in memory, and the retransmissions default shape from #2023 covered the full 7 days. Before this change both waited for the next 5 minute tick. ### Merge order with #2023 #2023 adds a tenth recomputer. Whichever of the two merges second has to add `recompRetransmissions` to `analyticsRecomputersLocked`, wire it to `loadedGate` instead of `LoadComplete`, and change 9 to 10 in `TestAnalyticsRecomputers_PostLoadOrder`; the retransmissions gate test then calls `signalStartupLoadDone()` instead of setting `loadComplete`. That resolution is what ran on staging above. ### Not verified - Repeater-enrich recomputer and the region/window TTL caches (hash-collisions region results have a 1 h TTL) may also keep partial results after the load; not changed here. - Recompute order is tested structurally, not with roles/clock-skew data. - Whether this reduces the #1724 stalls; not measured. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>