mirror of
https://github.com/Kpa-clawbot/meshcore-analyzer.git
synced 2026-09-25 22:23:50 +00:00
## Summary Adds `GET /api/analytics/retransmissions` and a "Retransmission Pressure (proxy)" chart on the Analytics Topology tab, implementing the metric agreed in #1699: for each flood, the number of distinct repeaters in the union of the paths of all its observations (`[A]`, `[A,B,C]`, `[A,D]` gives 4), averaged per time bucket. Topology is the tab that already shows hop counts and repeaters in paths, so the chart sits there instead of in a new tab. ## Definition - Flood routes only (`route_type` 0/1); TRACE excluded. Direct routes carry the route still to travel (firmware `src/Mesh.cpp:78-106,334-342`), zero-hop sends are direct (`src/Mesh.cpp:717-737`), TRACE path bytes are SNR values (`src/Mesh.cpp:59-61`, refused by `sendFlood` at `src/Mesh.cpp:637-641`). Firmware commit 0679dbef. - **Flood events, not hashes.** `transmissions.hash` is UNIQUE and the packet hash excludes the path (`src/Packet.cpp:41-50`), so when the same bytes flood again the observations land on the same transmission. Observations are sorted by time and split into events wherever two consecutive observations are more than 5 minutes apart. Each event is counted on its own and bucketed by its first observation. - Why 5 minutes: a node holds a flood for at most 32 s (`src/Dispatcher.cpp:11,243-251`) plus a random retransmit delay. On live over 7 days, 72,806 of 74,347 flood transmissions span 60 s or less, and of 1,372,283 consecutive observation gaps, 52 fall between 60 s and 300 s against 1,823 above 300 s. - Events that start before the store retention floor (now minus `retentionHours`) are left out for every request shape. The store keeps older observations only for hashes heard again recently, so they do not represent that period. Eviction of those transmissions is tracked in #2024. - A flood event heard only with an empty path counts as 0 repeaters. - **Prefixes are not resolved to nodes, and a prefix counts once per event**, whether it repeats across observations or inside one path. On live (7 days), a repeated 2-byte prefix inside one path occurs in 1.08% of flood transmissions and 6,178 of 6,596 such repeats match exactly one known node; for 3-byte it is 0.69% and 104 of 104. That is one node forwarding again after its 160-slot cyclic duplicate filter dropped the hash (`src/helpers/SimpleMeshTables.h:9,52-57`). A repeated 1-byte prefix (44.9% of 1-byte transmissions) is mostly two nodes; counting it once keeps the value a lower bound. `summary.one_byte_packets` reports how many events that affects. - Observations are stored once per observer and path per hash, so a later event of the same hash only holds pairs not stored before; its count is a lower bound too. On live these are 1,466 of 75,356 events (1.9%), and they are kept in the average. - Resolution was not used: on live, 1-byte observations nearly all have `resolved_path` NULL, and cold load refuses context-based resolution of history (`cmd/server/neighbor_persist.go:155-168`). - Buckets `5m|15m|1h|6h|1d`. `region` filters on observers like `/api/analytics/rf`, after the event split; a region with no known observers is not filtered, the same as the other analytics endpoints. `area` is not supported. ## Implementation - `cmd/server/retransmission_pressure.go:255` `addPath`: scans path JSON directly into a generation-stamped hash set, no allocation per observation. - `cmd/server/retransmission_pressure.go:367` `computeRetransmissionPressure`: one pass under `s.mu.RLock`. Per flood transmission it sorts the observations by cached parsed time into a reused scratch slice, splits events and counts each in `addEvent` (`:319`). O(T + O log k + H). - `cmd/server/retransmission_pressure.go:470` `GetRetransmissionPressure`: default shape from the recomputer (#1659 warm-up gate). Other shapes come from a typed TTL cache (max 64 entries) cleared on new paths and eviction (`cmd/server/store.go:2289,2336`); concurrent misses on one key share one compute through singleflight (`store.go:204`). - `cmd/server/retransmission_pressure.go:515` handler, `cmd/server/routes.go:331`, `cmd/server/openapi.go:108`, `docs/api-spec.md:1283`. - `public/analytics.js:761` card, `:855` `renderRetransmissionChart` (CSS variables only, lines break at missing buckets, caption states it is a proxy, names the observer coverage bias, the once-per-flood prefix rule and the 5 minute event split), `:829` loader with stale-response guard. ## Performance - `BenchmarkComputeRetransmissionPressure`, 50k transmissions x 20 observations, `-cpu 1`, i5-1335U: median 161 ms/op (132 ms/op before the event split); first pass after startup with timestamps not yet parsed 192 ms/op. About 22 KB and 281 allocations per op. - On staging the default shape is served from the recomputer in 0.3 s; the post-load recompute of this recomputer took 994 ms on a 121k-transmission store (log line quoted in #2025). A 336h store would be about twice that, every recompute interval, under the store read lock. ## Tests - `cmd/server/retransmission_pressure_test.go`: union counting (reporter example, overlaps, once per event for 1/2/3-byte, width/case, growth); route/TRACE/zero-hop filter, bucketing, window by event start; event split and the 5 minute settle gap (boundary, chained steps, unsorted input), retention floor; region filter, region applied after the split, unknown region, 1-byte share; recomputer read, TTL cache invalidation on new paths and on eviction, cache expiry, singleflight, recomputer gate wiring, handler, warm-up gate. - `test-issue-1699-retransmission-chart.js` (33 tests, registered in `test-all.sh` and `deploy.yml`). - Mutation-checked: 18 mutations of the event split, floor, prefix rule, bucketing, region order, cache clears, expiry, gate wiring and singleflight, all killed. - `go test ./...` in cmd/server passes; `scripts/check-css-vars.js` OK. ## Staging validation Build `c646310f` (this rework plus #2025 and the other review follow-ups), after a container restart and full load: default shape 74,974 flood events, average 27.08 repeaters, 169 hourly buckets from 2026-09-06 16:00 (the 168h floor) to the current hour, highest hourly average 53.3. Before the rework the same instance showed buckets back to 2026-07-18, averages up to 148, and for the first minutes after a restart only 5,911 packets. ## Merge order with #2025 #2025 fixes the recomputer startup for all analytics endpoints (the stale first snapshot seen here). Whichever of the two merges second has to add `recompRetransmissions` to `analyticsRecomputersLocked`, wire it to that PR's `loadedGate` instead of `LoadComplete`, bump the recomputer count in `TestAnalyticsRecomputers_PostLoadOrder` from 9 to 10, and make `TestStartAnalyticsRecomputers_RetransmissionsGatedOnLoadComplete` call `signalStartupLoadDone()`. That resolution is what ran on staging. ## Not verified - Recompute timing on a production-size (336h) store; only extrapolated. - Phone-width layout and dark theme of the reworked chart. - E2E Playwright suite. Fixes #1699 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
389 lines
13 KiB
Go
389 lines
13 KiB
Go
// Package main: analytics recomputer (issue #1240).
|
||
//
|
||
// Steady-state background recompute loop for expensive analytics
|
||
// endpoints. Reads always hit an atomic-pointer cache; compute runs
|
||
// on a fixed ticker in a goroutine. This eliminates the on-request
|
||
// compute-then-cache pattern where the first reader after expiry pays
|
||
// the full compute cost and blocks under writer contention.
|
||
//
|
||
// See issue #1240 and AGENTS.md "Performance is a feature".
|
||
package main
|
||
|
||
import (
|
||
"fmt"
|
||
"log"
|
||
"strings"
|
||
"sync"
|
||
"sync/atomic"
|
||
"time"
|
||
)
|
||
|
||
// analyticsRecomputer holds the latest snapshot of an analytics result
|
||
// in an atomic.Value, refreshed periodically by a background goroutine.
|
||
//
|
||
// Lifecycle:
|
||
// 1. Construct via newAnalyticsRecomputer(...)
|
||
// 2. Call Start() — runs initial compute synchronously, then launches
|
||
// the recompute goroutine. Initial compute is synchronous so the
|
||
// first Load() after Start returns never sees a nil cache.
|
||
// 3. Call Load() any number of times concurrently — never blocks
|
||
// beyond an atomic-pointer load.
|
||
// 4. Call Stop() to terminate the background goroutine cleanly.
|
||
//
|
||
// Compute func is called WITHOUT any lock held by this struct, so it
|
||
// may freely take any application-level locks it needs.
|
||
type analyticsRecomputer struct {
|
||
name string
|
||
interval time.Duration
|
||
compute func() interface{}
|
||
|
||
cache atomic.Value // holds interface{} — the latest snapshot
|
||
stop chan struct{}
|
||
done chan struct{}
|
||
recomputeReq chan chan struct{} // RecomputeNow → loop; the loop closes the inner channel when done
|
||
|
||
startOnce sync.Once
|
||
stopOnce sync.Once
|
||
|
||
// Stats (atomic).
|
||
computeRuns atomic.Int64
|
||
lastComputeNs atomic.Int64 // duration of last compute in nanoseconds
|
||
|
||
// Issue #1659 (PR #1688 r1) — warmup gate state, inlined here so
|
||
// hot-path readers (IsWarmingUp_1659) do lock-free atomic loads
|
||
// only (replaces the r0 package-level map + chanLock). See
|
||
// analytics_warmup_1659.go for full design notes.
|
||
firstPassDoneNs atomic.Int64
|
||
warmupStartedNs atomic.Int64
|
||
warmupReadyGate atomic.Value // *func() bool — gate must return true for markFirstPassDone to take effect
|
||
}
|
||
|
||
// newAnalyticsRecomputer constructs an unstarted recomputer.
|
||
// interval must be > 0; compute must be non-nil.
|
||
func newAnalyticsRecomputer(name string, interval time.Duration, compute func() interface{}) *analyticsRecomputer {
|
||
if interval <= 0 {
|
||
interval = 5 * time.Minute
|
||
}
|
||
return &analyticsRecomputer{
|
||
name: name,
|
||
interval: interval,
|
||
compute: compute,
|
||
stop: make(chan struct{}),
|
||
done: make(chan struct{}),
|
||
recomputeReq: make(chan chan struct{}),
|
||
}
|
||
}
|
||
|
||
// Start runs the initial compute synchronously (so the first Load
|
||
// after Start returns a populated snapshot, never nil), then launches
|
||
// a background goroutine to periodically recompute.
|
||
//
|
||
// Calling Start multiple times is a no-op after the first call.
|
||
func (r *analyticsRecomputer) Start() {
|
||
r.startOnce.Do(func() {
|
||
// Issue #1659 (#1688 munger #2): record warmup-start before
|
||
// the first compute, so IsWarmingUp_1659's fallback timeout
|
||
// is measured from "recomputer started" — not "first pass
|
||
// returned", which never happens if compute() hangs.
|
||
r.noteWarmupStart_1659()
|
||
// Initial synchronous compute — first read must NOT see empty
|
||
// or uninitialized data (acceptance criterion #1240).
|
||
r.runOnce()
|
||
go r.loop()
|
||
})
|
||
}
|
||
|
||
func (r *analyticsRecomputer) loop() {
|
||
defer close(r.done)
|
||
t := time.NewTicker(r.interval)
|
||
defer t.Stop()
|
||
for {
|
||
select {
|
||
case <-t.C:
|
||
r.runOnce()
|
||
case ack := <-r.recomputeReq:
|
||
r.runOnce()
|
||
t.Reset(r.interval)
|
||
close(ack)
|
||
case <-r.stop:
|
||
return
|
||
}
|
||
}
|
||
}
|
||
|
||
func (r *analyticsRecomputer) runOnce() {
|
||
if r.compute == nil {
|
||
return
|
||
}
|
||
defer func() {
|
||
// Don't let a compute panic kill the background goroutine.
|
||
// The previous snapshot remains valid. Even on panic, we
|
||
// still want IsWarmingUp_1659's fallback timeout to be the
|
||
// safety net (a perpetually panicking compute would never
|
||
// reach markFirstPassDone otherwise).
|
||
_ = recover()
|
||
}()
|
||
// Sample the #1659 readiness gate BEFORE computing: a pass that
|
||
// started on a partially loaded store must not end the warm-up,
|
||
// even if the load finishes while it runs.
|
||
ready := r.warmupReadyGateOpen_1659()
|
||
t0 := time.Now()
|
||
result := r.compute()
|
||
r.lastComputeNs.Store(int64(time.Since(t0)))
|
||
r.computeRuns.Add(1)
|
||
if result != nil {
|
||
r.cache.Store(result)
|
||
}
|
||
// Issue #1659: mark the first-pass clock so the warmup gate
|
||
// in GetAnalyticsRFWithWindow / Topology / Channels handlers
|
||
// can flip from 503-Retry-After to serving the cache.
|
||
//
|
||
// PR #1688 r1: called on EVERY successful pass (even nil
|
||
// result) so a compute that returns nil but doesn't panic
|
||
// still lifts the gate — banner-stuck-forever fix (munger #2).
|
||
if ready {
|
||
r.markFirstPassDone_1659()
|
||
}
|
||
}
|
||
|
||
// RecomputeNow has the loop goroutine run a compute that starts after
|
||
// this call, and waits for it to finish. The periodic ticker restarts
|
||
// from that compute, so the next periodic pass is a full interval
|
||
// later. Returns early if the recomputer is stopped; blocks until Start
|
||
// if it has not started yet.
|
||
func (r *analyticsRecomputer) RecomputeNow() {
|
||
ack := make(chan struct{})
|
||
select {
|
||
case r.recomputeReq <- ack:
|
||
case <-r.stop:
|
||
return
|
||
}
|
||
select {
|
||
case <-ack:
|
||
case <-r.stop:
|
||
}
|
||
}
|
||
|
||
// recomputeWhenLoaded waits for loaded to close, then recomputes each
|
||
// recomputer once, one at a time and in slice order. Sequential so the
|
||
// post-load passes do not all hold the store read lock at once, and so
|
||
// a recomputer that reads another one's snapshot can be placed after
|
||
// it. Returns when done or when stop closes.
|
||
func recomputeWhenLoaded(loaded, stop <-chan struct{}, rcs []*analyticsRecomputer) {
|
||
select {
|
||
case <-loaded:
|
||
case <-stop:
|
||
return
|
||
}
|
||
t0 := time.Now()
|
||
parts := make([]string, 0, len(rcs))
|
||
for _, rc := range rcs {
|
||
select {
|
||
case <-stop:
|
||
return
|
||
default:
|
||
}
|
||
rc.RecomputeNow()
|
||
parts = append(parts, fmt.Sprintf("%s=%s", rc.name, rc.LastComputeDuration().Round(time.Millisecond)))
|
||
}
|
||
log.Printf("[analytics-recompute] startup load done: recomputed %d snapshots in %s (%s)",
|
||
len(rcs), time.Since(t0).Round(time.Millisecond), strings.Join(parts, " "))
|
||
}
|
||
|
||
// Load returns the most recently computed snapshot, or nil if Start
|
||
// has not been called (or the very first compute returned nil).
|
||
// Never blocks beyond a single atomic load.
|
||
func (r *analyticsRecomputer) Load() interface{} {
|
||
v := r.cache.Load()
|
||
if v == nil {
|
||
return nil
|
||
}
|
||
return v
|
||
}
|
||
|
||
// Stop signals the background goroutine to exit and waits for it.
|
||
// Safe to call multiple times. Safe to call before Start (no-op).
|
||
func (r *analyticsRecomputer) Stop() {
|
||
r.stopOnce.Do(func() {
|
||
close(r.stop)
|
||
})
|
||
// Only wait if the goroutine was actually started.
|
||
select {
|
||
case <-r.done:
|
||
case <-time.After(5 * time.Second):
|
||
// Defensive timeout: shouldn't happen in practice.
|
||
}
|
||
}
|
||
|
||
// LastComputeDuration returns the duration of the most recent compute.
|
||
func (r *analyticsRecomputer) LastComputeDuration() time.Duration {
|
||
return time.Duration(r.lastComputeNs.Load())
|
||
}
|
||
|
||
// ComputeRuns returns the total number of compute invocations.
|
||
func (r *analyticsRecomputer) ComputeRuns() int64 {
|
||
return r.computeRuns.Load()
|
||
}
|
||
|
||
// AnalyticsRecomputeIntervals lets callers (main.go) override the
|
||
// per-endpoint recompute interval from config.json. Zero values fall
|
||
// back to the defaultInterval passed to StartAnalyticsRecomputers.
|
||
type AnalyticsRecomputeIntervals struct {
|
||
Topology time.Duration
|
||
RF time.Duration
|
||
Distance time.Duration
|
||
Channels time.Duration
|
||
HashCollisions time.Duration
|
||
HashSizes time.Duration
|
||
Roles time.Duration
|
||
ObserversClockSkew time.Duration
|
||
NodesClockSkew time.Duration
|
||
}
|
||
|
||
func pickInterval(override, def time.Duration) time.Duration {
|
||
if override > 0 {
|
||
return override
|
||
}
|
||
return def
|
||
}
|
||
|
||
// analyticsRecomputersLocked lists the analytics recomputers in the
|
||
// order they start and recompute after the startup load: the three
|
||
// warm-up-gated ones first, and roles after nodes-clock-skew because
|
||
// computeAnalyticsRoles reads that recomputer's snapshot
|
||
// (GetFleetClockSkew). Caller holds analyticsRecomputerMu.
|
||
func (s *PacketStore) analyticsRecomputersLocked() []*analyticsRecomputer {
|
||
return []*analyticsRecomputer{
|
||
s.recompRF, s.recompTopology, s.recompChannels,
|
||
s.recompDistance, s.recompHashCollisions, s.recompHashSizes,
|
||
s.recompObserversClockSkew, s.recompNodesClockSkew,
|
||
s.recompRoles,
|
||
s.recompRetransmissions,
|
||
}
|
||
}
|
||
|
||
// StartAnalyticsRecomputers wires each analytics endpoint to a
|
||
// background recompute goroutine. Each runs an initial compute
|
||
// synchronously (so the first read after startup is a cache hit, never
|
||
// cold) and then refreshes on a ticker.
|
||
//
|
||
// All recomputers serve the DEFAULT query shape only: region="" and
|
||
// zero-window (no ?since= / ?until= params). Region-keyed or windowed
|
||
// queries continue to use the legacy on-request compute + TTL cache —
|
||
// the recomputer count would explode if we maintained one per
|
||
// (endpoint × region × window) combination, and region filtering is
|
||
// fast read-time work anyway.
|
||
//
|
||
// Returns a stop closure that signals all goroutines and blocks until
|
||
// they exit. Safe to call once per PacketStore. Idempotent if called
|
||
// multiple times (subsequent calls return the first stop closure).
|
||
func (s *PacketStore) StartAnalyticsRecomputers(defaultInterval time.Duration, overrides ...AnalyticsRecomputeIntervals) func() {
|
||
if defaultInterval <= 0 {
|
||
defaultInterval = 5 * time.Minute
|
||
}
|
||
var ov AnalyticsRecomputeIntervals
|
||
if len(overrides) > 0 {
|
||
ov = overrides[0]
|
||
}
|
||
|
||
s.analyticsRecomputerMu.Lock()
|
||
if s.recompTopology != nil {
|
||
// Already started; return a no-op so the caller's defer is harmless.
|
||
s.analyticsRecomputerMu.Unlock()
|
||
return func() {}
|
||
}
|
||
|
||
// Each recomputer wraps the underlying compute* function with the
|
||
// default arguments. We use computeAnalytics* (not GetAnalytics*) to
|
||
// bypass the legacy TTL cache layer — the recomputer IS the cache.
|
||
s.recompTopology = newAnalyticsRecomputer(
|
||
"topology", pickInterval(ov.Topology, defaultInterval),
|
||
func() interface{} { return s.computeAnalyticsTopology("", "", TimeWindow{}) },
|
||
)
|
||
s.recompRF = newAnalyticsRecomputer(
|
||
"rf", pickInterval(ov.RF, defaultInterval),
|
||
func() interface{} { return s.computeAnalyticsRF("", "", TimeWindow{}) },
|
||
)
|
||
s.recompDistance = newAnalyticsRecomputer(
|
||
"distance", pickInterval(ov.Distance, defaultInterval),
|
||
func() interface{} { return s.computeAnalyticsDistance("", "") },
|
||
)
|
||
s.recompChannels = newAnalyticsRecomputer(
|
||
"channels", pickInterval(ov.Channels, defaultInterval),
|
||
func() interface{} { return s.computeAnalyticsChannels("", "", TimeWindow{}) },
|
||
)
|
||
s.recompHashCollisions = newAnalyticsRecomputer(
|
||
"hash-collisions", pickInterval(ov.HashCollisions, defaultInterval),
|
||
func() interface{} { return s.computeHashCollisions("", "") },
|
||
)
|
||
s.recompHashSizes = newAnalyticsRecomputer(
|
||
"hash-sizes", pickInterval(ov.HashSizes, defaultInterval),
|
||
func() interface{} { return s.computeAnalyticsHashSizesWithCapability("", "") },
|
||
)
|
||
s.recompRoles = newAnalyticsRecomputer(
|
||
"roles", pickInterval(ov.Roles, defaultInterval),
|
||
func() interface{} { return s.computeAnalyticsRoles() },
|
||
)
|
||
s.recompObserversClockSkew = newAnalyticsRecomputer(
|
||
"observers-clock-skew", pickInterval(ov.ObserversClockSkew, defaultInterval),
|
||
func() interface{} { return s.computeObserverCalibrations() },
|
||
)
|
||
s.recompNodesClockSkew = newAnalyticsRecomputer(
|
||
"nodes-clock-skew", pickInterval(ov.NodesClockSkew, defaultInterval),
|
||
func() interface{} { return s.computeFleetClockSkew() },
|
||
)
|
||
s.recompRetransmissions = newAnalyticsRecomputer(
|
||
"retransmissions", defaultInterval,
|
||
func() interface{} {
|
||
return s.computeRetransmissionPressure("", TimeWindow{}, retransmissionDefaultBucket)
|
||
},
|
||
)
|
||
all := s.analyticsRecomputersLocked()
|
||
s.analyticsRecomputerMu.Unlock()
|
||
|
||
// Issue #1659 (PR #1688 r1, munger #5): wire the loader readiness
|
||
// gate on the three warmup-gated recomputers (RF, Topology,
|
||
// Channels). Only a pass that STARTS after the whole startup load
|
||
// (hot window AND background fill) ends the warm-up. LoadComplete()
|
||
// is not enough: it flips at the end of the hot window, so the gate
|
||
// used to open on a snapshot that missed the background fill.
|
||
loaded := s.StartupLoadDone()
|
||
loadedGate := func() bool {
|
||
select {
|
||
case <-loaded:
|
||
return true
|
||
default:
|
||
return false
|
||
}
|
||
}
|
||
s.recompRF.setWarmupReadyGate_1659(loadedGate)
|
||
s.recompTopology.setWarmupReadyGate_1659(loadedGate)
|
||
s.recompChannels.setWarmupReadyGate_1659(loadedGate)
|
||
s.recompRetransmissions.setWarmupReadyGate_1659(loadedGate)
|
||
|
||
for _, rc := range all {
|
||
rc.Start()
|
||
}
|
||
|
||
// main.go starts the recomputers at the first load chunk, so the
|
||
// initial computes above only saw part of the data. Recompute as
|
||
// soon as the load is done instead of a full interval later.
|
||
stopPostLoad := make(chan struct{})
|
||
postLoadDone := make(chan struct{})
|
||
go func() {
|
||
defer close(postLoadDone)
|
||
recomputeWhenLoaded(loaded, stopPostLoad, all)
|
||
}()
|
||
|
||
var stopOnce sync.Once
|
||
return func() {
|
||
stopOnce.Do(func() {
|
||
close(stopPostLoad)
|
||
for _, rc := range all {
|
||
rc.Stop()
|
||
}
|
||
<-postLoadDone
|
||
})
|
||
}
|
||
}
|