Files
livekit/pkg/sfu/forwardstats.go
T
Raja SubramanianandClaude Opus 4.8 6ed445bad1 sfu: report forwarding latency as p90 instead of the mean (#4920)
* sfu: report forwarding latency as p90 instead of the mean

The forwarding-latency metric was the mean transit over all forwarded
packets. A mean is dominated by a few slow outliers, so a handful of
packets stalled on the forward path (e.g. goroutine scheduling latency)
inflated the whole node's reported latency even when nearly every packet
was forwarded promptly.

Report p90 instead: p90 rising means roughly a tenth of forwarded packets
are slow, a broad signal of systemic forwarding load rather than a sparse
tail. To read a percentile over the report window, the mergeable
per-interval summary now keeps a small power-of-two-bucket histogram of
transit instead of running moments (sum, sum-of-squares).

Drop the jitter (transit std dev) gauge: nothing consumed it, and any
spread is derivable from the forward-latency histogram. The protobuf
ForwardJitter field is left in place, now unset, to deprecate separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: clamp forwarding percentile to the observed [min, max]

Bucket interpolation assumes a uniform fill, so a single 20ms packet (or
uniform traffic) could report a p90 above every observed sample. Clamp the
interpolated value to the summary's already-tracked min/max, so a quantile
never falls outside the data. Exact for single samples and repeated
identical latencies.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: fix stale p50 comment in the percentile test

The reported metric is p90; the test comment still said p50 replaced the
mean. Reword it to reflect that a percentile, not the mean, is reported.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: quarter-octave forwarding-latency buckets for threshold resolution

Octave buckets are too coarse near the overload thresholds: 300us falls in
[256,512), so a p90 clustered at ~265us and one at ~500us interpolate to the
same value and would trip (or not) identically. Split each octave into four
linear sub-buckets so the two land in different buckets, on the correct side
of the threshold. Min/max clamping alone does not fix this once a node has a
high tail, since its max no longer bounds the interpolation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: interpolate percentiles within each bucket's observed range

Replace the quarter-octave split with plain octave buckets that also carry
the observed [min, max] of their samples, and interpolate a percentile
within that range instead of the bucket's nominal edges. This is exact when
a bucket's samples cluster, so a p90 near an overload threshold that falls
mid-bucket lands on the correct side of it regardless of bucket width -- no
threshold-aware boundaries needed. It subsumes the min/max clamp, since an
estimate can no longer leave the observed samples.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: cover the 500us cluster in the threshold test

Assert milos's full review example exactly: 265us and 500us clusters that
octave-nominal interpolation both read as ~341us now read 265us and 500us.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: make forwardSummary.addSample a pointer receiver

The summary grew from a small moments struct into a per-bucket histogram
(~700 bytes), so the value-receiver addSample copied the whole summary on
every drained sample in the flush loop. Mutate in place instead: ~28ns ->
~2ns per sample in the background fold, no behavior change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: eighth-octave buckets for accuracy near thresholds

Split each octave into 8 linear sub-buckets (was plain octaves). With the
per-bucket min/max interpolation this reports a tight p90 cluster exactly
even when it sits mid-octave: 850x265us + 150x410us now reads 410us (above a
400us threshold) instead of 395.5us, and a lognormal p90 lands within ~0.2us
of exact. Per-sample add cost is unchanged (~2ns); cost is ~5KB per summary
and a larger but per-report merge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-29 01:39:05 +05:30

368 lines
11 KiB
Go

package sfu
import (
"math/bits"
"sync"
"time"
"github.com/livekit/livekit-server/pkg/telemetry/prometheus"
"github.com/livekit/protocol/logger"
"go.uber.org/atomic"
)
const (
cHighForwardingLatency = 20 * time.Millisecond
cSkewFactor = 10
// forwardLatencyPercentile is the transit quantile reported as the node's
// forwarding latency and read by the overload controller. p90 tracks a broad
// slowdown but ignores the sparse tail (a few packets stalled on goroutine
// scheduling), which shedding cannot fix.
forwardLatencyPercentile = 0.90
)
const (
// A summary interval's worth of samples across all tracks must fit without
// dropping (ForwardStats is a singleton). Shard count spreads the per-packet
// atomic; shard capacity bounds memory (numShards*shardCap*16 bytes = 2MiB).
forwardSampleNumShards = 16
forwardSampleShardCap = 8192
forwardSampleShardMask = forwardSampleShardCap - 1
forwardSampleShardSel = forwardSampleNumShards - 1
)
// forwardSampleShard is a ring of transit samples with multiple producers and a
// single consumer. A producer reserves a slot, stores the value, then publishes
// the slot's epoch (reserved index + 1). The consumer reads a slot only once its
// epoch marks the value committed for that index.
type forwardSampleShard struct {
writeIdx atomic.Uint64 // advanced by producers to reserve a slot
readIdx uint64 // consumer-only cursor
ring [forwardSampleShardCap]atomic.Int64
seq [forwardSampleShardCap]atomic.Uint64 // per-slot publish epoch
}
// forwardSampleBuffer holds per-packet transit samples produced on the packet
// path and consumed by the background worker, which performs metric emission.
type forwardSampleBuffer struct {
shards [forwardSampleNumShards]forwardSampleShard
dropped atomic.Uint64
}
// push records a sample: reserve a slot, store the value, then publish the
// slot's epoch. The shard is selected from arrival time bits.
func (b *forwardSampleBuffer) push(arrival, transitNs int64) {
sh := &b.shards[(uint64(arrival)>>6)&forwardSampleShardSel]
i := sh.writeIdx.Add(1) - 1
slot := i & forwardSampleShardMask
sh.ring[slot].Store(transitNs)
sh.seq[slot].Store(i + 1)
}
// drain passes every committed sample to fn and advances the read cursor. Only
// the background worker calls this.
//
// A slot holds index r's value once its epoch equals r+1. If the slot at the
// cursor is still uncommitted (a producer reserved it but has not published),
// draining stops and resumes from there on the next call, so no sample is read
// stale or skipped. When producers get a shard's capacity ahead, or overwrite a
// slot before it is read, the affected samples are counted as dropped.
func (b *forwardSampleBuffer) drain(fn func(transitNs int64)) {
for si := range b.shards {
sh := &b.shards[si]
w := sh.writeIdx.Load()
r := sh.readIdx
if w-r > forwardSampleShardCap {
b.dropped.Add(w - r - forwardSampleShardCap)
r = w - forwardSampleShardCap
}
for r < w {
slot := r & forwardSampleShardMask
if sh.seq[slot].Load() < r+1 {
// reserved but not yet published; resume here next drain
break
}
v := sh.ring[slot].Load()
if sh.seq[slot].Load() != r+1 {
// overwritten by a newer sample during the read; original lost
b.dropped.Add(1)
r++
continue
}
fn(v)
r++
}
sh.readIdx = r
}
}
func (b *forwardSampleBuffer) takeDropped() uint64 {
return b.dropped.Swap(0)
}
// forwardSummary is a mergeable histogram of forwarding transit over an
// interval. Each power-of-two octave [2^e, 2^(e+1)) us is split into
// forwardHistSub linear sub-buckets, and each bucket keeps the observed
// [min, max] of the samples in it. A percentile interpolates within the
// crossing bucket's observed range rather than its nominal edges, so the
// estimate never leaves the samples and is exact when a bucket's samples
// cluster -- which keeps it accurate near an overload threshold that falls
// mid-bucket, without needing threshold-aware boundaries. Two summaries merge
// by adding their buckets, so the percentile covers the whole report window,
// and a few stalled packets (e.g. from goroutine scheduling latency) cannot
// drag it the way a mean does.
const (
forwardHistSub = 8 // linear sub-buckets per octave
forwardHistOctaves = 27 // up to 2^27 us (~134 s)
forwardHistBuckets = forwardHistOctaves * forwardHistSub // 216
)
// bucketStat is one bucket: how many samples fell in it and their observed
// transit range in nanoseconds.
type bucketStat struct {
count int64
minNs int64
maxNs int64
}
func (b *bucketStat) add(transitNs int64) {
if b.count == 0 {
b.minNs, b.maxNs = transitNs, transitNs
} else {
b.minNs = min(b.minNs, transitNs)
b.maxNs = max(b.maxNs, transitNs)
}
b.count++
}
func (b *bucketStat) mergeIn(o bucketStat) {
if o.count == 0 {
return
}
if b.count == 0 {
*b = o
return
}
b.count += o.count
b.minNs = min(b.minNs, o.minNs)
b.maxNs = max(b.maxNs, o.maxNs)
}
type forwardSummary struct {
count int64
minNs int64
maxNs int64
buckets [forwardHistBuckets]bucketStat
}
// forwardBucket returns the bucket index for a transit in nanoseconds: the
// octave floor(log2(us)) times forwardHistSub, plus the linear sub-bucket
// within the octave.
func forwardBucket(transitNs int64) int {
us := transitNs / 1000
if us <= 0 {
return 0
}
e := bits.Len64(uint64(us)) - 1 // floor(log2(us)), us >= 1 so e >= 0
if e >= forwardHistOctaves {
return forwardHistBuckets - 1
}
octave := int64(1) << e
return e*forwardHistSub + int((us-octave)*forwardHistSub/octave)
}
func (s *forwardSummary) addSample(transitNs int64) {
if s.count == 0 {
s.minNs, s.maxNs = transitNs, transitNs
} else {
s.minNs = min(s.minNs, transitNs)
s.maxNs = max(s.maxNs, transitNs)
}
s.count++
s.buckets[forwardBucket(transitNs)].add(transitNs)
}
func (s forwardSummary) merge(o forwardSummary) forwardSummary {
if o.count == 0 {
return s
}
if s.count == 0 {
return o
}
s.count += o.count
s.minNs = min(s.minNs, o.minNs)
s.maxNs = max(s.maxNs, o.maxNs)
for i := range s.buckets {
s.buckets[i].mergeIn(o.buckets[i])
}
return s
}
// percentile returns the p-quantile (0..1) of transit, interpolated within the
// observed [min, max] of the bucket the quantile falls in.
func (s forwardSummary) percentile(p float64) time.Duration {
if s.count == 0 {
return 0
}
target := p * float64(s.count)
var cum float64
for i := range s.buckets {
b := s.buckets[i]
c := float64(b.count)
if c == 0 {
continue
}
if cum+c >= target {
lo, hi := float64(b.minNs), float64(b.maxNs)
return time.Duration(lo + (hi-lo)*(target-cum)/c)
}
cum += c
}
return time.Duration(s.maxNs)
}
type ForwardStats struct {
samples forwardSampleBuffer
// ring of per-summary-interval summaries covering the report window.
// written by the background worker (flush) and read both by the worker
// (report) and by external callers (GetStats), so it is guarded by lock.
lock sync.Mutex
ring []forwardSummary
ringHead int
ringLen int
summaryInterval time.Duration
reportInterval time.Duration
closeCh chan struct{}
}
func NewForwardStats(summaryInterval, reportInterval, reportWindow time.Duration) *ForwardStats {
ringCap := int((reportWindow + summaryInterval - 1) / summaryInterval)
if ringCap < 1 {
ringCap = 1
}
s := &ForwardStats{
ring: make([]forwardSummary, ringCap),
summaryInterval: summaryInterval,
reportInterval: reportInterval,
closeCh: make(chan struct{}),
}
go s.run()
return s
}
// Update records a forwarded packet's transit latency. It buffers the sample
// and returns the transit and whether it exceeds the high-latency threshold.
// The sample is aggregated and emitted by the background worker.
func (s *ForwardStats) Update(arrival, left int64) (int64, bool) {
transit := left - arrival
s.samples.push(arrival, transit)
return transit, time.Duration(transit) > cHighForwardingLatency
}
func (s *ForwardStats) Stop() {
close(s.closeCh)
}
func (s *ForwardStats) run() {
summaryTicker := time.NewTicker(s.summaryInterval)
defer summaryTicker.Stop()
reportTicker := time.NewTicker(s.reportInterval)
defer reportTicker.Stop()
for {
select {
case <-s.closeCh:
return
case <-summaryTicker.C:
s.flush()
case <-reportTicker.C:
// the summary ticker keeps the window ring current to within one
// summary interval; report over it without advancing the ring.
s.report()
}
}
}
// flush drains the buffered samples, observes each into the Prometheus
// histogram, and folds the interval summary into the window ring used for the
// latency gauge.
func (s *ForwardStats) flush() {
var summ forwardSummary
s.samples.drain(func(transitNs int64) {
prometheus.RecordForwardLatencySample(transitNs)
summ.addSample(transitNs)
})
s.lock.Lock()
s.ring[s.ringHead] = summ
s.ringHead = (s.ringHead + 1) % len(s.ring)
if s.ringLen < len(s.ring) {
s.ringLen++
}
s.lock.Unlock()
}
// summarize merges the ring summaries covering the most recent window. A
// window <= 0 (or >= the report window) covers the entire ring.
func (s *ForwardStats) summarize(window time.Duration) forwardSummary {
s.lock.Lock()
defer s.lock.Unlock()
n := s.ringLen
if window > 0 && s.summaryInterval > 0 {
want := int((window + s.summaryInterval - 1) / s.summaryInterval)
if want < 1 {
want = 1
}
if want < n {
n = want
}
}
// walk backwards from the most recent entry (ringHead-1) over n entries.
var w forwardSummary
for i := 0; i < n; i++ {
idx := (s.ringHead - 1 - i + len(s.ring)) % len(s.ring)
w = w.merge(s.ring[idx])
}
return w
}
// GetStats returns the reported forwarding-latency percentile
// (forwardLatencyPercentile) over the most recent duration. The duration is
// rounded up to a whole number of summary intervals (the smallest bucket span
// that covers it). A duration <= 0, or one that meets/exceeds the report window,
// covers the full window.
func (s *ForwardStats) GetStats(duration time.Duration) time.Duration {
return s.summarize(duration).percentile(forwardLatencyPercentile)
}
func (s *ForwardStats) report() {
w := s.summarize(0)
p90 := w.percentile(forwardLatencyPercentile)
if dropped := s.samples.takeDropped(); dropped > 0 {
logger.Warnw("forward stats sample buffer overflow", nil, "dropped", dropped)
}
// a max far above p90 means a few packets stalled (e.g. goroutine scheduling
// latency) rather than a broad forwarding slowdown; p90 rides through it.
if w.count > 0 && w.maxNs > p90.Nanoseconds()*cSkewFactor {
logger.Infow(
"high spread in forwarding path",
"lowest", time.Duration(w.minNs),
"highest", time.Duration(w.maxNs),
"count", w.count,
"p90", p90,
)
}
prometheus.RecordForwardLatency(uint32(p90.Nanoseconds()))
}