mirror of
https://github.com/livekit/livekit.git
synced 2026-10-07 12:01:00 +00:00
* sfu: report forwarding latency as p90 instead of the mean The forwarding-latency metric was the mean transit over all forwarded packets. A mean is dominated by a few slow outliers, so a handful of packets stalled on the forward path (e.g. goroutine scheduling latency) inflated the whole node's reported latency even when nearly every packet was forwarded promptly. Report p90 instead: p90 rising means roughly a tenth of forwarded packets are slow, a broad signal of systemic forwarding load rather than a sparse tail. To read a percentile over the report window, the mergeable per-interval summary now keeps a small power-of-two-bucket histogram of transit instead of running moments (sum, sum-of-squares). Drop the jitter (transit std dev) gauge: nothing consumed it, and any spread is derivable from the forward-latency histogram. The protobuf ForwardJitter field is left in place, now unset, to deprecate separately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: clamp forwarding percentile to the observed [min, max] Bucket interpolation assumes a uniform fill, so a single 20ms packet (or uniform traffic) could report a p90 above every observed sample. Clamp the interpolated value to the summary's already-tracked min/max, so a quantile never falls outside the data. Exact for single samples and repeated identical latencies. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: fix stale p50 comment in the percentile test The reported metric is p90; the test comment still said p50 replaced the mean. Reword it to reflect that a percentile, not the mean, is reported. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: quarter-octave forwarding-latency buckets for threshold resolution Octave buckets are too coarse near the overload thresholds: 300us falls in [256,512), so a p90 clustered at ~265us and one at ~500us interpolate to the same value and would trip (or not) identically. Split each octave into four linear sub-buckets so the two land in different buckets, on the correct side of the threshold. Min/max clamping alone does not fix this once a node has a high tail, since its max no longer bounds the interpolation. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: interpolate percentiles within each bucket's observed range Replace the quarter-octave split with plain octave buckets that also carry the observed [min, max] of their samples, and interpolate a percentile within that range instead of the bucket's nominal edges. This is exact when a bucket's samples cluster, so a p90 near an overload threshold that falls mid-bucket lands on the correct side of it regardless of bucket width -- no threshold-aware boundaries needed. It subsumes the min/max clamp, since an estimate can no longer leave the observed samples. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: cover the 500us cluster in the threshold test Assert milos's full review example exactly: 265us and 500us clusters that octave-nominal interpolation both read as ~341us now read 265us and 500us. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: make forwardSummary.addSample a pointer receiver The summary grew from a small moments struct into a per-bucket histogram (~700 bytes), so the value-receiver addSample copied the whole summary on every drained sample in the flush loop. Mutate in place instead: ~28ns -> ~2ns per sample in the background fold, no behavior change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * sfu: eighth-octave buckets for accuracy near thresholds Split each octave into 8 linear sub-buckets (was plain octaves). With the per-bucket min/max interpolation this reports a tight p90 cluster exactly even when it sits mid-octave: 850x265us + 150x410us now reads 410us (above a 400us threshold) instead of 395.5us, and a lognormal p90 lands within ~0.2us of exact. Per-sample add cost is unchanged (~2ns); cost is ~5KB per summary and a larger but per-report merge. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
368 lines
11 KiB
Go
368 lines
11 KiB
Go
package sfu
|
|
|
|
import (
|
|
"math/bits"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/livekit/livekit-server/pkg/telemetry/prometheus"
|
|
"github.com/livekit/protocol/logger"
|
|
"go.uber.org/atomic"
|
|
)
|
|
|
|
const (
|
|
cHighForwardingLatency = 20 * time.Millisecond
|
|
cSkewFactor = 10
|
|
|
|
// forwardLatencyPercentile is the transit quantile reported as the node's
|
|
// forwarding latency and read by the overload controller. p90 tracks a broad
|
|
// slowdown but ignores the sparse tail (a few packets stalled on goroutine
|
|
// scheduling), which shedding cannot fix.
|
|
forwardLatencyPercentile = 0.90
|
|
)
|
|
|
|
const (
|
|
// A summary interval's worth of samples across all tracks must fit without
|
|
// dropping (ForwardStats is a singleton). Shard count spreads the per-packet
|
|
// atomic; shard capacity bounds memory (numShards*shardCap*16 bytes = 2MiB).
|
|
forwardSampleNumShards = 16
|
|
forwardSampleShardCap = 8192
|
|
forwardSampleShardMask = forwardSampleShardCap - 1
|
|
forwardSampleShardSel = forwardSampleNumShards - 1
|
|
)
|
|
|
|
// forwardSampleShard is a ring of transit samples with multiple producers and a
|
|
// single consumer. A producer reserves a slot, stores the value, then publishes
|
|
// the slot's epoch (reserved index + 1). The consumer reads a slot only once its
|
|
// epoch marks the value committed for that index.
|
|
type forwardSampleShard struct {
|
|
writeIdx atomic.Uint64 // advanced by producers to reserve a slot
|
|
readIdx uint64 // consumer-only cursor
|
|
ring [forwardSampleShardCap]atomic.Int64
|
|
seq [forwardSampleShardCap]atomic.Uint64 // per-slot publish epoch
|
|
}
|
|
|
|
// forwardSampleBuffer holds per-packet transit samples produced on the packet
|
|
// path and consumed by the background worker, which performs metric emission.
|
|
type forwardSampleBuffer struct {
|
|
shards [forwardSampleNumShards]forwardSampleShard
|
|
dropped atomic.Uint64
|
|
}
|
|
|
|
// push records a sample: reserve a slot, store the value, then publish the
|
|
// slot's epoch. The shard is selected from arrival time bits.
|
|
func (b *forwardSampleBuffer) push(arrival, transitNs int64) {
|
|
sh := &b.shards[(uint64(arrival)>>6)&forwardSampleShardSel]
|
|
i := sh.writeIdx.Add(1) - 1
|
|
slot := i & forwardSampleShardMask
|
|
sh.ring[slot].Store(transitNs)
|
|
sh.seq[slot].Store(i + 1)
|
|
}
|
|
|
|
// drain passes every committed sample to fn and advances the read cursor. Only
|
|
// the background worker calls this.
|
|
//
|
|
// A slot holds index r's value once its epoch equals r+1. If the slot at the
|
|
// cursor is still uncommitted (a producer reserved it but has not published),
|
|
// draining stops and resumes from there on the next call, so no sample is read
|
|
// stale or skipped. When producers get a shard's capacity ahead, or overwrite a
|
|
// slot before it is read, the affected samples are counted as dropped.
|
|
func (b *forwardSampleBuffer) drain(fn func(transitNs int64)) {
|
|
for si := range b.shards {
|
|
sh := &b.shards[si]
|
|
w := sh.writeIdx.Load()
|
|
r := sh.readIdx
|
|
if w-r > forwardSampleShardCap {
|
|
b.dropped.Add(w - r - forwardSampleShardCap)
|
|
r = w - forwardSampleShardCap
|
|
}
|
|
for r < w {
|
|
slot := r & forwardSampleShardMask
|
|
if sh.seq[slot].Load() < r+1 {
|
|
// reserved but not yet published; resume here next drain
|
|
break
|
|
}
|
|
v := sh.ring[slot].Load()
|
|
if sh.seq[slot].Load() != r+1 {
|
|
// overwritten by a newer sample during the read; original lost
|
|
b.dropped.Add(1)
|
|
r++
|
|
continue
|
|
}
|
|
fn(v)
|
|
r++
|
|
}
|
|
sh.readIdx = r
|
|
}
|
|
}
|
|
|
|
func (b *forwardSampleBuffer) takeDropped() uint64 {
|
|
return b.dropped.Swap(0)
|
|
}
|
|
|
|
// forwardSummary is a mergeable histogram of forwarding transit over an
|
|
// interval. Each power-of-two octave [2^e, 2^(e+1)) us is split into
|
|
// forwardHistSub linear sub-buckets, and each bucket keeps the observed
|
|
// [min, max] of the samples in it. A percentile interpolates within the
|
|
// crossing bucket's observed range rather than its nominal edges, so the
|
|
// estimate never leaves the samples and is exact when a bucket's samples
|
|
// cluster -- which keeps it accurate near an overload threshold that falls
|
|
// mid-bucket, without needing threshold-aware boundaries. Two summaries merge
|
|
// by adding their buckets, so the percentile covers the whole report window,
|
|
// and a few stalled packets (e.g. from goroutine scheduling latency) cannot
|
|
// drag it the way a mean does.
|
|
const (
|
|
forwardHistSub = 8 // linear sub-buckets per octave
|
|
forwardHistOctaves = 27 // up to 2^27 us (~134 s)
|
|
forwardHistBuckets = forwardHistOctaves * forwardHistSub // 216
|
|
)
|
|
|
|
// bucketStat is one bucket: how many samples fell in it and their observed
|
|
// transit range in nanoseconds.
|
|
type bucketStat struct {
|
|
count int64
|
|
minNs int64
|
|
maxNs int64
|
|
}
|
|
|
|
func (b *bucketStat) add(transitNs int64) {
|
|
if b.count == 0 {
|
|
b.minNs, b.maxNs = transitNs, transitNs
|
|
} else {
|
|
b.minNs = min(b.minNs, transitNs)
|
|
b.maxNs = max(b.maxNs, transitNs)
|
|
}
|
|
b.count++
|
|
}
|
|
|
|
func (b *bucketStat) mergeIn(o bucketStat) {
|
|
if o.count == 0 {
|
|
return
|
|
}
|
|
if b.count == 0 {
|
|
*b = o
|
|
return
|
|
}
|
|
b.count += o.count
|
|
b.minNs = min(b.minNs, o.minNs)
|
|
b.maxNs = max(b.maxNs, o.maxNs)
|
|
}
|
|
|
|
type forwardSummary struct {
|
|
count int64
|
|
minNs int64
|
|
maxNs int64
|
|
buckets [forwardHistBuckets]bucketStat
|
|
}
|
|
|
|
// forwardBucket returns the bucket index for a transit in nanoseconds: the
|
|
// octave floor(log2(us)) times forwardHistSub, plus the linear sub-bucket
|
|
// within the octave.
|
|
func forwardBucket(transitNs int64) int {
|
|
us := transitNs / 1000
|
|
if us <= 0 {
|
|
return 0
|
|
}
|
|
e := bits.Len64(uint64(us)) - 1 // floor(log2(us)), us >= 1 so e >= 0
|
|
if e >= forwardHistOctaves {
|
|
return forwardHistBuckets - 1
|
|
}
|
|
octave := int64(1) << e
|
|
return e*forwardHistSub + int((us-octave)*forwardHistSub/octave)
|
|
}
|
|
|
|
func (s *forwardSummary) addSample(transitNs int64) {
|
|
if s.count == 0 {
|
|
s.minNs, s.maxNs = transitNs, transitNs
|
|
} else {
|
|
s.minNs = min(s.minNs, transitNs)
|
|
s.maxNs = max(s.maxNs, transitNs)
|
|
}
|
|
s.count++
|
|
s.buckets[forwardBucket(transitNs)].add(transitNs)
|
|
}
|
|
|
|
func (s forwardSummary) merge(o forwardSummary) forwardSummary {
|
|
if o.count == 0 {
|
|
return s
|
|
}
|
|
if s.count == 0 {
|
|
return o
|
|
}
|
|
s.count += o.count
|
|
s.minNs = min(s.minNs, o.minNs)
|
|
s.maxNs = max(s.maxNs, o.maxNs)
|
|
for i := range s.buckets {
|
|
s.buckets[i].mergeIn(o.buckets[i])
|
|
}
|
|
return s
|
|
}
|
|
|
|
// percentile returns the p-quantile (0..1) of transit, interpolated within the
|
|
// observed [min, max] of the bucket the quantile falls in.
|
|
func (s forwardSummary) percentile(p float64) time.Duration {
|
|
if s.count == 0 {
|
|
return 0
|
|
}
|
|
target := p * float64(s.count)
|
|
var cum float64
|
|
for i := range s.buckets {
|
|
b := s.buckets[i]
|
|
c := float64(b.count)
|
|
if c == 0 {
|
|
continue
|
|
}
|
|
if cum+c >= target {
|
|
lo, hi := float64(b.minNs), float64(b.maxNs)
|
|
return time.Duration(lo + (hi-lo)*(target-cum)/c)
|
|
}
|
|
cum += c
|
|
}
|
|
return time.Duration(s.maxNs)
|
|
}
|
|
|
|
type ForwardStats struct {
|
|
samples forwardSampleBuffer
|
|
|
|
// ring of per-summary-interval summaries covering the report window.
|
|
// written by the background worker (flush) and read both by the worker
|
|
// (report) and by external callers (GetStats), so it is guarded by lock.
|
|
lock sync.Mutex
|
|
ring []forwardSummary
|
|
ringHead int
|
|
ringLen int
|
|
|
|
summaryInterval time.Duration
|
|
reportInterval time.Duration
|
|
|
|
closeCh chan struct{}
|
|
}
|
|
|
|
func NewForwardStats(summaryInterval, reportInterval, reportWindow time.Duration) *ForwardStats {
|
|
ringCap := int((reportWindow + summaryInterval - 1) / summaryInterval)
|
|
if ringCap < 1 {
|
|
ringCap = 1
|
|
}
|
|
|
|
s := &ForwardStats{
|
|
ring: make([]forwardSummary, ringCap),
|
|
summaryInterval: summaryInterval,
|
|
reportInterval: reportInterval,
|
|
closeCh: make(chan struct{}),
|
|
}
|
|
|
|
go s.run()
|
|
return s
|
|
}
|
|
|
|
// Update records a forwarded packet's transit latency. It buffers the sample
|
|
// and returns the transit and whether it exceeds the high-latency threshold.
|
|
// The sample is aggregated and emitted by the background worker.
|
|
func (s *ForwardStats) Update(arrival, left int64) (int64, bool) {
|
|
transit := left - arrival
|
|
s.samples.push(arrival, transit)
|
|
return transit, time.Duration(transit) > cHighForwardingLatency
|
|
}
|
|
|
|
func (s *ForwardStats) Stop() {
|
|
close(s.closeCh)
|
|
}
|
|
|
|
func (s *ForwardStats) run() {
|
|
summaryTicker := time.NewTicker(s.summaryInterval)
|
|
defer summaryTicker.Stop()
|
|
reportTicker := time.NewTicker(s.reportInterval)
|
|
defer reportTicker.Stop()
|
|
|
|
for {
|
|
select {
|
|
case <-s.closeCh:
|
|
return
|
|
|
|
case <-summaryTicker.C:
|
|
s.flush()
|
|
|
|
case <-reportTicker.C:
|
|
// the summary ticker keeps the window ring current to within one
|
|
// summary interval; report over it without advancing the ring.
|
|
s.report()
|
|
}
|
|
}
|
|
}
|
|
|
|
// flush drains the buffered samples, observes each into the Prometheus
|
|
// histogram, and folds the interval summary into the window ring used for the
|
|
// latency gauge.
|
|
func (s *ForwardStats) flush() {
|
|
var summ forwardSummary
|
|
s.samples.drain(func(transitNs int64) {
|
|
prometheus.RecordForwardLatencySample(transitNs)
|
|
summ.addSample(transitNs)
|
|
})
|
|
|
|
s.lock.Lock()
|
|
s.ring[s.ringHead] = summ
|
|
s.ringHead = (s.ringHead + 1) % len(s.ring)
|
|
if s.ringLen < len(s.ring) {
|
|
s.ringLen++
|
|
}
|
|
s.lock.Unlock()
|
|
}
|
|
|
|
// summarize merges the ring summaries covering the most recent window. A
|
|
// window <= 0 (or >= the report window) covers the entire ring.
|
|
func (s *ForwardStats) summarize(window time.Duration) forwardSummary {
|
|
s.lock.Lock()
|
|
defer s.lock.Unlock()
|
|
|
|
n := s.ringLen
|
|
if window > 0 && s.summaryInterval > 0 {
|
|
want := int((window + s.summaryInterval - 1) / s.summaryInterval)
|
|
if want < 1 {
|
|
want = 1
|
|
}
|
|
if want < n {
|
|
n = want
|
|
}
|
|
}
|
|
|
|
// walk backwards from the most recent entry (ringHead-1) over n entries.
|
|
var w forwardSummary
|
|
for i := 0; i < n; i++ {
|
|
idx := (s.ringHead - 1 - i + len(s.ring)) % len(s.ring)
|
|
w = w.merge(s.ring[idx])
|
|
}
|
|
return w
|
|
}
|
|
|
|
// GetStats returns the reported forwarding-latency percentile
|
|
// (forwardLatencyPercentile) over the most recent duration. The duration is
|
|
// rounded up to a whole number of summary intervals (the smallest bucket span
|
|
// that covers it). A duration <= 0, or one that meets/exceeds the report window,
|
|
// covers the full window.
|
|
func (s *ForwardStats) GetStats(duration time.Duration) time.Duration {
|
|
return s.summarize(duration).percentile(forwardLatencyPercentile)
|
|
}
|
|
|
|
func (s *ForwardStats) report() {
|
|
w := s.summarize(0)
|
|
|
|
p90 := w.percentile(forwardLatencyPercentile)
|
|
if dropped := s.samples.takeDropped(); dropped > 0 {
|
|
logger.Warnw("forward stats sample buffer overflow", nil, "dropped", dropped)
|
|
}
|
|
// a max far above p90 means a few packets stalled (e.g. goroutine scheduling
|
|
// latency) rather than a broad forwarding slowdown; p90 rides through it.
|
|
if w.count > 0 && w.maxNs > p90.Nanoseconds()*cSkewFactor {
|
|
logger.Infow(
|
|
"high spread in forwarding path",
|
|
"lowest", time.Duration(w.minNs),
|
|
"highest", time.Duration(w.maxNs),
|
|
"count", w.count,
|
|
"p90", p90,
|
|
)
|
|
}
|
|
|
|
prometheus.RecordForwardLatency(uint32(p90.Nanoseconds()))
|
|
}
|