Files
livekit/pkg/telemetry/prometheus/node.go
T
Raja SubramanianandClaude Opus 4.8 6ed445bad1 sfu: report forwarding latency as p90 instead of the mean (#4920)
* sfu: report forwarding latency as p90 instead of the mean

The forwarding-latency metric was the mean transit over all forwarded
packets. A mean is dominated by a few slow outliers, so a handful of
packets stalled on the forward path (e.g. goroutine scheduling latency)
inflated the whole node's reported latency even when nearly every packet
was forwarded promptly.

Report p90 instead: p90 rising means roughly a tenth of forwarded packets
are slow, a broad signal of systemic forwarding load rather than a sparse
tail. To read a percentile over the report window, the mergeable
per-interval summary now keeps a small power-of-two-bucket histogram of
transit instead of running moments (sum, sum-of-squares).

Drop the jitter (transit std dev) gauge: nothing consumed it, and any
spread is derivable from the forward-latency histogram. The protobuf
ForwardJitter field is left in place, now unset, to deprecate separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: clamp forwarding percentile to the observed [min, max]

Bucket interpolation assumes a uniform fill, so a single 20ms packet (or
uniform traffic) could report a p90 above every observed sample. Clamp the
interpolated value to the summary's already-tracked min/max, so a quantile
never falls outside the data. Exact for single samples and repeated
identical latencies.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: fix stale p50 comment in the percentile test

The reported metric is p90; the test comment still said p50 replaced the
mean. Reword it to reflect that a percentile, not the mean, is reported.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: quarter-octave forwarding-latency buckets for threshold resolution

Octave buckets are too coarse near the overload thresholds: 300us falls in
[256,512), so a p90 clustered at ~265us and one at ~500us interpolate to the
same value and would trip (or not) identically. Split each octave into four
linear sub-buckets so the two land in different buckets, on the correct side
of the threshold. Min/max clamping alone does not fix this once a node has a
high tail, since its max no longer bounds the interpolation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: interpolate percentiles within each bucket's observed range

Replace the quarter-octave split with plain octave buckets that also carry
the observed [min, max] of their samples, and interpolate a percentile
within that range instead of the bucket's nominal edges. This is exact when
a bucket's samples cluster, so a p90 near an overload threshold that falls
mid-bucket lands on the correct side of it regardless of bucket width -- no
threshold-aware boundaries needed. It subsumes the min/max clamp, since an
estimate can no longer leave the observed samples.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: cover the 500us cluster in the threshold test

Assert milos's full review example exactly: 265us and 500us clusters that
octave-nominal interpolation both read as ~341us now read 265us and 500us.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: make forwardSummary.addSample a pointer receiver

The summary grew from a small moments struct into a per-bucket histogram
(~700 bytes), so the value-receiver addSample copied the whole summary on
every drained sample in the flush loop. Mutate in place instead: ~28ns ->
~2ns per sample in the background fold, no behavior change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: eighth-octave buckets for accuracy near thresholds

Split each octave into 8 linear sub-buckets (was plain octaves). With the
per-bucket min/max interpolation this reports a tight p90 cluster exactly
even when it sits mid-octave: 850x265us + 150x410us now reads 410us (above a
400us threshold) instead of 395.5us, and a lognormal p90 lands within ~0.2us
of exact. Per-sample add cost is unchanged (~2ns); cost is ~5KB per summary
and a larger but per-report merge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-29 01:39:05 +05:30

310 lines
12 KiB
Go

// Copyright 2023 LiveKit, Inc.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package prometheus
import (
"time"
"github.com/prometheus/client_golang/prometheus"
"github.com/twitchtv/twirp"
"go.uber.org/atomic"
"github.com/livekit/protocol/livekit"
"github.com/livekit/protocol/rpc"
"github.com/livekit/protocol/utils/hwstats"
"github.com/livekit/protocol/webhook"
)
const (
livekitNamespace string = "livekit"
)
var (
initialized atomic.Bool
promMessageCounter *prometheus.CounterVec
promServiceOperationCounter *prometheus.CounterVec
promTwirpRequestStatusCounter *prometheus.CounterVec
promTwirpRequestLatency *prometheus.HistogramVec
sysPacketsStart uint32
sysDroppedPacketsStart uint32
promSysPacketGauge *prometheus.GaugeVec
cpuStats *hwstats.CPUStats
memoryStats *hwstats.MemoryStats
)
func Init(nodeID string, nodeType livekit.NodeType) error {
if initialized.Swap(true) {
return nil
}
promMessageCounter = prometheus.NewCounterVec(
prometheus.CounterOpts{
Namespace: livekitNamespace,
Subsystem: "node",
Name: "messages",
ConstLabels: prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()},
},
[]string{"type", "status", "direction"},
)
promServiceOperationCounter = prometheus.NewCounterVec(
prometheus.CounterOpts{
Namespace: livekitNamespace,
Subsystem: "node",
Name: "service_operation",
ConstLabels: prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()},
},
[]string{"type", "status", "error_type"},
)
promTwirpRequestStatusCounter = prometheus.NewCounterVec(
prometheus.CounterOpts{
Namespace: livekitNamespace,
Subsystem: "node",
Name: "twirp_request_status",
ConstLabels: prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()},
},
[]string{"service", "method", "status", "code"},
)
promTwirpRequestLatency = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Namespace: livekitNamespace,
Subsystem: "node",
Name: "twirp_request_latency_ms",
ConstLabels: prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()},
Buckets: []float64{5, 10, 25, 50, 100, 250, 500, 1000, 2500, 5000, 10000, 30000, 60000},
},
[]string{"service", "method", "status"},
)
promSysPacketGauge = prometheus.NewGaugeVec(
prometheus.GaugeOpts{
Namespace: livekitNamespace,
Subsystem: "node",
Name: "packet_total",
ConstLabels: prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()},
Help: "System level packet count. Count starts at 0 when service is first started.",
},
[]string{"type"},
)
prometheus.MustRegister(promMessageCounter)
prometheus.MustRegister(promServiceOperationCounter)
prometheus.MustRegister(promTwirpRequestStatusCounter)
prometheus.MustRegister(promTwirpRequestLatency)
prometheus.MustRegister(promSysPacketGauge)
sysPacketsStart, sysDroppedPacketsStart, _ = getTCStats()
initPacketStats(nodeID, nodeType)
initRoomStats(nodeID, nodeType)
rpc.InitPSRPCStats(prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()})
webhook.InitWebhookStats(prometheus.Labels{"node_id": nodeID, "node_type": nodeType.String()})
initQualityStats(nodeID, nodeType)
initDataPacketStats(nodeID, nodeType)
initDebugStats(nodeID, nodeType)
var err error
cpuStats, err = hwstats.NewCPUStats(nil)
if err != nil {
return err
}
memoryStats, err = hwstats.NewMemoryStats()
if err != nil {
return err
}
return nil
}
func GetNodeStats(nodeStartedAt int64, prevStats []*livekit.NodeStats, rateIntervals []time.Duration) (*livekit.NodeStats, error) {
loadAvg, err := getLoadAvg()
if err != nil {
return nil, err
}
// On MacOS, get "\"vm_stat\": executable file not found in $PATH" although it is in /usr/bin
// So, do not error out. Use the information if it is available.
memUsed, memTotal, _ := memoryStats.GetMemory()
sysPackets, sysDroppedPackets, _ := getTCStats()
promSysPacketGauge.WithLabelValues("out").Set(float64(sysPackets - sysPacketsStart))
promSysPacketGauge.WithLabelValues("dropped").Set(float64(sysDroppedPackets - sysDroppedPacketsStart))
stats := &livekit.NodeStats{
StartedAt: nodeStartedAt,
UpdatedAt: time.Now().Unix(),
NumRooms: roomCurrent.Load(),
NumClients: participantCurrent.Load(),
NumTracksIn: trackPublishedCurrent.Load(),
NumTracksOut: trackSubscribedCurrent.Load(),
NumTrackPublishAttempts: trackPublishAttempts.Load(),
NumTrackPublishSuccess: trackPublishSuccess.Load(),
NumTrackPublishCancels: trackPublishCancels.Load(),
NumTrackSubscribeAttempts: trackSubscribeAttempts.Load(),
NumTrackSubscribeSuccess: trackSubscribeSuccess.Load(),
NumTrackSubscribeCancels: trackSubscribeCancels.Load(),
BytesIn: bytesIn.Load(),
BytesOut: bytesOut.Load(),
PacketsIn: packetsIn.Load(),
PacketsOut: packetsOut.Load(),
RetransmitBytesOut: retransmitBytes.Load(),
RetransmitPacketsOut: retransmitPackets.Load(),
NackTotal: nackTotal.Load(),
ParticipantSignalConnected: participantSignalConnected.Load(),
ParticipantRtcInit: participantRTCInit.Load(),
ParticipantRtcConnected: participantRTCConnected.Load(),
ParticipantRtcCanceled: participantRTCCanceled.Load(),
ParticipantRtcActive: participantRTCActive.Load(),
ForwardLatency: forwardLatency.Load(),
NumCpus: uint32(cpuStats.NumCPU()), // this will round down to the nearest integer
CpuLoad: float32(cpuStats.GetCPULoad()),
MemoryTotal: memTotal,
MemoryUsed: memUsed,
LoadAvgLast1Min: float32(loadAvg.Loadavg1),
LoadAvgLast5Min: float32(loadAvg.Loadavg5),
LoadAvgLast15Min: float32(loadAvg.Loadavg15),
SysPacketsOut: sysPackets,
SysPacketsDropped: sysDroppedPackets,
}
for _, rateInterval := range rateIntervals {
for idx := len(prevStats) - 1; idx >= 0; idx-- {
prev := prevStats[idx]
if prev == nil {
continue
}
if stats.UpdatedAt-prev.UpdatedAt >= int64(rateInterval.Seconds()) {
if rate := getNodeStatsRate(append(prevStats[idx:], stats)); rate != nil {
stats.Rates = append(stats.Rates, rate)
}
break
}
}
}
return stats, nil
}
func getNodeStatsRate(statsHistory []*livekit.NodeStats) *livekit.NodeStatsRate {
if len(statsHistory) == 0 {
return nil
}
elapsed := statsHistory[len(statsHistory)-1].UpdatedAt - statsHistory[0].UpdatedAt
if elapsed <= 0 {
return nil
}
// time weighted averages
var cpuLoad, memoryUsed, memoryTotal, memoryLoad float32
for idx := len(statsHistory) - 1; idx > 0; idx-- {
stats := statsHistory[idx]
prevStats := statsHistory[idx-1]
if stats == nil || prevStats == nil {
continue
}
spanElapsed := stats.UpdatedAt - prevStats.UpdatedAt
if spanElapsed <= 0 {
continue
}
cpuLoad += stats.CpuLoad * float32(spanElapsed)
memoryUsed += float32(stats.MemoryUsed) * float32(spanElapsed)
memoryTotal += float32(stats.MemoryTotal) * float32(spanElapsed)
if stats.MemoryTotal > 0 {
memoryLoad += float32(stats.MemoryUsed) / float32(stats.MemoryTotal) * float32(spanElapsed)
}
}
earlier := statsHistory[0]
later := statsHistory[len(statsHistory)-1]
rate := &livekit.NodeStatsRate{
StartedAt: earlier.UpdatedAt,
EndedAt: later.UpdatedAt,
Duration: elapsed,
BytesIn: perSec(earlier.BytesIn, later.BytesIn, elapsed),
BytesOut: perSec(earlier.BytesOut, later.BytesOut, elapsed),
PacketsIn: perSec(earlier.PacketsIn, later.PacketsIn, elapsed),
PacketsOut: perSec(earlier.PacketsOut, later.PacketsOut, elapsed),
RetransmitBytesOut: perSec(earlier.RetransmitBytesOut, later.RetransmitBytesOut, elapsed),
RetransmitPacketsOut: perSec(earlier.RetransmitPacketsOut, later.RetransmitPacketsOut, elapsed),
NackTotal: perSec(earlier.NackTotal, later.NackTotal, elapsed),
ParticipantSignalConnected: perSec(earlier.ParticipantSignalConnected, later.ParticipantSignalConnected, elapsed),
ParticipantSignalFailed: perSec(earlier.ParticipantSignalFailed, later.ParticipantSignalFailed, elapsed),
ParticipantSignalValidationFailed: perSec(earlier.ParticipantSignalValidationFailed, later.ParticipantSignalValidationFailed, elapsed),
ParticipantRtcInit: perSec(earlier.ParticipantRtcInit, later.ParticipantRtcInit, elapsed),
ParticipantRtcConnected: perSec(earlier.ParticipantRtcConnected, later.ParticipantRtcConnected, elapsed),
ParticipantRtcCanceled: perSec(earlier.ParticipantRtcCanceled, later.ParticipantRtcCanceled, elapsed),
ParticipantRtcActive: perSec(earlier.ParticipantRtcActive, later.ParticipantRtcActive, elapsed),
SysPacketsOut: perSec(uint64(earlier.SysPacketsOut), uint64(later.SysPacketsOut), elapsed),
SysPacketsDropped: perSec(uint64(earlier.SysPacketsDropped), uint64(later.SysPacketsDropped), elapsed),
TrackPublishAttempts: perSec(uint64(earlier.NumTrackPublishAttempts), uint64(later.NumTrackPublishAttempts), elapsed),
TrackPublishSuccess: perSec(uint64(earlier.NumTrackPublishSuccess), uint64(later.NumTrackPublishSuccess), elapsed),
TrackPublishCancels: perSec(uint64(earlier.NumTrackPublishCancels), uint64(later.NumTrackPublishCancels), elapsed),
TrackSubscribeAttempts: perSec(uint64(earlier.NumTrackSubscribeAttempts), uint64(later.NumTrackSubscribeAttempts), elapsed),
TrackSubscribeSuccess: perSec(uint64(earlier.NumTrackSubscribeSuccess), uint64(later.NumTrackSubscribeSuccess), elapsed),
TrackSubscribeCancels: perSec(uint64(earlier.NumTrackSubscribeCancels), uint64(later.NumTrackSubscribeCancels), elapsed),
CpuLoad: cpuLoad / float32(elapsed),
MemoryLoad: memoryLoad / float32(elapsed),
MemoryUsed: memoryUsed / float32(elapsed),
MemoryTotal: memoryTotal / float32(elapsed),
}
return rate
}
func perSec(prev, curr uint64, secs int64) float32 {
return float32(curr-prev) / float32(secs)
}
func RecordSignalRequestSuccess() {
promMessageCounter.WithLabelValues("signal", "success", "request").Add(1)
}
func RecordSignalRequestFailure() {
promMessageCounter.WithLabelValues("signal", "failure", "request").Add(1)
}
func RecordSignalResponseSuccess() {
promMessageCounter.WithLabelValues("signal", "success", "response").Add(1)
}
func RecordSignalResponseFailure() {
promMessageCounter.WithLabelValues("signal", "failure", "response").Add(1)
}
func RecordServiceOperationSuccess(op string) {
promServiceOperationCounter.WithLabelValues(op, "success", "").Add(1)
}
func RecordServiceOperationError(op string, error string) {
promServiceOperationCounter.WithLabelValues(op, "error", error).Add(1)
}
func RecordTwirpRequestStatus(service string, method string, statusFamily string, code twirp.ErrorCode) {
promTwirpRequestStatusCounter.WithLabelValues(service, method, statusFamily, string(code)).Add(1)
}
func RecordTwirpRequestLatency(service, method string, duration time.Duration, statusFamily string) {
promTwirpRequestLatency.WithLabelValues(service, method, statusFamily).Observe(float64(duration.Milliseconds()))
}