* sfu: report forwarding latency as p90 instead of the mean
The forwarding-latency metric was the mean transit over all forwarded
packets. A mean is dominated by a few slow outliers, so a handful of
packets stalled on the forward path (e.g. goroutine scheduling latency)
inflated the whole node's reported latency even when nearly every packet
was forwarded promptly.
Report p90 instead: p90 rising means roughly a tenth of forwarded packets
are slow, a broad signal of systemic forwarding load rather than a sparse
tail. To read a percentile over the report window, the mergeable
per-interval summary now keeps a small power-of-two-bucket histogram of
transit instead of running moments (sum, sum-of-squares).
Drop the jitter (transit std dev) gauge: nothing consumed it, and any
spread is derivable from the forward-latency histogram. The protobuf
ForwardJitter field is left in place, now unset, to deprecate separately.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: clamp forwarding percentile to the observed [min, max]
Bucket interpolation assumes a uniform fill, so a single 20ms packet (or
uniform traffic) could report a p90 above every observed sample. Clamp the
interpolated value to the summary's already-tracked min/max, so a quantile
never falls outside the data. Exact for single samples and repeated
identical latencies.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: fix stale p50 comment in the percentile test
The reported metric is p90; the test comment still said p50 replaced the
mean. Reword it to reflect that a percentile, not the mean, is reported.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: quarter-octave forwarding-latency buckets for threshold resolution
Octave buckets are too coarse near the overload thresholds: 300us falls in
[256,512), so a p90 clustered at ~265us and one at ~500us interpolate to the
same value and would trip (or not) identically. Split each octave into four
linear sub-buckets so the two land in different buckets, on the correct side
of the threshold. Min/max clamping alone does not fix this once a node has a
high tail, since its max no longer bounds the interpolation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: interpolate percentiles within each bucket's observed range
Replace the quarter-octave split with plain octave buckets that also carry
the observed [min, max] of their samples, and interpolate a percentile
within that range instead of the bucket's nominal edges. This is exact when
a bucket's samples cluster, so a p90 near an overload threshold that falls
mid-bucket lands on the correct side of it regardless of bucket width -- no
threshold-aware boundaries needed. It subsumes the min/max clamp, since an
estimate can no longer leave the observed samples.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: cover the 500us cluster in the threshold test
Assert milos's full review example exactly: 265us and 500us clusters that
octave-nominal interpolation both read as ~341us now read 265us and 500us.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: make forwardSummary.addSample a pointer receiver
The summary grew from a small moments struct into a per-bucket histogram
(~700 bytes), so the value-receiver addSample copied the whole summary on
every drained sample in the flush loop. Mutate in place instead: ~28ns ->
~2ns per sample in the background fold, no behavior change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* sfu: eighth-octave buckets for accuracy near thresholds
Split each octave into 8 linear sub-buckets (was plain octaves). With the
per-bucket min/max interpolation this reports a tight p90 cluster exactly
even when it sits mid-octave: 850x265us + 150x410us now reads 410us (above a
400us threshold) instead of 395.5us, and a lognormal p90 lands within ~0.2us
of exact. Per-sample add cost is unchanged (~2ns); cost is ~5KB per summary
and a larger but per-report merge.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
getCPUStats has no callers: node CPU load comes from hwstats.CPUStats in
GetNodeStats, and nothing in the tree reads getCPUStats. Only getLoadAvg
is still used.
Remove the function from both the windows and non-windows variants along
with the state it kept. getLoadAvg is untouched, and go-osstat stays a
direct dependency through it.
Co-authored-by: XiaoShao <26596822+xiaoshao9704@users.noreply.github.com>
* Record subscribe stream start time in prometheus.
Adjust for mutes, i. e. take the last unmute time as the start point and
calculate time till the first byte is sent.
* close the tiny window of race
* Prevent long tail sample when publisher glitches.
Thanks to @milos-lk for this.
Publisher restarting would have reset the layer and would have caused a
sample with very high stream start time. We only need to capture when we
do a dummy start or when the state is seeded to a different node upon
migration.
* reduce a diff
* test
* changed the wrong thing, thank you Devin
Newer kernels report qdisc stats via TCA_STATS2 (Stats2) while older
kernels only populate the legacy TCA_STATS attribute (Stats). Whichever
is unused is left nil, so unconditionally dereferencing Stats crashed
with a SIGSEGV during telemetry init. Prefer Stats2 and fall back to
Stats, skipping when both are nil.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Leave out canceled attempts. Should make it easier to do percentages.
Not putting these in node stats yet. Will observe in prom before using
it in node stats.
* Prometheus metric for join latency.
Also including a couple of other failures in the signal connection path
and moving the signal connected to after all that.
Not doing counters for the new signal failure paths. I should not have
done for the other two I added a little while ago also (
validation failure and start participant failure) as those are not
scalable to keep adding to node stats. Will probably remove those two
from node stats later. Can add those counters if they are useful.
* deprecate signal failed counters
* Add prom metrics for peer connectino state.
By direction (PUBLISHER vs SUBSCRIBER) and state ("started" ->
"connected"). This gives a way to track peer connections failing to
finish establishment.
The RTC active count can be useful for primary peer connection, but not
for non-primary. This counter can be used to track any and can generally
be used to understand success/failure rate of peer connection
establishment.
* add a couple of more states
* clean up and avoid duplicate reporting fully established
* staticcheck
* Metrics for participant active, i. e. fully established.
- Egress stub for v2 API
- Fix the participant canceled counter 🤦
- Add active counter -> this is increment when a participant becomes
active, i. e. primary peer connection established. Can be used to
monitor node wise connection establishment issues.
- Add singnalling validation fail counter.
With this, we have
- signalling validation fail
- signalling failed --> this is when the `startSession` fails
- signalling connected -> signalling is succesful and can send back
joinResponse to client
on media connection side
- rtc_init -> start
- rtc_connected -> participant session created (joined)
- rtc_active -> primay peer connection established
- rtc_canceled -> could not proceed with RTC connection due to not being
able to resume.
* signalling counters deps
* revert pion/webrtc to 4.2.12 to get SCTP without interleaving
* go back to pion/webrtc 4.2.11 and sctp 1.9.5
* Forwarding latency measurement tweaks.
- prom transmission type public
- do not measure short term values as it is not used and saves some lock
contention time in packet path potentially. Adding a separate method
for that.
- Change latency/jitter summary reporting to `ns` also to match the
histogram.
* add GetShortStats
* Higher resolution forwarding latency histogram.
Was using the average latency/jitter of last second to populate
forwarding latency/jitter histogram. But, it is too coarse, i. e. the
average value of latency/jitter is very low and those summarised samples
end up in the lowest bucket always.
A few things to address it
- record per packet forwarding latency in histogram
- adjust histogram bins to include smaller values
- Drop jitter histogram
This is a per packet call, but prometheus histogram is supposedly
fast/light weight. Would be good to get better resolution histograms.
Hence doing this. Please let me know if there are performance concerns.
* typo
* one more typo
* Add prom histogram for forwarding latency and jitter.
Using short term stats for histogram.
An example setting is
1s - short term
1m - long term
Using the 1s (short term) data for histogram. In that 1 second, all
packet forwarding latencies are averaged for latency and std. dev. of
the collection is used as jitter.
* try different staticcheck
Currently, the signal requests are counted on media side and signal
responses are counted on controller side. This does not provide the
granularity to check how many response messages each media node is
sending.
Seeing some cases where track subscriptions are slow under load. This
would be good to see if the media node is doing a lot of signal response
messages.
* Rework node stats a bit.
Related protocol PR - https://github.com/livekit/protocol/pull/1023
- Make a config for node stats measurements. Wanted to put the config in
`routing` package, but a circular dependency forced me to put in
config.go
- Make rate calculations explicit, i. e. requested via config.
Previously, it had some odd checks to decide when to calculate rate
and it would have been calculating over different windows.
- Report signal/data channel bytes every 5 seconds to stats collection
module. Previously, it was doing it every 30 seconds and that meant
some windows could have had a large spike
NOTE: Still need to think about this for load calculations as a large
number of participants leaving could flush in a small window and that
could report a large spike in bytes/packets. Maybe need to ignore
signal bytes for load calculation?
* deps
* use default node stats config if given config is nil
* split out node stats into a struct for re-use
* update config
* De-centralize some configs to where they are used.
And make default variables.
Renaming a bit, but these are all internal config and have not been
added to documented config.
* Keep documented config as is.
* test
* typo
* Add counter for pub&sub time metrics
The pub&sub shows large value in migration related case like
muted/disabled migration, the subscription time depends on
the time when publisher unmute the track(sending rtp packet
after migration), add a counter to distinguish since we
can't control the time in such cases and the first subscription
attemps also is more meaningful than those cases.
* Add info log for high publish delay
* Record out-of-packet count/rate in prom.
Adding a field to AnalyticsStream to make this easier to report.
Let me know if adding to AnalyticsStream is not ok.
Will set up a protocol PR if it is okay.
* deps
* speed up track publication
Add metrics for track publication and subscription
Return EnabledCodecs in JoinResponse so client can
choose codec without server side codec fallback
Cache remote webrtc track without AddTrackRequest to
let client send publisher offer before AddTrackRequest response
* go mod
* clean code
Because we aren't able to get CPU count/load info on Windows, they are
stubbed out to return placeholders. This restores compatibility to run
on Windows.