Raja SubramanianandClaude Opus 4.8 6ed445bad1 sfu: report forwarding latency as p90 instead of the mean (#4920)
* sfu: report forwarding latency as p90 instead of the mean

The forwarding-latency metric was the mean transit over all forwarded
packets. A mean is dominated by a few slow outliers, so a handful of
packets stalled on the forward path (e.g. goroutine scheduling latency)
inflated the whole node's reported latency even when nearly every packet
was forwarded promptly.

Report p90 instead: p90 rising means roughly a tenth of forwarded packets
are slow, a broad signal of systemic forwarding load rather than a sparse
tail. To read a percentile over the report window, the mergeable
per-interval summary now keeps a small power-of-two-bucket histogram of
transit instead of running moments (sum, sum-of-squares).

Drop the jitter (transit std dev) gauge: nothing consumed it, and any
spread is derivable from the forward-latency histogram. The protobuf
ForwardJitter field is left in place, now unset, to deprecate separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: clamp forwarding percentile to the observed [min, max]

Bucket interpolation assumes a uniform fill, so a single 20ms packet (or
uniform traffic) could report a p90 above every observed sample. Clamp the
interpolated value to the summary's already-tracked min/max, so a quantile
never falls outside the data. Exact for single samples and repeated
identical latencies.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: fix stale p50 comment in the percentile test

The reported metric is p90; the test comment still said p50 replaced the
mean. Reword it to reflect that a percentile, not the mean, is reported.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: quarter-octave forwarding-latency buckets for threshold resolution

Octave buckets are too coarse near the overload thresholds: 300us falls in
[256,512), so a p90 clustered at ~265us and one at ~500us interpolate to the
same value and would trip (or not) identically. Split each octave into four
linear sub-buckets so the two land in different buckets, on the correct side
of the threshold. Min/max clamping alone does not fix this once a node has a
high tail, since its max no longer bounds the interpolation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: interpolate percentiles within each bucket's observed range

Replace the quarter-octave split with plain octave buckets that also carry
the observed [min, max] of their samples, and interpolate a percentile
within that range instead of the bucket's nominal edges. This is exact when
a bucket's samples cluster, so a p90 near an overload threshold that falls
mid-bucket lands on the correct side of it regardless of bucket width -- no
threshold-aware boundaries needed. It subsumes the min/max clamp, since an
estimate can no longer leave the observed samples.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: cover the 500us cluster in the threshold test

Assert milos's full review example exactly: 265us and 500us clusters that
octave-nominal interpolation both read as ~341us now read 265us and 500us.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: make forwardSummary.addSample a pointer receiver

The summary grew from a small moments struct into a per-bucket histogram
(~700 bytes), so the value-receiver addSample copied the whole summary on
every drained sample in the flush loop. Mutate in place instead: ~28ns ->
~2ns per sample in the background fold, no behavior change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* sfu: eighth-octave buckets for accuracy near thresholds

Split each octave into 8 linear sub-buckets (was plain octaves). With the
per-bucket min/max interpolation this reports a tight p90 cluster exactly
even when it sits mid-octave: 850x265us + 150x410us now reads 410us (above a
400us threshold) instead of 395.5us, and a lognormal p90 lands within ~0.2us
of exact. Per-sample add cost is unchanged (~2ns); cost is ~5KB per summary
and a larger but per-report merge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-29 01:39:05 +05:30
2026-09-16 22:11:45 -07:00
2023-01-11 14:49:50 -07:00
2026-09-14 10:20:40 +05:30
2026-09-14 10:20:40 +05:30
2026-09-28 16:39:55 +08:00
2026-09-28 16:39:55 +08:00
2021-06-03 23:22:19 -07:00
2023-07-27 16:43:19 -07:00
2026-09-08 15:48:01 -04:00

The LiveKit icon, the name of the repository and some sample code in the background.

LiveKit: Realtime infrastructure for voice, video, and AI agents

LiveKit is an open source platform for building voice, video, and physical AI agents. This repository is the LiveKit server: a scalable, distributed WebRTC SFU that moves realtime audio, video, and data between people, devices, and AI models. The SDKs, agents frameworks, and companion services are linked in the table at the bottom of this page.

LiveKit's server is written in Go, using the awesome Pion WebRTC implementation.

GitHub stars Slack community Twitter Follow Ask DeepWiki GitHub release (latest SemVer) GitHub Workflow Status License

Important

If you're building Voice AI, LiveKit Agents is the SDK for code-first realtime voice agents. STT, LLM, TTS, turn detection, expressive speech, keyterm accuracy, tool usage, and telephony all come bundled in the framework. It's available in both Python and Node.js.

# agent.py
from livekit import agents
from livekit.agents import Agent, AgentServer, AgentSession, STTContextOptions, TurnHandlingOptions, inference

server = AgentServer()


@server.rtc_session(agent_name="my-agent")
async def my_agent(ctx: agents.JobContext):
    session = AgentSession(
        stt=inference.STT(model="deepgram/nova-3", language="multi"),
        llm=inference.LLM(model="google/gemma-4-31b-it"),
        tts=inference.TTS(model="inworld/inworld-tts-2", voice="Ashley"),
        turn_handling=TurnHandlingOptions(turn_detection=inference.TurnDetector()),
        stt_context_options=STTContextOptions(keyterms=["LiveKit", "Acme Corp"]),
        expressive=True,
    )
    await session.start(room=ctx.room, agent=Agent(instructions="You are a helpful voice AI assistant."))
    await session.generate_reply(instructions="Greet the user and offer your assistance.")


if __name__ == "__main__":
    agents.cli.run_app(server)

Models come from LiveKit Inference with no per-provider API keys, and LiveKit Cloud handles deployment and observability. Visit the docs for more info at docs.livekit.io/agents.

Used in production by

LiveKit carries billions of calls a year for companies including Salesforce, Nvidia, Oracle, SAP, Deutsche Telekom, Spotify, Tinder, Coursera, Headspace, Skydio, Retell, Decagon, Cresta, and HeyGen. Read how Assort Health, Playback, and Polymath Robotics use it, or see more customers.

Features

Documentation & Guides

https://docs.livekit.io

Working with a coding agent? Give it the LiveKit Docs MCP server, or start with the coding agents guide.

Live Demos

Install

Tip

We recommend installing LiveKit CLI along with the server. It lets you access server APIs, create tokens, generate test traffic, and scaffold and deploy agents.

The following will install LiveKit's media server:

MacOS

brew install livekit

Linux

curl -sSL https://get.livekit.io | bash

Windows

Download the latest release here

Getting Started

Starting LiveKit

Start LiveKit in development mode by running livekit-server --dev. It'll use a placeholder API key/secret pair.

API Key: devkey
API Secret: secret

To customize your setup for production, refer to our deployment docs

Creating access token

A user connecting to a LiveKit room requires an access token. Access tokens (JWT) encode the user's identity and the room permissions they've been granted. You can generate a token with our CLI:

lk token create \
    --api-key devkey --api-secret secret \
    --join --room my-first-room --identity user1 \
    --valid-for 24h

Test with example app

Head over to our example app and enter a generated token to connect to your LiveKit server.

Once connected, your video and audio are now being published to your new LiveKit instance!

Simulating a test publisher

lk room join \
    --url ws://localhost:7880 \
    --api-key devkey --api-secret secret \
    --identity bot-user1 \
    --publish-demo \
    my-first-room

This command publishes a looped demo video to a room. Due to how the video clip was encoded (keyframes every 3s), there's a slight delay before the browser has sufficient data to begin rendering frames. This is an artifact of the simulation.

Adding an agent

Agents join rooms as participants, the same way a browser or a phone does. Follow the Voice AI quickstart to build one. An agent connects to a self-hosted server the same way it connects to LiveKit Cloud; when running without Cloud, use model plugins in place of LiveKit Inference.

Deployment

Use LiveKit Cloud

LiveKit Cloud is the fastest and most reliable way to run LiveKit. It runs in 19+ regions with 99.99% uptime and adds agent hosting, model inference, telephony, and observability on top of the server. The Build plan is free, with no credit card required.

Sign up for LiveKit Cloud.

Self-host

Read our deployment docs for more information. Official Docker images and Helm charts are available.

Building from source

Pre-requisites:

  • Go 1.26+ is installed
  • GOPATH/bin is in your PATH

Then run

git clone https://github.com/livekit/livekit
cd livekit
./bootstrap.sh
mage

Contributing

We welcome your contributions toward improving LiveKit! Please join us on Slack or in the Developer Community to discuss your ideas and/or PRs.

License

LiveKit server is licensed under Apache License v2.0.


LiveKit Ecosystem
Agents SDKsPython · Node.js
LiveKit SDKsBrowser · Swift · Android · Flutter · React Native · Rust · Node.js · Python · Unity · Unity (WebGL) · ESP32 · C++
Starter AppsPython Agent · TypeScript Agent · React App · SwiftUI App · Android App · Flutter App · React Native App · Web Embed
UI ComponentsReact · Android Compose · SwiftUI · Flutter
Server APIsNode.js · Golang · Ruby · Java/Kotlin · Python · Rust · PHP (community) · .NET (community)
ResourcesDocs · Docs MCP Server · CLI · LiveKit Cloud
LiveKit Server OSSLiveKit server · Egress · Ingress · SIP
CommunityDeveloper Community · Slack · X · YouTube

S
Description
No description provided
Readme Apache-2.0
204 MiB
Languages
Go 99.7%
Shell 0.2%
Dockerfile 0.1%