3869 Commits
Author SHA1 Message Date
b0e2d89826 Redact stream keys in UpdateStream API log fields (#4763)
* Redact stream keys in UpdateStream API log fields

UpdateStream add/remove urls were logged raw into the API request log,
including rtmp stream keys and mux/twitch shorthand keys. Redact them
with utils.RedactStreamKey before appending log fields.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Bump protocol for query-value redaction fallback

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 14:19:30 +02:00
Raja SubramanianandGitHub 035bef4111 Log invalid APIKey on API failures. (#4762)
Useful to understand which key is used.
2026-08-14 17:20:10 +05:30
68ecd38c00 Flush sequencer on stream restart; bound frame-integrity loops (#4760)
Flush the downtrack sequencer on stream restart (Resync, ReceiverRestart,
codec change) so NACK retransmissions can't use metadata that no longer
matches the resynced bucket. Add a defensive bounds guard on the RTX and
forward payload slicing.

Cap the PacketHistory and FrameIntegrityChecker catch-up loops to the ring
size so a large sequence/frame-number jump can't drive a big per-packet
iteration count.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-14 17:00:29 +05:30
cnderrauberandGitHub dbe06aa8d1 Experimental WARP (#4649)
* Experimental WARP

* fix panic

* go dep

* stats
2026-08-14 16:26:08 +08:00
Raja SubramanianandGitHub df20578a78 Remove auth token from log/being sent back to client on invalid token… (#4756)
* Remove auth token from log/being sent back to client on invalid token error

* actually remove API key also
2026-08-14 11:51:14 +05:30
7d612428f9 Process NACK retransmissions in a single worker per DownTrack (#4758)
Replace the per-NACK-packet goroutine spawn in DownTrack.handleRTCP with a
single long-lived worker that coalesces pending NACKs into one retransmit
pass. Pending sequence numbers are capped so a high NACK arrival rate cannot
grow goroutine count or memory unboundedly.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-14 11:51:03 +05:30
f72254ba6b Limit API request body size (#4757)
Bound the size of HTTP request bodies on the main API listener so large
messages cannot exhaust memory. Configurable via limit.max_api_request_body_size
(defaults to 10 MiB, 0 disables).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-14 10:36:40 +05:30
Raja SubramanianandGitHub cc6551d617 Check slice length before access in a couple of more places (#4752)
* Check slice length before access in a couple of more places

* min 🤦

* lint
2026-08-13 18:20:37 +05:30
cnderrauberandGitHub 7d556cfefe sample codec payload mismatch error log (#4751) 2026-08-13 19:32:21 +08:00
Raja SubramanianandGitHub 9d676e3a60 Limit number of pending tracks per participant. (#4750)
* Limit number of pending tracks per participant.

Prevents just a signalling connection adding tracks without actually
publishing them growing a large number.

* add to supervisor only if pending track is accepted
2026-08-13 16:48:49 +05:30
Raja SubramanianandGitHub 2561589868 Fail server start up on partial prom config. (#4749)
* Fail server start up on partial prom config.

* tweaking error message a bit
2026-08-13 13:49:17 +05:30
Raja SubramanianandGitHub d51533e25c Make subscription limit log Debugw as it could spam in a large room. (#4748) 2026-08-13 13:33:32 +05:30
Raja SubramanianandGitHub 2cd50a961f Record publish time on participant close for pending tracks. (#4738)
With https://github.com/livekit/livekit/pull/4706, there was a case of
some downstream component taking a long time while lock was held. While
the underlying cause of holding a lock while doing callback was removed
in that PR, to catch such cases, some publish side metric anomaly would
be useful to monitor and alert on.

Adding a publish time record for pending tracks on participant close.
That would inflate the publish time for participants not being able to
publish and can be alerted on as it will spike up the value at node
level and at cluster level if multiple nodes have the issue.
2026-08-12 21:41:52 +05:30
Raja SubramanianandGitHub 35fe831f1d Close web socket connections in all paths. (#4747)
* Close web socket connections in all paths.

There was a leak of WebSocket pingWorker if the initial response write
errored as it did not close the WebSocket connection.

* graceful close
2026-08-12 15:15:38 +05:30
Raja SubramanianandGitHub c432e49c1e Set relay quota per participant at 12 default for dual peer connection + resume scenarios (#4745) 2026-08-12 12:57:16 +05:30
3c6e56232e Add configurable read-message size limit on signalling WebSockets (#4743)
* Add configurable read-message size limit on signalling WebSockets

Set a read limit on both the client-facing (/rtc) and agent worker
WebSocket connections so an oversized frame is rejected by the transport
before being buffered. The limits are operator-tunable via
signal_message_size_limit and agent_signal_message_size_limit, both
defaulting to 2 MiB (0 disables).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Add tests for signalling WebSocket read-message size limit

Cover the configurable signal_message_size_limit added in the prior
commit:
- config: assert both limits default to 2 MiB and that a YAML override
  (including 0 to disable) is parsed correctly.
- full-path integration: a real client connects to /rtc on a single-node
  server and an oversized frame is rejected by the transport with a 1009
  close; a 0 limit leaves the connection unbounded and signalling
  proceeds.

Adds setupSingleNodeTestWithConfig so a single-node server can be started
with config overrides.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Bound decompressed size of signalling WebSocket messages

conn.SetReadLimit only accounts for the compressed bytes read off the
wire, and the client-facing /rtc upgrader negotiates permessage-deflate,
so a small compressed frame could still expand into a much larger buffer
once inflated. Enforce the same limit on the decompressed message by
reading through NextReader + io.LimitReader in WSSignalConnection instead
of the unbounded ReadMessage.

The transport-level SetReadLimit is kept as the cheap wire-level guard;
the new check is the decompressed-size backstop.

Adds NextReader to the WebsocketClient interface (regenerated fake) and a
unit test plus permessage-deflate integration tests covering the
amplification case.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-12 12:07:26 +05:30
cnderrauberandGitHub fad2cc4afe Use request id to make api idempotence on sdk retry (#4694)
* Use request id to make api idempotence on sdk retry

Derive resource id from request id
2026-08-12 09:28:39 +08:00
223587e140 Add per-participant concurrent TURN allocation quota (#4744)
* Add per-participant concurrent TURN allocation quota

The embedded TURN server authenticated each Allocate request but placed
no cap on how many relay allocations a single participant credential
could hold. One participant could reuse its credential across many client
5-tuples and open one relay socket/port per request, exhausting the
shared relay-port range for everyone else.

Add a configurable per-participant limit (turn.per_user_relay_allocation_limit,
default 4) wired to Pion's QuotaHandler, keyed by the participant ID from
HandleAuth. Slots are reserved before allocation and released when the
allocation ends, under a single lock, so concurrent Allocate bursts cannot
race past the limit; reservations are keyed by source address so retransmits
are idempotent. Over-quota requests receive 486 (Allocation Quota Reached).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Reclaim unconfirmed TURN allocation reservations

Allow reserved a quota slot before Pion built the relay, but the slot was
only released on the allocation-deleted event. An Allocate that passed the
quota check and then failed to create a relay (e.g. relay-port range
exhausted) emits no event, so the reservation leaked: after enough failures
a participant could lock itself out with 486, and the tracking map grew
without bound.

Reservations now start pending and are confirmed on allocation-created; each
pending reservation carries a reclaim timer that frees the slot after a TTL,
so a failed attempt cannot hold a slot forever while concurrent-burst safety
is preserved.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Make TURN reservation reclaim identity-aware

The reclaim timer captured only userID+key. Because Timer.Stop cannot cancel
a callback that has already fired and is waiting on the lock, a stale timer
could delete a replacement reservation created for the same userID+key after
the original was released, leaving a live allocation untracked and letting the
participant exceed its cap.

reclaimPending now captures the slot pointer and only removes the entry when
the map still holds that exact slot, so a stale timer is a no-op.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-11 23:12:07 +05:30
Raja SubramanianandGitHub 7167f91493 Validate TURN config to guard against invalid values (#4742) 2026-08-11 20:34:49 +05:30
c4c356f6ca Cover a couple of more cases on data track runt packet handling. (#4741)
* Cover a couple of more cases on data track runt packet handling.

* Guard data track header parser against extensions-size integer wraparound.

Widen the extensions-size arithmetic to int so a 0xFFFF wire value no
longer wraps in uint16, and reject any packet whose computed hdrSize
exceeds the buffer before slicing the payload.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-11 18:25:10 +05:30
Raja SubramanianandGitHub d279899b7c Fix publish track count on migration in. (#4740)
* Fix publish track count on migration in.

https://github.com/livekit/livekit/pull/4707 addressed the case of
publish tracks overcounting due to synthesised track publish on migrate
in. But, it introduced an issue where published tracks count could go
negative because unpublish subtracted the counter irrespective of the
track actually migrated in or not.

Fix it by keeping track of local publish.

Also, the older code was skipping publisher track count increase if the
synthesised publish was handled first. Address it by checking if the
track is actually new (i. e. fresh local publish) when the track was
already created in the migrate in path.

* fix pub time for tracks published after migration

* test

* prevent multiple track egresses
2026-08-11 16:13:57 +05:30
Raja SubramanianandGitHub 7f1c175a38 Check layer value in dependency descriptor and keep it in bounds. (#4739) 2026-08-11 12:04:19 +05:30
renovate[bot]GitHubrenovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
32368f79d2 Update actions/setup-go action to v7 (#4720)
Generated by renovateBot

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-08-09 00:44:18 -07:00
Kuba PodgórskiandGitHub 335990afa5 return psrpc.FailedPrecondition for "participant client version does not support moving" error (#4736) 2026-08-09 00:42:24 -07:00
4e921aa1b6 Expand room details in webhook events (#4730)
* pass room proto directly to telemetry events

* Keep telemetry analytics events on a minimal room, gate full room in webhooks

---------

Co-authored-by: Simon Beeli <simon.beeli@gmx.ch>
2026-08-07 15:29:41 +05:30
cnderrauberandGitHub 8e6077221c Return incompatible in SetCodecWithState if the codec PT changed (#4729) 2026-08-06 15:42:06 +08:00
Raja SubramanianandGitHub 3f9cf6bfc2 Do not report end time for participant if the participant is migrating (#4728)
out.
2026-08-06 00:59:37 +05:30
Raja SubramanianandGitHub 52ef3cd649 Include data track susbcriptions in WaitForSubscription. (#4727) 2026-08-05 13:28:48 +05:30
Nikita DavydovandGitHub 2a9bb36ee0 Apply ICE preference when switching to TCP on unstable UDP (#4703)
onMediaLossUpdate notified the participant handler directly, which only
sends a leave request with resume action. handleConnectionFailed that
actually switches the ICE preference to TCP/TLS was never called on
this path, so the client reconnected over UDP again and the fallback
kept firing every 30-60s without ever migrating.

Fixes livekit/livekit#4702
2026-08-04 13:04:36 +05:30
cnderrauberandGitHub 4d177cb01b Remove H.264 baseline (42001f) from default enabled codecs (#4723)
Users can explicitly enable this profile if they are certain that all device support it.
2026-08-04 11:12:42 +08:00
Raja SubramanianandGitHub 28468035d2 Check for pictureID existence in VP8 and VP9 (#4721)
* Check for pictureID existence in VP8 and VP9

* test
2026-08-03 15:22:59 +05:30
Felix-AyushandGitHub 4618b63eb7 Fix AgentHandler.DrainConnections deadlock on worker close. (#4710)
Snapshot workers and release h.mu before Close so HandleConnection can deregister without blocking drain.
2026-08-03 14:49:23 +05:30
3b9f118327 Release v1.13.5. (#4715)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v1.13.5
2026-07-31 12:28:21 +05:30
Raja SubramanianandGitHub ced94b8645 log high stream start latency. (#4714)
* log high stream start latency.

There is something wrong in measurement as audio is showing high p99
latency. Must be misattributing samples. So, logging for high latency to
understand this better.

* use correct variable

* time since create
2026-07-30 17:10:59 +05:30
Felix-AyushandGitHub 0759280cd0 Fix getRefLayerRTPTimestamp off-by-one that can panic on max layer index (#4712)
* Fix getRefLayerRTPTimestamp off-by-one that can panic on max layer index.

Reject ref/target layers with >= len(refInfos) so layer==len is an error instead of an out-of-range index.

* Remove historical comment from ref-layer bounds test.

Keep the regression coverage without referencing the old bounds check in source.
2026-07-30 15:24:50 +05:30
cnderrauberandGitHub 436a0cc3d3 Support more h264 profiles (#4708)
* Register h264 main profile if enabled explicitly

We don't support the h264 main profile for compatibility,
user can enabled it by set fmtp explicitly in codec config
to enable it if want to use it in special scenario.

* go mod
2026-07-28 16:18:19 +08:00
Raja SubramanianandGitHub b8a073cb68 A bit better counting for track publish. (#4707)
- Count a publish attempt on a migrating in tarck as there is no
  AddTrack for that.
- Add cancel publish only if the participant connection is canceled
- Do not add publish counter for synthetic publish attempts which
  happens for migrating in tracks. It will be counted on migrating in
  node when the track is actually published, i. e. negotiated/packets
  flowing.
2026-07-28 02:18:14 +05:30
Raja SubramanianandGitHub 25e3774cbb Do not call telemetry listener under pending track lock. (#4706)
* Do not call telemetry listener under pending track lock.

Fix the TrackPublishRequested call of telemetry listener.

Audited other callbacks to ensure that it is not under lock.

* missed some paths of recording it, thanks Devin
2026-07-27 23:52:21 +05:30
Raja SubramanianandGitHub 0aa296e126 Record subscribe stream start time in prometheus. (#4704)
* Record subscribe stream start time in prometheus.

Adjust for mutes, i. e. take the last unmute time as the start point and
calculate time till the first byte is sent.

* close the tiny window of race

* Prevent long tail sample when publisher glitches.

Thanks to @milos-lk for this.

Publisher restarting would have reset the layer and would have caused a
sample with very high stream start time. We only need to capture when we
do a dummy start or when the state is seeded to a different node upon
migration.

* reduce a diff

* test

* changed the wrong thing, thank you Devin
2026-07-27 20:37:44 +05:30
Raja SubramanianandGitHub f1e2eee2fe protocol update with webhook status logging (#4700) 2026-07-27 01:57:10 +05:30
Raja SubramanianandGitHub dfd3a3c4a5 protocol deps for logging webhook status (#4699) 2026-07-27 01:00:51 +05:30
Raja SubramanianandGitHub 1bbd4702b6 Spelling fixes (#4698) 2026-07-26 22:26:41 +05:30
renovate[bot]GitHubrenovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
a0d6e72017 Update go deps (#4645)
Generated by renovateBot

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-07-25 23:23:56 -07:00
He ChenandGitHub 93422be0c5 TELCU-1: send resolved ringing timeout (#4697) 2026-07-23 15:29:24 -07:00
Raja SubramanianandGitHub fc2fe3faf5 Use simulcast constructor for VP9 if simulcasted. (#4696)
Not active path, but noticed it while reviewing a bug report. Just
fixing it for correctness in code.
2026-07-23 13:15:01 +05:30
Ninad PundalikandGitHub afb9142c1e Add status code for twirp request latency prometheus metric (#4621)
* Add status code for twirp request latency prometheus metric

* Drop the twirp error codes as that adds a lot of coardinality
2026-07-22 15:52:30 +05:30
Raja SubramanianandGitHub 13e4aaec2b Tests for down stream packet push. (#4692)
* Tests for down stream packet push.

A recent issue (padding bit in RTP header) surfaced a gap which slipped
through due to lack of tests. Changes in pion/rtp were not adopted
properly.

So, adding some tests (thank you Claude for the heavy lifting) to test
the down stream packet path using the whole pion chain.

Split out some interfaces so it is easier to have it all in one place
and create fakes.

Will help adding more tests, for example include the upstream path also
in the integration test. May have to create more interfaces and make
things testable, but this is a start.

* missed file

* rtx specific test
2026-07-20 20:15:07 +05:30
Raja SubramanianandGitHub c684997c4f Add country to participant closing log (#4693) 2026-07-20 15:32:42 +05:30
David ZhaoandGitHub 3f59d0dd9e add mock for testing region pins with API (#4691)
we'll ensure that clients support redirection when making an API
to an unpinned region
2026-07-19 21:12:35 -07:00
Raja SubramanianandGitHub 366cadcd96 Fix padding bit in forwarded packet. (#4690)
Addresses https://github.com/livekit/livekit/issues/4689

We were probably missing a couple of bits with this
1. Not affected for regular traffic like from browsers as it does not
   add padding, but special clients were affected.
2. Probe packets were probably using wrong last byte as pion/rtp would
   have overwrriten with 0 because the header.PaddingSize for the newer
   versions were not set. That could have affected bandwidth estimation
   catch up.
2026-07-19 20:53:28 +05:30