Repository navigation
Conversation
🦋 Changeset detectedLatest commit: fd44b10 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
|
Nice, can you record some videos to show the impact in different network conditions ? You can share it in slack if that is easier. Thanks. |
|
@xianshijing-lk (2) On both real 4G network and WiFi network, we connect fast enough to avoid this path. |
Nice! I like the numbers. How about the false positive and false negative rates? I tried something similar this summer while I was in China behind a VPN, and sometimes I saw high latency under good network conditions and low latency under poor network conditions. I wouldn't think that is a blocker, but ideally it doesn't happen too frequently. |
Cap x-google-start-bitrate at 1 Mbps for connections that set up within 1.5 s, ramping linearly down to 300 kbps at 3.5 s or slower. Screen share is capped the same way once the cap is below 1 Mbps. Resumes keep their estimator; only a new peer connection (join or full reconnect) is measured.
…ports connect Room.connect returns after the signaling join, so an app that publishes video right away creates its first offer before the primary transport has connected and no setup time is recorded yet. The publisher now knows when the attempt began and uses the time elapsed so far for that offer, a lower bound on the eventual setup time, so it can only lean toward the 1 Mbps ceiling.
4e31c85 to
9ddd7d7
Compare
| if (oldVal != ConnectionState.RESUMING) { | ||
| recordConnectionSetupTime() | ||
| } |
There was a problem hiding this comment.
🟡 Early resume leaves video setup timer running
When signaling drops before initial ICE connection, RESUMING skips recording setup time after recovery. connectStartedAtMs keeps aging, so a later first video offer gets a 300 kbps hint regardless of the recovered link.
Learn more
A soft resume reuses the existing peer connection. If signaling closes during the initial connection, reconnect sets RESUMING while the publisher still holds the timestamp from setConnectStartedAt. On recovery, the connection-state callback skips recording, leaving both engine and publisher without a completed setup time. A later first video offer then uses the ever-growing elapsed time in connectionSetupTimeForOffer, even long after the connection stabilized.
Example: An initial join starts at 0 ms, signaling drops before ICE connects, and soft resume completes at 2 s. A camera first published at 30 s gets the 300 kbps hint rather than a value based on the completed connection.
Recommended fix: On RESUMING → CONNECTED, finalize a pending initial setup timestamp for the existing publisher, while continuing to skip timing for resumes of previously connected sessions.
Was this helpful? React with 👍 or 👎 to provide feedback.
detekt fails configure at its cyclomatic complexity threshold (15) with the extra branch. joinImpl calls configure before it negotiates, so setting the start time there is equally early.
|
Diffuse output: AARJAR |
…1031) * Scale the video start bitrate hint by connection setup time Cap x-google-start-bitrate at 1 Mbps for connections that set up within 1.5 s, ramping linearly down to 300 kbps at 3.5 s or slower. Screen share is capped the same way once the cap is below 1 Mbps. Resumes keep their estimator; only a new peer connection (join or full reconnect) is measured. * Use the elapsed setup time for a video offer created before the transports connect Room.connect returns after the signaling join, so an app that publishes video right away creates its first offer before the primary transport has connected and no setup time is recorded yet. The publisher now knows when the attempt began and uses the time elapsed so far for that offer, a lower bound on the eventual setup time, so it can only lean toward the 1 Mbps ceiling. * Hand the publisher its connect start time from joinImpl, not configure detekt fails configure at its cyclomatic complexity threshold (15) with the extra branch. joinImpl calls configure before it negotiates, so setting the start time there is equally early. (cherry picked from commit f297220)
Fork-local follow-up to livekit#1031 and a one-line extract from upstream livekit#1004. - `connectStartedAtMs` is written on the IO dispatcher in joinImpl and read and cleared on the RTC thread from the PeerConnectionStateListener, so mark it @volatile like the engine's other cross-thread flags. - close() clears the stamp alongside hasPublished so a cancelled attempt does not leave a stale start time behind. - joinImpl also negotiates the publisher when hasPublished is still set. The flag survives a full reconnect (closeResources does not reset it), and the reconnect loop then waits on publisher ICE; without fastPublish nothing produced a publisher offer before that wait, so it stalled for the ICE timeout and retried. Upstream fixed this inside the Data Tracks commit (livekit#1004), which the fork does not take. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Linear: [CLT-3380](https://linear.app/livekit/issue/CLT-3380/scale-the-video-start-bitrate-300-1000-kbps-by-connection-setup-time) Flutter counterpart of [rust-sdks#1467](livekit/rust-sdks#1467) and [client-sdk-android#1031](livekit/client-sdk-android#1031). Builds on #1218, which introduced the connection-level, write-once hint this change scales. ## Problem With #1218 the `x-google-start-bitrate` hint is min(90% of target, 1 Mbps) for camera tracks, and uncapped for screen share, chosen with no information about the network. On a constrained uplink 1 Mbps overshoots: the estimator has to recover, which shows up as a freeze and a retransmit storm over the first 5–15 s. libwebrtc cannot probe the path before a video sender exists, so connection setup time is the only network signal available before the first video offer. ## Change `Engine.connect` times each connection attempt from its start until the primary transport reports connected, and the publisher `Transport` scales the hint's cap by it: ``` cap_kbps = clamp(1000 − (setup_ms − 1500) × 700 / 2000, 300, 1000) ``` - ≤ 1.5 s → 1000 kbps (unchanged for healthy networks); 1.5–3.5 s → linear ramp; ≥ 3.5 s → 300 kbps, libwebrtc's own default starting estimate. - The existing 90%-of-target rule still applies on top, and the hint is still written at the floor rather than skipped. - Screen share stays exempt from both the 1 Mbps ceiling and the ramp, so its hint is 90% of its target however slow the setup. This deliberately differs from Rust and Android, which cap screen share once the ramp drops below 1 Mbps. - The publisher holds the attempt's `Stopwatch`. A video offer created before the primary transport connects (publishing during connect) uses the time elapsed so far, a lower bound on the setup time, so it can only err toward the ceiling. Once connected, the engine stops the stopwatch and the value is fixed. - Only a new peer connection is measured: the initial join and a full reconnect both go through `Engine.connect` and get a fresh stopwatch, and a retry against another region times itself from scratch. Resumes and ICE restarts never call `connect` and keep the estimator they have. The write-once latch from #1218 is unchanged. - Logs `connection setup took N ms` and `Applying x-google-start-bitrate=N kbps (connection setup N ms)` at info level, the same lines as Rust and Android. The measured window matches Android exactly (start of the join until the primary peer connection connects, dual peer connection). Rust uses a single peer connection by default, so its window likely excludes up to one signaling round trip that Flutter and Android include. ## Tests `start_bitrate_test.dart`: the ramp at the measured setup times and anchors (same values as Rust and Android), the 90%-of-target and floor rules under the ramp, screen share left uncapped by the ramp, the setup time applied to the connection-level value, the elapsed-time fallback before setup completes, and an engine-level test that the publisher holds a fixed setup time once connected. `flutter test`: 443 passed. Verified on an Android device against LiveKit Cloud that `connection setup took` is logged on connect.
Linear: CLT-3380
Rust counterpart: livekit/rust-sdks#1467
Builds on #973, which introduced the connection-level 1 Mbps cap this change scales.
Problem
With #973 the
x-google-start-bitratehint is min(90% of target, 1 Mbps) for camera tracks, and uncapped for screen share, chosen with no information about the network. On a constrained uplink 1 Mbps overshoots: the estimator has to recover, which shows up as a freeze and a retransmit storm over the first 5–15 s. On a good network 1 Mbps is right. libwebrtc cannot probe the path before a video sender exists, so the SDK has to pick the seed without a measurement; connection setup time is the only network signal available before the first video offer.Change
RTCEnginetimes each connection attempt from the top ofjoinImpluntil the primary transport reports connected, and hands it to the publisherPeerConnectionTransport, which scales the hint's cap by it:joinImpland build a new publisher, so each times itself from scratch. A resume (RESUMING → CONNECTED) keeps its peer connections and estimator and never records a time. fix: improve initial video quality by setting x-google-start-bitrate for all video codecs #973's write-once latch is unchanged. With no setup time recorded, the cap stays at 1 Mbps.Measurements and validation
Setup-time measurements on shaped links and the camera validation runs (LK2/LK3/LK4 profiles, fixed cap vs ramp) are in CLT-3380 and in livekit/rust-sdks#1467; the anchors are the same as in the Rust SDK. In short: on constrained links (500 kbps and 300 kbps uplinks) the ramp removes the opening overshoot, with 40–90% fewer retransmits and the frame rate held from the first frame; on a 1 Mbps link the two are within run-to-run variance.
Tests
SdpMungingTest: 7 pass (5 existing + 2 new covering the ramp at the measured setup times and anchors, and the screen-share cap).spotlessCheckclean.