Conversation
🦋 Changeset detectedLatest commit: f696c75 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
size-limit report 📦
|
d87bddc to
ca06b31
Compare
…estination One pipeline per process implements the SPEC backend contract in TypeScript: the Cloud ingest derived from the server URL, the observability grant and exp read from the token, the token re-supplied on every refresh, and the collector-answer table (partial success, 413 split, Retry-After, 404 and owner opt-out silences, 401 holds with refused tokens remembered by fingerprint until they expire, 5xx backoff with jitter, permanent 4xx dropped and counted). Every record is stamped with its owner and its collection decision when captured, and export reads only the stamp. Batches are written to TelemetryStorage before the network is tried and removed once the collector has taken them; a refused write lands in a bounded memory fallback, an unreadable batch is kept for the next pass, a failed deletion is remembered so a delivered batch is never resent, and every loss is counted into a self-report of deltas. The OTLP encoder is a small protobuf writer, so no OpenTelemetry package ships; JWT and bearer values are masked in log bodies and span statuses, and the test-only collector override never carries a token.
One scope per Room owns its session, the lk.connect, lk.reconnect, lk.publish and lk.subscribe spans with their checkpoints and cancelled/failed/reconnect_failed outcomes, join-time subscribe intents, a new connect ending a superseded reconnect, and the Room's custom attributes behind a 64-key limit. Span and window owners are fixed when they open. The poller issues one getStats() per peer connection at the cadence the pipeline sets, bounded per read, isolated per peer, invalidated when a connect is accepted, and re-checks the opt-out before every read and submit. Readings are folded per layer into 60 s windows the way the core does: counters summed across a track's layers, gauges from the window's own readings, windows closed on the pipeline clock, idle layers retired from the same timer.
The browser answers two device signals, page visibility and network information where Chromium exposes it; everything else arrives through the host seam. Any logger obtained through getLogger forwards warn and error records to the pipeline before the console level filters them, with every argument passed through unchanged; the forwarder survives setLevel, resetLevel and rebuild, and a log context carries its Room's scope through a weak reference so a line is attributed by identity, never by name.
disableTelemetry() takes effect when it returns, purges the pipeline's and the configured host's cache, closes every open span and keeps later Rooms silent; it is a TODO pending the token discussion. configureTelemetryHost() is the @internal seam a platform such as React Native uses to supply resource names, a storage and a device observer. Every entry point called from SDK code is guarded, and the guard's own diagnosis cannot throw, so telemetry never surfaces in an app. On a runtime without WeakRef the pipeline starts disabled.
Room installs the pipeline on construction, opens its scope on connect, forwards the join token and every refreshed token, records the required connect checkpoints in order, and ends the connect, reconnect, publish and subscribe spans with the outcome the engine and signal client report. Publishes and republishes share one instrumented path, so a full reconnect's republishes are spans and warnings during a publish carry its id. Room.emitTelemetryEvent() and Room.setTelemetryAttribute() are the two public additions; the local participant reaches its Room's scope through a WeakMap so the emitted declarations expose nothing else.
Unit tests cover the collector-answer table, record ownership across projects and reconnects, cache bounds, fallbacks and losses, delete markers, masking, the opt-out before and after install, the poller gate and generation, window closure and cadence relief, and the per-layer fold. The e2e suite gains the telemetry story in Chromium: two Rooms on livekit-server --dev, mock media, every span lifecycle including failed and cancelled connects and republishes, reconnects, a refreshed token, a custom event and attribute, device events, disconnect and the opt-out, with every batch posted to an OTLP collector. It skips locally when a service is unreachable and fails on CI, where both run as services.
The UMD size limit moves from 130 to 140 kB for the pipeline and encoder; TELEMETRY.md is a short map of the module for maintainers.
c0a5b20 to
f696c75
Compare
| } | ||
| if (!exempt && !this.floodOk()) return; | ||
| record.host ??= owner ? owner.host : this.latestHost; | ||
| this.queue.push({ record, scope, owner }); |
There was a problem hiding this comment.
🔴 Pre-connect events retain abandoned Rooms
When an unconnected Room emits an event, queue keeps its scope until that Room selects a server. The scope's poller callback retains the Room, so abandoned Rooms remain in memory until queue eviction or opt-out.
Learn more
A Room can call emitTelemetryEvent before connect. emit submits that event without a host, so persist keeps it in the queue until the Room connects. The queue entry retains the scope strongly. StatsPoller installs a callback on that scope and holds source closures that capture the Room. An application that drops a never-connected Room therefore cannot release the Room, its participant, or its engine while this entry remains in the bounded queue.
Example: Create a Room, call room.emitTelemetryEvent('opened'), and discard the Room without calling connect. Its event stays unassigned and its scope keeps the entire Room alive, rather than letting it be collected.
Recommended fix: Do not strongly retain unassigned scopes in queued entries. Use a weak owner reference with a cleanup path for orphaned records, while preserving the first-server adoption behavior for live Rooms.
Was this helpful? React with 👍 or 👎 to provide feedback.
| this.setAndEmitConnectionState(ConnectionState.Connected); | ||
| this.telemetry.connectStep('room_connected'); | ||
| this.telemetry.connectEnded('ok'); | ||
| this.statsPoller.start(); |
There was a problem hiding this comment.
🟡 RTC polling continues without an ingest
When a self-hosted or ungranted Room connects, statsPoller.start() still schedules regular and fast getStats() calls. record drops self-hosted samples, while persist discards ungranted batches.
Learn more
The Room starts its telemetry poller after every successful connection, regardless of whether its server has an ingest or its project granted telemetry. A self-hosted scope has collects === false, so its samples are dropped at capture. An ungranted Cloud project's samples are discarded when the queue is persisted. The poller still invokes getStats on each peer every 30 seconds, and every second while awaiting a publication or subscription, alongside the existing stats logger.
Example: A self-hosted Room publishes a microphone. The poller requests stats every second initially, even though no telemetry record from the Room can reach a collector.
Recommended fix: Gate StatsPoller.start and its wake-ups on collection eligibility, and restart polling when a valid grant becomes available if the project's policy permits that transition.
Was this helpful? React with 👍 or 👎 to provide feedback.
| }; | ||
|
|
||
| private onMediaDevicesError = (e: Error, kind?: MediaDeviceKind) => { | ||
| telemetry.captureFailed(e, kind === 'videoinput' ? 'camera' : 'microphone'); |
There was a problem hiding this comment.
🟡 Screen-share failures masquerade as microphone failures
When screen capture fails, sourceToKind(ScreenShare) supplies no device kind to this handler. The fallback records a microphone failure instead, leaving screen-share failures misclassified.
Learn more
The capture error event uses sourceToKind(source) for camera, microphone and screen sharing. sourceToKind returns undefined for screen sharing. This handler then treats that undefined kind as a microphone, so the lk.device.capture.failed event identifies the wrong capture device.
Example: A user denies a screen-share prompt. setScreenShareEnabled(true) emits a media-device error with no kind, and telemetry records lk.device.capture.device=microphone rather than screen_share.
Recommended fix: Carry the original track source or an explicit screen-share device identifier through the capture-error event. Map that value to screen_share without treating every absent kind as microphone.
Was this helpful? React with 👍 or 👎 to provide feedback.
| report = await Promise.race([ | ||
| pending, | ||
| new Promise<undefined>((resolve) => unref(setTimeout(resolve, READ_TIMEOUT_MS))), | ||
| ]); |
There was a problem hiding this comment.
🟡 Successful stats reads leave timers running
When getStats() resolves quickly, its five-second race timer remains active. Fast polling leaves several pending timers per Room, wasting resources until each timer fires.
Learn more
Each getStats call creates a five-second timeout as the other branch of Promise.race. Winning the race does not cancel the losing timeout. During the one-second polling cadence, each peer accumulates up to five active timeout callbacks even if every read completes promptly.
Example: With publisher and subscriber stats resolving in 10 ms, five fast polls create ten timeout callbacks that remain scheduled for up to five seconds instead of being cleared as reads complete.
Recommended fix: Keep the timeout handle and clear it in a finally around each individual stats read, including rejection and timeout cases.
Was this helpful? React with 👍 or 👎 to provide feedback.
| project.granted ??= claims.granted; | ||
| if (project.granted && claims.granted && !claims.expired && !this.refused(project, token)) { | ||
| project.token = token; | ||
| project.expiresAt = claims.expiresAt; | ||
| } |
There was a problem hiding this comment.
| key.length > 0 && utf8(key) <= MAX_KEY_BYTES && !key.startsWith('lk.') && key !== 'session.id' | ||
| ); |
There was a problem hiding this comment.
🟥 Custom attributes overwrite trusted event names
A caller can supply otel.event.name through emitTelemetryEvent or setTelemetryAttribute. The spread of custom attributes overrides the trusted event name, letting caller data impersonate SDK events.
Was this helpful? React with 👍 or 👎 to provide feedback.
|
|
||
| function onLog(level: LogLevel, message: unknown, context?: object, logger?: string) { | ||
| if (pipeline.disabled || level < LogLevel.warn) return; | ||
| const body = mask(String(message)); |
There was a problem hiding this comment.
Client telemetry for JavaScript, implementing the shared contract (livekit/rust-sdks#1396,
livekit-telemetry/SPEC.md) in TypeScript, since a browser cannot link the Rust core. Every Room on LiveKit Cloud reports its spans, RTC statistics, SDK warnings/errors and device state to its own project (only when the token carries the observability grant), for about 2k lines of TypeScript, no new dependency and no new public types;@livekit/react-nativereuses the pipeline through one internal hook.Public API
disableTelemetry()room.emitTelemetryEvent(name, attributes = {})room.setTelemetryAttribute(key, value)nullremovesNothing else is public: the emitted declarations add exactly these three calls plus the
@internalhost seam. The pipeline, transport, storage and device code stay internal;configureTelemetryHost()(resource names, a file-backed cache, device state and events, the collector override for local runs) is exported@internalfor React Native; every tuning value (60 s export, 60 s windows, stats poll interval) is the SPEC default.Platform code
What this platform adds — here, everything: the browser has no Rust core, so destination, token handling, retries, cache, holds, stats mapping and span state are TypeScript too.
Nine files, 2,413 lines; 2,706 lines in total including the wiring in existing files.
Files and responsibilities
telemetry/pipeline.tsRetry-Afterbounded to 24 h, 404 and owner opt-out silences, 401 token retirement, 5xx full-jitter backoff, 10 s export timeout); device state → events, the cadence read live (a CPU-limited encoder stretches at once; one pipeline timer times the relief, and a Room leaving re-derives every remaining Room's deadlines, so relief never waits for a tick), a flush on background;lk.telemetry.reportdeltas riding in the next logs batch; a batch the host's store refuses kept in a bounded memory fallback and counted as a write error; a batch the store cannot read right now kept for the next pass and counted (cache.read_errors); refused tokens remembered by fingerprint (two hashes plus the length, no token material) until they expire and never forgotten before — one withoutexpfor the process; a batch the collector took but the store would not delete is remembered and never sent again from this launch (cache.delete_errors; markers retire only when the store's listing or eviction list shows the batch gone — an unavailable listing reconciles and sends nothing,cache.list_errors— never to make room: at the cap of 1,024, the store-plus-fallback bound, nothing new is sent until a deletion goes through; cleared at opt-out; an uploaded batch leaving the store is never a loss); the collector override never sends a bearer; the synchronous opt-out and purge (nothing is written to the cache afterwards, a late 413 split included; every credential dropped)telemetry/scope.tslk.connect,lk.reconnect(one span per cycle, attempts as checkpoints),lk.publishparenting,lk.subscribe(intent,subscribed,first_media, 30 stimed_out, cancelled on unsubscribe/unpublish/disconnect); RTC windows per track and direction, closed at the 60 s boundary on their own timer (re-derived when the cadence changes, relief included), on background, on an attribute change, at disconnect and when a new connect is accepted; simulcast layers kept per track for cumulative counters only (a vanished layer keeps counting, retired after 3 min idle by the window timer, no reading needed, the timer gone at disconnect), gauges from the window's own readings; required connect checkpoints kept in SPEC order (early_pc_createdfor a peer connection ahead of the join); a reconnect superseded by a new connect ends cancelled; CPU-limited detection; the two app calls with SPEC limits;lk.room.disconnected; nothing accepted after the opt-out (no span event, attribute, subscribe timer, server, token or identity; open spans closed and emptied, identity and credentials dropped); the guard that keeps every call from throwing into the SDKtelemetry/otlp.tsExportLogsServiceRequest/ExportTraceServiceRequest; a field-awarepartial_successreader; the JWT/bearer masktelemetry/transport.tsexp), answer classification,Retry-After, backoff, the gzippedfetchwithPriority: u=7that never follows a redirecttelemetry/webrtc.tsgetStats()report → SPEC readings per layer (codec; RTT from the remote-inbound stream byremoteIdorlocalId, or the selected / nominated candidate pair; gauges unrounded) and the core's fold (cumulative counters summed across every layer seen, limitation durations the maximum; gauges only from the layers in the reading); the min/max/avg windowtelemetry/device.tslk.device.*events, cadence factor, soft holds; the browser observer (visibility, online/offline,navigator.connection); the capture-failure taxonomytelemetry/index.tsdisableTelemetry()(purges the host's cache from earlier launches too); the host hook; SDK warn/error capture (tokens masked), attributed to its Room by identity (the Room's loggers carry their scope in the structured context, non-enumerably); a Room-shaped context with no scope behind it is dropped and counted (dropped.unattributed), never the process's or another Room's; only a line with no Room context is the process's; the guard around every entry pointtelemetry/storage.tsreplacefor splits; a store without it never splits — the batch stays whole and waits for the project's backoff, so React Native's file store must addreplaceto split), the bounded memory store (a lone batch over the bytes goes too), the fallback that keeps a host store's refused batches in memory, a shape check on cached bodies; batch ids carry time, count, kind and projecttelemetry/poller.tsgetStats()per peer connection: every second while a subscribe or a new publication (30 s at most) awaits its first reading; one connection's failing read never costs the other its reading, else every 30 s; the opt-out gate before every read; one generation per poll, invalidated by a stop or by a connect accepted meanwhile, so neither an answer nor a further peer read lands under a later connection; a 5 s read bound; wake-ups only move the deadline closerChanges in existing code
Wiring only: every existing public API, the console logging and the per-track statistics behave as on main. Every telemetry entry point called from SDK code is guarded — the install in the Room constructor, scope creation, every method of the scope and the poller, capture failures: a telemetry failure — in a call or in a scheduled timer — is dropped (a debug line that may itself fail silently) and never skips a connect, publish or teardown step; a unit test with a throwing pipeline proves it.
Changed files
RoomWeakRef— none of the supported browsers, Hermes has it — telemetry is disabled before anything is registered or captured); the Room's loggers carry the scope for warning attribution; hands the scope the URL and token whenconnect()accepts the attempt, on every token refresh and on a room move;setRoomat join, room update and move; onelk.connectper connect —ws_open/signalat the socket upgrade,join_recv,pc_created,pc_connected,room_connected— ended before the disconnect handling when a connect fails or is aborted; subscribe intents for join-time tracks (also after a full reconnect), autoSubscribe publishes andsetSubscribed(true); subscribed / failed / unsubscribed / unpublished; onelk.reconnectper cycle from the engine's resume and restart events; the stats poller runs fromroom_connectedto the disconnect;lk.room.disconnectedwith the protocol reason,reconnect_failedwhen the engine gave up; the two public methodsRTCEngine,SignalClient,events.tsResuming/Restarting(an optional parameter); aSignalOpenedengine event from the WebSocket upgrade (ws_open); the engine's log context carries its Room's scope (looked up lazily)LocalParticipant.publishTracklk.publishspan per attempt, republishes after a full reconnect included, nested under the open connect/reconnect span when any; sid and source set on successlogger.tssetLevel,resetLevelandrebuild); every call reaches the console with its original arguments; a Room's loggers carry the Room's telemetry scope as a non-enumerable, weak context entry, so console objects and log extensions see the same context and a retained context never keeps a Room alivesetLogExtensionE2eeManagerindex.tsdisableTelemetry()public;configureTelemetryHostand its types exported@internalfor React Native.size-limit.cjspackage.jsonvitest.e2e.config.mts,.github/workflows/e2e.yaml/__otlpto the collector (no CORS set-up on the collector); CI startslivekit-serverand an OTLP collector as services; locally the story skips when either is unreachable, on CI a missing service fails the jobTELEMETRY.mdsrc/telemetry/Events
11 of 19 SPEC signals fully covered, 1 partially, 7 skipped (not exposed by browsers — React Native supplies three of them, thermal, low power and memory, through the host hook — or smoke-test only).
error.typeis the thrown error'sname(ConnectionError), or the reason name when the server ended the call (signal_close,reconnect_failed).Event coverage table
lk.connectspan (+ checkpoints)Room.connect:ws_open·signal(stamped together at the socket upgrade) ·join_recv·pc_created(the optionalearly_pc_createdwhen the peer connection precedes the join response, as on Swift) ·pc_connected·room_connected; cancelled bydisconnect(), failed with the error otherwise (both exercised by the story)lk.reconnectspanattempt <n> quick|fullper attempt (a resume that becomes a restart is attempt 2 of the same span); reason from the engine'sReconnectReasonlk.publishspanLocalParticipant.publishTrackand every republish (full reconnect), nested under the open connect/reconnect span; warnings during a publish point at its span (the newest open publish wins over the connect or reconnect it nests under)lk.subscribespansetSubscribed(true); subscribed / failed (SubscriptionErrorname) / unsubscribed / unpublished;first_mediafrom the 1 s polls;timed_outafter 30 slk.rtc.stats.samplegetStats()per peer connection (publisher + subscriber), paced by the scope; simulcast layers folded as the core does (a layer that stops reporting keeps counting); windows close at the 60 s boundary on the pipeline's clock, on background, on an attribute change and at disconnectlk.room.disconnectedDisconnectReason,reconnect_failedwhen the engine gave uplk.telemetry.reportcustom.<name>Room.emitTelemetryEventgetLogger(the SDK'slivekit-*, a host package's too), whatever the console level; JWT-shaped substrings masked; attributed to its Room by identity through its loggers' context (a Room-shaped context with no scope: dropped and counted), else the process; the browser exposes no WebRTC-internal logslk.device.thermal.changedlk.device.low_power.changedlk.device.app_state.changeddocument.visibilityState(hidden = background, which also flushes)lk.device.memory.changedlk.device.network.changednavigator.connection(Chromium): wifi / cell / wired / bluetooth / other,saveDataas constrained, cellular as expensive;online/offline→unavailable; VPN not told apartlk.device.battery.changedlk.device.audio_route.changedlk.device.audio.interruptionlk.device.capture.failedMediaDevicesErrorfrom track creation during publish: theDOMExceptionname → permission_denied / not_found / in_use / otherlk.pingLocal testing
Try it against a local OTel backend (LGTM) in a few minutes: the e2e story runs a whole call in Chromium and prints its Rooms' trace ids.
Commands
Then open http://localhost:3000 → Explore:
lk.connect(with its checkpoints),lk.reconnect,lk.publishon the publisher;lk.connect,lk.subscribe(first_media) on the subscriber.{service_name="livekit-client-js"} | app_call_id="<marker>"for the publisher's stats windows (lk.rtc.stats.sample),lk.room.disconnected,custom.e2e.checkpointand SDK warnings;{service_name="livekit-client-js"} | otel_event_name=~"lk.device.+"for the device events.pnpm test:e2eruns the same story next to the signal-connection suite (CI starts the server and a collector as services, and fails the job if one is missing); locally, without a reachable server or collector, the story skips and prints why.