Final piece of #20. The EncodeOutcome::SkippedDuplicate variant was
introduced in #19 but its count was invisible — silent Ok(_) arm in
encode_thread_loop. Now exposed as a stat.
Changes:
- stats.rs: PipelineStats gains duplicate_frames_skipped (window delta)
and prev_duplicate_frames_skipped (running total). Snapshot field
added. Display format places it after over_budget (both are counters).
Reset clears window delta but preserves running total (same pattern
as pipewire_dropped).
- state_portal.rs: EncodeThread struct gains duplicate_count:
Arc<AtomicU64>. Cloned for encode_thread_loop, stored for main-thread
reads. encode_thread_loop now explicitly matches SkippedDuplicate and
increments with Ordering::Relaxed. Stats snapshot code reads atomic
and calls set_duplicate_frames_skipped after timing drain.
What this enables:
Diagnosing encoded_fps health. Examples:
- capture_fps=60 encoded_fps=30 duplicate_frames_skipped=30
→ healthy: encoder at 30fps target, 30 frames were true duplicates
- capture_fps=60 encoded_fps=5 duplicate_frames_skipped=0
→ problem: frames not being dedup'd but encoder can't keep up
- capture_fps=1.7 encoded_fps=1.7 duplicate_frames_skipped=0
→ healthy static: low fps because KWin damage-driven delivery
Original #20 issues status:
- 'encoded_fps stuck at ~30': FIXED via #19 (EncodeOutcome enum), now
tracks capture_fps when below 30
- 'filler masking real fps': FIXED via #15/#18 (filler deleted)
- 'duplicate count invisible': FIXED via this commit
- 'unique_encoded_fps / delivered_fps': not implemented, deemed
unnecessary now that the core metrics are trustworthy
Tests:
- cargo build --release: 0 new warnings (19 baseline preserved)
- cargo test: 96 lib + 3 integration, 0 failed
- SAFETY comments preserved verbatim
- 2 files changed, +45/-6 lines
Cleanup pass after the latency investigation concluded:
- state_portal.rs: remove TODO(#19) comment block (17 lines). The Option A
upgrade path was tentatively documented during the #19 fix, but the user
decided to accept current latency behavior. The TODO is now noise.
- state.rs:605: remove 'let _fps = self.args.fps as i64;' dead binding left
over from #25 time_base change. The variable became unused when PTS formula
switched from fps-multiplier to literal 90_000, and was renamed to _fps to
silence the warning. Removing it entirely is cleaner.
No behavior change. Build clean (0 new warnings). All 96+96+3 tests pass.
Final piece of the latency puzzle. Server-side frame_age was 6ms but browser
jitterBufferDelay still spiked to 500+ ms during active periods.
Root cause found via librarian investigation of str0m 0.20 source:
str0m LeakyBucketPacer (active when BWE enabled) limits send rate to
BWE_estimate * 1.1. For a 100KB IDR frame at 8Mbps pacing, the pacer
queue adds ~100ms of send-side latency. Multiple frames stack during
burst encode, causing receiver jitter buffer to grow.
Video streams are PACED BY DEFAULT in str0m (audio is unpaced).
Evidence: str0m/src/streams/send.rs:1116 default unpaced logic.
Fix: call stream_tx.set_unpaced(true) on the video stream when media
is added. BWE remains enabled for bitrate adaptation / TWCC feedback,
but the pacer no longer throttles our video egress.
Why this is the right call for wl-webrtc:
- LAN scenario: dedicated link, no competing flows to smooth against
- Encoder cap (8 Mbps) + VBV buffer (250ms) already provide rate control
- The pacer's smoothing benefit (avoid bursts that compete with TCP) is
irrelevant for our use case
- BWE adaptation still works (we still receive EgressBitrateEstimate events
and adapt encoder bitrate via BitrateCommand::UpdateBitrate)
Verification expectations:
- Active-period jitterBufferDelay: 500+ ms -> < 200 ms
- 'Stop mouse, client continues 10s' symptom: should disappear entirely
- Server-side frame_age: unchanged (~6ms)
- Bitrate/resolution adaptation: unchanged (still uses BWE)
- MP4 mode: unaffected (different code path)
Tests:
- cargo build --release: 0 new warnings (19 baseline preserved)
- cargo test: 96 lib + 3 integration, 0 failed
- SAFETY comments preserved verbatim
- 1 file changed, 7 insertions(+), 1 deletion(-)
Closes the latency investigation started in #23 / #24 / #25.
Add quantitative capture-to-send latency measurement so we can diagnose
remaining latency sources after #24/#25 PTS fixes.
Previously frame_age_p95 was always 0.0ms because the WebRTC code path
never propagated capture timestamps, even though stats.rs already
supported the metric. The infrastructure existed but was disconnected.
Changes:
- avhw.rs: Add capture_time: Instant field to CpuNv12Frame (set when
PipeWire delivers frame) and EncodedH264Frame (propagated through
encode thread via new last_capture_time side-channel on SwEncEncode).
- state_portal.rs: Change sent_gap channel type from Sender<f64> to
Sender<(f64, Option<f64>)> so WebRTC thread can send pre-computed
age_ms = capture_time.elapsed() at the exact send moment (not at
stats drain time, which would inflate the measurement by ~1s).
- stats.rs: record_send_from_thread now accepts Option<f64> age_ms
and pushes to frame_age_ms Vec when Some.
After this commit:
- stats: log lines show real frame_age_p95 / frame_age_max in ms
- Expected range: 5-30ms (import + scale + encode + channel send)
- If much higher: server pipeline has queueing issue
- If low but user still sees latency: confirms bottleneck is network
or browser-side (jitter buffer, decode queue)
Scope notes:
- Only Portal/PipeWire path is instrumented. wlr-screencopy path uses
different code path (EncState, not SwEncState) and will continue to
report frame_age=0.0ms. Adding wlr instrumentation is separate scope.
- This is diagnostic only — does NOT change user-visible behavior.
No encoding, sending, or stats output format changes.
Tests:
- cargo build --release: 0 new warnings (19 baseline preserved)
- cargo test: 96 lib + 3 integration, 0 failed
- SAFETY comments preserved verbatim
- 3 files changed, +39/-10 lines
Refs #20.
Root cause:
After #24 fixed PTS propagation, browser jitter buffer still accumulated
to 10+ seconds during active mouse movement. User reported: stop moving
mouse, client continues showing motion for ~10 seconds.
The encoder time_base was 1/fps (33ms granularity at 30fps). When KWin
delivers frames at 60fps (16.7ms apart), compute_capture_pts integer
math mapped multiple captures to the same tick. The monotonicity guard
then bumped them to sequential ticks (0, 1, 2, 3, ...).
Result: 60 captures in 1 real second produced 60 sequential RTP
timestamps spanning 60 * 33ms = 1.98 seconds of RTP time. Browser
played at RTP rate (half real speed), buffer accumulated.
Math verification from test8 log:
- 2067 frames * 3000 RTP jumps = 68s of active RTP time
- 61 frames * 54000 RTP jumps = 37s of static RTP time
- Total RTP time 105s vs real time 98.6s (7% inflation in short session;
long active sessions amplify to 50%+ inflation matching user-reported
10-second trailing).
Fix (per Oracle round review):
Change WebRTC encoder time_base from 1/fps to 1/90000 (90kHz). This
matches the RTP video clock directly, providing 11us PTS granularity.
Captures 16.7ms apart now produce distinct ticks (~1500 each), no
quantization, RTP timestamps accurately reflect real time.
Oracle-required revisions incorporated:
1. **set_frame_rate alongside time_base** — libx264 infers fps from
time_base when not explicit. With 1/90000 time_base and no explicit
framerate, x264 would assume ~90000fps and VBV rate control would
break. Setting framerate=fps/1 preserves real frame semantics while
using 90kHz PTS precision.
2. **rtp_timestamp_from_pts_ticks returns u64 not u32** — MediaTime::new
takes u64. Returning u32 would truncate at 13.25 hours and create
backwards MediaTime. str0m handles RTP u32 wrap internally; we feed
it full u64.
3. **wlr-screencopy path also updated** — state.rs:605 used fps-based
PTS formula. Changed to 90kHz ticks so wlr path matches Portal path
unit. Without this, wlr-screencopy users would have wrong PTS after
the time_base change.
4. **MP4 path (create_software_h264_muxer) UNCHANGED** — verified at
avhw.rs:1693-1789, keeps 1/fps time_base, no set_frame_rate added.
File output doesn't need real-time PTS.
Implementation:
- src/avhw.rs: WEBRTC_RTP_CLOCK_HZ=90_000 const; create_software_h264_encoder
uses 1/90000 time_base + explicit framerate
- src/state_portal.rs: compute_capture_pts uses WEBRTC_RTP_CLOCK_HZ for
tick conversion (was fps multiplier)
- src/state.rs: wlr PTS formula uses 90_000 (was fps multiplier)
- src/webrtc.rs: rtp_timestamp_from_pts_ticks simplified to identity
function (pts_ticks.max(0) as u64), drop fps parameter; write_h264_frame
signature drops fps (was only used for rtp conversion); 5 unit tests
updated to assert 90kHz identity (0->0, 1500->1500, 90000->90000)
Verification expectations:
- Active period jitterBufferDelay: 1000+ ms -> < 100 ms
- 'Stop mouse, client continues 10s' symptom: should disappear
- Static period behavior: unchanged (was already correct)
- MP4 file output: unchanged
- VBV-constrained IDR sizes: unchanged (framerate explicit preserves
rate control semantics)
- All prior fixes (#19, #23, #15, #18, #24) preserved
Out of scope (Oracle noted, not blocking):
- build_swenc_filter_graph still uses 1/fps time_base at avhw.rs:1603/1620
(semantic mismatch but no functional impact since scale_vaapi passes
PTS integers through)
- Runtime VBV update on bitrate change (separate pre-existing issue)
Tests:
- cargo build --release: 0 new warnings (23 baseline preserved)
- cargo test: 96 lib + 3 integration, 0 failed
- 5 rtp_timestamp_* tests updated for 90kHz identity
- SAFETY comments preserved verbatim
- 4 files changed, +48/-35 lines
Previous commit 079611a claimed to fix#24 but the gating condition was
wrong, making the entire fix dead code:
src/state_portal.rs:467
- let pts = if self.webrtc.is_some() { // ALWAYS false here
+ let pts = if self.webrtc_thread.is_some() { // correct lifecycle check
Why self.webrtc was wrong:
WebRtcState lifecycle in Portal path:
1. StatePortal::new() sets self.webrtc = Some(...) if args.port > 0
2. First frame arrives -> WaitingForFormat branch
3. state_portal.rs:274 does self.webrtc.take() and moves WebRtcState
into the webrtc thread
4. Subsequent frames -> Streaming branch -> handle_pw_frame
5. By this point self.webrtc is None
So my gating check 'if self.webrtc.is_some()' at handle_pw_frame ALWAYS
returned false, and compute_capture_pts was NEVER called. Confirmed by
debug instrumentation showing 0 invocations across a 174s WebRTC session.
Net effect: #24's PTS fix was completely inert. RTP timestamps were
still computed from sequential frame counter (old broken behavior).
User reports of 'latency got worse' were due to other test conditions,
not the dead code.
The correct check is self.webrtc_thread.is_some() because:
- webrtc_thread is set AFTER WebRtcState is moved into it (line 301)
- webrtc_thread stays Some for the entire WebRTC session
- webrtc_thread is None for MP4 mode (no thread spawned)
So this check correctly distinguishes WebRTC mode from MP4 mode at the
point where PTS is computed for each frame in handle_pw_frame.
Lesson learned:
- Oracle review (rounds 1 and 2) verified the design and code structure
but did not catch the lifecycle issue because they reasoned about the
code statically.
- Runtime verification via debug instrumentation was needed to confirm
the function was never called.
- This is why the user ran the test BEFORE I committed - their feedback
that latency got worse was the canary that exposed the dead code.
Verification plan (next user test):
- Run with --stats and confirm rtp= field in write_h264 debug logs
shows VARIABLE jumps (not uniform 3000 increments)
- During static periods (capture_fps < 5), rtp should jump by 30000+
- During active periods (capture_fps > 30), rtp increments may still
look sequential due to 1/fps time_base quantization (acceptable)
- Browser jitter buffer should stabilize at < 500ms
Root cause (3-stage bug found via Oracle round 1+2 review):
Browser WebRTC clients accumulated 2-3 seconds jitter buffer under
damage-driven variable frame rate (KWin Portal/PipeWire). User moved
mouse, saw action 2-3 seconds later on client.
Three coordinated bugs formed a chain that defeated any single-point fix:
1. state_portal.rs:455 used sequential frame counter as PTS instead of
real capture time. (Portal path only — wlr-screencopy already correct.)
2. avhw.rs Channel output sent only Vec<u8>, DISCARDING AVPacket PTS.
Even with correct encoder PTS, timing metadata was thrown away.
3. webrtc.rs:695 computed RTP timestamp as 'frame_number * 90000 / fps'
from a counter, ignoring any real PTS. Browser saw uniform 33ms RTP
spacing regardless of actual 1.6-57fps variable delivery, growing
jitter buffer to compensate for perceived 'network jitter'.
Fix (7-step ordered implementation per Oracle round 2):
1. EncodedH264Frame struct in avhw.rs carries data + pts_ticks
2. FrameOutput::Channel type changed from Sender<Vec<u8>> to
Sender<EncodedH264Frame>; both new_webrtc signatures updated
3. Channel drain in avhw.rs extracts pkt.pts(), normalizes via
'p - start_ts' (mirrors existing Muxer branch logic). Drops
packets with missing PTS instead of silently emitting zero.
4. write_h264_frame signature: frame_number:u64 -> pts_ticks:i64
(both WebRtcState and WebRtcInner layers). Extracted pure function
rtp_timestamp_from_pts_ticks(pts_ticks, fps) with 5 unit tests
covering zero, one-frame, one-second, negative clamp, fps=0.
5. state_portal.rs receiver loop consumes EncodedH264Frame, passes
.data and .pts_ticks to write_h264_frame.
6. state.rs (wlr-screencopy) receiver loop updated for compile
compatibility — its existing real-time PTS computation at state.rs:606
was already correct, now properly propagates through new channel type.
7. state_portal.rs:467 PTS computation GATED on output mode:
- WebRTC branch: compute_capture_pts() uses PipeWire's ns timestamp
(or Instant fallback), normalizes to first-frame-origin, converts
to encoder time_base units with i128 intermediate math, enforces
monotonicity via safe checked_add pattern.
- MP4 branch: KEEPS self.frames_encoded as i64 (sequential counter).
File output does not need real-time PTS; changing it would alter
file playback speed during static periods.
Oracle round 2 critical revisions incorporated:
- Single-point PTS normalization (only in avhw Channel drain), NOT at
source. Avoids double-subtraction with existing Muxer logic.
- MP4 path explicitly preserved — real PTS only applied to WebRTC branch.
- Safe Rust monotonicity guard (no unsafe pointer tricks).
- i64::try_from(ticks_i128) instead of broken i128::try_from(...).unwrap_or(i64::MAX).
- pkt.pts() missing -> log + drop, not silent unwrap_or(0).
- Updates span 4 files (avhw.rs, webrtc.rs, state_portal.rs, state.rs)
because channel type change ripples through both Portal and wlr paths.
Verification expectations:
- Browser jitter buffer should stabilize at 100-500ms (typical) instead
of growing to 2-3 seconds under damage-driven delivery.
- chrome://webrtc-internals: jitterBufferDelay / jitterBufferEmittedCount
ratio should drop significantly.
- Server-side metrics (output_bps, frame rate, IDR size) unchanged.
- MP4 file output (--output mode) behavior unchanged.
Out of scope (deferred):
- frame_age metric fix (Oracle: separate commit to isolate behavioral
change from observability change)
- VFR encoder redesign (1/fps time_base sufficient for this fix)
- MP4 VFR recording (semantic change, separate decision)
- Existing client jitter buffers may not auto-shrink; reconnect may be
required for users with already-accumulated latency
Tests:
- cargo build --release: clean, 0 new warnings (19 pre-existing)
- cargo test: 96 lib + 96 bin + 3 integration, 0 failed
- 5 new RTP unit tests covering edge cases
- SAFETY comments preserved verbatim
- 4 files changed, +157/-27 lines
Paradigm shift: stop treating KWin's damage-driven frame delivery as an
anomaly. Static content = no new frames is correct Wayland behavior.
Root causes (combined #15 + #18):
#15: stall detection threshold was 100ms (max(100ms, 3*frame_interval)).
KWin/PipeWire damage-driven delivery meant normal static periods triggered
WARN 'compositor frame delivery stalled' continuously. Test3 data showed
55% stall rate during active streaming (38 stalls in 69s). User perceived
severe stutter pattern 'cardboard-effect freeze-release-freeze'.
#18: filler mechanism (maybe_send_filler_frame) cloned 3MB NV12 data and
sent to encode thread during stalls. Encode thread hashed Y plane, found
duplicate, skipped via dedup. Net: wasted CPU + channel bandwidth with
zero visual benefit (the dedup path was already catching it).
Oracle review revealed the two issues are causally linked: filler is the
failed response to the false stall alarm. Removing both together is correct.
Fixes (all in state_portal.rs):
1. Remove filler mechanism entirely:
- Remove fields: last_fillable_frame, next_filler_at, filler_frames_sent
- Remove method: maybe_send_filler_frame (49 lines)
- Remove constant: MAX_FILLER_DURATION
- Remove fillable_frame clone cascade in handle_pw_frame (10 lines
of 3MB NV12 cloning per frame, the largest CPU/memory win)
- Remove filler_frames_sent from stats output
2. Redefine stall as idle (Oracle-revised):
- Rename: stall_start -> idle_log_start (semantic clarity)
- Change threshold: 100ms -> 5s (CAPTURE_IDLE_LOG_THRESHOLD)
- Change log level: WARN -> DEBUG
- Change wording: 'compositor frame delivery stalled' ->
'portal capture idle; no damage frames received (normal Wayland behavior)'
- One-shot log per idle episode (not repeated every second)
- Use last_capture_arrival as idle start for accurate elapsed duration
(old code set stall_start=now at first detection, undercounting by
threshold value)
Explicit product decision (Oracle flagged trade-off):
Static-content PLI repair is deferred. When WebRTC client sends PLI
during static content:
- Server sets force_keyframe_pending in encode thread
- Encode thread blocks on input_rx.recv() (no frames coming)
- Client may send more PLIs (all rate-limited by #23 to 1/sec)
- When user interacts -> KWin delivers frame -> encode thread produces IDR
This means during fully static content, client may wait for next damage
to receive keyframe. Acceptable because static content is by definition
unchanged - the last received frame is still visually accurate. If user
reports unacceptable PLI latency on static screens, follow-up with
event-driven one-shot IDR mechanism (Option B per Oracle).
Verification expectation:
- Zero WARN 'stalled' messages during normal session
- Optional DEBUG 'portal capture idle' after 5s of no frames
- Optional DEBUG 'portal capture resumed after idle period' on recovery
- capture_fps will still vary with content activity (this is correct)
- encoded_fps will only count real frames (filler no longer inflates it)
Out of scope:
- Event-driven cached-frame IDR (Option B): follow-up if needed
- PipeWire CursorFromCache negotiation: separate enhancement
- capture_fps expectations documentation: defer to #20 stats rework
Tests:
- cargo build --release: clean, no new warnings
- cargo test: 91 passed + 3 passed + 0 failed
- SAFETY comments preserved verbatim
- Net change: -80 lines (20 insertions, 100 deletions)
Closes#15, closes#18.
Root causes (3, discovered via Oracle review + x264 source verification):
A. PLI storm: WebRtcInner.force_keyframe_to_encode is a single bool with
no rate limit. Client PLI storms (5 PLIs in 845ms observed) produced
5 back-to-back IDRs (~1.5MB burst), swamping the network.
B. Bitrate runaway: BWE feedback had no upper bound. Observed bitrate
escalating from 5 Mbps to 9.9 Mbps in 12 seconds, all applied to
encoder. Combined with A, made each IDR grow 4-5x.
C. VBV effectively disabled (NEW finding not in original issue):
x264's vbv-maxrate and vbv-bufsize x264opts expect kbit/s and kbit
(confirmed via x264 source ratecontrol.c:658-661 which multiplies by
1000 at use site). The code passed bps, making VBV interpret 5.5 Mbps
as 5.5 Gbps (clipped to 2 Gbps). Buffer was 173MB instead of 172KB,
so VBV never constrained anything. This is why IDRs could balloon to
336KB after runtime bitrate increases.
Fixes (3, all in this commit per Oracle review):
Fix 0 (avhw.rs): divide bitrate by 1000 when formatting x264opts string.
Updated existing vbv_x264opts_format and vbv_bufsize_is_quarter_of_maxrate
tests to assert correct kbit/s values. Old tests passed but asserted
wrong values - classic 'tests covered the wrong implementation'.
Fix 1 (webrtc.rs): split keyframe trigger into two paths.
- set_need_keyframe() (internal: connect, resolution change) remains
unthrottled but updates last_forced_keyframe_at timestamp.
- request_keyframe_from_viewer() (external PLI) rate-limited to
Duration::from_secs(1), checked against last_forced_keyframe_at.
- Key insight from Oracle: track ALL keyframe production time, not
just PLI time, to prevent 'connect -> immediate PLI -> duplicate IDR'.
Fix 2 (args.rs + state_portal.rs + avhw.rs):
- New --max-bitrate CLI flag, default 8 Mbps.
- Primary clamp in webrtc_thread_loop (policy layer): clamp BWE via
variable shadowing so all downstream code (bitrate_tx, resolution
adaptation) uses clamped value.
- Defensive guardrail in encode_cpu_frame UpdateBitrate handler:
50 Mbps hard ceiling in case future callers bypass policy layer.
- Per Oracle: flat 8 Mbps default, no auto-scaling
(max(8M, 2*initial) would have failed the observed case).
Verification (82.1s session, 34.7 fps avg vs previous 18.2):
PLI storm absorption:
- 15 PLIs received from viewer
- 10 PLIs throttled (67%)
- 6 IDRs produced total (1 connect + 5 honored)
- 3 distinct PLI storms (625ms, 640ms, 858ms duration) all absorbed
VBV constraint working:
- First IDR: 65KB (was 67KB)
- Largest IDR: 169KB (was 336KB)
- All IDRs under VBV buffer bound of 172KB
- No more runaway IDR growth
Bitrate cap working:
- 5 bitrate updates applied (was 8)
- Peak bitrate: 7.4 Mbps (was 9.9 Mbps)
- 52952 BWE readings clamped (>8 Mbps filtered out)
Overall quality:
- Frame rate: 34.7 fps (was 18.2, 1.9x improvement, exceeds 30 fps target)
- Average bitrate: 4451 kb/s (was 5616, lower and more stable)
Out of scope (deferred to future work):
- Runtime VBV reconfiguration on bitrate change (Fix 3): encoder
recreation is expensive, observe whether VBV mismatch becomes a
quality issue after cap is in place
- Asymmetric BWE filter (Fix 4): rise=15%/fall=5% thresholds; nice
tuning but cap is sufficient for now
- Compositor stalls (#15): still causes some stutter at session end
but no longer compounds into latency
Tests:
- cargo test: 91 passed + 3 passed + 0 failed
- vbv_x264opts_format and vbv_bufsize_is_quarter_of_maxrate updated
and passing with new kbit/s assertions
- SAFETY comments preserved verbatim
- cargo build --release: clean, no new warnings
Root cause (verified via code reading + Oracle design review):
- webrtc_paused IS correctly initialized to true in StatePortal::new()
- encode_cpu_frame DOES respect paused via early return
- BUT encode_thread_loop unconditionally sent timing on every Ok(()),
causing stats.record_encode_thread to tick encoded_frames even for
paused-dropped frames -> phantom encoded_fps=29.7 during idle
- Real waste: handle_pw_frame imports DMA-BUF + VAAPI scale + NV12 clone
+ crossbeam send at 60fps even when no WebRTC client is connected
Fix (Option B - surgical bugfix):
- Add EncodeOutcome enum (Encoded/SkippedPaused/SkippedDisconnected/
SkippedDuplicate) to encode_cpu_frame return type
- encode_thread_loop only reports timing on Ok(Encoded), not on skips
-> encoded_fps naturally stays at 0 during idle (also helps #20)
- handle_pw_frame entry: early return on paused, skipping ALL frame
processing (DMA-BUF import, VAAPI scale, NV12 clone, channel send)
- MP4 mode unchanged (webrtc_paused is None, gate is no-op)
- Side benefit: last_fillable_frame stays None during initial idle,
so maybe_send_filler_frame also early-returns -> no filler waste
during the pre-connect idle window (partial mitigation for #18)
Out of scope (TODO comment added at state_portal.rs:178):
- Encoder still initializes on first PipeWire frame (one-time ~50ms
SwEncEncode::new_webrtc cost). Full deferral (Option A) requires
splitting WebRTC signaling lifecycle from media lifecycle - deferred
until startup cost becomes user-perceptible
- Bitrate formula unchanged (5*W*H*fps/100) - tracked by #21
Verification (34s idle + 60s connected session):
- 0 'skipping duplicate frame' events during idle (was ~30/sec before)
- 0 BWE bitrate updates during idle
- 0 IDR production during idle (was 180 frames into the void)
- First IDR produced 178ms after connect (ForceKeyframe -> IDR in 10ms)
- cargo test transform/fps_limit/backend_detect: 34 passed
- SAFETY comments preserved verbatim
shutdown() fired twice on exit (explicit call in main.rs + Drop impl), causing "Total: N frames" and "StatePortal shutdown complete" to log twice 23us apart. Add `shutdown_started: bool` guard at function entry.
Plain bool (not AtomicBool) because &mut self already grants exclusive access. Guard is set BEFORE cleanup so Drop re-entry during panic unwinding is suppressed.
Includes scripts/test_shutdown_idempotency.sh as a live regression test (requires Wayland session). Verified PASS on KWin: both lines print exactly once.
encode_thread_loop hardcoded sws_us=0 and output_bytes=0, making the
stats dashboard show zeros for those columns. The previous encode_us
was measured around the entire encode_cpu_frame call, mixing sws +
encode work into one bucket.
Add SwEncodeTiming { sws_us, encode_us, output_bytes } to SwEncEncode.
- drain_encoder returns Result<usize> (total encoded bytes), accumulated
before the Muxer/Channel match so branch duplication and multi-packet
drain are both handled correctly. output_bytes = bytes produced by
libavcodec even if downstream delivery drops them.
- encode_cpu_frame resets last_timing to Default at entry (so early
returns from disconnect/pause/dedup never report stale prior values),
measures sws_us around sws_scale, measures encode_us around
avcodec_send_frame + drain, then stores the complete snapshot.
- take_timing() uses mem::take to return and clear in one step.
- flush() ignores the byte count from drain_encoder.
- state_portal encode_thread_loop calls take_timing() for real values.
WebRTC client bandwidth estimate now drives both encoder bitrate and
resolution tier selection, replacing the previous static-target encoder.
- webrtc.rs: enable str0m BWE (seeded at 5 Mbps), surface
EgressBitrateEstimate + KeyframeRequest events, expose
get_bwe_estimate() / set_need_keyframe()
- state_portal.rs: wire bitrate/resolution channels between the WebRTC
thread and the encode thread; tier ladder [1440p, 1080p, 720p] with
downscale at 60% budget and upscale hysteresis (120% sustained 10s)
- avhw.rs: SwEncImport::poll_resolution_commands() rebuilds the import
filter graph on UpdateResolution; SwEncEncode::recreate_encoder()
rebuilds sws/enc_video/yuv_frame atomically; hash_sampled_y_plane()
skips duplicate frames; VBV x264opts cap IDR bursts; H.264 level 4.0
(muxer) / 4.2 (WebRTC)
- state.rs: sync wlr-screencopy GOP to fps*2 max 20 for parity
- fix: drain bitrate_rx + resolution_rx BEFORE the stride check in
encode_cpu_frame() so the new (smaller-stride) frame produced after
a resolution change does not hit the stale (larger) enc_width and
crash the encode thread
- WebRTC GOP widened to fps*2 max 20 (was fps/2 max 10)
- Move WebRTC send to dedicated wl-webrtc-webrtc thread (was inline in main loop)
- Reduce frame_rx 16→1, input_tx 2→1 (drop-on-full), webrtc_tx 32→2
- Recv_timeout 10ms→2ms to reduce pipeline latency
- Fix sent_gap_p95 stats bug: compute gap at actual send time in WebRTC
thread instead of batch-draining at snapshot time (was always 0.0ms)
- High profile via AVCodecContext.profile, veryfast preset, 5x bitrate
- Stats drain via sent_gap channel with record_send_from_thread()
- Shutdown: drop input_tx → join encode → drop webrtc_tx → join webrtc
Add lightweight per-second pipeline statistics for stutter diagnosis:
- --stats CLI flag enables structured stats logging
- PipelineStats tracks capture/encode/send timing with p95/pmax
- FrameTimings records import/scale/transfer/sws/encode per-frame
- StatsSnapshot produces one structured log line per second
The 500 error response previously included the raw error message {e}
in the body, potentially leaking internal implementation details (SDP
parse errors, ICE candidate info) to clients.
The detailed error is already logged server-side via tracing::error!,
so the response body is now a fixed generic string with a proper
HTTP/1.1 status line.
poll_rtc() always returned Ok(false), preventing WebRtcState from
clearing self.inner on disconnect. This leaked the UDP socket, Rtc
instance, and 65KB buffer permanently if the client never reconnected.
Closes#10
shutdown() calls enc.flush() → drain_encoder() → tx.send() on a
crossbeam bounded(32) channel. If the channel is full and the
receiver (webrtc_rx) is alive but not being drained, send() blocks
forever — a self-deadlock since both ends belong to the same struct.
Two-layer fix:
- avhw.rs: replace tx.send() with tx.try_send(); handle Full (drop
frame) and Disconnected (set flag) separately.
- state_portal.rs: drop webrtc_rx before flushing in shutdown() so
try_send returns Disconnected immediately.
Regression tests added for the channel semantics.
- Replace 'let _ = tx.send()' with proper error handling: log warning,
set webrtc_disconnected flag, and break drain loop on SendError
- Add Arc<AtomicBool> webrtc_paused shared between State/StatePortal
and SwEncState, synced from wrtc.is_connected() in poll_webrtc()
- Skip encoding in encode_filtered_frame() when paused or disconnected
- Drain and discard stale channel frames on disconnect
- Resume encoding automatically on WebRTC reconnection
- Add src/webrtc.rs: HTTP signaling server + str0m Sans-IO WebRTC transport
with H.264 Annex-B → RTP packetization and key-frame request handling
- avhw: introduce FrameOutput enum (Muxer | Channel) so SwEncState can
output to either MP4 muxer or crossbeam channel for WebRTC
- cap_portal: support portal session restore tokens (PersistMode::ExplicitlyRevoked)
to skip re-authorization dialog; add --no-persist flag to force fresh dialog
- args: make --output optional when --port is used for WebRTC mode
- state_portal: integrate WebRTC pipeline (encoder channel → RTP forwarding)
with shorter GOP for WebRTC (fps/2, min 10)
- main: redirect tracing to stderr; validate --output or --port required
- Add dependencies: str0m 0.20, serde_json 1, dirs 6
The old FpsLimit compared timestamps between CONSECUTIVE frames.
When PipeWire delivers at 60fps (16ms intervals) and target is 30fps
(33ms min_interval), the gap between consecutive frames is always
16ms < 33ms, so EVERY frame was rejected after the first.
Fix: track last_output_time and compare against that instead of the
previous frame's timestamp. Now frames pass when enough time has
elapsed since the last OUTPUT, not since the last INPUT.
Also adds PipeWire process callback counter logging and frame
diagnostic STATS in state_portal.rs for debugging.
ashpd caches zbus::Connection in a global OnceLock. When check_portal_available()
created a Screencast proxy, the connection was cached there. When the function
returned and its tokio Runtime dropped, the cached connection became dead.
Subsequent setup_portal() calls reused this dead connection and hung forever.
Fix: replace ashpd Screencast proxy with direct zbus D-Bus interface check,
which does not touch the ashpd global connection cache.
Add examples/test_portal.rs for minimal Portal ScreenCast testing.
- Rename PwEvent to PwCtrlEvent, separate frame data into its own channel
- Add null chunk check to prevent crash on malformed PipeWire buffer
- Remove redundant inline comments and signal handlers
- Use try_send for error events to avoid blocking on full channel
- Add Drop impl for StatePortal to flush encoder on drop (bug #2)
- Use enc.take() in shutdown() to prevent double-flush of write_trailer
- Null out data[0] after Box::from_raw recovery to avoid dangling pointer
- Extract compute_pts() for testable PTS calculation
- Add 8 tests: PTS calculation, DRM device resolution, descriptor building
BUG-2 (HIGH): SHM Buffer event caused permanent hang
In the ZwlrScreencopyFrameV1 dispatcher, receiving a SHM Buffer event
left in_flight_surface stuck at AllocQueued forever, preventing
queue_alloc_frame() from requesting new frames.
Fix: treat Buffer as a metadata offer (v3 protocol), wait for
BufferDone to decide failure, and add AllocQueued state guard to
LinuxDmabuf handler.
BUG-3 (MEDIUM): Portal backend picked wrong GPU on multi-GPU systems
state_portal.rs hardcoded /dev/dri/renderD128 then renderD129, which
selects the wrong GPU when PipeWire uses a different device.
Fix: extract find_drm_render_nodes() as shared utility; defer DRM
device selection to first PipeWire frame; test each candidate with
av_hwframe_transfer_data to find the GPU that can actually import
the DMA-BUF frame.
BUG-4 (LOW): VAAPI device context created twice unnecessarily
try_finalize_output() created an AvHwDevCtx stored in EverythingButFmt,
but negotiate_format() discarded it (_hw_device_ctx) and EncState::new
created a new one.
Fix: thread the existing hw_device_ctx through negotiate_format() and
create_encoder() to EncState::new() which reuses it when provided.
pw::init() is guarded by an internal OnceCell (process-global one-shot).
pw::deinit() is unsafe and requires 'only called once per process lifetime
after all PipeWire use has permanently stopped'. Since CapPortal can be
created/destroyed multiple times, calling deinit() from a function-local
scope would prevent re-initialization (OnceCell already consumed) and
violate the unsafe contract.
The 5 early-return error paths in pipewire_thread() that previously
leaked global state are now consistent with the success path — neither
calls pw::deinit(). Process exit reclaims global PipeWire state.
Add comprehensive Chinese documentation comments to cap_portal,
main, and state_portal modules covering architecture, lifecycle,
and data flow for each component.
Add check_portal_available() and check_screencopy_available() to
probe each backend independently before committing. This enables
smarter fallback logic and better diagnostics when no backend is
found. Includes Chinese documentation comments.