Files
voice-cat/docs/voice.md
Talon 8c90e250f0 feat(apple): screen-audio sharing -- macOS ScreenCaptureKit, iOS ReplayKit
Implement system/desktop audio sharing on the Apple clients, feeding the
existing SCREEN_AUDIO Opus -> AEAD -> UDP path via vc_stream_feed_pcm. No
C++/protocol/codec changes -- the core was already ready (the Windows-only
loopback is #ifdef VOICECAT_HAS_LOOPBACK; off Windows the stream just waits
for fed PCM). Audio only; video is dropped.

macOS (in-process):
- ScreenAudioCapture.swift drives an audio-only SCStream
  (excludesCurrentProcessAudio), converts Float32 -> int16 in the channel's
  mono/stereo mode, and calls feedPcm. Capture starts on the self
  .streamStarted event (effective config known then). Wired into
  MainWindowController.screenAudioClicked().

iOS (forward-to-host, single session):
- VoiceCatBroadcast: a ReplayKit Broadcast Upload Extension consumes
  .audioApp only, resamples to 48kHz int16 stereo (AVAudioConverter), and
  writes a shared App Group SPSC ring (BroadcastAudioRing.swift). It does
  not link libvoicecat.
- Host BroadcastAudioPump drains the ring (reacting to the extension's
  Darwin notifications) and feeds the SCREEN_AUDIO stream it owns, downmixing
  to mono when the channel is mono. Screen audio appears as a second stream
  of the same user; no credentials persisted. UI is RPSystemBroadcastPicker
  View in VoiceControlsView. Removes the speculative BroadcastCredentials.

Docs: voice.md s9, CLAUDE.md status, PROGRESS.md.
2026-06-21 00:14:31 +02:00

20 KiB
Raw Blame History

Voice & Media

Real-time audio runs over UDP, secured per security.md. The control channel (TCP/TLS) handles signaling — announcing streams, channel membership, talk state — while UDP carries only the encoded audio frames. This split keeps media latency low and independent of TCP head-of-line blocking.

1. The multi-stream model

A user publishes one or more streams. Each stream is an independent audio source with its own encoder, its own stream_id (unique per user) and ssrc (media-plane id assigned by the server), and is independently mutable/mutable at the receiver.

        User "Alex"                              Receiver "Sam"
   ┌────────────────────┐                   ┌──────────────────────────┐
   │ mic       → enc ───┼──ssrc 1001──▶     │ jitter(1001)→dec→┐        │
   │ desktop   → enc ───┼──ssrc 1002──▶     │ jitter(1002)→dec→┤        │
   │ 2nd mic   → enc ───┼──ssrc 1003──▶     │ jitter(1003)→dec→┴─mix──▶ out
   └────────────────────┘                   └──────────────────────────┘

Stream kinds (v1): MIC, SCREEN_AUDIO (system/desktop audio for listening together), AUX_DEVICE (a second capture device). Receivers can set, per incoming stream: gain, mute, and noise reduction (see §10) — so Sam can turn down Alex's desktop audio while keeping the mic, and independently apply noise suppression to a third user who has a loud fan. The mixer sums all active streams from all users in the channel into the local playback device. All of these receiver-side controls are local to the listener and carry no protocol traffic.

SCREEN_AUDIO capture is platform-specific and covered in §9 — it is supported on Windows, macOS, and iOS (via a ReplayKit broadcast extension).

2. Voice frame format (UDP payload, inside the media AEAD)

A fixed binary header — no protobuf on the RT path. Multi-byte fields are big-endian.

 0      1      2      3      4      5      6      7      8 ...
┌──────┬──────┬──────┬──────┬──────┬──────┬──────┬──────┬───────────────┐
│ type │flags │  codec      │     ssrc (u32)                            │
├──────┴──────┴──────┴──────┼──────┬──────┬──────┬──────┬──────────────┤
│  seq (u16)  │       timestamp (u32, in samples @48k)   │  payload ... │
└─────────────┴──────────────────────────────────────────┴──────────────┘

type    u8   1 = VOICE, 2 = KEEPALIVE, 3 = UDP_BINDING (handshake)
flags   u8   bit0 marker (start of talkspurt) · bit1 FEC-present
             bit2 DTX/comfort-noise · bit3 last-frame-before-stop
codec   u16  0 = OPUS  (room for future codecs)
ssrc    u32  media-plane stream id. Client sends its own ssrc; the server
             validates it against the bound session and relays unchanged.
seq     u16  per-ssrc sequence number, wraps; drives loss detection + reorder
timestamp u32 RTP-style sample clock @48 kHz; drives the jitter buffer
payload      one Opus packet (the encoder's output for one frame)

This is intentionally RTP-shaped (familiar semantics: ssrc/seq/timestamp) without RTP's full machinery. The server relays the payload unmodified — it only reads the header to route by ssrc→channel and may restamp nothing (the client's ssrc is globally unique once assigned at StreamAnnounce). No server-side decode.

Why client-sends-ssrc is safe

The UDP 5-tuple is bound to an authenticated session (protocol.md §4). The server checks that the ssrc in each frame belongs to a stream that session announced; spoofed ssrcs are dropped. So identity is anchored by the session binding + transport encryption, not by trusting the header.

3. Per-channel audio configuration

Opus is configured per channel and pushed to clients in JoinChannelResult.audio / StreamAnnounceResult.effective_audio. All members of a channel encode with mutually decodable parameters.

message AudioConfig {
  uint32 codec = 1;               // 0 = OPUS
  ChannelMode mode = 2;           // MONO / STEREO
  uint32 sample_rate = 3;         // 8000/12000/16000/24000/48000 (48000 recommended)
  uint32 bitrate_bps = 4;         // e.g. 24000 (speech) … 128000 (music/stereo)
  uint32 frame_ms = 5;            // 2.5/5/10/20/40/60 (20 default)
  OpusApplication application = 6;// VOIP / AUDIO / LOWDELAY
  bool   fec = 7;                 // in-band forward error correction
  uint32 expected_packet_loss = 8;// %, tunes FEC aggressiveness
  bool   dtx = 9;                 // discontinuous transmission (silence suppression)
  uint32 complexity = 10;         // 0..10 encoder complexity
}

Guidance baked into defaults / docs:

  • Sample rate: always run Opus at 48 kHz internally. Opus resamples internally anyway; 48 kHz avoids surprises. The sample_rate field mainly constrains capture/narrowband modes for very low bitrate channels. Default 48000.
  • Frame size: 20 ms default. Smaller (10 ms) lowers latency at the cost of more per-packet overhead and CPU; larger (40/60 ms) improves efficiency and loss resilience at the cost of latency. Expose it per channel for "low-latency talk" vs "stable music" rooms.
  • Mode/bitrate: speech channels → MONO, VOIP, 2432 kbps, DTX on, FEC on. Music/screen-audio channels → STEREO, AUDIO, 96128 kbps, DTX off, FEC optional.
  • application: VOIP for talk, AUDIO for music/screen-share, LOWDELAY for monitoring use cases.

4. Packet-loss resilience (Opus 1.6)

Layered, all configurable per channel:

  1. In-band FEC — Opus embeds a low-bitrate copy of the previous frame; the decoder recovers a lost packet from the next one (costs one frame of latency on recovery). Tuned by expected_packet_loss.
  2. PLC (packet loss concealment) — decoder synthesizes a plausible frame for an unrecovered loss; always on, free.
  3. DTX — sender stops transmitting during silence and sends sparse comfort-noise updates; cuts bandwidth and is bandwidth-friendly on busy channels.
  4. DRED (Deep REDundancy, per-channel toggle) — Opus 1.6's ML redundancy: the encoder embeds 20 ms of acoustic features in every packet (bool dred in AudioConfig, off by default). When a packet is lost, the receiver peeks at the next already-buffered packet, parses its DRED extension (opus_dred_parse), and reconstructs the lost frame with opus_decoder_dred_decode — producing significantly better audio than PLC comfort noise for single-frame gaps. Falls back silently to PLC if the next packet has not arrived yet or if the sender did not embed DRED. Heavier CPU on the encoder (~510 % at 24 kbps); minimal overhead on the decoder (parse is a fast header check on non-DRED packets).

5. Jitter buffer

Each receiver keeps an adaptive jitter buffer per ssrc.

  • Frames are inserted by timestamp; playback reads in order at the device callback rate.
  • Target depth adapts to observed network jitter between a configurable min/max latency (channel-level "stability vs latency" knob). A "low-latency" channel runs a shallow buffer; a "stable" channel runs deeper.
  • Late frames past the playout point are dropped; gaps are filled by FEC (if the next frame arrived) or PLC.
  • The marker flag (start of talkspurt) lets the buffer resynchronize cleanly after silence/DTX without accumulating drift.
 incoming (out of order) ──▶ [ reorder by ts | adaptive depth ] ──▶ Opus decode ──▶ mixer
                                     ▲
                             jitter estimate feeds depth

6. UDP keepalive & NAT

  • A KEEPALIVE (type 2) frame flows both directions on the media channel every ~5 s to hold NAT bindings and measure media-path RTT/loss independent of TCP. The frame is plaintext (14-byte header, no payload, no AEAD) — the server identifies the sender by its already-verified UDP endpoint (established during the UdpBinding handshake). On receipt the server bumps the sender's last_seen (so media activity defers the TCP reaper independently of control-channel traffic) and echoes the frame back so the client can measure media-path RTT.
  • If the media path dies but TCP is alive, the client surfaces a "voice disconnected" state and attempts UDP re-binding (re-derive media keys + fresh UdpBinding) without dropping the control session.
  • No ICE/STUN/TURN. The expectation matches TeamSpeak/Mumble: the server is reachable (public IP or port-forward); clients sit behind NAT and initiate, so their bindings are created by their outbound first packet.

7. Talk-state signaling

"Who is talking" can be derived two ways; we use both:

  • Implicit: presence of recent voice frames for an ssrc → that stream is "active". The receiver drives talk indicators from the jitter buffer, so they're accurate and need no extra messages.
  • Explicit (optional): StreamStateUpdate on TCP for coarse UI state (muted, hold) and for users not currently subscribed to the media. Server-side mute/deafen is authoritative and always signaled on TCP.

8. Capture/playback pipeline (inside the core)

 mic device ─(miniaudio capture, 48k, mono)→ resample? → send-side VAD/PTT gate
        → Opus encode → frame header → AEAD → UDP send

 screen audio ─(WASAPI loopback, 48k, mono or stereo per channel mode)→ Opus encode
        → frame header → AEAD → UDP send

 UDP recv → AEAD open → parse header → jitter(ssrc) → Opus decode
         → per-stream recv-side NS (optional, per user) → per-stream gain/mute
         → mixer (sum all ssrc, stereo; mono streams upmixed L=R) → (miniaudio playback,
           48k, stereo) → device
  • Capture and playback run on miniaudio's real-time callbacks (WASAPI / CoreAudio / ALSA). Playback is genuinely stereo end-to-end. Mic capture is mono by default; stereo mic capture is supported via vc_set_capture_channels(stream_id, 2) — when enabled, the capture device opens in stereo (interleaved L/R) and the encoder receives real stereo PCM (no upmix). A mono mic frame on a stereo channel is upmixed L=R before encoding so the Opus bitstream is still spec-correct stereo. Screen-audio (SCREEN_AUDIO) loopback captures in the channel's mode — stereo when the channel is stereo (real interleaved L/R, no downmix), mono when the channel is mono — so a stereo music/screen-share channel gets genuine stereo end-to-end. See §9 for the platform-specific loopback mechanism.
  • iOS mic capture: all iOS audio routing is driven from Swift via AVAudioSession by the IOSAudioRouter singleton before the core (miniaudio) opens its device — miniaudio does NOT touch AVAudioSession on iOS. Input port selection (availableInputs), built-in mic orientation (setPreferredDataSource: front/back/top/bottom), polar patterns (setPreferredPolarPattern: omni/cardioid/subcardioid/bidirectional), mic processing mode (.voiceChat = Standard with AEC/AGC/HPF, or .measurement = Raw/Studio with all processing off), Bluetooth mode (.allowBluetoothHFP HFP voice vs .allowBluetoothA2DP stereo output vs neither), and stereo capture (.stereo polar pattern + setPreferredInput + setInputDataSourcevc_set_capture_channels) are all set from Swift. The core then opens whatever route AVAudioSession has established. When the user changes audio settings mid-session, IOSAudioRouter suspends the core's devices (vc_audio_suspend), reconfigures AVAudioSession, then restarts the devices (vc_audio_restart) so they reopen against the new route — mirroring TeamTalk5's closeSoundDevices/initSoundInputDevice/ initSoundOutputDevice pattern.
  • DSP engine: see §11. The original plan was webrtc-audio-processing (AEC + NS + AGC + VAD in one tuned module, BSD-licensed) — but it has no working Windows/MSVC build upstream (confirmed via its own issue tracker: GCC-only Meson build, MinGW support unfinished, hard abseil-cpp dependency, Linux-tested only — gitlab.freedesktop.org/pulseaudio/webrtc-audio-processing#1). v1 ships a lightweight, dependency-free energy/RMS VAD instead (§11); there is no AEC, NS, or AGC implementation at all yet — not just a deferred VAD, the whole APM is unbuilt. Real webrtc-audio-processing stays a tracked future swap, behind the same ApmProcessor interface (core/src/audio/apm_processor.h), revisit if/when a Linux build target exists or upstream Windows support matures.
  • The mixer sums decoded streams; clipping is handled by soft limiting on the master bus.

10. Noise reduction — two-sided

Noise reduction can be applied at the sender, at the listener, or both — they are independent.

  • Sender-side (the talker's choice): the publishing client runs APM noise suppression on its mic before encoding, controlled by that user's own settings. This cleans the signal for everyone and saves bitrate.
  • Listener-side, per user (the listener's choice): on the receive path, after decoding each stream and before mixing, the listener can enable an additional NS pass on a specific sender's stream. So even if Alex chose not to denoise his mic, Sam can locally suppress Alex's background noise without affecting how anyone else hears Alex.

Implementation: a per-ssrc APM NS instance on the receive path, instantiated lazily only for streams the listener has flagged. State lives entirely on the listener's machine; toggling it is a local UI action with no protocol message and no effect on other listeners. Because each receive stream is decoded independently before the mixer (voice.md §1), per-user receive NS is a clean drop-in on that per-stream stage.

All three receive-side controls (gain, mute, NR) are queryable via vc_get_remote_stream — the counterpart to vc_set_remote_stream — so a UI can reopen its per-stream mix controls at the listener's actual current settings (defaults: gain 1.0, unmuted, NR off). Like the setter, it carries no protocol traffic.

11. Input activation — VAD and PTT (client-configurable)

Whether the mic transmits is decided locally by the input gate, and the client supports both modes, switchable per client (vc_set_input_mode):

  • Voice activation (VAD): v1 implements this as a lightweight, dependency-free energy/RMS-threshold VAD (EnergyVadProcessor, core/src/audio/apm_processor.cpp) — no external DSP dependency, since real webrtc-audio-processing has no working Windows/MSVC build (see §8). It opens the gate when a frame's RMS exceeds a configurable threshold (default ~0.025, normalized to int16 range), with a configurable hang-time (default 300 ms, matching the talk-indicator hangover so "talking" and "gate open" agree) to avoid clipping word tails. DTX naturally complements this — when the gate is closed nothing (or only comfort noise) is sent. This implementation has no AEC — a real limitation versus the originally-planned APM, not just a deferred VAD.
  • Push-to-talk (PTT): vc_set_push_to_talk(active) opens/closes the gate directly. The UI exposes a configurable keybind; the core just receives gate open/close.

Gating applies to the MIC stream onlySCREEN_AUDIO/AUX_DEVICE always bypass it (gating a desktop-audio share on the user's own voice activity would silently drop shared music/video audio whenever the user isn't talking, which defeats the feature).

This is purely a send-side, client-local concern — it gates what gets encoded and sent. It needs no protocol support; remote talk indicators are still derived from the presence of received frames (§7), so they work identically under VAD or PTT.

9. System / screen audio capture (SCREEN_AUDIO)

"Listen together" needs to capture the audio another app is playing. The capture mechanism differs per OS, but it always feeds the same Opus-encode → media-AEAD → UDP path as a normal stream; only the source is platform-specific.

Platform Mechanism Notes
Windows WASAPI loopback capture of the default render endpoint (via miniaudio's loopback mode) Implemented. Captures in the channel's mode — stereo (interleaved L/R) when the channel is stereo, mono when the channel is mono — so a stereo music/screen-share channel gets genuine stereo end-to-end (no downmix). Whole-device capture, not process-specific — it inherently captures this app's own incoming voice mix along with everything else playing (an accepted self-echo-loop characteristic of desktop-audio capture, not a bug). Windows 10 2004+'s process-specific loopback (AUDIOCLIENT_ACTIVATION_PARAMS) would avoid this but miniaudio doesn't expose it — a future enhancement.
macOS ScreenCaptureKit system-audio capture (macOS 13+) Implemented (clients/apple/macOS/VoiceCatMac/Audio/ScreenAudioCapture.swift). OS requires screen-recording permission; capture happens in the main app. An SCStream with capturesAudio + excludesCurrentProcessAudio delivers audio CMSampleBuffers; Swift converts Float32 → int16 (in the channel's mono/stereo mode) and calls vc_stream_feed_pcm — no miniaudio loopback device involved (VOICECAT_HAS_LOOPBACK is Windows-only).
iOS ReplayKit Broadcast Upload Extension (the Discord mechanism) Implemented. See below — separate process, App Group, ~50 MB cap (fine for audio-only).

iOS detail

The extension captures, the host app sends. Unlike a self-connecting extension, this keeps a single session — the screen-audio share appears as a second stream of the same user (exactly like macOS/Windows), and no credentials are ever persisted to disk.

  • The user starts a broadcast from Control Center's screen-record button; we surface it via RPSystemBroadcastPickerView from inside the app (VoiceControlsView) for one-tap start.
  • The Broadcast Upload Extension (clients/apple/iOS/VoiceCatBroadcast/SampleHandler.swift) receives RPSampleBufferType.audioApp (system/app audio), .audioMic, and .video. We consume .audioApp only and drop video + mic — video is what blows the ~50 MB extension memory budget, so an audio-only consumer stays comfortably inside it. The extension does not link libvoicecat.
  • The extension converts each chunk to the core's canonical format (48 kHz int16 stereo, via AVAudioConverter) and writes it into a lock-free single-producer/single-consumer ring in a shared App Group mmap'd file (clients/apple/iOS/Shared/BroadcastAudioRing.swift). It posts Darwin notifications on start/stop so the host reacts promptly.
  • The host app owns the stream: its BroadcastAudioPump announces the SCREEN_AUDIO stream over the control channel (StreamAnnounce), drains the ring, and calls vc_stream_feed_pcm (the external PCM feed API — see architecture.md §4) to drive the Opus encode + AEAD + send path. It downmixes to mono when the channel's effective config is mono.
  • Mic + voice also run in the host app. When the broadcast stops (broadcastFinished), the extension clears the ring's active flag (and posts a Darwin notification); the host stops feeding and emits StreamStop. The host must be alive to relay — always true while in a call (the app declares the audio background mode).