feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as out of scope: - Device enumeration (vc_list_devices) + input device selection (vc_set_input_device), backed by AudioEngine::enumerate_devices() via miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded ma_device_id strings. - VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk). webrtc-audio-processing (the originally-planned APM) has no working Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard abseil-cpp dependency), so VAD is a new lightweight, dependency-free energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it. - True stereo playback: AudioEngine's mixer and output device now carry stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded stereo streams to mono before mixing. - Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via miniaudio's loopback device type), replacing test-only injection as the production capture path. Also: vccli gains --list-devices, --input-device, --input-mode, and --share-screen-audio flags, plus a stdin command loop (ptt on/off, mode vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all four items (ABI-level + a white-box AudioEngine stereo-mix check). Docs updated to match: voice.md, roadmap.md (decision-log entry superseding the original webrtc-audio-processing choice), tech-stack.md, README.md, architecture.md, CLAUDE.md, PROGRESS.md. Still explicitly out of scope, documented not silently dropped: real webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS SCREEN_AUDIO capture, process-specific loopback, and a pre-existing RT-thread rule violation in the capture path that predates this work. Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive standalone runs; manually verified live (vccli --list-devices against real hardware, vccli --voice --input-mode vad streaming without incident). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -159,20 +159,27 @@ Each receiver keeps an **adaptive jitter buffer per ssrc**.
|
||||
## 8. Capture/playback pipeline (inside the core)
|
||||
|
||||
```
|
||||
device ─(miniaudio capture, 48k)→ resample? → send-side APM
|
||||
(AEC + NS + AGC + VAD/PTT gate) → Opus encode → frame header → AEAD → UDP send
|
||||
device ─(miniaudio capture, 48k, mono)→ resample? → send-side VAD/PTT gate
|
||||
→ Opus encode → frame header → AEAD → UDP send
|
||||
|
||||
UDP recv → AEAD open → parse header → jitter(ssrc) → Opus decode
|
||||
→ per-stream recv-side NS (optional, per user) → per-stream gain/mute
|
||||
→ mixer (sum all ssrc) → (miniaudio playback, 48k) → device
|
||||
→ mixer (sum all ssrc, stereo; mono streams upmixed L=R) → (miniaudio playback,
|
||||
48k, stereo) → device
|
||||
```
|
||||
|
||||
- Capture and playback run on miniaudio's real-time callbacks (WASAPI / CoreAudio / ALSA).
|
||||
- **DSP engine: webrtc-audio-processing (APM)** — the "better one". It provides
|
||||
high-quality **AEC** (acoustic echo cancellation, essential for speaker users), **noise
|
||||
suppression**, **AGC**, and a **VAD** in one tuned module, BSD-licensed. speexdsp is kept
|
||||
only for resampling and as a lightweight jitter-buffer reference. AEC is **in from the
|
||||
start**, not deferred.
|
||||
Playback is genuinely stereo end-to-end; capture stays mono (no stereo mic in v1).
|
||||
- **DSP engine: see §11.** The original plan was `webrtc-audio-processing` (AEC + NS + AGC +
|
||||
VAD in one tuned module, BSD-licensed) — but it has no working Windows/MSVC build upstream
|
||||
(confirmed via its own issue tracker: GCC-only Meson build, MinGW support unfinished, hard
|
||||
`abseil-cpp` dependency, Linux-tested only —
|
||||
[gitlab.freedesktop.org/pulseaudio/webrtc-audio-processing#1](https://gitlab.freedesktop.org/pulseaudio/webrtc-audio-processing/-/issues/1)).
|
||||
v1 ships a lightweight, dependency-free energy/RMS VAD instead (§11); there is **no AEC, NS,
|
||||
or AGC implementation at all yet** — not just a deferred VAD, the whole APM is unbuilt. Real
|
||||
`webrtc-audio-processing` stays a tracked future swap, behind the same `ApmProcessor`
|
||||
interface (`core/src/audio/apm_processor.h`), revisit if/when a Linux build target exists or
|
||||
upstream Windows support matures.
|
||||
- The mixer sums decoded streams; clipping is handled by soft limiting on the master bus.
|
||||
|
||||
## 10. Noise reduction — two-sided
|
||||
@@ -197,13 +204,23 @@ NS is a clean drop-in on that per-stream stage.
|
||||
## 11. Input activation — VAD and PTT (client-configurable)
|
||||
|
||||
Whether the mic transmits is decided locally by the **input gate**, and the client supports
|
||||
**both** modes, switchable per client (and ideally per input device):
|
||||
**both** modes, switchable per client (`vc_set_input_mode`):
|
||||
|
||||
- **Voice activation (VAD):** the APM VAD opens the gate when speech is detected, with a
|
||||
configurable threshold and hang-time to avoid clipping word tails. DTX naturally
|
||||
complements this — when the gate is closed nothing (or only comfort noise) is sent.
|
||||
- **Push-to-talk (PTT):** a held key/button opens the gate. The UI exposes a configurable
|
||||
keybind; the core just receives gate open/close.
|
||||
- **Voice activation (VAD):** v1 implements this as a lightweight, dependency-free
|
||||
energy/RMS-threshold VAD (`EnergyVadProcessor`, `core/src/audio/apm_processor.cpp`) — no
|
||||
external DSP dependency, since real `webrtc-audio-processing` has no working Windows/MSVC
|
||||
build (see §8). It opens the gate when a frame's RMS exceeds a configurable threshold
|
||||
(default ~0.025, normalized to int16 range), with a configurable hang-time (default 300 ms,
|
||||
matching the talk-indicator hangover so "talking" and "gate open" agree) to avoid clipping
|
||||
word tails. DTX naturally complements this — when the gate is closed nothing (or only
|
||||
comfort noise) is sent. This implementation has **no AEC** — a real limitation versus the
|
||||
originally-planned APM, not just a deferred VAD.
|
||||
- **Push-to-talk (PTT):** `vc_set_push_to_talk(active)` opens/closes the gate directly. The UI
|
||||
exposes a configurable keybind; the core just receives gate open/close.
|
||||
|
||||
Gating applies to the **MIC stream only** — `SCREEN_AUDIO`/`AUX_DEVICE` always bypass it
|
||||
(gating a desktop-audio share on the user's own voice activity would silently drop shared
|
||||
music/video audio whenever the user isn't talking, which defeats the feature).
|
||||
|
||||
This is purely a send-side, client-local concern — it gates what gets encoded and sent. It
|
||||
needs **no protocol support**; remote talk indicators are still derived from the presence of
|
||||
@@ -217,7 +234,7 @@ normal stream; only the *source* is platform-specific.
|
||||
|
||||
| Platform | Mechanism | Notes |
|
||||
|----------|-----------|-------|
|
||||
| **Windows** | **WASAPI loopback** capture of the default render endpoint (via miniaudio's loopback mode) | Cleanest case; no extra process. Can capture system mix or a specific endpoint. |
|
||||
| **Windows** | **WASAPI loopback** capture of the default render endpoint (via miniaudio's loopback mode) | **Implemented.** Whole-device capture, not process-specific — it inherently captures this app's own incoming voice mix along with everything else playing (an accepted self-echo-loop characteristic of desktop-audio capture, not a bug). Windows 10 2004+'s process-specific loopback (`AUDIOCLIENT_ACTIVATION_PARAMS`) would avoid this but miniaudio doesn't expose it — a future enhancement. |
|
||||
| **macOS** | **ScreenCaptureKit** system-audio capture (macOS 13+), or a virtual audio device fallback on older OSes | OS requires screen-recording permission; capture happens in the main app. |
|
||||
| **iOS** | **ReplayKit Broadcast Upload Extension** (the Discord mechanism) | See below — separate process, App Group, ~50 MB cap (fine for audio-only). |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user