feat(audio): real noise suppression via vendored RNNoise (send + receive)
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled

The two-sided NR plumbing (RemoteStream::recv_ns + the per-listener
vc_set_remote_stream noise_reduction toggle) was wired but inert:
ApmProcessor::create() returned a no-op passthrough, because the
originally-planned webrtc-audio-processing has no working Windows/macOS
build. Drop in RNNoise as the real backend behind the same ApmProcessor
interface, lighting up both NR paths.

- Vendor RNNoise (BSD-3 + CC0) at third_party/rnnoise/ — the vcpkg port
  is !windows !arm, so it can't cover our primary targets. Shrunk int8
  model (78MB -> 11.7MB via upstream scripts/shrink_model.sh), built as a
  standalone C static lib with no RTCD (portable scalar path on x86,
  auto-NEON on arm64) under -DDISABLE_DEBUG_FLOAT. Model is baked in
  (rnnoise_create(NULL)); no runtime file.
- New RnnoiseProcessor (core/src/audio/apm_processor.cpp) selected by
  ApmProcessor::create() when VOICECAT_HAS_NS. Mono/48kHz/480-sample;
  our clock is fixed 48kHz and Opus frame sizes are multiples of 480, so
  no resampling. RT-safe: allocates at construction, lock-free in the
  capture/playback callbacks.
- Receive-side: lit up via the factory; gated to mono streams (a stereo
  stream is a screen-audio share, not voice).
- Send-side (new): vc_set_input_noise_reduction(client, enable) ABI +
  vc_client::mic_ns_, run before input gain/VAD in on_capture_frame. A
  stereo mic is downmixed to mono ONLY when NR is on — with NR off a
  stereo mic keeps full stereo (never collapse mic quality unasked).
- Enable C as a project language for the vendored lib.
- New noise_suppression test: white noise through ApmProcessor::create()
  drops ~99.9% RMS. ctest --preset dev green, 28/28. windows-client DLL
  builds clean with vc_set_input_noise_reduction exported, system-only deps.
- Docs synced: voice.md §10, tech-stack.md §1/§5, third_party/README.md,
  vcpkg.json note, PROGRESS.md, CLAUDE.md.

Client on/off UI toggles (Windows/macOS/iOS) are the remaining follow-up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-06-23 13:30:54 +02:00
parent 7249a8fd30
commit bad9c7533a
50 changed files with 381395 additions and 23 deletions

View File

@@ -287,17 +287,37 @@ Each receiver keeps an **adaptive jitter buffer per ssrc** with **bounded-depth
Noise reduction can be applied **at the sender, at the listener, or both** — they are
independent.
- **Sender-side** (the talker's choice): the publishing client runs APM noise suppression on
its mic before encoding, controlled by that user's own settings. This cleans the signal for
*everyone* and saves bitrate.
- **Sender-side** (the talker's choice): the publishing client runs noise suppression on its
mic before the input gain and the VAD/PTT gate, controlled by that user's own settings
(`vc_set_input_noise_reduction`). This cleans the signal for *everyone* in one pass and helps
bitrate/VAD. MIC stream only.
- **Listener-side, per user** (the listener's choice): on the receive path, *after* decoding
each stream and *before* mixing, the listener can enable an **additional** NS pass on a
**specific** sender's stream. So even if Alex chose not to denoise his mic, Sam can locally
suppress Alex's background noise without affecting how anyone else hears Alex.
**specific** sender's stream (`vc_set_remote_stream(..., noise_reduction)`). So even if Alex
chose not to denoise his mic, Sam can locally suppress Alex's background noise without
affecting how anyone else hears Alex.
Implementation: a per-`ssrc` APM NS instance on the receive path, instantiated lazily only
for streams the listener has flagged. State lives entirely on the listener's machine; toggling
it is a local UI action with **no protocol message** and no effect on other listeners. Because
**Backend: RNNoise** (vendored in [`third_party/rnnoise/`](../third_party/rnnoise), BSD-3 + CC0).
The original plan was WebRTC's APM, but `webrtc-audio-processing` has no working Windows/MSVC
build (see §8). RNNoise is a small, dependency-free C library — a hybrid DSP/RNN speech denoiser
that runs ~60× faster than real time. Both NR paths share one `ApmProcessor` implementation
(`RnnoiseProcessor`, `core/src/audio/apm_processor.cpp`), selected by `ApmProcessor::create()`
when the core is built with `VOICECAT_HAS_NS` (a no-op `ApmPassthrough` otherwise). Allocation
happens at construction; `process_capture()` runs lock-free on the RT thread (architecture.md §3).
RNNoise is a **mono, 48 kHz, 480-sample (10 ms)** denoiser. Our engine clock is fixed at 48 kHz
and every Opus frame size (480/960/1920/2880) is a multiple of 480, so frames are processed as
whole 480-sample chunks with no resampling. Because it's mono-only:
- **Send-side:** a stereo mic is downmixed to mono **only when NR is enabled** — with NR off a
stereo mic keeps full stereo (we never collapse mic quality unless asked).
- **Receive-side:** NR is skipped on stereo streams (a stereo stream is a screen-audio share,
not voice).
Implementation: a per-`ssrc` NS instance (`RemoteStream::recv_ns`) on the receive path,
instantiated lazily only for streams the listener has flagged; the send-side instance
(`vc_client::mic_ns_`) is built once with the MIC stream and gated by an atomic flag so toggling
never allocates on the capture callback. State lives entirely on the local machine; toggling
either is a local UI action with **no protocol message** and no effect on other users. Because
each receive stream is decoded independently before the mixer (voice.md §1), per-user receive
NS is a clean drop-in on that per-stream stage.