feat(audio): real noise suppression via vendored RNNoise (send + receive)
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled

The two-sided NR plumbing (RemoteStream::recv_ns + the per-listener
vc_set_remote_stream noise_reduction toggle) was wired but inert:
ApmProcessor::create() returned a no-op passthrough, because the
originally-planned webrtc-audio-processing has no working Windows/macOS
build. Drop in RNNoise as the real backend behind the same ApmProcessor
interface, lighting up both NR paths.

- Vendor RNNoise (BSD-3 + CC0) at third_party/rnnoise/ — the vcpkg port
  is !windows !arm, so it can't cover our primary targets. Shrunk int8
  model (78MB -> 11.7MB via upstream scripts/shrink_model.sh), built as a
  standalone C static lib with no RTCD (portable scalar path on x86,
  auto-NEON on arm64) under -DDISABLE_DEBUG_FLOAT. Model is baked in
  (rnnoise_create(NULL)); no runtime file.
- New RnnoiseProcessor (core/src/audio/apm_processor.cpp) selected by
  ApmProcessor::create() when VOICECAT_HAS_NS. Mono/48kHz/480-sample;
  our clock is fixed 48kHz and Opus frame sizes are multiples of 480, so
  no resampling. RT-safe: allocates at construction, lock-free in the
  capture/playback callbacks.
- Receive-side: lit up via the factory; gated to mono streams (a stereo
  stream is a screen-audio share, not voice).
- Send-side (new): vc_set_input_noise_reduction(client, enable) ABI +
  vc_client::mic_ns_, run before input gain/VAD in on_capture_frame. A
  stereo mic is downmixed to mono ONLY when NR is on — with NR off a
  stereo mic keeps full stereo (never collapse mic quality unasked).
- Enable C as a project language for the vendored lib.
- New noise_suppression test: white noise through ApmProcessor::create()
  drops ~99.9% RMS. ctest --preset dev green, 28/28. windows-client DLL
  builds clean with vc_set_input_noise_reduction exported, system-only deps.
- Docs synced: voice.md §10, tech-stack.md §1/§5, third_party/README.md,
  vcpkg.json note, PROGRESS.md, CLAUDE.md.

Client on/off UI toggles (Windows/macOS/iOS) are the remaining follow-up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-06-23 13:30:54 +02:00
parent 7249a8fd30
commit bad9c7533a
50 changed files with 381395 additions and 23 deletions

View File

@@ -1,9 +1,14 @@
#include "audio/apm_processor.h"
#include <algorithm>
#include <atomic>
#include <chrono>
#include <cmath>
#ifdef VOICECAT_HAS_NS
#include "rnnoise.h"
#endif
namespace voicecat::audio {
namespace {
@@ -57,12 +62,55 @@ class EnergyVadProcessor final : public ApmProcessor {
int64_t last_voice_ms_ = 0; // epoch start -> gate begins closed until first loud frame
};
#ifdef VOICECAT_HAS_NS
// ── RnnoiseProcessor ─────────────────────────────────────────────────────────
// Real noise suppression via vendored RNNoise (third_party/rnnoise; docs/voice.md §10-11).
// RNNoise is a mono, 48 kHz, fixed 480-sample (10 ms) speech denoiser; our engine clock is fixed
// at 48 kHz and every Opus frame size (480/960/1920/2880) is a multiple of 480, so we process
// whole 480-sample chunks with no resampling and no cross-call carry. Mono only — callers gate
// on a single channel (a stereo screen-audio share is never voice and isn't denoised).
//
// RT-safety (docs/architecture.md §3): the DenoiseState and the float scratch are allocated in the
// ctor; process_capture() does no allocation/locking. The state is owned per-stream (recv) or
// per-mic (send) so it persists across calls, which is exactly what RNNoise's overlap needs.
class RnnoiseProcessor final : public ApmProcessor {
public:
RnnoiseProcessor() : st_(rnnoise_create(nullptr)) {}
~RnnoiseProcessor() override {
if (st_) rnnoise_destroy(st_);
}
void process_render(const int16_t*, int, int) override {} // NS needs no AEC reference
bool process_capture(int16_t* pcm, int samples, int sample_rate) override {
// RNNoise is 48 kHz only; anything else passes through untouched (our clock is 48 kHz, so
// this guard never trips in practice — it's just a correctness backstop).
if (!st_ || sample_rate != 48000) return true;
for (int off = 0; off + kFrame <= samples; off += kFrame) {
for (int i = 0; i < kFrame; ++i) in_[i] = static_cast<float>(pcm[off + i]);
rnnoise_process_frame(st_, out_, in_);
for (int i = 0; i < kFrame; ++i) {
int32_t v = static_cast<int32_t>(std::lround(out_[i]));
pcm[off + i] = static_cast<int16_t>(std::clamp(v, -32768, 32767));
}
}
return true; // NS doesn't gate; the send path's VAD stays a separate stage
}
private:
static constexpr int kFrame = 480; // rnnoise_get_frame_size()
DenoiseState* st_;
float in_[kFrame];
float out_[kFrame];
};
#endif // VOICECAT_HAS_NS
std::unique_ptr<ApmProcessor> ApmProcessor::create() {
#ifdef VOICECAT_HAS_APM
// TODO: return std::make_unique<WebrtcApmProcessor>(); — see create_vad()'s doc comment for
// why this isn't wired up yet (no working Windows/MSVC build upstream).
#endif
#ifdef VOICECAT_HAS_NS
return std::make_unique<RnnoiseProcessor>();
#else
return std::make_unique<ApmPassthrough>();
#endif
}
std::unique_ptr<ApmProcessor> ApmProcessor::create_vad(float rms_threshold,