feat(M3): multi-stream & per-channel tuning

Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC +
SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise-
reduction, talk indicators, and enforced per-channel Opus configurability
(mono/stereo, bitrate, frame size, FEC/DTX, application).

Bugs fixed along the way (found while implementing, not pre-existing scope):
- Server hard-coded stream_id=1 for every announce, so a second stream from
  the same user silently overwrote the first in SessionRegistry::set_user_stream.
  Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop
  validates against announced_stream_ids_ before clearing.
- Client dropped mode/dtx/complexity/application from effective_audio even for
  the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever
  applied to OpusParams. Fixed on both the send (handle_stream_announce_result)
  and receive (sync_remote_streams) paths via a shared
  opus_params_from_audio_config() helper.
- OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application
  and wired it through.
- on_playback's per-stream decode passed the wrong frame_size to opus_decode
  (total samples instead of samples-per-channel), which would have overflowed
  the decode buffer for any stereo stream.
- teardown_voice() raced when called concurrently from run_io()'s own cleanup
  and from disconnect() on a different thread -- both could see
  udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the
  same std::thread (intermittent std::system_error under ctest). Fixed with a
  teardown_mu_ guard instead of carrying the flake forward.

New:
- Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/
  FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX);
  handle_stream_announce enforces the channel's config, clamping (not
  overriding) bitrate_bps to its ceiling.
- core/src/core/client.h/.cpp: local-stream state is now a
  std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with
  request_id-correlated announce/result handling (request_id already
  round-tripped on the wire; just wasn't read before). on_capture_frame is
  kind-aware and upmixes mono capture to stereo when a stream's config calls
  for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired
  through set_remote_stream. New run_talk_timer() thread emits
  VC_EVENT_TALK_STATE from both remote and local edge detection.
- core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps
  (inject_capture), stereo-to-mono downmix at the decode/mix boundary,
  RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled
  and last_voice_ms/talking; new set_stream_noise_reduction() and
  poll_talk_transitions().
- core/src/session/session.h/.cpp: Stream now carries the full AudioConfig,
  not just sample_rate/frame_ms.
- New additive C ABI (core/include/voicecat.h): vc_audio_config +
  vc_get_stream_audio_config (effective Opus config for any stream you own or
  a peer's); vc_test_inject_capture (test-only synthetic PCM injection,
  clearly marked, mirrors AudioEngine::inject_capture).
- tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI
  (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two
  concurrent local streams, independent gain/mute/NS control, per-channel
  config divergence via vc_get_stream_audio_config, talk indicators.

Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently
dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback
capture for SCREEN_AUDIO (synthetic injection only); true stereo playback
output (AudioEngine's mixer/output device stays mono -- Opus itself is fully
stereo-correct on the wire).

ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive
full-suite runs plus 8 standalone runs of the new test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-06-16 14:12:37 +02:00
parent c693cab35c
commit 867557eda1
17 changed files with 1102 additions and 133 deletions

View File

@@ -31,6 +31,8 @@
#include "codec/opus_codec.h"
#endif
#include "audio/apm_processor.h"
namespace voicecat::audio {
// ── JitterBuffer ─────────────────────────────────────────────────────────────
@@ -83,8 +85,12 @@ struct AudioParams {
// Owns miniaudio capture/playback, per-ssrc jitter buffers + Opus decoders, and the mixer.
class AudioEngine {
public:
// Callback type for encoded capture frames ready to be sent.
using CaptureCallback = std::function<void(const int16_t* pcm, int samples)>;
// Callback type for encoded capture frames ready to be sent. `kind` identifies which
// local stream this PCM belongs to (a vc_stream_kind value; 0 = MIC for the real capture
// device, which is always the "primary" tap). M3: multiple concurrent local streams are
// possible (e.g. MIC + SCREEN_AUDIO), each fed via its own injection tap (see
// inject_capture) since there is only one real hardware capture device.
using CaptureCallback = std::function<void(int kind, const int16_t* pcm, int samples)>;
AudioEngine();
~AudioEngine();
@@ -100,8 +106,11 @@ class AudioEngine {
bool running() const { return running_.load(std::memory_order_acquire); }
// Inject synthetic PCM directly into the capture pipeline (bypasses real device).
// Thread-safe; can be called from any thread including tests.
void inject_capture(const int16_t* pcm, size_t n);
// Thread-safe; can be called from any thread including tests. `kind` selects which local
// stream's injection tap to feed (each gets its own ring buffer); the 2-arg overload
// targets kind 0 (MIC) for source compatibility with existing callers.
void inject_capture(int kind, const int16_t* pcm, size_t n);
void inject_capture(const int16_t* pcm, size_t n) { inject_capture(0, pcm, n); }
// Called by the net thread when a decoded voice frame arrives for a remote stream.
void push_recv_frame(uint32_t ssrc, JitterBuffer::Frame f);
@@ -109,8 +118,17 @@ class AudioEngine {
// Per-stream receive-side controls (safe from any thread).
void set_stream_gain(uint32_t ssrc, float gain); // 0.02.0, default 1.0
void set_stream_mute(uint32_t ssrc, bool mute);
// Listener-chosen, local-only noise reduction on a specific remote stream (docs/voice.md
// §10) — lazily instantiates an ApmProcessor on first enable, frees it on disable.
void set_stream_noise_reduction(uint32_t ssrc, bool enable);
void remove_stream(uint32_t ssrc);
// Edge-triggered talk-state transitions since the last call (docs/voice.md §7: talk state
// is derived from recent frame arrival, no protocol message). Call from a lightweight
// poller, not the audio callback thread. Returns {ssrc, now_talking} for each stream whose
// state flipped.
std::vector<std::pair<uint32_t, bool>> poll_talk_transitions();
// Get stats for a remote stream's jitter buffer.
uint32_t stream_packets_lost(uint32_t ssrc) const;
uint32_t stream_target_depth_ms(uint32_t ssrc) const;
@@ -138,12 +156,16 @@ class AudioEngine {
CaptureCallback capture_cb_;
std::atomic<bool> running_{false};
// Inject ring: stores raw int16 PCM written by inject_capture().
// The encode thread reads from this (no real capture device needed in tests).
std::mutex inject_mu_;
std::vector<int16_t> inject_ring_; // circular, size = frame_samples_
std::atomic<size_t> inject_write_{0};
std::atomic<size_t> inject_read_{0};
// Inject ring(s): stores raw int16 PCM written by inject_capture(), one ring per local
// stream kind so e.g. MIC and SCREEN_AUDIO can each be fed independently in tests.
// The encode thread reads from these (no real capture device needed in tests).
struct InjectTap {
std::vector<int16_t> ring; // circular, size = kInjectCapSamples
std::atomic<size_t> write{0};
std::atomic<size_t> read{0};
};
std::mutex inject_mu_;
std::unordered_map<int, std::unique_ptr<InjectTap>> inject_taps_;
static constexpr size_t kInjectCapSamples = 48000 * 2; // 2 s @48 kHz mono
// Per remote stream (protected by streams_mu_).
@@ -155,11 +177,24 @@ class AudioEngine {
float gain = 1.0f;
bool mute = false;
uint32_t playout_ts = 0;
// M3: listener-chosen, local-only noise reduction (docs/voice.md §10). Lazily
// created only when enabled — bounded by how many remote streams this listener
// subscribes to, so no separate instance cap is needed.
bool noise_reduction_enabled = false;
std::unique_ptr<ApmProcessor> recv_ns;
// M3: talk-indicator edge detection (docs/voice.md §7) — updated by push_recv_frame
// (already off the real-time audio thread), polled by poll_talk_transitions().
std::atomic<int64_t> last_voice_ms{0};
bool talking = false;
};
mutable std::mutex streams_mu_;
std::unordered_map<uint32_t, RemoteStream> streams_;
int frame_samples_ = 960; // 20 ms @48 kHz
static constexpr int64_t kTalkHangoverMs = 300;
};
} // namespace voicecat::audio