feat(M3): multi-stream & per-channel tuning

Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC +
SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise-
reduction, talk indicators, and enforced per-channel Opus configurability
(mono/stereo, bitrate, frame size, FEC/DTX, application).

Bugs fixed along the way (found while implementing, not pre-existing scope):
- Server hard-coded stream_id=1 for every announce, so a second stream from
  the same user silently overwrote the first in SessionRegistry::set_user_stream.
  Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop
  validates against announced_stream_ids_ before clearing.
- Client dropped mode/dtx/complexity/application from effective_audio even for
  the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever
  applied to OpusParams. Fixed on both the send (handle_stream_announce_result)
  and receive (sync_remote_streams) paths via a shared
  opus_params_from_audio_config() helper.
- OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application
  and wired it through.
- on_playback's per-stream decode passed the wrong frame_size to opus_decode
  (total samples instead of samples-per-channel), which would have overflowed
  the decode buffer for any stereo stream.
- teardown_voice() raced when called concurrently from run_io()'s own cleanup
  and from disconnect() on a different thread -- both could see
  udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the
  same std::thread (intermittent std::system_error under ctest). Fixed with a
  teardown_mu_ guard instead of carrying the flake forward.

New:
- Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/
  FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX);
  handle_stream_announce enforces the channel's config, clamping (not
  overriding) bitrate_bps to its ceiling.
- core/src/core/client.h/.cpp: local-stream state is now a
  std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with
  request_id-correlated announce/result handling (request_id already
  round-tripped on the wire; just wasn't read before). on_capture_frame is
  kind-aware and upmixes mono capture to stereo when a stream's config calls
  for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired
  through set_remote_stream. New run_talk_timer() thread emits
  VC_EVENT_TALK_STATE from both remote and local edge detection.
- core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps
  (inject_capture), stereo-to-mono downmix at the decode/mix boundary,
  RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled
  and last_voice_ms/talking; new set_stream_noise_reduction() and
  poll_talk_transitions().
- core/src/session/session.h/.cpp: Stream now carries the full AudioConfig,
  not just sample_rate/frame_ms.
- New additive C ABI (core/include/voicecat.h): vc_audio_config +
  vc_get_stream_audio_config (effective Opus config for any stream you own or
  a peer's); vc_test_inject_capture (test-only synthetic PCM injection,
  clearly marked, mirrors AudioEngine::inject_capture).
- tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI
  (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two
  concurrent local streams, independent gain/mute/NS control, per-channel
  config divergence via vc_get_stream_audio_config, talk indicators.

Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently
dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback
capture for SCREEN_AUDIO (synthetic injection only); true stereo playback
output (AudioEngine's mixer/output device stays mono -- Opus itself is fully
stereo-correct on the wire).

ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive
full-suite runs plus 8 standalone runs of the new test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-06-16 14:12:37 +02:00
parent c693cab35c
commit 867557eda1
17 changed files with 1102 additions and 133 deletions

View File

@@ -58,6 +58,14 @@ struct vc_client {
vc_result list_devices(vc_device_kind kind, vc_device_list* out);
// M3: effective Opus config for a (user_id, stream_id) — our own pending/active local
// streams, or any peer's broadcast StreamInfo.audio.
vc_result get_stream_audio_config(uint32_t user_id, uint32_t stream_id,
vc_audio_config* out);
// TEST-ONLY (see voicecat.h) — inject synthetic PCM into a local stream's encode pipeline.
vc_result test_inject_capture(uint32_t stream_id, const int16_t* pcm, size_t samples);
vc_connection_state state() const {
#ifdef VOICECAT_HAS_NET
return state_net_.load(std::memory_order_acquire);
@@ -122,20 +130,43 @@ struct vc_client {
std::unique_ptr<voicecat::crypto::SodiumMediaCrypto> media_recv_crypto_;
voicecat::audio::AudioEngine audio_engine_;
voicecat::codec::OpusEncoder local_encoder_;
std::atomic<bool> local_stream_active_{false};
bool local_stream_pending_{false}; // announced, waiting for result
uint32_t local_stream_id_{0};
uint32_t local_ssrc_{0};
uint32_t next_local_stream_id_{1};
uint32_t local_timestamp_{0};
uint16_t local_frame_samples_{960};
// M3: one LocalStream per concurrently-active stream kind (MIC/SCREEN_AUDIO/AUX_DEVICE
// are each singletons for a given client — see docs/roadmap.md §M3). Replaces the M2
// single-stream fields (local_encoder_/local_stream_active_/etc).
struct LocalStream {
voicecat::codec::OpusEncoder encoder;
std::atomic<bool> active{false};
bool pending{false}; // announced, awaiting StreamAnnounceResult
uint32_t stream_id{0};
uint32_t ssrc{0};
uint32_t timestamp{0};
uint16_t frame_samples{960};
uint64_t pending_request_id{0};
// Full effective OpusParams from the last StreamAnnounceResult — retained so
// vc_get_stream_audio_config() has something to read back for our own streams.
voicecat::codec::OpusParams effective_params;
// Talk-indicator edge detection (docs/voice.md §7) — updated in on_capture_frame.
std::atomic<int64_t> last_capture_ms{0};
bool talking = false;
};
mutable std::mutex local_streams_mu_;
std::unordered_map<int, LocalStream> local_streams_; // keyed by vc_stream_kind
std::unordered_map<uint64_t, int> pending_announce_kind_; // request_id -> kind
uint32_t next_local_stream_id_{1};
// ssrc → (user_id, stream_id) for remote streams already wired into audio_engine_.
mutable std::mutex remote_streams_mu_;
std::unordered_map<uint32_t, std::pair<uint32_t, uint32_t>> remote_streams_;
// M3: talk-indicator polling thread (separate from udp_thread_ / the miniaudio callback
// thread — see architecture.md §3 real-time rule).
std::thread talk_timer_thread_;
std::atomic<bool> talk_timer_stop_{false};
static constexpr int64_t kTalkPollMs = 100;
static constexpr int64_t kTalkHangoverMs = 300;
// Local UDP destination (server media endpoint), resolved once during binding.
uint32_t udp_dest_addr_{0}; // network byte order
uint16_t udp_dest_port_{0}; // network byte order
@@ -143,6 +174,14 @@ struct vc_client {
std::atomic<bool> self_mic_muted_{false};
std::atomic<bool> self_deafened_{false};
// teardown_voice() is called both from run_io()'s own cleanup (on the io_thread_, when
// the read loop exits) and from disconnect() (on the caller's thread) -- without
// serializing those two call sites, both can see udp_thread_/talk_timer_thread_ as
// joinable() at the same time and race to join() the same std::thread object (UB; an
// intermittent "No such process" std::system_error on Windows is the typical symptom).
// This mutex makes teardown_voice() idempotent under concurrent calls.
std::mutex teardown_mu_;
// ── io_thread_ entry point ──────────────────────────────────────────────────
void run_io(std::string host, uint16_t port);
@@ -155,7 +194,8 @@ struct vc_client {
void handle_text_message(const voicecat::v1::TextMessage& msg);
void handle_disconnect(const voicecat::v1::Disconnect& msg);
void handle_udp_binding_ack(const voicecat::v1::UdpBinding& msg);
void handle_stream_announce_result(const voicecat::v1::StreamAnnounceResult& msg);
void handle_stream_announce_result(uint64_t req_id,
const voicecat::v1::StreamAnnounceResult& msg);
// ── M2: UDP / media helpers ──────────────────────────────────────────────────
// Kicks off TCP UdpBinding request; called once after a successful AuthResult.
@@ -164,8 +204,9 @@ struct vc_client {
void finish_udp_binding();
// udp_thread_ entry point: recv loop, AEAD-open, decode, push to audio_engine_.
void run_udp_recv();
// capture_cb passed to audio_engine_.start(): encode + seal + send one frame.
void on_capture_frame(const int16_t* pcm, int samples);
// capture_cb passed to audio_engine_.start(): encode + seal + send one frame for the
// given local stream `kind` (M3: multiple concurrent local streams are possible).
void on_capture_frame(int kind, const int16_t* pcm, int samples);
// Inspect a User proto's streams and wire up any new remote ssrc into audio_engine_,
// emitting VC_EVENT_STREAM_STARTED/STOPPED as streams appear/disappear.
void sync_remote_streams(const voicecat::v1::User& user);
@@ -174,6 +215,12 @@ struct vc_client {
// Joins udp_thread_, stops audio_engine_, clears media crypto/remote-stream state.
// Safe to call multiple times. Called both from run_io()'s cleanup and disconnect().
void teardown_voice();
// talk_timer_thread_ entry point (M3): polls audio_engine_ for remote talk-state edges
// and local capture activity, emitting VC_EVENT_TALK_STATE. Never the audio RT thread.
void run_talk_timer();
// Find a LocalStream by its client-assigned stream_id (held under local_streams_mu_ by
// the caller, or taken internally). Returns nullptr if not found/not active.
LocalStream* find_local_stream_by_id(uint32_t stream_id);
// ── Helpers (io_thread_ and caller threads) ─────────────────────────────────
// Queue an encoded envelope to be sent on io_thread_.