feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -11,9 +11,16 @@ bool OpusEncoder::init(const OpusParams& p) {
|
||||
channels_ = p.stereo ? 2 : 1;
|
||||
frame_samples_ = opus_frame_samples(p);
|
||||
|
||||
int opus_application = OPUS_APPLICATION_VOIP;
|
||||
switch (p.application) {
|
||||
case OpusApplication::Audio: opus_application = OPUS_APPLICATION_AUDIO; break;
|
||||
case OpusApplication::LowDelay: opus_application = OPUS_APPLICATION_RESTRICTED_LOWDELAY; break;
|
||||
case OpusApplication::Voip: default: opus_application = OPUS_APPLICATION_VOIP; break;
|
||||
}
|
||||
|
||||
int err = 0;
|
||||
enc_ = opus_encoder_create(static_cast<opus_int32>(p.sample_rate), channels_,
|
||||
OPUS_APPLICATION_VOIP, &err);
|
||||
opus_application, &err);
|
||||
if (err != OPUS_OK || !enc_) {
|
||||
err_ = opus_strerror(err);
|
||||
return false;
|
||||
|
||||
@@ -16,15 +16,24 @@
|
||||
|
||||
namespace voicecat::codec {
|
||||
|
||||
// Mirrors voicecat::v1::OpusApplication (core/proto/voicecat.proto) without depending on
|
||||
// generated protobuf headers from this low-level codec module.
|
||||
enum class OpusApplication {
|
||||
Voip = 0, // speech, optimized for low-rate intelligibility
|
||||
Audio = 1, // music/screen-audio, optimized for fidelity
|
||||
LowDelay = 2, // monitoring, minimal algorithmic delay
|
||||
};
|
||||
|
||||
struct OpusParams {
|
||||
uint32_t sample_rate = 48000;
|
||||
uint32_t bitrate_bps = 24000;
|
||||
uint32_t frame_ms = 20;
|
||||
bool stereo = false;
|
||||
bool fec = true;
|
||||
bool dtx = false;
|
||||
uint32_t complexity = 10;
|
||||
uint32_t expected_packet_loss = 0; // % 0..100
|
||||
uint32_t sample_rate = 48000;
|
||||
uint32_t bitrate_bps = 24000;
|
||||
uint32_t frame_ms = 20;
|
||||
bool stereo = false;
|
||||
bool fec = true;
|
||||
bool dtx = false;
|
||||
uint32_t complexity = 10;
|
||||
uint32_t expected_packet_loss = 0; // % 0..100
|
||||
OpusApplication application = OpusApplication::Voip;
|
||||
};
|
||||
|
||||
// Returns frame_samples for a given sample_rate + frame_ms.
|
||||
|
||||
Reference in New Issue
Block a user