feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -2,6 +2,7 @@
|
||||
|
||||
#ifdef VOICECAT_HAS_NET
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
@@ -329,15 +330,31 @@ void ConnSession::handle_udp_binding(uint64_t req_id, const voicecat::v1::UdpBin
|
||||
void ConnSession::handle_stream_announce(uint64_t req_id,
|
||||
const voicecat::v1::StreamAnnounce& msg) {
|
||||
uint32_t ssrc = registry_->assign_ssrc(session_id_);
|
||||
uint32_t stream_id = next_stream_id_++;
|
||||
|
||||
auto env = make_env(req_id);
|
||||
auto* res = env.mutable_stream_announce_result();
|
||||
res->set_ok(true);
|
||||
res->set_stream_id(1);
|
||||
res->set_stream_id(stream_id);
|
||||
res->set_ssrc(ssrc);
|
||||
|
||||
// Per-channel AudioConfig is authoritative (docs/voice.md §3): the channel's mode/
|
||||
// frame_ms/application/fec/dtx/complexity/expected_packet_loss apply to every stream
|
||||
// announced into it, regardless of kind. bitrate_bps is clamped (not overridden) to the
|
||||
// channel's ceiling so a client may still request less. sample_rate stays
|
||||
// client-requested-or-48000 — everything runs at 48kHz internally per voice.md §3.
|
||||
auto* eff = res->mutable_effective_audio();
|
||||
if (msg.has_requested_audio()) {
|
||||
auto chan_cfg = registry_->channel_audio_config(registry_->user_channel(user_id_.load()));
|
||||
uint32_t requested_bps =
|
||||
msg.has_requested_audio() ? msg.requested_audio().bitrate_bps() : 0;
|
||||
uint32_t requested_rate =
|
||||
msg.has_requested_audio() ? msg.requested_audio().sample_rate() : 0;
|
||||
if (chan_cfg) {
|
||||
*eff = *chan_cfg;
|
||||
eff->set_bitrate_bps(requested_bps > 0 ? std::min(requested_bps, chan_cfg->bitrate_bps())
|
||||
: chan_cfg->bitrate_bps());
|
||||
eff->set_sample_rate(requested_rate > 0 ? requested_rate : 48000);
|
||||
} else if (msg.has_requested_audio()) {
|
||||
*eff = msg.requested_audio();
|
||||
} else {
|
||||
eff->set_codec(0); // OPUS
|
||||
@@ -350,10 +367,10 @@ void ConnSession::handle_stream_announce(uint64_t req_id,
|
||||
if (eff->bitrate_bps() == 0) eff->set_bitrate_bps(24000);
|
||||
if (eff->frame_ms() == 0) eff->set_frame_ms(20);
|
||||
|
||||
announced_stream_id_ = res->stream_id();
|
||||
announced_stream_ids_.push_back(stream_id);
|
||||
|
||||
voicecat::v1::StreamInfo info;
|
||||
info.set_stream_id(res->stream_id());
|
||||
info.set_stream_id(stream_id);
|
||||
info.set_ssrc(ssrc);
|
||||
info.set_kind(msg.kind());
|
||||
*info.mutable_audio() = *eff;
|
||||
@@ -372,8 +389,12 @@ void ConnSession::handle_stream_announce(uint64_t req_id,
|
||||
}
|
||||
|
||||
void ConnSession::handle_stream_stop(const voicecat::v1::StreamStop& msg) {
|
||||
auto it = std::find(announced_stream_ids_.begin(), announced_stream_ids_.end(),
|
||||
msg.stream_id());
|
||||
if (it == announced_stream_ids_.end()) return; // not ours — ignore (no spoofed stops)
|
||||
announced_stream_ids_.erase(it);
|
||||
|
||||
auto updated = registry_->clear_user_stream(user_id_.load(), msg.stream_id());
|
||||
if (msg.stream_id() == announced_stream_id_) announced_stream_id_ = 0;
|
||||
if (updated) {
|
||||
auto bcast = make_env();
|
||||
auto* ue = bcast.mutable_user_event();
|
||||
|
||||
@@ -121,8 +121,11 @@ class ConnSession : public std::enable_shared_from_this<ConnSession> {
|
||||
std::unique_ptr<voicecat::crypto::SodiumMediaCrypto> send_crypto_;
|
||||
std::unique_ptr<voicecat::crypto::SodiumMediaCrypto> recv_crypto_;
|
||||
|
||||
// M2: locally-announced stream (single MIC stream per session for now)
|
||||
uint32_t announced_stream_id_{0};
|
||||
// M2/M3: locally-announced streams. The server assigns the stream_id (unique per
|
||||
// session), so a per-session counter + the set of currently-active ids is enough to
|
||||
// support multiple concurrent streams (MIC + SCREEN_AUDIO + AUX_DEVICE) per user.
|
||||
uint32_t next_stream_id_{1};
|
||||
std::vector<uint32_t> announced_stream_ids_;
|
||||
};
|
||||
|
||||
} // namespace voicecat::server
|
||||
|
||||
@@ -17,7 +17,44 @@ void SessionRegistry::init_default_channels() {
|
||||
lobby.proto.set_name("Lobby");
|
||||
lobby.proto.set_type(voicecat::v1::CHANNEL_PERMANENT);
|
||||
lobby.proto.set_order(0);
|
||||
{
|
||||
// Speech profile: mono, low bitrate, FEC+DTX on for resilience/silence-suppression.
|
||||
auto* a = lobby.proto.mutable_audio();
|
||||
a->set_codec(0);
|
||||
a->set_mode(voicecat::v1::MODE_MONO);
|
||||
a->set_sample_rate(48000);
|
||||
a->set_bitrate_bps(24000);
|
||||
a->set_frame_ms(20);
|
||||
a->set_application(voicecat::v1::OPUS_VOIP);
|
||||
a->set_fec(true);
|
||||
a->set_expected_packet_loss(10);
|
||||
a->set_dtx(true);
|
||||
a->set_complexity(5);
|
||||
}
|
||||
channels_[1] = std::move(lobby);
|
||||
|
||||
ChannelEntry music;
|
||||
music.proto.set_id(2);
|
||||
music.proto.set_name("Music Room");
|
||||
music.proto.set_type(voicecat::v1::CHANNEL_PERMANENT);
|
||||
music.proto.set_order(1);
|
||||
{
|
||||
// Music/screen-audio profile: stereo, high bitrate, FEC/DTX off (continuous signal).
|
||||
auto* a = music.proto.mutable_audio();
|
||||
a->set_codec(0);
|
||||
a->set_mode(voicecat::v1::MODE_STEREO);
|
||||
a->set_sample_rate(48000);
|
||||
a->set_bitrate_bps(128000);
|
||||
a->set_frame_ms(20);
|
||||
a->set_application(voicecat::v1::OPUS_AUDIO);
|
||||
a->set_fec(false);
|
||||
a->set_expected_packet_loss(0);
|
||||
a->set_dtx(false);
|
||||
a->set_complexity(8);
|
||||
}
|
||||
channels_[2] = std::move(music);
|
||||
|
||||
next_channel_id_ = 3; // 1 and 2 are now reserved (Lobby, Music Room)
|
||||
}
|
||||
|
||||
uint64_t SessionRegistry::register_session(std::weak_ptr<ConnSession> session) {
|
||||
@@ -203,6 +240,14 @@ uint32_t SessionRegistry::user_channel(uint32_t user_id) const {
|
||||
return (it == users_.end()) ? 0 : it->second.proto.channel_id();
|
||||
}
|
||||
|
||||
std::optional<voicecat::v1::AudioConfig> SessionRegistry::channel_audio_config(
|
||||
uint32_t channel_id) const {
|
||||
std::shared_lock lk(mu_);
|
||||
auto it = channels_.find(channel_id);
|
||||
if (it == channels_.end()) return std::nullopt;
|
||||
return it->second.proto.audio();
|
||||
}
|
||||
|
||||
} // namespace voicecat::server
|
||||
|
||||
#endif // VOICECAT_HAS_NET
|
||||
|
||||
@@ -112,6 +112,12 @@ class SessionRegistry {
|
||||
// Return the channel_id of a user (0 if not found).
|
||||
uint32_t user_channel(uint32_t user_id) const;
|
||||
|
||||
// Return a channel's authoritative AudioConfig (M3 per-channel Opus tuning), or nullopt
|
||||
// if the channel doesn't exist. There is no per-id Channel getter today otherwise —
|
||||
// channel_snapshot() copies every channel, which callers needing just one config should
|
||||
// avoid.
|
||||
std::optional<voicecat::v1::AudioConfig> channel_audio_config(uint32_t channel_id) const;
|
||||
|
||||
private:
|
||||
mutable std::shared_mutex mu_;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user