Files
voice-cat/server/src/conn_session.cpp

419 lines
15 KiB
C++
Raw Normal View History

2026-06-15 23:48:44 +02:00
#include "conn_session.h"
#ifdef VOICECAT_HAS_NET
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
#include <algorithm>
2026-06-15 23:48:44 +02:00
#include <chrono>
#include <cstdio>
#include <cstring>
#include <sodium.h>
2026-06-15 23:48:44 +02:00
#include "db.h"
#include "session_registry.h"
#include "core/worker_pool.h"
#include "protocol/envelope.h"
namespace voicecat::server {
static voicecat::v1::Envelope make_env(uint64_t req_id = 0) {
voicecat::v1::Envelope e;
e.set_request_id(req_id);
return e;
}
ConnSession::ConnSession(std::shared_ptr<Database> db,
std::shared_ptr<SessionRegistry> registry,
std::shared_ptr<voicecat::WorkerPool> workers,
const std::array<uint8_t, 32>& server_fp,
bool allow_guests,
uint16_t udp_media_port)
2026-06-15 23:48:44 +02:00
: db_(std::move(db)),
registry_(std::move(registry)),
workers_(std::move(workers)),
server_fp_(server_fp),
allow_guests_(allow_guests),
udp_media_port_(udp_media_port) {
randombytes_buf(udp_token_.data(), udp_token_.size());
}
2026-06-15 23:48:44 +02:00
void ConnSession::set_io(SendFn send_fn, CloseFn close_fn) {
send_fn_ = std::move(send_fn);
close_fn_ = std::move(close_fn);
}
void ConnSession::begin() {
// Nothing to do at TCP level — wait for ClientHello
}
void ConnSession::on_frame(std::vector<uint8_t> frame) {
voicecat::v1::Envelope env;
if (!protocol::decode_envelope(frame, env)) return;
auto st = state_.load(std::memory_order_acquire);
switch (env.body_case()) {
case voicecat::v1::Envelope::kClientHello:
if (st == State::WaitingHello)
handle_client_hello(env.request_id(), env.client_hello());
break;
case voicecat::v1::Envelope::kAuthRequest:
if (st == State::WaitingAuth)
handle_auth_request(env.request_id(), env.auth_request());
break;
case voicecat::v1::Envelope::kJoinChannel:
if (st == State::Authenticated)
handle_join_channel(env.request_id(), env.join_channel());
break;
case voicecat::v1::Envelope::kTextMessage:
if (st == State::Authenticated)
handle_text_message(env.text_message());
break;
case voicecat::v1::Envelope::kPing:
handle_ping(env.ping());
break;
case voicecat::v1::Envelope::kLeaveChannel:
if (st == State::Authenticated)
registry_->set_user_channel(user_id_.load(), 1);
break;
case voicecat::v1::Envelope::kUdpBinding:
if (st == State::Authenticated)
handle_udp_binding(env.request_id(), env.udp_binding());
break;
case voicecat::v1::Envelope::kStreamAnnounce:
if (st == State::Authenticated)
handle_stream_announce(env.request_id(), env.stream_announce());
break;
case voicecat::v1::Envelope::kStreamStop:
if (st == State::Authenticated)
handle_stream_stop(env.stream_stop());
break;
2026-06-15 23:48:44 +02:00
default:
break;
}
}
void ConnSession::on_disconnect() { close(); }
void ConnSession::send_envelope(const voicecat::v1::Envelope& env) {
if (!send_fn_ || closed_.load()) return;
// Serialize to raw protobuf bytes; send_fn_ (→ TcpServerConn::send_frame)
// adds the [4-byte len] framing, so we must NOT pre-frame here.
std::string bytes;
if (!env.SerializeToString(&bytes)) return;
std::vector<uint8_t> raw(bytes.begin(), bytes.end());
send_fn_(std::move(raw));
}
void ConnSession::close() {
if (closed_.exchange(true)) return;
state_.store(State::Disconnecting, std::memory_order_release);
uint32_t uid = user_id_.load();
if (uid) registry_->remove_user(uid);
if (session_id_) registry_->unregister_session(session_id_);
if (close_fn_) close_fn_();
}
// ── M2: media crypto ─────────────────────────────────────────────────────────
void ConnSession::set_media_crypto(
std::unique_ptr<voicecat::crypto::SodiumMediaCrypto> send,
std::unique_ptr<voicecat::crypto::SodiumMediaCrypto> recv) {
std::lock_guard lk(crypto_mu_);
send_crypto_ = std::move(send);
recv_crypto_ = std::move(recv);
}
voicecat::crypto::SodiumMediaCrypto* ConnSession::send_crypto() {
std::lock_guard lk(crypto_mu_);
return send_crypto_.get();
}
voicecat::crypto::SodiumMediaCrypto* ConnSession::recv_crypto() {
std::lock_guard lk(crypto_mu_);
return recv_crypto_.get();
}
// ── M2: UDP endpoint ─────────────────────────────────────────────────────────
void ConnSession::set_udp_endpoint(asio::ip::udp::endpoint ep) {
{
std::lock_guard lk(udp_ep_mu_);
udp_ep_ = ep;
}
has_udp_ep_.store(true, std::memory_order_release);
registry_->register_udp_endpoint(ep, session_id_);
}
asio::ip::udp::endpoint ConnSession::udp_endpoint() const {
std::lock_guard lk(udp_ep_mu_);
return udp_ep_;
}
// ── Handlers ─────────────────────────────────────────────────────────────────
2026-06-15 23:48:44 +02:00
void ConnSession::handle_client_hello(uint64_t req_id, const voicecat::v1::ClientHello& msg) {
if (msg.proto_version() != 1) {
send_disconnect_and_close(1, "unsupported protocol version");
return;
}
auto env = make_env(req_id);
auto* hello = env.mutable_server_hello();
hello->set_proto_version(1);
hello->set_server_name("VoiceCat Server");
hello->set_server_version("0.1.0");
if (allow_guests_) hello->add_auth_methods("guest");
hello->add_auth_methods("password");
hello->set_server_identity_fingerprint(server_fp_.data(), server_fp_.size());
if (udp_media_port_) hello->set_udp_port(udp_media_port_);
2026-06-15 23:48:44 +02:00
send_envelope(env);
state_.store(State::WaitingAuth, std::memory_order_release);
}
void ConnSession::handle_auth_request(uint64_t req_id, const voicecat::v1::AuthRequest& msg) {
if (msg.has_guest()) {
finish_guest_auth(msg.guest(), req_id);
} else if (msg.has_password()) {
finish_password_auth(msg.password().username(), msg.password().password(), req_id);
} else {
auto env = make_env(req_id);
env.mutable_auth_result()->set_ok(false);
env.mutable_auth_result()->set_error("unknown auth method");
send_envelope(env);
}
}
void ConnSession::finish_guest_auth(const voicecat::v1::GuestAuth& guest, uint64_t req_id) {
if (!allow_guests_) {
auto env = make_env(req_id);
env.mutable_auth_result()->set_ok(false);
env.mutable_auth_result()->set_error("guest login not permitted");
send_envelope(env);
return;
}
voicecat::v1::User user;
user.set_nickname(guest.nickname().empty() ? "Guest" : guest.nickname());
user.set_is_guest(true);
user.set_channel_id(1);
uint32_t uid = registry_->add_user(session_id_, user);
user.set_id(uid);
user_id_.store(uid, std::memory_order_relaxed);
state_.store(State::Authenticated, std::memory_order_release);
registry_->register_udp_token(udp_token_, session_id_);
2026-06-15 23:48:44 +02:00
{
auto env = make_env(req_id);
auto* res = env.mutable_auth_result();
res->set_ok(true);
res->set_session_id(session_id_);
*res->mutable_self() = user;
res->set_udp_token(udp_token_.data(), udp_token_.size());
2026-06-15 23:48:44 +02:00
send_envelope(env);
}
broadcast_user_joined(user);
send_state_snapshot();
}
void ConnSession::finish_password_auth(const std::string& username,
const std::string& password, uint64_t req_id) {
// Argon2id runs on the worker pool (deliberately slow).
auto self = shared_from_this();
workers_->post([self, username, password, req_id] {
auto acc = self->db_->authenticate(username, password);
if (!acc) {
auto env = make_env(req_id);
env.mutable_auth_result()->set_ok(false);
env.mutable_auth_result()->set_error("invalid credentials");
self->send_envelope(env);
return;
}
voicecat::v1::User user;
user.set_nickname(acc->username);
user.set_is_guest(false);
user.set_channel_id(1);
uint32_t uid = self->registry_->add_user(self->session_id_, user);
user.set_id(uid);
self->user_id_.store(uid, std::memory_order_relaxed);
self->state_.store(State::Authenticated, std::memory_order_release);
self->registry_->register_udp_token(self->udp_token_, self->session_id_);
2026-06-15 23:48:44 +02:00
{
auto env = make_env(req_id);
auto* res = env.mutable_auth_result();
res->set_ok(true);
res->set_session_id(self->session_id_);
*res->mutable_self() = user;
auto* perms = res->mutable_permissions();
perms->set_is_admin(acc->is_admin);
res->set_udp_token(self->udp_token_.data(), self->udp_token_.size());
2026-06-15 23:48:44 +02:00
self->send_envelope(env);
}
self->broadcast_user_joined(user);
self->send_state_snapshot();
});
}
void ConnSession::send_state_snapshot() {
auto env = make_env();
auto* snap = env.mutable_server_state();
for (auto& ch : registry_->channel_snapshot()) *snap->add_channels() = ch;
for (auto& u : registry_->user_snapshot()) *snap->add_users() = u;
send_envelope(env);
}
void ConnSession::broadcast_user_joined(const voicecat::v1::User& user) {
auto bcast = make_env();
auto* ue = bcast.mutable_user_event();
ue->set_kind(voicecat::v1::UserEvent::JOINED);
*ue->mutable_user() = user;
registry_->broadcast(bcast, session_id_);
}
void ConnSession::handle_join_channel(uint64_t req_id,
const voicecat::v1::JoinChannelRequest& msg) {
bool ok = registry_->set_user_channel(user_id_.load(), msg.channel_id());
auto env = make_env(req_id);
auto* res = env.mutable_join_channel_result();
res->set_ok(ok);
if (!ok) res->set_error("channel not found");
else res->set_channel_id(msg.channel_id());
send_envelope(env);
}
void ConnSession::handle_text_message(const voicecat::v1::TextMessage& msg) {
using namespace std::chrono;
int64_t now_ms = duration_cast<milliseconds>(
system_clock::now().time_since_epoch()).count();
voicecat::v1::TextMessage relay = msg;
relay.set_sender_id(user_id_.load(std::memory_order_relaxed));
relay.set_sent_at_unix_ms(now_ms);
voicecat::v1::Envelope fwd;
*fwd.mutable_text_message() = relay;
auto targets = registry_->resolve_text_targets(session_id_, msg.scope(), msg.target_id());
for (auto& t : targets) t->send_envelope(fwd);
// Ack
auto env = make_env();
auto* ack = env.mutable_text_message_ack();
ack->set_client_msg_id(msg.client_msg_id());
ack->set_ok(true);
send_envelope(env);
}
void ConnSession::handle_ping(const voicecat::v1::Ping& msg) {
auto env = make_env();
env.mutable_pong()->set_nonce(msg.nonce());
send_envelope(env);
}
void ConnSession::handle_udp_binding(uint64_t req_id, const voicecat::v1::UdpBinding& msg) {
if (msg.ack()) return; // server→client direction; ignore if echoed back
const std::string& tok = msg.udp_token();
if (tok.size() != 16 || std::memcmp(tok.data(), udp_token_.data(), 16) != 0) {
// Bad token — silently ignore (don't leak timing information)
return;
}
// Ack over TCP; MediaRelay will set the UDP endpoint when the UDP binding packet arrives.
auto env = make_env(req_id);
env.mutable_udp_binding()->set_ack(true);
send_envelope(env);
}
void ConnSession::handle_stream_announce(uint64_t req_id,
const voicecat::v1::StreamAnnounce& msg) {
uint32_t ssrc = registry_->assign_ssrc(session_id_);
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
uint32_t stream_id = next_stream_id_++;
auto env = make_env(req_id);
auto* res = env.mutable_stream_announce_result();
res->set_ok(true);
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
res->set_stream_id(stream_id);
res->set_ssrc(ssrc);
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
// Per-channel AudioConfig is authoritative (docs/voice.md §3): the channel's mode/
// frame_ms/application/fec/dtx/complexity/expected_packet_loss apply to every stream
// announced into it, regardless of kind. bitrate_bps is clamped (not overridden) to the
// channel's ceiling so a client may still request less. sample_rate stays
// client-requested-or-48000 — everything runs at 48kHz internally per voice.md §3.
auto* eff = res->mutable_effective_audio();
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
auto chan_cfg = registry_->channel_audio_config(registry_->user_channel(user_id_.load()));
uint32_t requested_bps =
msg.has_requested_audio() ? msg.requested_audio().bitrate_bps() : 0;
uint32_t requested_rate =
msg.has_requested_audio() ? msg.requested_audio().sample_rate() : 0;
if (chan_cfg) {
*eff = *chan_cfg;
eff->set_bitrate_bps(requested_bps > 0 ? std::min(requested_bps, chan_cfg->bitrate_bps())
: chan_cfg->bitrate_bps());
eff->set_sample_rate(requested_rate > 0 ? requested_rate : 48000);
} else if (msg.has_requested_audio()) {
*eff = msg.requested_audio();
} else {
eff->set_codec(0); // OPUS
eff->set_sample_rate(48000);
eff->set_bitrate_bps(24000);
eff->set_frame_ms(20);
eff->set_fec(true);
}
if (eff->sample_rate() == 0) eff->set_sample_rate(48000);
if (eff->bitrate_bps() == 0) eff->set_bitrate_bps(24000);
if (eff->frame_ms() == 0) eff->set_frame_ms(20);
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
announced_stream_ids_.push_back(stream_id);
voicecat::v1::StreamInfo info;
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
info.set_stream_id(stream_id);
info.set_ssrc(ssrc);
info.set_kind(msg.kind());
*info.mutable_audio() = *eff;
info.set_label(msg.label());
send_envelope(env);
auto updated = registry_->set_user_stream(user_id_.load(), info);
if (updated) {
auto bcast = make_env();
auto* ue = bcast.mutable_user_event();
ue->set_kind(voicecat::v1::UserEvent::UPDATED);
*ue->mutable_user() = *updated;
registry_->broadcast(bcast, session_id_);
}
}
void ConnSession::handle_stream_stop(const voicecat::v1::StreamStop& msg) {
feat(M3): multi-stream & per-channel tuning Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
auto it = std::find(announced_stream_ids_.begin(), announced_stream_ids_.end(),
msg.stream_id());
if (it == announced_stream_ids_.end()) return; // not ours — ignore (no spoofed stops)
announced_stream_ids_.erase(it);
auto updated = registry_->clear_user_stream(user_id_.load(), msg.stream_id());
if (updated) {
auto bcast = make_env();
auto* ue = bcast.mutable_user_event();
ue->set_kind(voicecat::v1::UserEvent::UPDATED);
*ue->mutable_user() = *updated;
registry_->broadcast(bcast, session_id_);
}
}
2026-06-15 23:48:44 +02:00
void ConnSession::send_disconnect_and_close(uint32_t code, const std::string& reason) {
auto env = make_env();
auto* d = env.mutable_disconnect();
d->set_code(code);
d->set_reason(reason);
send_envelope(env);
close();
}
} // namespace voicecat::server
#endif // VOICECAT_HAS_NET