feat(M1): TCP/TLS control plane -- auth, channels, ephemeral text
Implements the full M1 milestone. Two clients authenticate over TLS 1.3
(guest + Argon2id password) and exchange channel + private text messages
through a real server. All five ctest --preset m1-dev tests pass in ~1 s.
Key components added:
- vcpkg baseline + m1-dev preset (protobuf/mbedTLS/libsodium/asio/sqlite3)
- FrameCodec feed+emit, encode/decode_envelope, protobuf codegen
- TcpServerConn with blocking TLS handshake thread + tls_read_loop
- TlsContext (mbedTLS 1.3, ECDSA-P256 self-signed cert, TOFU on client)
- WorkerPool (3 threads, used for Argon2id)
- Database: SQLite + libsodium Argon2id, account lifecycle, bootstrap admin
- ServerIdentityManager: Ed25519 key + cert generate/persist/fingerprint
- ConnSession state machine: WaitingHello -> WaitingAuth -> Authenticated
- SessionRegistry: channel tree, user map, text routing, broadcast
- vc_client full M1 C ABI: connect/TLS/handshake/auth/text/disconnect
- voicecat-admin CLI: account add/reset/del/list
- test_m1_integration: M1 exit criterion, verified green
Bug fixed: double-framing in ConnSession::send_envelope -- encode_envelope
was adding the [4-byte len] prefix, then TcpServerConn::send_frame added
a second one, causing the client to parse [len][proto] as protobuf (silent
failure). Fixed by serializing raw protobuf bytes in send_envelope and
letting send_frame apply the single length prefix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 23:48:44 +02:00
|
|
|
#include "session_registry.h"
|
|
|
|
|
|
|
|
|
|
#ifdef VOICECAT_HAS_NET
|
|
|
|
|
|
feat(M2): UDP voice/media plane -- SFU relay, Opus, AEAD, jitter buffer
Adds the full voice pipeline: 14-byte binary frame header, ChaCha20-Poly1305
AEAD keyed from the TLS exporter, libopus encode/decode with FEC/PLC/DTX,
an adaptive per-ssrc jitter buffer, a miniaudio capture/playback engine, an
APM passthrough stub, and the UdpBinding/StreamAnnounce signaling chain
wired through ConnSession/SessionRegistry into a new server-side SFU
(MediaRelay) that decrypts and re-encrypts frames per channel member.
Exit criterion verified: test_m2_voice — two headless clients relay 50
encrypted Opus frames through the server; ctest --preset m1-dev is 9/9
green. Also corrects protocol.md's UdpBinding diagram, which described the
UDP-side binding packet as AEAD-sealed when it is in fact a plaintext
bootstrap frame (separate from the TCP/TLS UdpBinding ack).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 01:31:14 +02:00
|
|
|
#include <atomic>
|
feat(M1): TCP/TLS control plane -- auth, channels, ephemeral text
Implements the full M1 milestone. Two clients authenticate over TLS 1.3
(guest + Argon2id password) and exchange channel + private text messages
through a real server. All five ctest --preset m1-dev tests pass in ~1 s.
Key components added:
- vcpkg baseline + m1-dev preset (protobuf/mbedTLS/libsodium/asio/sqlite3)
- FrameCodec feed+emit, encode/decode_envelope, protobuf codegen
- TcpServerConn with blocking TLS handshake thread + tls_read_loop
- TlsContext (mbedTLS 1.3, ECDSA-P256 self-signed cert, TOFU on client)
- WorkerPool (3 threads, used for Argon2id)
- Database: SQLite + libsodium Argon2id, account lifecycle, bootstrap admin
- ServerIdentityManager: Ed25519 key + cert generate/persist/fingerprint
- ConnSession state machine: WaitingHello -> WaitingAuth -> Authenticated
- SessionRegistry: channel tree, user map, text routing, broadcast
- vc_client full M1 C ABI: connect/TLS/handshake/auth/text/disconnect
- voicecat-admin CLI: account add/reset/del/list
- test_m1_integration: M1 exit criterion, verified green
Bug fixed: double-framing in ConnSession::send_envelope -- encode_envelope
was adding the [4-byte len] prefix, then TcpServerConn::send_frame added
a second one, causing the client to parse [len][proto] as protobuf (silent
failure). Fixed by serializing raw protobuf bytes in send_envelope and
letting send_frame apply the single length prefix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 23:48:44 +02:00
|
|
|
#include <mutex>
|
|
|
|
|
#include <shared_mutex>
|
|
|
|
|
|
|
|
|
|
#include "conn_session.h"
|
|
|
|
|
|
|
|
|
|
namespace voicecat::server {
|
|
|
|
|
|
|
|
|
|
void SessionRegistry::init_default_channels() {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
ChannelEntry lobby;
|
|
|
|
|
lobby.proto.set_id(1);
|
|
|
|
|
lobby.proto.set_name("Lobby");
|
|
|
|
|
lobby.proto.set_type(voicecat::v1::CHANNEL_PERMANENT);
|
|
|
|
|
lobby.proto.set_order(0);
|
feat(M4): Windows WinForms client, TOFU identity pinning, VAD threshold + always-on mode
Core ABI extensions (voicecat.h):
- vc_list_channels / vc_list_users / vc_list_user_streams — pull-based snapshot getters
for the channel-tree and user-list UI; session_model_mu_ guards cross-thread reads
- VC_EVENT_JOIN_RESULT / vc_join_channel — channel join with optional password
- VC_EVENT_SERVER_IDENTITY + vc_confirm_server_identity — TOFU gate that blocks io_thread_
until the UI approves or rejects; pins TLS leaf-cert SHA-256 (not declared Ed25519)
- vc_get_server_identity_display — Ed25519 fingerprint for human-readable display only
- VC_INPUT_ALWAYS_ON = 2 in vc_input_mode — transmit unconditionally, no VAD gate
- vc_set_vad_threshold — live RMS threshold update (0.0–1.0); EnergyVadProcessor stores
it atomically so the audio RT path reads without a lock
C++ implementation:
- SessionModel::apply_snapshot / apply_channel_event fixed to populate parent_id,
password_protected, and max_users (were permanently zeroed)
- TlsContext::peer_cert_fingerprint — SHA-256 of peer leaf cert DER via mbedTLS
- TofuStore split into peek (read-only) + pin (write) so first-connect only persists
after user approval; tofu_store_path in vc_config for per-user pin file location
- TcpAcceptor uses dual-stack IPv6+IPv4 fallback (fixes localhost → ::1 on Windows)
- windows-client CMake preset: Release shared DLL, static MinGW runtime, no tools/tests
- New C++ tests: test_channel_user_list_abi, test_tofu_flow (14/14 green)
Windows client (clients/windows/ — .NET 10 WinForms):
- VoiceCat.Interop: LibraryImport P/Invoke surface, UnmanagedCallersOnly callbacks,
Channel<VoiceCatEvent> event delivery drained by 30ms WinForms Timer
- VoiceCat.App: ConnectDialog (saved servers, DPAPI password storage), ServerIdentity-
Dialog (TOFU first-connect / mismatch warning), MainForm (channel TreeView, user
ListBox, RichTextBox chat, voice controls, device pickers, VAD/PTT/always-on mode,
per-user gain/mute/NR tuning, VAD sensitivity TrackBar, level meter ProgressBar)
- PttKeyCaptureDialog — focus-scoped PTT key capture (documented limitation)
- PerUserTuningDialog — real-time gain/mute/NR applied to all of a user's streams
- Accessibility: explicit AccessibleName/Description on every control, & mnemonics,
Activity log ListBox as durable screen-reader record, AutomationNotification for
curated live announcements
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 00:35:16 +02:00
|
|
|
// Non-zero so vc_list_channels/SessionModel round-trip this field for real (a regression
|
|
|
|
|
// test for the M4 SessionModel field-population fix needs at least one non-default value).
|
|
|
|
|
lobby.proto.set_max_users(20);
|
feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC +
SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise-
reduction, talk indicators, and enforced per-channel Opus configurability
(mono/stereo, bitrate, frame size, FEC/DTX, application).
Bugs fixed along the way (found while implementing, not pre-existing scope):
- Server hard-coded stream_id=1 for every announce, so a second stream from
the same user silently overwrote the first in SessionRegistry::set_user_stream.
Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop
validates against announced_stream_ids_ before clearing.
- Client dropped mode/dtx/complexity/application from effective_audio even for
the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever
applied to OpusParams. Fixed on both the send (handle_stream_announce_result)
and receive (sync_remote_streams) paths via a shared
opus_params_from_audio_config() helper.
- OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application
and wired it through.
- on_playback's per-stream decode passed the wrong frame_size to opus_decode
(total samples instead of samples-per-channel), which would have overflowed
the decode buffer for any stereo stream.
- teardown_voice() raced when called concurrently from run_io()'s own cleanup
and from disconnect() on a different thread -- both could see
udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the
same std::thread (intermittent std::system_error under ctest). Fixed with a
teardown_mu_ guard instead of carrying the flake forward.
New:
- Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/
FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX);
handle_stream_announce enforces the channel's config, clamping (not
overriding) bitrate_bps to its ceiling.
- core/src/core/client.h/.cpp: local-stream state is now a
std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with
request_id-correlated announce/result handling (request_id already
round-tripped on the wire; just wasn't read before). on_capture_frame is
kind-aware and upmixes mono capture to stereo when a stream's config calls
for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired
through set_remote_stream. New run_talk_timer() thread emits
VC_EVENT_TALK_STATE from both remote and local edge detection.
- core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps
(inject_capture), stereo-to-mono downmix at the decode/mix boundary,
RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled
and last_voice_ms/talking; new set_stream_noise_reduction() and
poll_talk_transitions().
- core/src/session/session.h/.cpp: Stream now carries the full AudioConfig,
not just sample_rate/frame_ms.
- New additive C ABI (core/include/voicecat.h): vc_audio_config +
vc_get_stream_audio_config (effective Opus config for any stream you own or
a peer's); vc_test_inject_capture (test-only synthetic PCM injection,
clearly marked, mirrors AudioEngine::inject_capture).
- tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI
(mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two
concurrent local streams, independent gain/mute/NS control, per-channel
config divergence via vc_get_stream_audio_config, talk indicators.
Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently
dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback
capture for SCREEN_AUDIO (synthetic injection only); true stereo playback
output (AudioEngine's mixer/output device stays mono -- Opus itself is fully
stereo-correct on the wire).
ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive
full-suite runs plus 8 standalone runs of the new test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
|
|
|
{
|
|
|
|
|
// Speech profile: mono, low bitrate, FEC+DTX on for resilience/silence-suppression.
|
|
|
|
|
auto* a = lobby.proto.mutable_audio();
|
|
|
|
|
a->set_codec(0);
|
|
|
|
|
a->set_mode(voicecat::v1::MODE_MONO);
|
|
|
|
|
a->set_sample_rate(48000);
|
|
|
|
|
a->set_bitrate_bps(24000);
|
|
|
|
|
a->set_frame_ms(20);
|
|
|
|
|
a->set_application(voicecat::v1::OPUS_VOIP);
|
|
|
|
|
a->set_fec(true);
|
|
|
|
|
a->set_expected_packet_loss(10);
|
|
|
|
|
a->set_dtx(true);
|
|
|
|
|
a->set_complexity(5);
|
|
|
|
|
}
|
feat(M1): TCP/TLS control plane -- auth, channels, ephemeral text
Implements the full M1 milestone. Two clients authenticate over TLS 1.3
(guest + Argon2id password) and exchange channel + private text messages
through a real server. All five ctest --preset m1-dev tests pass in ~1 s.
Key components added:
- vcpkg baseline + m1-dev preset (protobuf/mbedTLS/libsodium/asio/sqlite3)
- FrameCodec feed+emit, encode/decode_envelope, protobuf codegen
- TcpServerConn with blocking TLS handshake thread + tls_read_loop
- TlsContext (mbedTLS 1.3, ECDSA-P256 self-signed cert, TOFU on client)
- WorkerPool (3 threads, used for Argon2id)
- Database: SQLite + libsodium Argon2id, account lifecycle, bootstrap admin
- ServerIdentityManager: Ed25519 key + cert generate/persist/fingerprint
- ConnSession state machine: WaitingHello -> WaitingAuth -> Authenticated
- SessionRegistry: channel tree, user map, text routing, broadcast
- vc_client full M1 C ABI: connect/TLS/handshake/auth/text/disconnect
- voicecat-admin CLI: account add/reset/del/list
- test_m1_integration: M1 exit criterion, verified green
Bug fixed: double-framing in ConnSession::send_envelope -- encode_envelope
was adding the [4-byte len] prefix, then TcpServerConn::send_frame added
a second one, causing the client to parse [len][proto] as protobuf (silent
failure). Fixed by serializing raw protobuf bytes in send_envelope and
letting send_frame apply the single length prefix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 23:48:44 +02:00
|
|
|
channels_[1] = std::move(lobby);
|
feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC +
SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise-
reduction, talk indicators, and enforced per-channel Opus configurability
(mono/stereo, bitrate, frame size, FEC/DTX, application).
Bugs fixed along the way (found while implementing, not pre-existing scope):
- Server hard-coded stream_id=1 for every announce, so a second stream from
the same user silently overwrote the first in SessionRegistry::set_user_stream.
Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop
validates against announced_stream_ids_ before clearing.
- Client dropped mode/dtx/complexity/application from effective_audio even for
the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever
applied to OpusParams. Fixed on both the send (handle_stream_announce_result)
and receive (sync_remote_streams) paths via a shared
opus_params_from_audio_config() helper.
- OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application
and wired it through.
- on_playback's per-stream decode passed the wrong frame_size to opus_decode
(total samples instead of samples-per-channel), which would have overflowed
the decode buffer for any stereo stream.
- teardown_voice() raced when called concurrently from run_io()'s own cleanup
and from disconnect() on a different thread -- both could see
udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the
same std::thread (intermittent std::system_error under ctest). Fixed with a
teardown_mu_ guard instead of carrying the flake forward.
New:
- Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/
FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX);
handle_stream_announce enforces the channel's config, clamping (not
overriding) bitrate_bps to its ceiling.
- core/src/core/client.h/.cpp: local-stream state is now a
std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with
request_id-correlated announce/result handling (request_id already
round-tripped on the wire; just wasn't read before). on_capture_frame is
kind-aware and upmixes mono capture to stereo when a stream's config calls
for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired
through set_remote_stream. New run_talk_timer() thread emits
VC_EVENT_TALK_STATE from both remote and local edge detection.
- core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps
(inject_capture), stereo-to-mono downmix at the decode/mix boundary,
RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled
and last_voice_ms/talking; new set_stream_noise_reduction() and
poll_talk_transitions().
- core/src/session/session.h/.cpp: Stream now carries the full AudioConfig,
not just sample_rate/frame_ms.
- New additive C ABI (core/include/voicecat.h): vc_audio_config +
vc_get_stream_audio_config (effective Opus config for any stream you own or
a peer's); vc_test_inject_capture (test-only synthetic PCM injection,
clearly marked, mirrors AudioEngine::inject_capture).
- tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI
(mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two
concurrent local streams, independent gain/mute/NS control, per-channel
config divergence via vc_get_stream_audio_config, talk indicators.
Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently
dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback
capture for SCREEN_AUDIO (synthetic injection only); true stereo playback
output (AudioEngine's mixer/output device stays mono -- Opus itself is fully
stereo-correct on the wire).
ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive
full-suite runs plus 8 standalone runs of the new test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
|
|
|
|
|
|
|
|
ChannelEntry music;
|
|
|
|
|
music.proto.set_id(2);
|
|
|
|
|
music.proto.set_name("Music Room");
|
|
|
|
|
music.proto.set_type(voicecat::v1::CHANNEL_PERMANENT);
|
|
|
|
|
music.proto.set_order(1);
|
|
|
|
|
{
|
|
|
|
|
// Music/screen-audio profile: stereo, high bitrate, FEC/DTX off (continuous signal).
|
|
|
|
|
auto* a = music.proto.mutable_audio();
|
|
|
|
|
a->set_codec(0);
|
|
|
|
|
a->set_mode(voicecat::v1::MODE_STEREO);
|
|
|
|
|
a->set_sample_rate(48000);
|
|
|
|
|
a->set_bitrate_bps(128000);
|
|
|
|
|
a->set_frame_ms(20);
|
|
|
|
|
a->set_application(voicecat::v1::OPUS_AUDIO);
|
|
|
|
|
a->set_fec(false);
|
|
|
|
|
a->set_expected_packet_loss(0);
|
|
|
|
|
a->set_dtx(false);
|
|
|
|
|
a->set_complexity(8);
|
|
|
|
|
}
|
|
|
|
|
channels_[2] = std::move(music);
|
|
|
|
|
|
|
|
|
|
next_channel_id_ = 3; // 1 and 2 are now reserved (Lobby, Music Room)
|
feat(M1): TCP/TLS control plane -- auth, channels, ephemeral text
Implements the full M1 milestone. Two clients authenticate over TLS 1.3
(guest + Argon2id password) and exchange channel + private text messages
through a real server. All five ctest --preset m1-dev tests pass in ~1 s.
Key components added:
- vcpkg baseline + m1-dev preset (protobuf/mbedTLS/libsodium/asio/sqlite3)
- FrameCodec feed+emit, encode/decode_envelope, protobuf codegen
- TcpServerConn with blocking TLS handshake thread + tls_read_loop
- TlsContext (mbedTLS 1.3, ECDSA-P256 self-signed cert, TOFU on client)
- WorkerPool (3 threads, used for Argon2id)
- Database: SQLite + libsodium Argon2id, account lifecycle, bootstrap admin
- ServerIdentityManager: Ed25519 key + cert generate/persist/fingerprint
- ConnSession state machine: WaitingHello -> WaitingAuth -> Authenticated
- SessionRegistry: channel tree, user map, text routing, broadcast
- vc_client full M1 C ABI: connect/TLS/handshake/auth/text/disconnect
- voicecat-admin CLI: account add/reset/del/list
- test_m1_integration: M1 exit criterion, verified green
Bug fixed: double-framing in ConnSession::send_envelope -- encode_envelope
was adding the [4-byte len] prefix, then TcpServerConn::send_frame added
a second one, causing the client to parse [len][proto] as protobuf (silent
failure). Fixed by serializing raw protobuf bytes in send_envelope and
letting send_frame apply the single length prefix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 23:48:44 +02:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
uint64_t SessionRegistry::register_session(std::weak_ptr<ConnSession> session) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
uint64_t id = next_session_id_++;
|
|
|
|
|
sessions_[id] = std::move(session);
|
|
|
|
|
return id;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void SessionRegistry::unregister_session(uint64_t session_id) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
sessions_.erase(session_id);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
uint32_t SessionRegistry::add_user(uint64_t session_id, const voicecat::v1::User& user) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
uint32_t uid = next_user_id_++;
|
|
|
|
|
UserEntry entry;
|
|
|
|
|
entry.proto = user;
|
|
|
|
|
entry.proto.set_id(uid);
|
|
|
|
|
entry.proto.set_channel_id(1); // start in Lobby
|
|
|
|
|
entry.session_id = session_id;
|
|
|
|
|
users_[uid] = std::move(entry);
|
|
|
|
|
return uid;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void SessionRegistry::remove_user(uint32_t user_id) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
users_.erase(user_id);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
bool SessionRegistry::set_user_channel(uint32_t user_id, uint32_t channel_id) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
auto ch_it = channels_.find(channel_id);
|
|
|
|
|
if (ch_it == channels_.end()) return false;
|
|
|
|
|
auto user_it = users_.find(user_id);
|
|
|
|
|
if (user_it == users_.end()) return false;
|
|
|
|
|
user_it->second.proto.set_channel_id(channel_id);
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<voicecat::v1::Channel> SessionRegistry::channel_snapshot() const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
std::vector<voicecat::v1::Channel> result;
|
|
|
|
|
result.reserve(channels_.size());
|
|
|
|
|
for (auto& [id, entry] : channels_) result.push_back(entry.proto);
|
|
|
|
|
return result;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<voicecat::v1::User> SessionRegistry::user_snapshot() const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
std::vector<voicecat::v1::User> result;
|
|
|
|
|
result.reserve(users_.size());
|
|
|
|
|
for (auto& [id, entry] : users_) result.push_back(entry.proto);
|
|
|
|
|
return result;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<std::shared_ptr<ConnSession>> SessionRegistry::resolve_text_targets(
|
|
|
|
|
uint64_t sender_session_id, voicecat::v1::TextScope scope, uint32_t target_id) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
std::vector<std::shared_ptr<ConnSession>> targets;
|
|
|
|
|
|
|
|
|
|
if (scope == voicecat::v1::TEXT_CHANNEL) {
|
|
|
|
|
// Find channel_id of the target, then all users in that channel
|
|
|
|
|
for (auto& [uid, entry] : users_) {
|
|
|
|
|
if (entry.proto.channel_id() != target_id) continue;
|
|
|
|
|
if (entry.session_id == sender_session_id) continue;
|
|
|
|
|
auto sit = sessions_.find(entry.session_id);
|
|
|
|
|
if (sit == sessions_.end()) continue;
|
|
|
|
|
if (auto sess = sit->second.lock()) targets.push_back(sess);
|
|
|
|
|
}
|
|
|
|
|
} else if (scope == voicecat::v1::TEXT_PRIVATE) {
|
|
|
|
|
// target_id is user_id
|
|
|
|
|
auto user_it = users_.find(target_id);
|
|
|
|
|
if (user_it != users_.end()) {
|
|
|
|
|
auto sit = sessions_.find(user_it->second.session_id);
|
|
|
|
|
if (sit != sessions_.end()) {
|
|
|
|
|
if (auto sess = sit->second.lock()) targets.push_back(sess);
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
return targets;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void SessionRegistry::broadcast(const voicecat::v1::Envelope& env,
|
|
|
|
|
uint64_t exclude_session_id) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
for (auto& [sid, weak] : sessions_) {
|
|
|
|
|
if (sid == exclude_session_id) continue;
|
|
|
|
|
if (auto sess = weak.lock()) sess->send_envelope(env);
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
feat(M2): UDP voice/media plane -- SFU relay, Opus, AEAD, jitter buffer
Adds the full voice pipeline: 14-byte binary frame header, ChaCha20-Poly1305
AEAD keyed from the TLS exporter, libopus encode/decode with FEC/PLC/DTX,
an adaptive per-ssrc jitter buffer, a miniaudio capture/playback engine, an
APM passthrough stub, and the UdpBinding/StreamAnnounce signaling chain
wired through ConnSession/SessionRegistry into a new server-side SFU
(MediaRelay) that decrypts and re-encrypts frames per channel member.
Exit criterion verified: test_m2_voice — two headless clients relay 50
encrypted Opus frames through the server; ctest --preset m1-dev is 9/9
green. Also corrects protocol.md's UdpBinding diagram, which described the
UDP-side binding packet as AEAD-sealed when it is in fact a plaintext
bootstrap frame (separate from the TCP/TLS UdpBinding ack).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 01:31:14 +02:00
|
|
|
// ── M2: UDP / media ──────────────────────────────────────────────────────────
|
|
|
|
|
|
|
|
|
|
void SessionRegistry::register_udp_token(const std::array<uint8_t, 16>& token,
|
|
|
|
|
uint64_t session_id) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
udp_tokens_[token] = session_id;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::shared_ptr<ConnSession> SessionRegistry::find_by_udp_token(
|
|
|
|
|
const std::array<uint8_t, 16>& token) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
auto it = udp_tokens_.find(token);
|
|
|
|
|
if (it == udp_tokens_.end()) return nullptr;
|
|
|
|
|
auto sit = sessions_.find(it->second);
|
|
|
|
|
if (sit == sessions_.end()) return nullptr;
|
|
|
|
|
return sit->second.lock();
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void SessionRegistry::register_udp_endpoint(asio::ip::udp::endpoint ep,
|
|
|
|
|
uint64_t session_id) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
udp_endpoints_[ep] = session_id;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::shared_ptr<ConnSession> SessionRegistry::find_by_udp_endpoint(
|
|
|
|
|
const asio::ip::udp::endpoint& ep) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
auto it = udp_endpoints_.find(ep);
|
|
|
|
|
if (it == udp_endpoints_.end()) return nullptr;
|
|
|
|
|
auto sit = sessions_.find(it->second);
|
|
|
|
|
if (sit == sessions_.end()) return nullptr;
|
|
|
|
|
return sit->second.lock();
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
uint32_t SessionRegistry::assign_ssrc(uint64_t session_id) {
|
|
|
|
|
uint32_t ssrc = next_ssrc_.fetch_add(1, std::memory_order_relaxed);
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
ssrc_to_session_[ssrc] = session_id;
|
|
|
|
|
return ssrc;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<std::shared_ptr<ConnSession>> SessionRegistry::find_channel_sessions(
|
|
|
|
|
uint32_t channel_id, uint64_t exclude_session_id) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
std::vector<std::shared_ptr<ConnSession>> result;
|
|
|
|
|
for (auto& [uid, entry] : users_) {
|
|
|
|
|
if (entry.proto.channel_id() != channel_id) continue;
|
|
|
|
|
if (entry.session_id == exclude_session_id) continue;
|
|
|
|
|
auto sit = sessions_.find(entry.session_id);
|
|
|
|
|
if (sit == sessions_.end()) continue;
|
|
|
|
|
if (auto sess = sit->second.lock()) result.push_back(sess);
|
|
|
|
|
}
|
|
|
|
|
return result;
|
|
|
|
|
}
|
|
|
|
|
|
2026-06-16 02:12:50 +02:00
|
|
|
std::optional<voicecat::v1::User> SessionRegistry::set_user_stream(
|
|
|
|
|
uint32_t user_id, const voicecat::v1::StreamInfo& info) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
auto it = users_.find(user_id);
|
|
|
|
|
if (it == users_.end()) return std::nullopt;
|
|
|
|
|
auto* streams = it->second.proto.mutable_streams();
|
|
|
|
|
for (int i = 0; i < streams->size(); ++i) {
|
|
|
|
|
if (streams->Get(i).stream_id() == info.stream_id()) {
|
|
|
|
|
*streams->Mutable(i) = info;
|
|
|
|
|
return it->second.proto;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
*streams->Add() = info;
|
|
|
|
|
return it->second.proto;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::optional<voicecat::v1::User> SessionRegistry::clear_user_stream(uint32_t user_id,
|
|
|
|
|
uint32_t stream_id) {
|
|
|
|
|
std::unique_lock lk(mu_);
|
|
|
|
|
auto it = users_.find(user_id);
|
|
|
|
|
if (it == users_.end()) return std::nullopt;
|
|
|
|
|
auto* streams = it->second.proto.mutable_streams();
|
|
|
|
|
for (int i = 0; i < streams->size(); ++i) {
|
|
|
|
|
if (streams->Get(i).stream_id() == stream_id) {
|
|
|
|
|
streams->erase(streams->begin() + i);
|
|
|
|
|
break;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
return it->second.proto;
|
|
|
|
|
}
|
|
|
|
|
|
feat(M2): UDP voice/media plane -- SFU relay, Opus, AEAD, jitter buffer
Adds the full voice pipeline: 14-byte binary frame header, ChaCha20-Poly1305
AEAD keyed from the TLS exporter, libopus encode/decode with FEC/PLC/DTX,
an adaptive per-ssrc jitter buffer, a miniaudio capture/playback engine, an
APM passthrough stub, and the UdpBinding/StreamAnnounce signaling chain
wired through ConnSession/SessionRegistry into a new server-side SFU
(MediaRelay) that decrypts and re-encrypts frames per channel member.
Exit criterion verified: test_m2_voice — two headless clients relay 50
encrypted Opus frames through the server; ctest --preset m1-dev is 9/9
green. Also corrects protocol.md's UdpBinding diagram, which described the
UDP-side binding packet as AEAD-sealed when it is in fact a plaintext
bootstrap frame (separate from the TCP/TLS UdpBinding ack).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 01:31:14 +02:00
|
|
|
uint32_t SessionRegistry::user_channel(uint32_t user_id) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
auto it = users_.find(user_id);
|
|
|
|
|
return (it == users_.end()) ? 0 : it->second.proto.channel_id();
|
|
|
|
|
}
|
|
|
|
|
|
feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC +
SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise-
reduction, talk indicators, and enforced per-channel Opus configurability
(mono/stereo, bitrate, frame size, FEC/DTX, application).
Bugs fixed along the way (found while implementing, not pre-existing scope):
- Server hard-coded stream_id=1 for every announce, so a second stream from
the same user silently overwrote the first in SessionRegistry::set_user_stream.
Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop
validates against announced_stream_ids_ before clearing.
- Client dropped mode/dtx/complexity/application from effective_audio even for
the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever
applied to OpusParams. Fixed on both the send (handle_stream_announce_result)
and receive (sync_remote_streams) paths via a shared
opus_params_from_audio_config() helper.
- OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application
and wired it through.
- on_playback's per-stream decode passed the wrong frame_size to opus_decode
(total samples instead of samples-per-channel), which would have overflowed
the decode buffer for any stereo stream.
- teardown_voice() raced when called concurrently from run_io()'s own cleanup
and from disconnect() on a different thread -- both could see
udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the
same std::thread (intermittent std::system_error under ctest). Fixed with a
teardown_mu_ guard instead of carrying the flake forward.
New:
- Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/
FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX);
handle_stream_announce enforces the channel's config, clamping (not
overriding) bitrate_bps to its ceiling.
- core/src/core/client.h/.cpp: local-stream state is now a
std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with
request_id-correlated announce/result handling (request_id already
round-tripped on the wire; just wasn't read before). on_capture_frame is
kind-aware and upmixes mono capture to stereo when a stream's config calls
for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired
through set_remote_stream. New run_talk_timer() thread emits
VC_EVENT_TALK_STATE from both remote and local edge detection.
- core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps
(inject_capture), stereo-to-mono downmix at the decode/mix boundary,
RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled
and last_voice_ms/talking; new set_stream_noise_reduction() and
poll_talk_transitions().
- core/src/session/session.h/.cpp: Stream now carries the full AudioConfig,
not just sample_rate/frame_ms.
- New additive C ABI (core/include/voicecat.h): vc_audio_config +
vc_get_stream_audio_config (effective Opus config for any stream you own or
a peer's); vc_test_inject_capture (test-only synthetic PCM injection,
clearly marked, mirrors AudioEngine::inject_capture).
- tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI
(mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two
concurrent local streams, independent gain/mute/NS control, per-channel
config divergence via vc_get_stream_audio_config, talk indicators.
Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently
dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback
capture for SCREEN_AUDIO (synthetic injection only); true stereo playback
output (AudioEngine's mixer/output device stays mono -- Opus itself is fully
stereo-correct on the wire).
ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive
full-suite runs plus 8 standalone runs of the new test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00
|
|
|
std::optional<voicecat::v1::AudioConfig> SessionRegistry::channel_audio_config(
|
|
|
|
|
uint32_t channel_id) const {
|
|
|
|
|
std::shared_lock lk(mu_);
|
|
|
|
|
auto it = channels_.find(channel_id);
|
|
|
|
|
if (it == channels_.end()) return std::nullopt;
|
|
|
|
|
return it->second.proto.audio();
|
|
|
|
|
}
|
|
|
|
|
|
feat(M1): TCP/TLS control plane -- auth, channels, ephemeral text
Implements the full M1 milestone. Two clients authenticate over TLS 1.3
(guest + Argon2id password) and exchange channel + private text messages
through a real server. All five ctest --preset m1-dev tests pass in ~1 s.
Key components added:
- vcpkg baseline + m1-dev preset (protobuf/mbedTLS/libsodium/asio/sqlite3)
- FrameCodec feed+emit, encode/decode_envelope, protobuf codegen
- TcpServerConn with blocking TLS handshake thread + tls_read_loop
- TlsContext (mbedTLS 1.3, ECDSA-P256 self-signed cert, TOFU on client)
- WorkerPool (3 threads, used for Argon2id)
- Database: SQLite + libsodium Argon2id, account lifecycle, bootstrap admin
- ServerIdentityManager: Ed25519 key + cert generate/persist/fingerprint
- ConnSession state machine: WaitingHello -> WaitingAuth -> Authenticated
- SessionRegistry: channel tree, user map, text routing, broadcast
- vc_client full M1 C ABI: connect/TLS/handshake/auth/text/disconnect
- voicecat-admin CLI: account add/reset/del/list
- test_m1_integration: M1 exit criterion, verified green
Bug fixed: double-framing in ConnSession::send_envelope -- encode_envelope
was adding the [4-byte len] prefix, then TcpServerConn::send_frame added
a second one, causing the client to parse [len][proto] as protobuf (silent
failure). Fixed by serializing raw protobuf bytes in send_envelope and
letting send_frame apply the single length prefix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 23:48:44 +02:00
|
|
|
} // namespace voicecat::server
|
|
|
|
|
|
|
|
|
|
#endif // VOICECAT_HAS_NET
|