Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
18 KiB
PROGRESS — VoiceCat
Living status. Update this file in the same commit as your work so the next agent picks up instantly. Newest status at the top.
- Date convention: ISO (YYYY-MM-DD).
- Statuses:
[ ]not started ·[~]in progress ·[x]done.
▶ Where we left off / next action
- Done: M3 — multi-stream & per-channel tuning ✓ complete (2026-06-16). See the M3
section below for the full file-by-file change list.
ctest --test-dir build/m1-dev— 11/11 tests green (3 consecutive full-suite runs), including the newtest_m3_multistream(realvc_clients, not raw sockets — same lesson as M2: ABI-level coverage is what proves the client library, not just the wire protocol). Two items intentionally still open (carried forward, not silently dropped):vc_set_input_device/vc_set_input_mode/vc_set_push_to_talk/vc_list_devices(device enumeration + VAD/PTT input gate) remainVC_ERR_NOT_IMPLEMENTED— explicitly scoped out of this M3 pass; revisit in a future milestone.- Stereo Opus is now wire-correct end-to-end (a channel configured
MODE_STEREOreally encodes/decodes 2-channel Opus packets), butAudioEngine's playback mixer/output device stays mono internally — stereo streams are downmixed (avg L/R) immediately after decode, before mixing. True stereo playback output is a follow-up, not part of M3.
- Next: M4 — native clients (Windows C#, macOS/iOS Swift). See
docs/roadmap.md §M4.
Milestones (see docs/roadmap.md for full detail)
- M0 — Scaffolding ✓ complete
- M1 — Control plane ✓ complete (2026-06-15)
- M2 — Voice, single stream ✓ complete (2026-06-16)
- M3 — Multi-stream & per-channel tuning ✓ complete (2026-06-16)
- M4 — Native clients (Windows C#, macOS/iOS Swift) ← next
- M5 — Moderation, polish, beyond (perms, bans, DRED; then file transfer, E2EE, …)
M0 — Scaffolding ✓ (completed)
- Repo layout (
core/ server/ tools/ clients/ tests/), CMake + presets, vcpkg manifest. - C ABI header
core/include/voicecat.h(full surface, stubbed). - Protocol source-of-truth
core/proto/voicecat.proto(matches docs/protocol.md). - Core stubs for all six subsystems (net/crypto/codec/protocol/session/audio) +
vc_client. voicecat-server(arg parsing, config, stub run) andvccli(drives the C ABI).- CTest smoke test asserting the C ABI contract (not just "it compiles").
.gitattributes(LF),.gitignore,.clang-format, onboarding docs.- Verified:
cmake --preset dev && cmake --build --preset dev && ctest --preset dev→ green.
M1 — Control plane ✓ (completed 2026-06-15)
Exit criterion: ✓ test_m1_integration — two clients authenticate over TLS 1.3 (guest
- Argon2id password), exchange channel and private text messages. Passes in ~1 s.
- vcpkg baseline +
m1-devpreset;find_packagefor protobuf/mbedTLS/libsodium/asio/sqlite3. FrameCodecfeed + emit;encode_envelope/decode_envelope.- Asio TCP acceptor +
TcpServerConn(TLS path: blocking handshake thread +tls_read_loop). TlsContext(mbedTLS 1.3, server cert/identity, ECDSA-P256 self-signed, TOFU on client).WorkerPool(3 threads, used for Argon2id).Database— SQLite, Argon2id via libsodium,create_account/authenticate/ bootstrap admin.voicecat-admin— account add/reset/del/list against live DB file.ServerIdentityManager— generate/persist Ed25519 key + cert; fingerprint display.ConnSession— WaitingHello → WaitingAuth → Authenticated state machine; full protocol relay.SessionRegistry— channel tree, user map, broadcast, text routing.vc_client(client.cpp) — full M1 C ABI: connect/TLS/ClientHello/AuthRequest/text/disconnect.Server::run()— io_context, acceptor, worker pool, signal handling,on_readycallback.test_m1_integration— M1 exit criterion. Verified green 2026-06-15.
Key bug fixed: double-framing in ConnSession::send_envelope — encode_envelope was
pre-framing the protobuf, then TcpServerConn::send_frame re-framed it. Fixed by serializing
raw protobuf bytes directly and letting send_frame add the single [4-byte len] prefix.
M2 — Voice, single stream ✓ (completed 2026-06-16)
Exit criterion: ✓ test_m2_voice — two headless clients authenticate over TLS, bind UDP,
announce a MIC stream, send 50 encrypted Opus frames; server SFU relay re-encrypts + forwards
to the second client; B receives ≥ 25 frames and all decrypt correctly. Passes in ~4 s.
m2-devpreset (inheritsvcpkg-base, binaryDirbuild/m2-dev);m1-devalso builds all M2 code.core/CMakeLists.txt—find_package(Opus),find_path(MINIAUDIO_INCLUDE_DIR).core/src/net/voice_frame.h— 14-byte UDP header (type/flags/codec/ssrc/seq/ts), serialize/parse,make_udp_binding_packet.SodiumMediaCrypto— ChaCha20-Poly1305 AEAD; counter-nonce; 64-bit sliding-window anti-replay;derive_send/recvfrom TLS RFC 5705 exporter.OpusEncoder/OpusDecoder— libopus 1.6, FEC, DTX, PLC (free; nullptr → decoder extrapolates).UdpMediaChannel— async UDP socket (asio); thread-safesend_to; async recv loop.JitterBuffer— per-ssrc, EWMA jitter estimation, adaptive depth 20–200 ms, late-drop at 500 ms.AudioEngine— miniaudio capture+playback;inject_capture()bypass for headless tests; per-ssrc RemoteStream with OpusDecoder + JitterBuffer.ApmProcessor—ApmPassthroughstub (VAD always open); WebRTC APM deferred until M3.on_tls_readycallback inTcpChannelCallbacks— server derives and stores media AEAD keys immediately after TLS handshake.ConnSessionM2 —udp_tokengenerated at construction; included inAuthResult;handle_udp_binding(verifies token, TCP ack);handle_stream_announce(assigns SSRC via registry);udp_media_portinServerHello.SessionRegistryM2 —register_udp_token,find_by_udp_token,register_udp_endpoint,find_by_udp_endpoint,assign_ssrc,find_channel_sessions,user_channel.MediaRelay— SFU UDP relay;kFrameUdpBinding→ endpoint binding;kFrameVoice→ decrypt/re-encrypt/forward to channel members.Server::run()— creates and bindsMediaRelay; passes media port toConnSession; wireson_tls_readyto derive per-connection media AEAD keys.test_voice_frame— header round-trip, big-endian layout, binding packet format.test_media_aead— seal/open round-trip, anti-replay, tamper detection, multi-packet sequence.test_opus_codec— encode/decode round-trip energy check (within 3 dB), PLC, frame-samples helper.test_m2_voice— M2 exit criterion (raw-socket harness). Verified green 2026-06-16.
Follow-up (same day): the above made test_m2_voice pass, but vc_client's public voice
methods were still stubs — the actual M2 exit criterion ("two vccli/early-GUI clients talk")
wasn't met. Closed the gap:
core/src/core/client.cpp— realstream_start/stream_stop/set_self_mute/set_remote_stream; UDP-binding handshake (start_udp_binding/handle_udp_binding_ack/finish_udp_binding); media key derivation fromtls_(RFC 5705 exporter);run_udp_recv(AEAD-open →JitterBuffer::Frame→audio_engine_.push_recv_frame);on_capture_frame(encode → seal →sendto);sync_remote_streams(diffs aUserproto'sstreamsagainstremote_streams_, wiring upOpusDecoders and emittingSTREAM_STARTED/STOPPED).set_input_device/set_input_mode/set_push_to_talk/list_devicesremainVC_ERR_NOT_IMPLEMENTED— no device-enumeration backend yet; scoped to M3 (VAD/PTT).core/src/session/session.cpp/h—SessionModel::find_user,find_user_by_ssrc,Stream{stream_id, ssrc, kind, label, sample_rate, frame_ms}.server/src/conn_session.cpp/h—handle_stream_announce/handle_stream_stopnow broadcast viaSessionRegistry::set_user_stream/clear_user_stream→UserEvent::UPDATED.server/src/session_registry.cpp/h—set_user_stream/clear_user_stream(mutate a user'sStreamInfolist, return the updatedUserproto for broadcast).tests/test_voice_client_abi.cpp— drives two realvc_clientinstances throughvc_connect/vc_authenticate_guest/vc_stream_start/vc_stream_stop; asserts client B observes client A'sSTREAM_STARTED/STOPPEDevents. Verified green 2026-06-16.tools/vccli/src/main.cpp— argv parsing (--host/--port/--nick/--channel/--voice/ --mute/--text);--voicestarts a MIC stream and blocks on SIGINT, printingon_eventcallbacks live (unbuffered stdout — MinGW/MSVCRT treat_IOLBFas full buffering for non-console streams). Dropped the originally-planned--voice-loopbackand thetx=N rx=M lost=K jitter=Jstats line:voicecat.hexposes no PCM-injection hook or jitter/loss stats getter publicly, onlyon_event+on_level(RMS). Manually verified: twovccli --voiceinstances see each other's stream start in real time.
M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)
Exit criterion: ✓ test_m3_multistream — a real vc_client (A) runs two concurrent local
streams (MIC + SCREEN_AUDIO) with distinct stream ids; a second client (B) sees both as
separate STREAM_STARTED events and a VC_EVENT_TALK_STATE talking edge for A's MIC stream;
B independently sets gain/mute/noise-reduction on each of A's streams without one call
affecting the other; A then joins "Music Room" (channel 2: stereo/128kbps/OPUS_AUDIO/no
DTX) and announces a fresh MIC stream there, while B stays in "Lobby" (channel 1: mono/24kbps/
OPUS_VOIP/DTX on) — vc_get_stream_audio_config shows their effective Opus config differs
exactly as the server enforces per channel. Passes in ~2.4s; verified across 8 consecutive
standalone runs + 3 consecutive full-suite runs with no flakes.
Exploration before implementing turned up several bugs/gaps where the wire format already supported this milestone but the client/server logic didn't — these were fixed as part of M3, not treated as pre-existing-and-out-of-scope:
- Server
stream_idbug —handle_stream_announcealways wrotestream_id=1, so a second stream from the same user silently overwrote the first inSessionRegistry::set_user_stream's replace-by-id logic. Fixed with a per-session counter (ConnSession::next_stream_id_) +announced_stream_ids_(also now validated inhandle_stream_stop, rejecting stops for ids the session never announced). - Per-channel
AudioConfigwas modeled but never populated/enforced.SessionRegistry::init_default_channels()now seeds Lobby (id=1: mono, 24kbps,OPUS_VOIP, FEC+DTX on) and a new "Music Room" (id=2: stereo, 128kbps,OPUS_AUDIO, FEC+DTX off) with realAudioConfigs; newSessionRegistry::channel_audio_config(channel_id)accessor (there was no per-id channel getter before, onlychannel_snapshot()).handle_stream_announcenow treats the channel's config as authoritative (mode/frame_ms/application/fec/dtx/ complexity), clamping (not overriding)bitrate_bpsto the channel's ceiling. - Client silently dropped
mode/dtx/complexity/applicationfromeffective_audioeven for the single M2 stream —handle_stream_announce_resultandsync_remote_streamsonly copiedsample_rate/bitrate_bps/frame_ms/fecintoOpusParams. New sharedopus_params_from_audio_config()helper (client.cpp) fixes both the send and receive paths. core/src/codec/opus_codec.h/.cpp— newOpusApplicationenum +OpusParams::applicationfield;OpusEncoder::initnow honors it instead of hardcodingOPUS_APPLICATION_VOIP.core/src/session/session.h/.cpp—Streamstruct extended with the fullAudioConfig(mode/bitrate_bps/application/fec/expected_packet_loss/dtx/complexity), not just sample_rate/frame_ms;copy_streams()now copies all of it.core/src/core/client.h/.cpp— local-stream state is now astd::unordered_map<int, LocalStream>keyed byvc_stream_kind(one active stream per kind — MIC/SCREEN_AUDIO/ AUX_DEVICE are each singletons for a client), replacing the M2 single-stream fields.StreamAnnounce/StreamAnnounceResultround-trips are now correlated byrequest_id(already round-tripped on the wire; just wasn't read) viapending_announce_kind_, so multiple concurrent announces from one client resolve to the rightLocalStream.on_capture_frametakes akindparameter and upmixes mono capture to stereo (duplicate L=R) when a stream's channel config calls for it.vc_set_self_mute'smic_mutedonly gates theMICkind — a concurrentSCREEN_AUDIOshare keeps playing while muted.set_remote_streamnow actually wiresnoise_reductionthrough (previously parsed and discarded). Newrun_talk_timer()(a small dedicated thread, started alongside the UDP media path, never the miniaudio callback thread) polls both remote talk-state edges (AudioEngine::poll_talk_transitions()) and local capture-activity edges, emittingVC_EVENT_TALK_STATE.- Fixed a thread-join race in
teardown_voice()— it's called both fromrun_io()'s own cleanup and fromdisconnect(), on different threads; without serialization both could seeudp_thread_/talk_timer_thread_asjoinable()simultaneously and race tojoin()the samestd::thread(UB; surfaced as an intermittentstd::system_error: No such processunderctest). Added ateardown_mu_guard around the whole function. This pre-existed forudp_thread_alone (likely the same root cause as thetest_m1_integration/test_m2_voicecleanup-path flake noted in the M2 section above) — addingtalk_timer_thread_'s join just made it surface more often, so it was fixed properly here rather than carried forward again. core/src/audio/audio_engine.h/.cpp—CaptureCallbackgained akindparameter (the real miniaudio capture device is always taggedkind=0/MIC; a second concurrent local stream is fed via its owninject_capture(kind, ...)ring buffer —inject_taps_, keyed by kind — since there is only one real hardware capture device in M3). Fixed a buffer-sizing bug inon_playback's per-stream decode (opus_decode'sframe_sizeparameter is samples-per-channel, not total samples — the old code passedframes * params_.channels, which would have overflowed the decode buffer for any stereo stream). Stereo decoder output is downmixed (avg L/R) into the engine's mono mix accumulator immediately after decode.RemoteStreamgainedrecv_ns/noise_reduction_enabled(lazyApmProcessorinstantiation — freed on disable, so no separate instance cap is needed per the roadmap's guidance) andlast_voice_ms/talking(talk-indicator edge state, updated inpush_recv_frame); newset_stream_noise_reduction()andpoll_talk_transitions(). Note: untilVOICECAT_HAS_APMis wired to a real WebRTC APM build, the NS toggle is plumbed end-to-end but behaviorally a passthrough no-op (ApmPassthroughdoesn't touch PCM) — same situation send-side APM has been in since M2; M3's job was the plumbing, not the DSP backend.- New C ABI surface (
core/include/voicecat.h, additive only):vc_audio_configstruct +vc_get_stream_audio_config(c, user_id, stream_id, out)— the effective Opus config for a stream you own or a peer's, reading from the (now richer)LocalStream/session::Stream.vc_test_inject_capture(c, stream_id, pcm, samples)— clearly-marked test-only, forwards toAudioEngine::inject_capture, sotest_m3_multistreamcan drive two concurrent synthetic-audio streams through the real ABI without a microphone. tests/test_m3_multistream.cpp— the M3 exit criterion (ABI-level, mirrorstest_voice_client_abi.cpp's approach per the M2 lesson). Registered intests/CMakeLists.txt.
Explicitly out of scope for this pass (confirmed with the user before implementing):
vc_set_input_device/vc_set_input_mode/vc_set_push_to_talk/vc_list_devices(device enumeration + VAD/PTT input gate) — stillVC_ERR_NOT_IMPLEMENTED. These were mentioned as "scoped to M3" in the M2 follow-up notes above, but docs/roadmap.md's M3 bullets never actually listed them — deferred again, now tracked explicitly rather than implicitly.- Real WASAPI desktop-audio loopback capture for
SCREEN_AUDIO— the engine now supports feeding a second concurrent local stream viainject_capture, but only synthetic PCM is wired up; a real loopback capture device is a follow-up. - True stereo playback output —
AudioEngine's mixer/output device stays mono; stereo streams are downmixed after decode (see above). The Opus wire format itself is fully stereo-correct.
Decisions log
All architecture/scope decisions are settled and recorded in
docs/roadmap.md §2 "Resolved decisions" and reflected across docs/.
If you make a new decision, record it there and link it here.
How to update this file
- Check off tasks as you complete them; flip a milestone to
[x]only when its exit criterion test passes. - Keep the "Where we left off / next action" block at the top accurate — it's the first thing the next agent reads.
- When you start a milestone, copy its task list from
docs/roadmap.mdinto a section here.