Files
voice-cat/PROGRESS.md
Talon 867557eda1 feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC +
SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise-
reduction, talk indicators, and enforced per-channel Opus configurability
(mono/stereo, bitrate, frame size, FEC/DTX, application).

Bugs fixed along the way (found while implementing, not pre-existing scope):
- Server hard-coded stream_id=1 for every announce, so a second stream from
  the same user silently overwrote the first in SessionRegistry::set_user_stream.
  Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop
  validates against announced_stream_ids_ before clearing.
- Client dropped mode/dtx/complexity/application from effective_audio even for
  the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever
  applied to OpusParams. Fixed on both the send (handle_stream_announce_result)
  and receive (sync_remote_streams) paths via a shared
  opus_params_from_audio_config() helper.
- OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application
  and wired it through.
- on_playback's per-stream decode passed the wrong frame_size to opus_decode
  (total samples instead of samples-per-channel), which would have overflowed
  the decode buffer for any stereo stream.
- teardown_voice() raced when called concurrently from run_io()'s own cleanup
  and from disconnect() on a different thread -- both could see
  udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the
  same std::thread (intermittent std::system_error under ctest). Fixed with a
  teardown_mu_ guard instead of carrying the flake forward.

New:
- Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/
  FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX);
  handle_stream_announce enforces the channel's config, clamping (not
  overriding) bitrate_bps to its ceiling.
- core/src/core/client.h/.cpp: local-stream state is now a
  std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with
  request_id-correlated announce/result handling (request_id already
  round-tripped on the wire; just wasn't read before). on_capture_frame is
  kind-aware and upmixes mono capture to stereo when a stream's config calls
  for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired
  through set_remote_stream. New run_talk_timer() thread emits
  VC_EVENT_TALK_STATE from both remote and local edge detection.
- core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps
  (inject_capture), stereo-to-mono downmix at the decode/mix boundary,
  RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled
  and last_voice_ms/talking; new set_stream_noise_reduction() and
  poll_talk_transitions().
- core/src/session/session.h/.cpp: Stream now carries the full AudioConfig,
  not just sample_rate/frame_ms.
- New additive C ABI (core/include/voicecat.h): vc_audio_config +
  vc_get_stream_audio_config (effective Opus config for any stream you own or
  a peer's); vc_test_inject_capture (test-only synthetic PCM injection,
  clearly marked, mirrors AudioEngine::inject_capture).
- tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI
  (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two
  concurrent local streams, independent gain/mute/NS control, per-channel
  config divergence via vc_get_stream_audio_config, talk indicators.

Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently
dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback
capture for SCREEN_AUDIO (synthetic injection only); true stereo playback
output (AudioEngine's mixer/output device stays mono -- Opus itself is fully
stereo-correct on the wire).

ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive
full-suite runs plus 8 standalone runs of the new test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 14:12:37 +02:00

18 KiB
Raw Blame History

PROGRESS — VoiceCat

Living status. Update this file in the same commit as your work so the next agent picks up instantly. Newest status at the top.

  • Date convention: ISO (YYYY-MM-DD).
  • Statuses: [ ] not started · [~] in progress · [x] done.

▶ Where we left off / next action

  • Done: M3 — multi-stream & per-channel tuning ✓ complete (2026-06-16). See the M3 section below for the full file-by-file change list. ctest --test-dir build/m1-dev11/11 tests green (3 consecutive full-suite runs), including the new test_m3_multistream (real vc_clients, not raw sockets — same lesson as M2: ABI-level coverage is what proves the client library, not just the wire protocol). Two items intentionally still open (carried forward, not silently dropped):
    • vc_set_input_device/vc_set_input_mode/vc_set_push_to_talk/vc_list_devices (device enumeration + VAD/PTT input gate) remain VC_ERR_NOT_IMPLEMENTED — explicitly scoped out of this M3 pass; revisit in a future milestone.
    • Stereo Opus is now wire-correct end-to-end (a channel configured MODE_STEREO really encodes/decodes 2-channel Opus packets), but AudioEngine's playback mixer/output device stays mono internally — stereo streams are downmixed (avg L/R) immediately after decode, before mixing. True stereo playback output is a follow-up, not part of M3.
  • Next: M4 — native clients (Windows C#, macOS/iOS Swift). See docs/roadmap.md §M4.

Milestones (see docs/roadmap.md for full detail)

  • M0 — Scaffolding ✓ complete
  • M1 — Control plane ✓ complete (2026-06-15)
  • M2 — Voice, single stream ✓ complete (2026-06-16)
  • M3 — Multi-stream & per-channel tuning ✓ complete (2026-06-16)
  • M4 — Native clients (Windows C#, macOS/iOS Swift) ← next
  • M5 — Moderation, polish, beyond (perms, bans, DRED; then file transfer, E2EE, …)

M0 — Scaffolding ✓ (completed)

  • Repo layout (core/ server/ tools/ clients/ tests/), CMake + presets, vcpkg manifest.
  • C ABI header core/include/voicecat.h (full surface, stubbed).
  • Protocol source-of-truth core/proto/voicecat.proto (matches docs/protocol.md).
  • Core stubs for all six subsystems (net/crypto/codec/protocol/session/audio) + vc_client.
  • voicecat-server (arg parsing, config, stub run) and vccli (drives the C ABI).
  • CTest smoke test asserting the C ABI contract (not just "it compiles").
  • .gitattributes (LF), .gitignore, .clang-format, onboarding docs.
  • Verified: cmake --preset dev && cmake --build --preset dev && ctest --preset dev → green.

M1 — Control plane ✓ (completed 2026-06-15)

Exit criterion:test_m1_integration — two clients authenticate over TLS 1.3 (guest

  • Argon2id password), exchange channel and private text messages. Passes in ~1 s.
  • vcpkg baseline + m1-dev preset; find_package for protobuf/mbedTLS/libsodium/asio/sqlite3.
  • FrameCodec feed + emit; encode_envelope / decode_envelope.
  • Asio TCP acceptor + TcpServerConn (TLS path: blocking handshake thread + tls_read_loop).
  • TlsContext (mbedTLS 1.3, server cert/identity, ECDSA-P256 self-signed, TOFU on client).
  • WorkerPool (3 threads, used for Argon2id).
  • Database — SQLite, Argon2id via libsodium, create_account / authenticate / bootstrap admin.
  • voicecat-admin — account add/reset/del/list against live DB file.
  • ServerIdentityManager — generate/persist Ed25519 key + cert; fingerprint display.
  • ConnSession — WaitingHello → WaitingAuth → Authenticated state machine; full protocol relay.
  • SessionRegistry — channel tree, user map, broadcast, text routing.
  • vc_client (client.cpp) — full M1 C ABI: connect/TLS/ClientHello/AuthRequest/text/disconnect.
  • Server::run() — io_context, acceptor, worker pool, signal handling, on_ready callback.
  • test_m1_integration — M1 exit criterion. Verified green 2026-06-15.

Key bug fixed: double-framing in ConnSession::send_envelopeencode_envelope was pre-framing the protobuf, then TcpServerConn::send_frame re-framed it. Fixed by serializing raw protobuf bytes directly and letting send_frame add the single [4-byte len] prefix.



M2 — Voice, single stream ✓ (completed 2026-06-16)

Exit criterion:test_m2_voice — two headless clients authenticate over TLS, bind UDP, announce a MIC stream, send 50 encrypted Opus frames; server SFU relay re-encrypts + forwards to the second client; B receives ≥ 25 frames and all decrypt correctly. Passes in ~4 s.

  • m2-dev preset (inherits vcpkg-base, binaryDir build/m2-dev); m1-dev also builds all M2 code.
  • core/CMakeLists.txtfind_package(Opus), find_path(MINIAUDIO_INCLUDE_DIR).
  • core/src/net/voice_frame.h — 14-byte UDP header (type/flags/codec/ssrc/seq/ts), serialize/parse, make_udp_binding_packet.
  • SodiumMediaCrypto — ChaCha20-Poly1305 AEAD; counter-nonce; 64-bit sliding-window anti-replay; derive_send/recv from TLS RFC 5705 exporter.
  • OpusEncoder / OpusDecoder — libopus 1.6, FEC, DTX, PLC (free; nullptr → decoder extrapolates).
  • UdpMediaChannel — async UDP socket (asio); thread-safe send_to; async recv loop.
  • JitterBuffer — per-ssrc, EWMA jitter estimation, adaptive depth 20200 ms, late-drop at 500 ms.
  • AudioEngine — miniaudio capture+playback; inject_capture() bypass for headless tests; per-ssrc RemoteStream with OpusDecoder + JitterBuffer.
  • ApmProcessorApmPassthrough stub (VAD always open); WebRTC APM deferred until M3.
  • on_tls_ready callback in TcpChannelCallbacks — server derives and stores media AEAD keys immediately after TLS handshake.
  • ConnSession M2 — udp_token generated at construction; included in AuthResult; handle_udp_binding (verifies token, TCP ack); handle_stream_announce (assigns SSRC via registry); udp_media_port in ServerHello.
  • SessionRegistry M2 — register_udp_token, find_by_udp_token, register_udp_endpoint, find_by_udp_endpoint, assign_ssrc, find_channel_sessions, user_channel.
  • MediaRelay — SFU UDP relay; kFrameUdpBinding → endpoint binding; kFrameVoice → decrypt/re-encrypt/forward to channel members.
  • Server::run() — creates and binds MediaRelay; passes media port to ConnSession; wires on_tls_ready to derive per-connection media AEAD keys.
  • test_voice_frame — header round-trip, big-endian layout, binding packet format.
  • test_media_aead — seal/open round-trip, anti-replay, tamper detection, multi-packet sequence.
  • test_opus_codec — encode/decode round-trip energy check (within 3 dB), PLC, frame-samples helper.
  • test_m2_voice — M2 exit criterion (raw-socket harness). Verified green 2026-06-16.

Follow-up (same day): the above made test_m2_voice pass, but vc_client's public voice methods were still stubs — the actual M2 exit criterion ("two vccli/early-GUI clients talk") wasn't met. Closed the gap:

  • core/src/core/client.cpp — real stream_start/stream_stop/set_self_mute/ set_remote_stream; UDP-binding handshake (start_udp_binding/handle_udp_binding_ack/ finish_udp_binding); media key derivation from tls_ (RFC 5705 exporter); run_udp_recv (AEAD-open → JitterBuffer::Frameaudio_engine_.push_recv_frame); on_capture_frame (encode → seal → sendto); sync_remote_streams (diffs a User proto's streams against remote_streams_, wiring up OpusDecoders and emitting STREAM_STARTED/STOPPED). set_input_device/set_input_mode/set_push_to_talk/list_devices remain VC_ERR_NOT_IMPLEMENTED — no device-enumeration backend yet; scoped to M3 (VAD/PTT).
  • core/src/session/session.cpp/hSessionModel::find_user, find_user_by_ssrc, Stream{stream_id, ssrc, kind, label, sample_rate, frame_ms}.
  • server/src/conn_session.cpp/hhandle_stream_announce/handle_stream_stop now broadcast via SessionRegistry::set_user_stream/clear_user_streamUserEvent::UPDATED.
  • server/src/session_registry.cpp/hset_user_stream/clear_user_stream (mutate a user's StreamInfo list, return the updated User proto for broadcast).
  • tests/test_voice_client_abi.cpp — drives two real vc_client instances through vc_connect/vc_authenticate_guest/vc_stream_start/vc_stream_stop; asserts client B observes client A's STREAM_STARTED/STOPPED events. Verified green 2026-06-16.
  • tools/vccli/src/main.cpp — argv parsing (--host/--port/--nick/--channel/--voice/ --mute/--text); --voice starts a MIC stream and blocks on SIGINT, printing on_event callbacks live (unbuffered stdout — MinGW/MSVCRT treat _IOLBF as full buffering for non-console streams). Dropped the originally-planned --voice-loopback and the tx=N rx=M lost=K jitter=J stats line: voicecat.h exposes no PCM-injection hook or jitter/loss stats getter publicly, only on_event + on_level (RMS). Manually verified: two vccli --voice instances see each other's stream start in real time.

M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)

Exit criterion:test_m3_multistream — a real vc_client (A) runs two concurrent local streams (MIC + SCREEN_AUDIO) with distinct stream ids; a second client (B) sees both as separate STREAM_STARTED events and a VC_EVENT_TALK_STATE talking edge for A's MIC stream; B independently sets gain/mute/noise-reduction on each of A's streams without one call affecting the other; A then joins "Music Room" (channel 2: stereo/128kbps/OPUS_AUDIO/no DTX) and announces a fresh MIC stream there, while B stays in "Lobby" (channel 1: mono/24kbps/ OPUS_VOIP/DTX on) — vc_get_stream_audio_config shows their effective Opus config differs exactly as the server enforces per channel. Passes in ~2.4s; verified across 8 consecutive standalone runs + 3 consecutive full-suite runs with no flakes.

Exploration before implementing turned up several bugs/gaps where the wire format already supported this milestone but the client/server logic didn't — these were fixed as part of M3, not treated as pre-existing-and-out-of-scope:

  • Server stream_id bughandle_stream_announce always wrote stream_id=1, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream's replace-by-id logic. Fixed with a per-session counter (ConnSession::next_stream_id_) + announced_stream_ids_ (also now validated in handle_stream_stop, rejecting stops for ids the session never announced).
  • Per-channel AudioConfig was modeled but never populated/enforced. SessionRegistry::init_default_channels() now seeds Lobby (id=1: mono, 24kbps, OPUS_VOIP, FEC+DTX on) and a new "Music Room" (id=2: stereo, 128kbps, OPUS_AUDIO, FEC+DTX off) with real AudioConfigs; new SessionRegistry::channel_audio_config(channel_id) accessor (there was no per-id channel getter before, only channel_snapshot()). handle_stream_announce now treats the channel's config as authoritative (mode/frame_ms/application/fec/dtx/ complexity), clamping (not overriding) bitrate_bps to the channel's ceiling.
  • Client silently dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream — handle_stream_announce_result and sync_remote_streams only copied sample_rate/bitrate_bps/frame_ms/fec into OpusParams. New shared opus_params_from_audio_config() helper (client.cpp) fixes both the send and receive paths.
  • core/src/codec/opus_codec.h/.cpp — new OpusApplication enum + OpusParams::application field; OpusEncoder::init now honors it instead of hardcoding OPUS_APPLICATION_VOIP.
  • core/src/session/session.h/.cppStream struct extended with the full AudioConfig (mode/bitrate_bps/application/fec/expected_packet_loss/dtx/complexity), not just sample_rate/frame_ms; copy_streams() now copies all of it.
  • core/src/core/client.h/.cpp — local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind (one active stream per kind — MIC/SCREEN_AUDIO/ AUX_DEVICE are each singletons for a client), replacing the M2 single-stream fields. StreamAnnounce/StreamAnnounceResult round-trips are now correlated by request_id (already round-tripped on the wire; just wasn't read) via pending_announce_kind_, so multiple concurrent announces from one client resolve to the right LocalStream. on_capture_frame takes a kind parameter and upmixes mono capture to stereo (duplicate L=R) when a stream's channel config calls for it. vc_set_self_mute's mic_muted only gates the MIC kind — a concurrent SCREEN_AUDIO share keeps playing while muted. set_remote_stream now actually wires noise_reduction through (previously parsed and discarded). New run_talk_timer() (a small dedicated thread, started alongside the UDP media path, never the miniaudio callback thread) polls both remote talk-state edges (AudioEngine::poll_talk_transitions()) and local capture-activity edges, emitting VC_EVENT_TALK_STATE.
  • Fixed a thread-join race in teardown_voice() — it's called both from run_io()'s own cleanup and from disconnect(), on different threads; without serialization both could see udp_thread_/talk_timer_thread_ as joinable() simultaneously and race to join() the same std::thread (UB; surfaced as an intermittent std::system_error: No such process under ctest). Added a teardown_mu_ guard around the whole function. This pre-existed for udp_thread_ alone (likely the same root cause as the test_m1_integration/test_m2_voice cleanup-path flake noted in the M2 section above) — adding talk_timer_thread_'s join just made it surface more often, so it was fixed properly here rather than carried forward again.
  • core/src/audio/audio_engine.h/.cppCaptureCallback gained a kind parameter (the real miniaudio capture device is always tagged kind=0/MIC; a second concurrent local stream is fed via its own inject_capture(kind, ...) ring buffer — inject_taps_, keyed by kind — since there is only one real hardware capture device in M3). Fixed a buffer-sizing bug in on_playback's per-stream decode (opus_decode's frame_size parameter is samples-per-channel, not total samples — the old code passed frames * params_.channels, which would have overflowed the decode buffer for any stereo stream). Stereo decoder output is downmixed (avg L/R) into the engine's mono mix accumulator immediately after decode. RemoteStream gained recv_ns/noise_reduction_enabled (lazy ApmProcessor instantiation — freed on disable, so no separate instance cap is needed per the roadmap's guidance) and last_voice_ms/talking (talk-indicator edge state, updated in push_recv_frame); new set_stream_noise_reduction() and poll_talk_transitions(). Note: until VOICECAT_HAS_APM is wired to a real WebRTC APM build, the NS toggle is plumbed end-to-end but behaviorally a passthrough no-op (ApmPassthrough doesn't touch PCM) — same situation send-side APM has been in since M2; M3's job was the plumbing, not the DSP backend.
  • New C ABI surface (core/include/voicecat.h, additive only): vc_audio_config struct + vc_get_stream_audio_config(c, user_id, stream_id, out) — the effective Opus config for a stream you own or a peer's, reading from the (now richer) LocalStream/session::Stream. vc_test_inject_capture(c, stream_id, pcm, samples) — clearly-marked test-only, forwards to AudioEngine::inject_capture, so test_m3_multistream can drive two concurrent synthetic-audio streams through the real ABI without a microphone.
  • tests/test_m3_multistream.cpp — the M3 exit criterion (ABI-level, mirrors test_voice_client_abi.cpp's approach per the M2 lesson). Registered in tests/CMakeLists.txt.

Explicitly out of scope for this pass (confirmed with the user before implementing):

  • vc_set_input_device/vc_set_input_mode/vc_set_push_to_talk/vc_list_devices (device enumeration + VAD/PTT input gate) — still VC_ERR_NOT_IMPLEMENTED. These were mentioned as "scoped to M3" in the M2 follow-up notes above, but docs/roadmap.md's M3 bullets never actually listed them — deferred again, now tracked explicitly rather than implicitly.
  • Real WASAPI desktop-audio loopback capture for SCREEN_AUDIO — the engine now supports feeding a second concurrent local stream via inject_capture, but only synthetic PCM is wired up; a real loopback capture device is a follow-up.
  • True stereo playback outputAudioEngine's mixer/output device stays mono; stereo streams are downmixed after decode (see above). The Opus wire format itself is fully stereo-correct.

Decisions log

All architecture/scope decisions are settled and recorded in docs/roadmap.md §2 "Resolved decisions" and reflected across docs/. If you make a new decision, record it there and link it here.


How to update this file

  1. Check off tasks as you complete them; flip a milestone to [x] only when its exit criterion test passes.
  2. Keep the "Where we left off / next action" block at the top accurate — it's the first thing the next agent reads.
  3. When you start a milestone, copy its task list from docs/roadmap.md into a section here.