Files
voice-cat/PROGRESS.md
Talon b2af1a3001 feat(macos): validate dev + apple-dev presets on macOS, fix 3 cross-platform bugs
macOS port groundwork — core, server, tools, and tests now build and run on
macOS 26.5 / Apple Silicon. ctest --preset dev green 21/21 (2 consecutive runs).
apple-dev produces valid arm64 libvoicecat.a + XCFramework for the Swift Package.

Three real cross-platform bugs found and fixed (all latent on Windows/Linux):

1. test_m2_voice.cpp POSIX branch missing <netdb.h> — Linux glibc transitively
   includes it, macOS doesn't. Would fail on any strict POSIX system.

2. SIGPIPE killing processes on macOS — writing to a closed TCP socket raises
   SIGPIPE by default (doesn't exist on Windows, benign on Linux). Fixed by
   ignoring SIGPIPE in both core client init and server startup (POSIX-only,
   #ifndef _WIN32). Production fix, not just tests.

3. Use-after-free of Asio's kqueue reactor on server shutdown — the
   deterministic test_tofu_flow segfault. TcpServerConn's tls_read_loop runs on
   a blocking-I/O thread; when Server::run() returned, io_context was destroyed
   while those threads were still running. On macOS kqueue the reactor pointer
   is null'd immediately -> segfault in socket.close(). Latent on Windows IOCP
   and Linux epoll. Fix: TcpAcceptor now tracks connections; new shutdown()
   closes all + joins threads before io is destroyed; Server::stop() now closes
   acceptor + media_relay too (was just io.stop()).

Verified: dev + apple-dev presets build green, 21/21 tests pass, server starts
+ two vccli text chat over TLS (M1 on Mac), vccli --voice starts MIC stream via
CoreAudio (M2 protocol-level), vccli --list-devices enumerates CoreAudio
devices, xcodebuild -create-xcframework produces valid VoiceCatCore.xcframework.

No ABI or proto changes. Docs updated: building.md, clients/apple/README.md,
PROGRESS.md, CLAUDE.md status line.
2026-06-18 13:24:42 +02:00

60 KiB
Raw Blame History

PROGRESS — VoiceCat

Living status. Update this file in the same commit as your work so the next agent picks up instantly. Newest status at the top.

  • Date convention: ISO (YYYY-MM-DD).
  • Statuses: [ ] not started · [~] in progress · [x] done.

▶ Where we left off / next action

  • Done: macOS port — dev + apple-dev presets validated (2026-06-18). The core, server, tools, and tests now build and run on macOS (Apple Silicon, macOS 26.5, Apple clang 21). This lays the groundwork for the macOS/iOS Swift client. Three real bugs found and fixed (all were cross-platform issues that manifested on macOS but were latent on Windows/Linux):

    1. Missing <netdb.h> in test_m2_voice.cpp POSIX branch — the raw-socket test's #else branch included <arpa/inet.h>/<netinet/in.h>/<sys/socket.h>/<unistd.h> but not <netdb.h> (needed for addrinfo/getaddrinfo/freeaddrinfo). On Linux glibc these headers transitively include <netdb.h>; on macOS they don't. Fixed by adding # include <netdb.h> to the POSIX branch (mirrors core/src/core/client.cpp:16 which already had it). Real latent bug — would fail on any strict POSIX system.
    2. SIGPIPE killing processes on macOS — on macOS, writing to a closed TCP socket raises SIGPIPE by default (unlike Windows where it doesn't exist, or Linux where it's often benign). This killed test_tofu_flow (intermittent SIGPIPE/SEGFAULT) and would also kill voicecat-server and vccli in production when a peer dropped mid-write. Fixed by ignoring SIGPIPE (std::signal(SIGPIPE, SIG_IGN)) in both the core client init (core/src/core/client.cpp POSIX branch of the #ifdef _WIN32 WSAStartup block) and the server startup (server/src/server.cpp before the asio::signal_set). Both are POSIX-only (#ifndef _WIN32), process-global, and idempotent. The server's asio::signal_set(SIGINT, SIGTERM) is unaffected (independent signals).
    3. Use-after-free of Asio's kqueue reactor on server shutdown — the root cause of the deterministic test_tofu_flow segfault (EXC_BAD_ACCESS in kqueue_reactor::deregister_descriptor(this=0x0000000000000000)). TcpServerConn's tls_read_loop runs on a dedicated blocking-I/O thread (not async on io_context). When Server::stop()io.stop()Server::run() returned, the local asio::io_context was destroyed while tls_read_loop threads were still running. When a thread detected the disconnect and called TcpServerConn::close()socket.close() → Asio tried to deregister from the kqueue reactor — but the reactor (owned by io_context) was already destroyed, and on macOS kqueue the reactor pointer is null'd immediately. Latent on Windows (IOCP) and Linux (epoll) where the timing is more forgiving. Fix: TcpAcceptor now tracks its connections (new conns_ vector + conns_mu_); TcpAcceptor::stop() closes all tracked connections while io_context is still alive; new TcpAcceptor::shutdown() method calls stop() then wait_closed() on each connection (new TcpServerConn::wait_closed() joins tls_thread_); Server::run() calls acceptor.shutdown() after io.run() returns and before io is destroyed; Server::stop()'s stop_fn_ now calls acceptor.stop() + media_relay->stop() + io.stop() (was just io.stop()). After acceptor.shutdown(), when tls_read_loop threads exit and call on_disconnectedConnSession::close()TcpServerConn::close(), the close() is a no-op (closing_.exchange(true) returns true) — no reactor access occurs after io is destroyed.
    • Environment setup: vcpkg cloned to ~/code/vcpkg + bootstrapped. VCPKG_ROOT must be set. Homebrew autoconf-archive is required (vcpkg's libsodium port needs it for autoreconf — brew install autoconf-archive). autoconf/automake/libtool were already installed; glibtoolize (Homebrew's macOS name for libtoolize) is handled by vcpkg automatically.
    • Verified: cmake --preset dev + cmake --build --preset dev green (21 binaries). ctest --preset dev --parallel 121/21 green (2 consecutive runs, 64s each). cmake --preset apple-dev + cmake --build --preset apple-dev green → valid 1.9 MB arm64 libvoicecat.a (167 exported C ABI symbols, correct visibility). xcodebuild -create-xcframework → valid VoiceCatCore.xcframework (macOS-arm64 slice with voicecat.h headers). Server runtime: voicecat-server starts, generates identity, SQLite, Lobby, binds TCP+UDP, clean SIGINT shutdown. Two vccli text chat over TLS (M1 exit criterion on Mac). vccli --voice --input-mode vad starts a MIC stream via CoreAudio (M2 protocol-level on Mac — ear test pending). vccli --list-devices enumerates 4 CoreAudio input + 3 output devices with correct defaults.
    • Apple framework linking: NOT needed — modern macOS ld (Xcode 26.5) auto-discovers CoreAudio/CoreFoundation frameworks in /System/Library/Frameworks without explicit -framework flags. miniaudio's MINIAUDIO_IMPLEMENTATION compiles the CoreAudio calls inline, and the linker resolves them automatically. No if(APPLE) CMake block was added.
    • Docs updated: docs/building.md (apple-dev row + §6 updated from "scaffolding" to "validated"), clients/apple/README.md (macOS slice build confirmed, iOS slices still scaffolding), PROGRESS.md (this entry).
    • Still deferred (per scope): macOS SCREEN_AUDIO loopback via ScreenCaptureKit (stub returns false — per docs/voice.md §9); iOS cross-compile presets (apple-ios/ apple-ios-sim — scaffolding); vc_audio_suspend/vc_audio_resume ABI hooks (defer to iOS client milestone, keep ABI stable).
  • Done: CMake preset cleanup + cross-platform build config (2026-06-18). The preset set was a mess — dev (never used), m1-dev (the one everyone used), m2-dev (cache-identical to m1-dev, never used), no optimized+tests preset, no stripping. Cleaned up to a sensible set + added cross-platform triplet auto-resolution + Apple platform scaffolding:

    • Renames: devskeleton (no-deps stub smoke — accurately named now); m1-devdev (the default development preset — milestone-named presets were misleading since the project is past M5); m2-dev dropped (cache-identical to m1-dev).
    • New presets: release (optimized Release + tests on, symbols kept — run the suite against optimized code or profile); apple-dev / apple-ios / apple-ios-sim (scaffolding — static libvoicecat.a slices for the Swift Package / XCFramework; marked "not yet CI-validated, build on macOS to verify").
    • server-release now strips binaries (CMAKE_EXE_LINKER_FLAGS=-s + CMAKE_SHARED_LINKER_FLAGS=-s) — smaller executables for deployment.
    • Cross-platform triplet auto-resolution: new cmake/voicecat-toolchain.cmake wraps vcpkg's toolchain and resolves VCPKG_TARGET_TRIPLET / VCPKG_HOST_TRIPLET from CMAKE_HOST_SYSTEM_NAME + CMAKE_HOST_SYSTEM_PROCESSORx64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon. The main presets (dev, release, server-release) now work on all three platforms without per-OS variants. Cross-compile presets (apple-ios, apple-ios-sim) override VCPKG_TARGET_TRIPLET explicitly. Hidden base preset renamed vcpkg-basevcpkg-common (now points at the wrapper toolchain instead of vcpkg's toolchain directly).
    • No C++ source changes — the core was already portable (every #ifdef _WIN32 in client.cpp already had a POSIX #else; voicecat.h's export macro already handled GCC visibility; transport.cpp is pure Asio; miniaudio abstracts WASAPI/CoreAudio/ALSA).
    • Docs updated: docs/building.md (full rewrite — 8-preset table, platform matrix, preset history mapping old names to new, Apple scaffolding section), CLAUDE.md (build section reframed around dev as default, status line updated to M5-done), README.md (stale "M0 skeleton" framing replaced with current M5 reality), AGENTS.md (build section updated, stale "placeholder builtin-baseline" sentence deleted), docs/deployment.md (server-release now stripped + cross-platform note), docs/tech-stack.md §4 (triplet auto-resolution + Apple scaffolding note), clients/apple/README.md (new "Building the core for Apple platforms" section with XCFramework workflow), clients/windows/README.md (m1-dev→dev).
    • Code comments updated: core/include/voicecat.h, core/src/net/transport.h, core/src/voicecat.cpp, tests/CMakeLists.txt — preset name references updated.
    • Historical PROGRESS.md entries left intactm1-dev/m2-dev mentions in older entries are a true record of what was run; rewriting them would falsify history. The preset history table in docs/building.md §1 maps old names to new.
    • Verified: cmake --list-presets shows all 8 presets. skeleton + dev configure and build green on Windows. Apple presets are scaffolding (won't build on Windows — expected; they require macOS).
  • Done: Disconnect, timeout & keepalive system (2026-06-18). Three reported bugs traced to one root cause + two missing designed features, all fixed:

    1. Stale users after disconnect + eternal PLC hiss (root cause): ConnSession::close() silently erased dropped users from the registry without broadcasting UserEvent::LEFT. Peers never learned the user left → their user lists stayed stale AND their audio engines never called remove_stream → Opus PLC synthesized comfort noise forever (the "soft hissing that never goes away"). Fix: SessionRegistry::broadcast_left() helper (mirrors kick_user's first half); ConnSession::close() now broadcasts LEFT before erasing. Regression test test_disconnect_left (tests/).
    2. PLC cap (defense-in-depth): AudioEngine::on_playback now caps pure-PLC at ~2 s (kPlcCapSamples); after that it emits digital silence instead of more comfort noise, so a stale stream can never hiss forever even if remove_stream is never called. Resets automatically when fresh packets arrive. Test test_plc_cap (tests/).
    3. No timeout / no ping (missing feature): the client never sent Ping, the server had no last_seen / reaper, and half-open connections (NAT timeout, wifi loss, sleep) left ghost users forever. Fix: client sends Ping every 15 s from the io thread (RTT measured from Pong nonce); ConnSession::last_seen bumped on every inbound TCP/UDP frame; asio::steady_timer reaper sweeps every 15 s and drops sessions older than 45 s (configurable via server::Config::reaper_timeout_ms/reaper_sweep_ms). Each drop broadcasts LEFT via fix #1. Test test_reaper_timeout (tests/, 2 s timeout for fast turnaround).
    4. UDP KEEPALIVE (missing feature): client sends a plaintext KEEPALIVE frame every 5 s (voice_frame.h::kFrameKeepalive); server bumps last_seen + echoes back. Keeps NAT bindings alive and lets media activity defer the reaper independently of TCP.
    5. Graceful client disconnect: vc_disconnect() now sends Disconnect{code=0} when fully authenticated (VC_STATE_CONNECTED): it queues the envelope, sets a graceful_disconnect_pending_ flag, and joins the io thread — the io thread's drain_sends() sends the Disconnect, sees the flag, sets io_stop_, and exits naturally (its cleanup handles teardown_voice() + socket close). No main-thread socket close → no double-close race. Server handles client-sent Disconnect (handle_client_disconnectclose() → immediate LEFT broadcast, no reaper/EOF wait). In non-authenticated states, the original force-close path runs (with TOFU unblock).
    • Docs updated: docs/protocol.md §6/§7 (graceful disconnect, pinned N=3, reaper, last_seen), docs/voice.md §6 (KEEPALIVE plaintext + echo + last_seen), docs/architecture.md §5 (reaper timer), server::Config (reaper fields). ctest --preset m2-dev21/21 green (3 new tests: disconnect_left, plc_cap, reaper_timeout).
  • Done: Stereo screen-audio loopback capture on Windows (2026-06-17). The WASAPI loopback path (start_loopback_capture) used to hardcode cfg.capture.channels = 1, downmixing the system's stereo mix to mono before the encoder ever saw it — so even on a stereo channel, SCREEN_AUDIO was effectively mono (the encoder then upmixed L=R to produce a fake stereo bitstream). Now the loopback device opens in the channel's mode: stereo (interleaved L/R) when the channel is configured stereo, mono when mono. Real stereo flows end-to-end through loopback → Opus encode → decode → stereo playback mixer.

    • audio_engine.hCaptureCallback gained an int channels parameter (the encoder needs to know whether the PCM is real stereo or mono to avoid upmixing real stereo). start_loopback_capture(int kind)start_loopback_capture(int kind, int channels). New loopback_channels_ member; new feed_loopback_for_test test hook (self-contained, works on headless CI where the real WASAPI device can't init).
    • audio_engine.cppon_loopback accumulator is now channel-aware (sized to frame_samples_*loopback_channels_); start_loopback_capture sizes the accumulator off the RT thread before ma_device_start, opens the device with channels, and falls back to mono if the render endpoint rejects stereo (mirrors the playback path's fallback). on_capture/inject_capture/feed_capture_for_test forward channels through the callback (mic path always 1; loopback path 1 or 2).
    • client.cppon_capture_frame takes channels; for channels==2 (real stereo loopback PCM) it encodes directly with no upmix; for channels==1 on a stereo channel it keeps the existing L=R upmix (mic stays mono in v1). handle_stream_announce_result reads effective_params.stereo under the lock and passes 2 or 1 to start_loopback_capture.
    • test_vad_ptt_devices.cpp — new test_loopback_stereo_capture behavior test: feeds a loud-L / silent-R stereo signal through feed_loopback_for_test, encodes (as on_capture_frame now does for channels==2), decodes, mixes, and asserts L≠R across the frame (total_diff ~8.2M, well above the 960k threshold). A mono-downmixed-then- upmixed bitstream would have L==R. Existing test lambda updated for the 4-arg callback.
    • docs/voice.md §8/§9 — diagram + notes updated: mic stays mono; SCREEN_AUDIO loopback captures stereo when the channel is stereo.
    • cmake --build --preset m2-dev + ctest --preset m2-dev --parallel 118/18 green. The vad_ptt_devices VAD-gate sub-test has a pre-existing parallel-run timing flake (passes serially and in isolation); unrelated to this change (VAD gate logic is unchanged for channels==1, the only path the MIC uses).
  • Done: Screen-audio sharing wired into the Windows WinForms client (2026-06-17). The core already fully supported SCREEN_AUDIO capture on Windows (post-M3 WASAPI loopback via VOICECAT_HAS_LOOPBACK, always on for the windows-client preset — core/CMakeLists.txt:85; loopback start/stop at client.cpp:1093/audio_engine.cpp:529; VAD/PTT/self-mute/server-mute correctly bypassed for non-MIC kinds at client.cpp:862-879) and the C# Interop layer was already complete (VcStreamKind.ScreenAudio, StartStream/StopStream/SetRemoteStream/ListUserStreams all generic). The gap was purely UI wiring. No core, proto, or C ABI changes were needed — confirming the "the core should support it already" assessment.

    • MainForm.Designer.cs — new btnScreenShareToggle button in the voice panel top row (flpVoiceTop), right after btnMicToggle, with full AccessibleName/ AccessibleDescription per the existing accessibility convention.
    • MainForm.cs — new _screenStreamId field; BtnScreenShareToggle_Click handler mirroring BtnMicToggle_Click but with no device picker / VAD / PTT / mode / mute (screen audio bypasses all of those in the core). Independent of mic — can share without joining voice and vice versa. HandleDisconnected now resets _screenStreamId and disables the screen toggle. OnFormClosed now explicitly stops both mic and screen streams before Disconnect() (clean StreamStop messages go out before the control channel closes). HandleStreamStarted already labeled ScreenAudio => "screen audio"; PerUserTuningDialog.ApplySettings already iterates all of a peer's streams — both unchanged, peers can independently volume-tune a screen-audio stream vs that user's mic.
    • VoiceCatClientSmokeTests.cs — new ScreenAudioStream_Starts_And_Stops test: connects + TOFU + guest auth, StartStream(ScreenAudio), asserts VC_EVENT_STREAM_STARTED arrives with matching StreamId, StopStream, asserts VC_EVENT_STREAM_STOPPED. Passes headless (the StreamAnnounce succeeds regardless of whether the loopback device initializes on a CI box).
    • dotnet build — 0 warnings/errors. dotnet test4/4 tests green (3 existing + 1 new). Event trace confirms the full StreamStarted → UserUpdated → StreamStopped path through P/Invoke against a live voicecat-server.exe.
    • Not yet confirmed audible by ear — pending manual two-instance live test (one shares screen audio while something plays on the default render endpoint, the other hears it). This is the observable-behavior exit criterion per AGENTS.md.
    • Documented caveat (docs/voice.md §9, unchanged): whole-device WASAPI loopback inherently re-captures this app's own incoming voice mix (self-echo loop) — accepted characteristic, not a bug. Process-specific loopback (Windows 10 2004+ AUDIOCLIENT_ACTIVATION_PARAMS) is a future enhancement; miniaudio doesn't expose it.
  • In progress: M5 — moderation & admin (2026-06-17). Server-side and C ABI are implemented and tested: permissions, kick/ban/move/server-mute, channel CRUD, in-app account management. Four new tests pass: test_m5_permissions, test_m5_kick_ban_move_mute, test_m5_admin_accounts, test_m5_channel_crud. vccli now exposes all M5 operations via CLI flags (--kick, --ban, --move, --server-mute/-unmute/-deafen/-undeafen, --set-permission, --create-channel, --edit-channel, --delete-channel, --create-account, --reset-password, --delete-account, --list-accounts) plus --username/--password for account auth and --self-mute/--self-deafen. Docs updated: docs/protocol.md (envelope tags for ServerMuteRequest/ListAccountsResult, User.server_deafened, GenericResult usage), docs/security.md (BLAKE2b channel passwords, bans schema). ctest --preset m1-dev18/18 green. Windows WinForms UI now exposes all M5 operations: channel CRUD with full per-channel Opus audio config, user moderation (kick/ban/move/server mute/server deafen/set permissions), and server account management. dotnet test of the Windows solution passes. Still to do: DRED/audio-quality polish.

  • Done: Fixed multi-user voice — relayed frames failed AEAD decryption (nonce desync) (2026-06-17, reported live: with 2+ people in a channel, audio was one-directional — "I can hear them but they can't hear me" — and a 3rd joiner heard nobody). Root cause was in the SFU relay (server/src/media_relay.cpp). The media AEAD nonce is an implicit per-direction monotonic counter; open() reconstructs it from the 14-byte header's seq field (the AAD), so the wire contract is header.seq == the counter seal() used (the client honors this at client.cpp:907). The relay decrypted each inbound frame with the sender's key, then re-sealed with the recipient's send_crypto (its own counter) but forwarded the sender's header verbatim — so header.seq carried the sender's counter, not the recipient's. The recipient's open() rebuilt the wrong nonce → every relayed frame failed auth and was silently dropped. It only "worked" while the sender's counter coincidentally equalled the server→recipient counter (a single first-ever sender into a fresh recipient), which is exactly why the first/sole talker was heard but reverse/3rd-party audio was not. Fix: before re-sealing, the relay rewrites the outgoing header's seq (bytes [8..9]) to the recipient's peek_send_counter(), so each server→client direction is one contiguous monotonic counter and the nonce always matches (the anti-replay window also stops seeing false replays from interleaved senders). Safe because the jitter buffer orders by timestamp, not seq (audio_engine.h); seq exists only to carry the AEAD counter. No wire-format/proto/ABI change. Regression test added in tests/test_media_aead.cpp (test_relay_interleaved_reseal): two senders interleaved into one recipient all decrypt with the fix, and the verbatim-seq path is asserted to fail. ctest --preset m1-dev18/18 green (run via PowerShell; Git Bash can't resolve the runtime DLLs. vad_ptt_devices is timing-flaky over loopback — passes on re-run — unrelated to this fix). Latent, separate: the secondary "3rd joiner sometimes can't see other users" report is a control-plane (TCP snapshot/UserEvent) issue, not this AEAD bug — re-verify after live testing before investigating. Also still latent: the 16-bit seq wraps after 65536 frames per direction (faster on a busy relay) with no ROC, so the implicit counter desyncs on long continuous sessions (crypto.cpp open() TODO).

  • Done: Fixed "randomly bumped to Lobby" in the Windows client — actors were excluded from their own state-change broadcasts (2026-06-17, reported live: a connected client would intermittently snap from its joined channel back to Lobby in the UI). Root cause was a design inconsistency, not a disconnect: the server delivered self-initiated state changes (channel join/leave, stream announce/stop) only as a private *Result to the actor and broadcast the authoritative UserEvent::UPDATED to everyone else. The core never applied the join result to its SessionModel, so vc_list_users() kept self in the old channel; the Windows HandleUserUpdated rebuilds _currentChannelId from vc_list_users() on any user's UPDATED event, so the next unrelated event (someone joining, announcing/stopping a stream, being muted) surfaced the stale self-channel → "bumped to Lobby." Flaky because it depended on other users' activity. Fix (broad, per the actor-sees-own-changes principle): the server now broadcasts these UserEvent::UPDATEDs to all clients including the actor (server/src/conn_session.cpp, broadcast(…, /*exclude*/ 0)), and text fan-out now includes the sender (resolve_text_targets), so every client converges via one authoritative path. The *Result is now purely ack/correlation/actor-private payload; response handlers no longer mutate the local model. Windows client drops its optimistic text echo (the relay comes back) and renders the sender's own message via HandleTextMessage. Documented the response-vs-broadcast contract in docs/protocol.md §6. Registry-level admin broadcasts (move/mute/kick/channel CRUD) already used exclude=0 and were correct. ctest --test-dir build/m1-dev18/18 green (PowerShell); Windows VoiceCat.App builds 0 warnings. Latent, not fixed: the client never sends the Ping keepalive that docs/protocol.md §7 describes (only the server answers pings) — unrelated to this bug, noted for later. Resolved (2026-06-18): full keepalive/timeout/disconnect system implemented — see "disconnect, timeout & keepalive" entry below.

  • Done: Fixed a second silent-playback bug — the playout clock free-ran and drifted off the stream (2026-06-17, reported live: both vccli and the Windows client showed talking=1/0 correctly on VAD/PTT, mic + screen-share were recognized by peers, but nothing was audible). Root cause: RemoteStream::playout_ts was only ever seeded to 0 and then advanced one Opus frame per playback callback via the PLC path too (core/src/audio/audio_engine.cpp on_playback), so it free-ran at ~1× wall-clock regardless of whether the sender was transmitting. The sender's frame timestamps only advance while it actually sends (the VAD/PTT gate in core/src/core/client.cpp returns before ls.timestamp += samples). Across a late join or any VAD/PTT silence gap the two clocks diverged without bound; once past the jitter buffer's 500 ms late-drop window, every real frame was dropped-as-late (clock ahead) or never-due (clock behind) → permanent silence, while the talk indicator (driven by push_recv_frame, independent of the jitter buffer) stayed lit. The M3 E2E test missed it because clients there talked continuously right after joining, keeping the clocks aligned. Fix: JitterBuffer gained peek_front_ts() (try-lock, RT-safe); on_playback now seeds/re-syncs playout_ts to the earliest buffered frame on the first frame and whenever it has drifted past ±200/500 ms (kResyncAheadSamples/kResyncBehindSamples), which both seeds startup and recovers after every silence gap. New regression test test_playout_resync (tests/test_vad_ptt_devices.cpp): free-runs the clock ~2 s past the drop window, pushes a ts=0 frame, asserts audible output — verified to fail (energy=0) with the fix disabled, pass (energy≈15M) with it. ctest --test-dir build/m1-dev14/14 green (run via PowerShell; Git Bash exec gotcha for these binaries, see docs/building.md). Not yet confirmed audible by ear — pending the user re-running their live test.

  • Done: Fixed silent-playback bug in AudioEngine::on_playback (2026-06-16, found via live manual test: two vccli --voice clients, control-plane events and VAD all correct, but zero audible output). Root cause: opus_decode()'s max_samples was being passed the hardware playback callback's frame count (miniaudio's own choice, frequently smaller than one Opus frame — e.g. ~480 samples on default low-latency WASAPI periods), instead of the decoder's fixed frame size (960 @ 20ms/48kHz). Since the real packet almost always decodes to more samples than that, opus_decode returned OPUS_BUFFER_TOO_SMALL on nearly every callback — frames were correctly received/decrypted/jitter-buffered, just never decoded into audible PCM. mix_for_test()'s white-box test masked this because it always called on_playback with frames == frame_samples, the one case where the bug is invisible. Fix: RemoteStream (core/src/audio/audio_engine.h) gained a small ring buffer (init_ring/push_ring/pop_ring) that decouples decode cadence from playback-callback cadence — on_playback (core/src/audio/audio_engine.cpp) now tops the ring up by decoding whole Opus frames (decoder.frame_samples(), never the hardware frames) and drains exactly frames samples-per-channel from it each callback, silence-padding (PLC) on underrun. Also fixes a latent playout_ts bug: it now advances by the actual decoded sample count per Opus frame, not by the hardware callback's (unrelated) frame count, which was the wrong unit for jitter-buffer timestamp comparisons. ctest --test-dir build/m1-dev — 12/12 green (run via PowerShell; Git Bash exec gotcha for these binaries, see docs/building.md). Not yet confirmed audible by ear — pending the user re-running their live two-vccli test.

  • Done: Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI loopback ✓ complete (2026-06-16). Closes all three items M3 explicitly carried forward as out of scope (see the dated section below for the full file-by-file change list). ctest --test-dir build/m1-dev12/12 tests green (3 consecutive full-suite runs), including the new test_vad_ptt_devices (real vc_clients against a real server, plus a white-box AudioEngine stereo-mix check — same ABI-level-coverage lesson as M2/M3). Manually verified live: vccli --list-devices against real hardware, and vccli --voice --input-mode vad connecting/streaming without incident. Still explicitly out of scope (carried forward, not silently dropped):

    • Real webrtc-audio-processing/AEC — no working Windows/MSVC build upstream; v1 ships a lightweight energy/RMS VAD instead (see docs/roadmap.md §2, docs/voice.md §8/§11). There is no AEC, NS, or AGC implementation at all, not just a deferred VAD. vc_set_remote_stream(..., noise_reduction)'s per-stream NS toggle is unaffected by this pass and stays exactly as inert as it was after M3 (ApmPassthrough, no PCM modification).
    • macOS/iOS SCREEN_AUDIO capture (ScreenCaptureKit / ReplayKit) — this pass is Windows-only for real loopback capture; other platforms keep vc_test_inject_capture as the only way to feed SCREEN_AUDIO.
    • Process-specific WASAPI loopback — miniaudio's loopback mode captures the whole render endpoint (including this app's own incoming voice mix), not a single process.
    • The pre-existing RT-thread rule violation in on_capture_frame/AudioEngine::on_capture (mutex lock, heap allocation, blocking sendto on the miniaudio real-time callback thread) — predates this work, documented but not fixed; fixing it needs the lock-free ring-buffer hand-off docs/architecture.md §3 specifies, a separate, larger refactor.
  • Done: M4 — Windows WinForms C# client ✓ complete (2026-06-17). Full details in the M4 section below. ctest --preset m1-dev14/14 tests green. dotnet build — 0 warnings/errors across all three C# projects. Manually verified: saved servers, TOFU first-connect dialog, channel tree, join, voice (VAD/PTT/always-on + sensitivity slider + per-user gain/mute/NR), text chat (channel + private), device pickers, level meter.

  • Next: macOS/iOS Swift client (M4 continued) and/or M5 (admin UI, moderation, kick/ ban). The C ABI is complete and stable through M4 — both directions are unblocked. See docs/roadmap.md §M4M5.


Milestones (see docs/roadmap.md for full detail)

  • M0 — Scaffolding ✓ complete
  • M1 — Control plane ✓ complete (2026-06-15)
  • M2 — Voice, single stream ✓ complete (2026-06-16)
  • M3 — Multi-stream & per-channel tuning ✓ complete (2026-06-16)
  • M4 — Native clients — Windows WinForms ✓ (2026-06-17); macOS/iOS Swift pending
  • [~] M5 — Moderation, polish, beyond (perms, bans, DRED; then file transfer, E2EE, …)

M0 — Scaffolding ✓ (completed)

  • Repo layout (core/ server/ tools/ clients/ tests/), CMake + presets, vcpkg manifest.
  • C ABI header core/include/voicecat.h (full surface, stubbed).
  • Protocol source-of-truth core/proto/voicecat.proto (matches docs/protocol.md).
  • Core stubs for all six subsystems (net/crypto/codec/protocol/session/audio) + vc_client.
  • voicecat-server (arg parsing, config, stub run) and vccli (drives the C ABI).
  • CTest smoke test asserting the C ABI contract (not just "it compiles").
  • .gitattributes (LF), .gitignore, .clang-format, onboarding docs.
  • Verified: cmake --preset dev && cmake --build --preset dev && ctest --preset dev → green.

M1 — Control plane ✓ (completed 2026-06-15)

Exit criterion:test_m1_integration — two clients authenticate over TLS 1.3 (guest

  • Argon2id password), exchange channel and private text messages. Passes in ~1 s.
  • vcpkg baseline + m1-dev preset; find_package for protobuf/mbedTLS/libsodium/asio/sqlite3.
  • FrameCodec feed + emit; encode_envelope / decode_envelope.
  • Asio TCP acceptor + TcpServerConn (TLS path: blocking handshake thread + tls_read_loop).
  • TlsContext (mbedTLS 1.3, server cert/identity, ECDSA-P256 self-signed, TOFU on client).
  • WorkerPool (3 threads, used for Argon2id).
  • Database — SQLite, Argon2id via libsodium, create_account / authenticate / bootstrap admin.
  • voicecat-admin — account add/reset/del/list against live DB file.
  • ServerIdentityManager — generate/persist Ed25519 key + cert; fingerprint display.
  • ConnSession — WaitingHello → WaitingAuth → Authenticated state machine; full protocol relay.
  • SessionRegistry — channel tree, user map, broadcast, text routing.
  • vc_client (client.cpp) — full M1 C ABI: connect/TLS/ClientHello/AuthRequest/text/disconnect.
  • Server::run() — io_context, acceptor, worker pool, signal handling, on_ready callback.
  • test_m1_integration — M1 exit criterion. Verified green 2026-06-15.

Key bug fixed: double-framing in ConnSession::send_envelopeencode_envelope was pre-framing the protobuf, then TcpServerConn::send_frame re-framed it. Fixed by serializing raw protobuf bytes directly and letting send_frame add the single [4-byte len] prefix.



M2 — Voice, single stream ✓ (completed 2026-06-16)

Exit criterion:test_m2_voice — two headless clients authenticate over TLS, bind UDP, announce a MIC stream, send 50 encrypted Opus frames; server SFU relay re-encrypts + forwards to the second client; B receives ≥ 25 frames and all decrypt correctly. Passes in ~4 s.

  • m2-dev preset (inherits vcpkg-base, binaryDir build/m2-dev); m1-dev also builds all M2 code.
  • core/CMakeLists.txtfind_package(Opus), find_path(MINIAUDIO_INCLUDE_DIR).
  • core/src/net/voice_frame.h — 14-byte UDP header (type/flags/codec/ssrc/seq/ts), serialize/parse, make_udp_binding_packet.
  • SodiumMediaCrypto — ChaCha20-Poly1305 AEAD; counter-nonce; 64-bit sliding-window anti-replay; derive_send/recv from TLS RFC 5705 exporter.
  • OpusEncoder / OpusDecoder — libopus 1.6, FEC, DTX, PLC (free; nullptr → decoder extrapolates).
  • UdpMediaChannel — async UDP socket (asio); thread-safe send_to; async recv loop.
  • JitterBuffer — per-ssrc, EWMA jitter estimation, adaptive depth 20200 ms, late-drop at 500 ms.
  • AudioEngine — miniaudio capture+playback; inject_capture() bypass for headless tests; per-ssrc RemoteStream with OpusDecoder + JitterBuffer.
  • ApmProcessorApmPassthrough stub (VAD always open); WebRTC APM deferred until M3.
  • on_tls_ready callback in TcpChannelCallbacks — server derives and stores media AEAD keys immediately after TLS handshake.
  • ConnSession M2 — udp_token generated at construction; included in AuthResult; handle_udp_binding (verifies token, TCP ack); handle_stream_announce (assigns SSRC via registry); udp_media_port in ServerHello.
  • SessionRegistry M2 — register_udp_token, find_by_udp_token, register_udp_endpoint, find_by_udp_endpoint, assign_ssrc, find_channel_sessions, user_channel.
  • MediaRelay — SFU UDP relay; kFrameUdpBinding → endpoint binding; kFrameVoice → decrypt/re-encrypt/forward to channel members.
  • Server::run() — creates and binds MediaRelay; passes media port to ConnSession; wires on_tls_ready to derive per-connection media AEAD keys.
  • test_voice_frame — header round-trip, big-endian layout, binding packet format.
  • test_media_aead — seal/open round-trip, anti-replay, tamper detection, multi-packet sequence.
  • test_opus_codec — encode/decode round-trip energy check (within 3 dB), PLC, frame-samples helper.
  • test_m2_voice — M2 exit criterion (raw-socket harness). Verified green 2026-06-16.

Follow-up (same day): the above made test_m2_voice pass, but vc_client's public voice methods were still stubs — the actual M2 exit criterion ("two vccli/early-GUI clients talk") wasn't met. Closed the gap:

  • core/src/core/client.cpp — real stream_start/stream_stop/set_self_mute/ set_remote_stream; UDP-binding handshake (start_udp_binding/handle_udp_binding_ack/ finish_udp_binding); media key derivation from tls_ (RFC 5705 exporter); run_udp_recv (AEAD-open → JitterBuffer::Frameaudio_engine_.push_recv_frame); on_capture_frame (encode → seal → sendto); sync_remote_streams (diffs a User proto's streams against remote_streams_, wiring up OpusDecoders and emitting STREAM_STARTED/STOPPED). set_input_device/set_input_mode/set_push_to_talk/list_devices remain VC_ERR_NOT_IMPLEMENTED — no device-enumeration backend yet; scoped to M3 (VAD/PTT).
  • core/src/session/session.cpp/hSessionModel::find_user, find_user_by_ssrc, Stream{stream_id, ssrc, kind, label, sample_rate, frame_ms}.
  • server/src/conn_session.cpp/hhandle_stream_announce/handle_stream_stop now broadcast via SessionRegistry::set_user_stream/clear_user_streamUserEvent::UPDATED.
  • server/src/session_registry.cpp/hset_user_stream/clear_user_stream (mutate a user's StreamInfo list, return the updated User proto for broadcast).
  • tests/test_voice_client_abi.cpp — drives two real vc_client instances through vc_connect/vc_authenticate_guest/vc_stream_start/vc_stream_stop; asserts client B observes client A's STREAM_STARTED/STOPPED events. Verified green 2026-06-16.
  • tools/vccli/src/main.cpp — argv parsing (--host/--port/--nick/--channel/--voice/ --mute/--text); --voice starts a MIC stream and blocks on SIGINT, printing on_event callbacks live (unbuffered stdout — MinGW/MSVCRT treat _IOLBF as full buffering for non-console streams). Dropped the originally-planned --voice-loopback and the tx=N rx=M lost=K jitter=J stats line: voicecat.h exposes no PCM-injection hook or jitter/loss stats getter publicly, only on_event + on_level (RMS). Manually verified: two vccli --voice instances see each other's stream start in real time.

M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)

Exit criterion:test_m3_multistream — a real vc_client (A) runs two concurrent local streams (MIC + SCREEN_AUDIO) with distinct stream ids; a second client (B) sees both as separate STREAM_STARTED events and a VC_EVENT_TALK_STATE talking edge for A's MIC stream; B independently sets gain/mute/noise-reduction on each of A's streams without one call affecting the other; A then joins "Music Room" (channel 2: stereo/128kbps/OPUS_AUDIO/no DTX) and announces a fresh MIC stream there, while B stays in "Lobby" (channel 1: mono/24kbps/ OPUS_VOIP/DTX on) — vc_get_stream_audio_config shows their effective Opus config differs exactly as the server enforces per channel. Passes in ~2.4s; verified across 8 consecutive standalone runs + 3 consecutive full-suite runs with no flakes.

Exploration before implementing turned up several bugs/gaps where the wire format already supported this milestone but the client/server logic didn't — these were fixed as part of M3, not treated as pre-existing-and-out-of-scope:

  • Server stream_id bughandle_stream_announce always wrote stream_id=1, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream's replace-by-id logic. Fixed with a per-session counter (ConnSession::next_stream_id_) + announced_stream_ids_ (also now validated in handle_stream_stop, rejecting stops for ids the session never announced).
  • Per-channel AudioConfig was modeled but never populated/enforced. SessionRegistry::init_default_channels() now seeds Lobby (id=1: mono, 24kbps, OPUS_VOIP, FEC+DTX on) and a new "Music Room" (id=2: stereo, 128kbps, OPUS_AUDIO, FEC+DTX off) with real AudioConfigs; new SessionRegistry::channel_audio_config(channel_id) accessor (there was no per-id channel getter before, only channel_snapshot()). handle_stream_announce now treats the channel's config as authoritative (mode/frame_ms/application/fec/dtx/ complexity), clamping (not overriding) bitrate_bps to the channel's ceiling.
  • Client silently dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream — handle_stream_announce_result and sync_remote_streams only copied sample_rate/bitrate_bps/frame_ms/fec into OpusParams. New shared opus_params_from_audio_config() helper (client.cpp) fixes both the send and receive paths.
  • core/src/codec/opus_codec.h/.cpp — new OpusApplication enum + OpusParams::application field; OpusEncoder::init now honors it instead of hardcoding OPUS_APPLICATION_VOIP.
  • core/src/session/session.h/.cppStream struct extended with the full AudioConfig (mode/bitrate_bps/application/fec/expected_packet_loss/dtx/complexity), not just sample_rate/frame_ms; copy_streams() now copies all of it.
  • core/src/core/client.h/.cpp — local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind (one active stream per kind — MIC/SCREEN_AUDIO/ AUX_DEVICE are each singletons for a client), replacing the M2 single-stream fields. StreamAnnounce/StreamAnnounceResult round-trips are now correlated by request_id (already round-tripped on the wire; just wasn't read) via pending_announce_kind_, so multiple concurrent announces from one client resolve to the right LocalStream. on_capture_frame takes a kind and channels parameter; for mono capture (channels==1) on a stereo channel it upmixes L=R, and for real stereo capture (channels==2, the SCREEN_AUDIO loopback path on a stereo channel) it encodes directly with no upmix. vc_set_self_mute's mic_muted only gates the MIC kind — a concurrent SCREEN_AUDIO share keeps playing while muted. set_remote_stream now actually wires noise_reduction through (previously parsed and discarded). New run_talk_timer() (a small dedicated thread, started alongside the UDP media path, never the miniaudio callback thread) polls both remote talk-state edges (AudioEngine::poll_talk_transitions()) and local capture-activity edges, emitting VC_EVENT_TALK_STATE.
  • Fixed a thread-join race in teardown_voice() — it's called both from run_io()'s own cleanup and from disconnect(), on different threads; without serialization both could see udp_thread_/talk_timer_thread_ as joinable() simultaneously and race to join() the same std::thread (UB; surfaced as an intermittent std::system_error: No such process under ctest). Added a teardown_mu_ guard around the whole function. This pre-existed for udp_thread_ alone (likely the same root cause as the test_m1_integration/test_m2_voice cleanup-path flake noted in the M2 section above) — adding talk_timer_thread_'s join just made it surface more often, so it was fixed properly here rather than carried forward again.
  • core/src/audio/audio_engine.h/.cppCaptureCallback gained a kind parameter (the real miniaudio capture device is always tagged kind=0/MIC; a second concurrent local stream is fed via its own inject_capture(kind, ...) ring buffer — inject_taps_, keyed by kind — since there is only one real hardware capture device in M3). Fixed a buffer-sizing bug in on_playback's per-stream decode (opus_decode's frame_size parameter is samples-per-channel, not total samples — the old code passed frames * params_.channels, which would have overflowed the decode buffer for any stereo stream). Stereo decoder output is downmixed (avg L/R) into the engine's mono mix accumulator immediately after decode. RemoteStream gained recv_ns/noise_reduction_enabled (lazy ApmProcessor instantiation — freed on disable, so no separate instance cap is needed per the roadmap's guidance) and last_voice_ms/talking (talk-indicator edge state, updated in push_recv_frame); new set_stream_noise_reduction() and poll_talk_transitions(). Note: until VOICECAT_HAS_APM is wired to a real WebRTC APM build, the NS toggle is plumbed end-to-end but behaviorally a passthrough no-op (ApmPassthrough doesn't touch PCM) — same situation send-side APM has been in since M2; M3's job was the plumbing, not the DSP backend.
  • New C ABI surface (core/include/voicecat.h, additive only): vc_audio_config struct + vc_get_stream_audio_config(c, user_id, stream_id, out) — the effective Opus config for a stream you own or a peer's, reading from the (now richer) LocalStream/session::Stream. vc_test_inject_capture(c, stream_id, pcm, samples) — clearly-marked test-only, forwards to AudioEngine::inject_capture, so test_m3_multistream can drive two concurrent synthetic-audio streams through the real ABI without a microphone.
  • tests/test_m3_multistream.cpp — the M3 exit criterion (ABI-level, mirrors test_voice_client_abi.cpp's approach per the M2 lesson). Registered in tests/CMakeLists.txt.

Explicitly out of scope for this pass (confirmed with the user before implementing):

  • vc_set_input_device/vc_set_input_mode/vc_set_push_to_talk/vc_list_devices (device enumeration + VAD/PTT input gate) — still VC_ERR_NOT_IMPLEMENTED. These were mentioned as "scoped to M3" in the M2 follow-up notes above, but docs/roadmap.md's M3 bullets never actually listed them — deferred again, now tracked explicitly rather than implicitly.
  • Real WASAPI desktop-audio loopback capture for SCREEN_AUDIO — the engine now supports feeding a second concurrent local stream via inject_capture, but only synthetic PCM is wired up; a real loopback capture device is a follow-up.
  • True stereo playback outputAudioEngine's mixer/output device stays mono; stereo streams are downmixed after decode (see above). The Opus wire format itself is fully stereo-correct.

Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI loopback ✓ (completed 2026-06-16)

Closes all three items the M3 section above explicitly carried forward as out of scope.

Exit verification: ctest --test-dir build/m1-dev12/12 tests green (3 consecutive full-suite runs), including the new test_vad_ptt_devices (device enumeration + VAD/PTT gate through real vc_clients against a real server, plus a white-box AudioEngine stereo-mix check — no audio hardware needed for that last part). Also verified 5 consecutive standalone runs of the new test alone, no flakes. Manually verified live on Windows: vccli --list-devices against real hardware (3 input / 4 output devices, correct is_default flags), and vccli --voice --input-mode vad connecting + streaming without incident.

  • Device enumeration (vc_list_devices) — AudioEngine::enumerate_devices(bool capture) (static, works without a running engine — inits a throwaway ma_context via ma_context_get_devices). device_id/vc_device.id is an opaque hex-encoded raw ma_device_id (not the device name — names aren't guaranteed unique); documented as an internal contract callers must round-trip, never construct by hand. vc_client::list_devices works in any connection state (no VC_STATE_CONNECTED gate) since device pickers need to populate pre-connect. vc_free_device_list is now a real free (was a no-op stub).
  • Input device selection (vc_set_input_device) — stores the device id on the targeted LocalStream (new field); for the MIC stream, if the engine is already running, restarts it (stop() + ensure_audio_running()) to pick up the new device. Simplified: it restarts unconditionally rather than trying to detect whether the device id actually changed (AudioEngine has no getter for "current device").
  • VAD/PTT input gate (vc_set_input_mode, vc_set_push_to_talk) — new EnergyVadProcessor (core/src/audio/apm_processor.cpp) implementing the existing ApmProcessor interface: energy/RMS threshold (default ~0.025 normalized) + hang-time (default 300 ms, matching kTalkHangoverMs). New factory ApmProcessor::create_vad(), kept separate from create() (which recv-side per-stream NS still uses, unaffected by this pass). vc_client gained current_input_mode_/ptt_active_/mic_vad_; the gate is inserted in on_capture_frame, MIC-onlySCREEN_AUDIO/AUX_DEVICE always bypass it (gating a desktop-audio share on the user's own voice activity would silently drop shared music/video audio). last_capture_ms (drives the talk indicator) is now updated after the gate check, not before, so a VAD/PTT-closed frame never shows as "talking". mic_vad_ is constructed once the MIC stream's StreamAnnounceResult lands (on io_thread_), not lazily inside the capture path.
  • True stereo playbackAudioParams::channels split into capture_channels (stays
    1. and playback_channels (now 2, unconditionally). AudioEngine::on_playback no longer downmixes decoded stereo streams to mono before mixing — stereo decode output is mixed directly (L→L, R→R); mono decode output is upmixed (duplicated into both channels). Falls back to a 1-channel playback device once if the 2-channel ma_device_init fails (unusual hardware). New test-only AudioEngine::mix_for_test() exposes the mixer for white-box testing without a real ma_device.
  • WASAPI loopback capture for SCREEN_AUDIO — new VOICECAT_HAS_LOOPBACK macro (core/CMakeLists.txt, Windows-only). AudioEngine gained a separate loopback_device_ (own lifecycle, decoupled from the mic capture/playback devices) with start_loopback_capture()/stop_loopback_capture(), using miniaudio's ma_device_type_loopback against the default render endpoint. Its callback feeds capture_cb_ directly (same pattern as the real mic capture device), not through inject_capture()'s test-only ring. Wired into vc_client::handle_stream_announce_result (start, alongside ensure_audio_running()) and stream_stop (stop) for VC_STREAM_SCREEN_AUDIO. Non-Windows builds keep vc_test_inject_capture as the only way to feed SCREEN_AUDIO.
  • tools/vccli/src/main.cpp — new flags --list-devices, --input-device, --input-mode vad|ptt, --share-screen-audio; while --voice is running, a background stdin-reader thread accepts ptt on/ptt off/mode vad/mode ptt (the most portable way to drive PTT interactively from a headless CLI — no SIGUSR1 equivalent on Windows). Also prints VC_EVENT_TALK_STATE. Known minor caveat: on Windows the stdin-reader thread is detached (not joined) on exit, since std::getline can't be interrupted from another thread — a vc_client* use-after-free is theoretically possible if a command line arrives in the brief window between teardown and process exit; acceptable for a headless test/dev tool.
  • tests/test_vad_ptt_devices.cpp — new test covering all four items above; registered in tests/CMakeLists.txt. tests/test_smoke.cpp's device-list assertion is now conditional on VOICECAT_HAS_AUDIO (was a hard VC_ERR_NOT_IMPLEMENTED assertion) — VC_OK only, never count > 0 (a headless CI build agent may legitimately report zero audio devices).

Still explicitly out of scope (carried forward, not silently dropped):

  • Real webrtc-audio-processing/AEC — no working Windows/MSVC build upstream (see docs/roadmap.md §2's superseding decision-log entry). There is no AEC, NS, or AGC implementation at all, not just a deferred VAD. The per-stream NS toggle (vc_set_remote_stream(..., noise_reduction)) is unaffected by this pass and stays exactly as inert as it was after M3 (ApmPassthrough, no PCM modification) — don't mistake this pass for having fixed it.
  • macOS/iOS SCREEN_AUDIO capture (ScreenCaptureKit / ReplayKit) — Windows-only loopback in this pass.
  • Process-specific WASAPI loopback — whole-device capture only; inherently captures this app's own incoming voice mix along with everything else playing.
  • The pre-existing RT-thread rule violation in on_capture_frame/AudioEngine::on_capture (mutex lock, heap allocation for the stereo-upmix path, blocking sendto, all on the miniaudio real-time callback thread) — predates this work (was already present in M2/M3); documented here explicitly rather than silently carried forward again. Fixing it properly needs the lock-free ring-buffer hand-off docs/architecture.md §3 specifies — a separate, larger refactor, out of scope for this pass.

M4 — Windows WinForms C# client ✓ (completed 2026-06-17)

Exit criterion:ctest --preset m1-dev14/14 tests green (existing 12 + 2 new C++ tests: test_channel_user_list_abi, test_tofu_flow). dotnet build — 0 warnings/errors. Manually verified: connect, TOFU first-connect dialog, channel tree, join, voice, text, device pickers, level meter. Accessibility: explicit AccessibleName/AccessibleDescription on every control, & mnemonics on every button, activity-log ListBox as screen-reader record.

New C++ ABI surface (all additive, backward-compatible):

  • vc_list_channels / vc_list_users / vc_list_user_streams — pull-based snapshot getters for the channel tree + user list; session_model_mu_ added for cross-thread safety. SessionModel::apply_snapshot/apply_channel_event fixed to populate parent_id, password_protected, max_users (were permanently zeroed despite the struct declaring them).
  • vc_join_channel(channel_id, password) — join with optional password; server replies via new VC_EVENT_JOIN_RESULT.
  • VC_EVENT_SERVER_IDENTITY + vc_confirm_server_identity(accept) — TOFU gate that blocks io_thread_ until the UI approves or rejects. Pins the TLS leaf-cert SHA-256 fingerprint (verifiable directly at handshake), not the declared Ed25519 value (see docs/security.md §1.1 for why — the TLS cert and Ed25519 key are generated independently, no binding). vc_get_server_identity_display exposes the Ed25519 fingerprint for human-readable display.
  • vc_config::tofu_store_path — optional per-user pin file path; defaults to a relative "./voicecat_tofu_pins.txt" so existing tests need no change.
  • TcpAcceptor now dual-stacks (IPv6 + IPv4 fallback) so localhost::1 on Windows connects correctly without forcing users to type 127.0.0.1.
  • VC_INPUT_ALWAYS_ON = 2 in vc_input_mode — transmit unconditionally (no VAD gate).
  • vc_set_vad_threshold(float) — live VAD threshold update; EnergyVadProcessor stores it atomically so the RT capture path reads without a lock or allocation.

New C++ tests:

  • test_channel_user_list_abi — snapshot getters, parent_id/password_protected/max_users regression, per-user stream list, invalid-user-id error, double-free idempotency.
  • test_tofu_flow — first-connect blocks until confirmed; reject doesn't persist; reconnect to same identity reports MATCHED; rotated identity reports MISMATCH; vc_confirm_* with nothing pending returns an error.

windows-client CMake preset — Release, VOICECAT_BUILD_SHARED=ON, static MinGW runtime (-static-libgcc -static-libstdc++ -static -lwinpthread), no tools/tests. Outputs build/windows-client/bin/voicecat.dll with zero MinGW DLL dependencies (only Windows system DLLs remain — verified via objdump -p).

C# solution (clients/windows/, .NET 10 LTS net10.0-windows):

  • VoiceCat.Interop[LibraryImport] P/Invoke surface, [UnmanagedCallersOnly] callbacks, System.Threading.Channels.Channel<VoiceCatEvent> event delivery drained by 30ms WinForms Timer; VoiceCatClientHandle : SafeHandle guarantees vc_client_destroy.
  • VoiceCat.App — WinForms UI:
    • ConnectDialog — saved-server ListBox, Add/Remove/Edit; servers persisted to %AppData%\VoiceCat\servers.json; passwords DPAPI-encrypted (ProtectedData, opt-in).
    • ServerIdentityDialog — shown only on FIRST_CONNECT/MISMATCH (never MATCHED); mismatch text and button ordering are starkly different ("WARNING" framing, Cancel default).
    • MainFormTreeView channel tree, ListBox user list, RichTextBox chat, scope ComboBox (Channel/Private), activity-log ListBox, voice panel with mic toggle, mute/deafen checkboxes, VAD/PTT/Always-On radio group, VAD sensitivity TrackBar (1100, hidden for non-VAD modes), device ComboBox + refresh, level ProgressBar.
    • PerUserTuningDialog — real-time gain TrackBar + mute/NR checkboxes; applied to all of a user's streams immediately (no OK/Cancel round-trip).
    • PttKeyCaptureDialog — focus-scoped PTT key capture.
  • VoiceCat.Interop.Tests — xunit smoke test: connect → TOFU → guest auth → list channels purely via P/Invoke against a live voicecat-server.exe.

Explicitly out of scope for this pass:

  • macOS/iOS Swift client — pending.
  • Admin/moderation UI (kick/ban/permissions/account provisioning) — server-side dispatch for these messages is M5's job; building the UI now would require building the server side too.
  • PTT hotkey is focus-scoped only (works while VoiceCat window has focus). A system-wide WH_KEYBOARD_LL hook would require escalated permissions and risk AV flagging — documented limitation, not silently omitted.
  • Receive-side noise reduction (vc_set_remote_stream(..., noise_reduction)) is end-to-end plumbed but behaviorally a passthrough no-op (ApmPassthrough, no PCM modification) — same as before M4. The per-user NR checkbox in PerUserTuningDialog is labeled accordingly.

M5 — Moderation, polish, and beyond [~] (in progress 2026-06-17)

Exit criterion: four ABI-level tests green (test_m5_permissions, test_m5_kick_ban_move_mute, test_m5_admin_accounts, test_m5_channel_crud); vccli can drive all moderation/admin/channel operations against a live server.

  • Server-side moderation & permissions:
    • server/src/session_registry.h/.cpp — per-session Permissions, permission helpers (can_kick, can_ban, etc.), kick/ban/move/server-mute, channel CRUD, DB-backed channel tree load/save, in-memory channel state.
    • server/src/conn_session.cpp — M5 dispatch handlers, permission checks, channel-password
      • max_users enforcement, UserEvent::UPDATED broadcast on join/leave.
    • core/proto/voicecat.protoServerMuteRequest, UserEvent.reason, User.server_deafened, ListAccountsResult, AccountEntry.
  • C ABI / client-side:
    • core/include/voicecat.hvc_permissions, vc_channel_info, vc_kick_user, vc_ban_user, vc_set_permission, vc_set_server_mute, vc_move_user, vc_create_channel, vc_edit_channel, vc_delete_channel, vc_create_account, vc_reset_password, vc_delete_account, vc_list_accounts, vc_get_permissions; new events VC_EVENT_GENERIC_RESULT and VC_EVENT_ACCOUNT_LIST.
    • core/src/voicecat.cpp, core/src/core/client.h/.cpp — implementations + server-mute/deafen gating on the client.
  • Database: server/src/db.h/.cpp schema v2 (channels, bans), Argon2id accounts, BLAKE2b channel passwords, migrations.
  • Tests: four new M5 tests registered in tests/CMakeLists.txt:
    • test_m5_permissions — grant/revoke permissions, verify enforcement.
    • test_m5_kick_ban_move_mute — kick, ban, move, server-mute/deafen.
    • test_m5_admin_accounts — create/reset/delete/list accounts.
    • test_m5_channel_crud — create/edit/delete channels, password + max_users enforcement.
  • vccli (tools/vccli/src/main.cpp) — all M5 operations exposed via flags; account auth via --username/--password; async VC_EVENT_GENERIC_RESULT/VC_EVENT_ACCOUNT_LIST handling.
  • Docs kept in sync: docs/protocol.md, docs/security.md, PROGRESS.md.

Key bug fixed: test_m5_channel_crud failed because SessionRegistry::create_channel broadcast ChannelEvent::CREATED from a moved-from entry.proto after channels_[id] = std::move(entry). Fixed by building the event before moving into the map.

Still to do:

  • DRED/audio-quality polish.
  • macOS/iOS Swift client (carried from M4).

Decisions log

All architecture/scope decisions are settled and recorded in docs/roadmap.md §2 "Resolved decisions" and reflected across docs/. If you make a new decision, record it there and link it here.


How to update this file

  1. Check off tasks as you complete them; flip a milestone to [x] only when its exit criterion test passes.
  2. Keep the "Where we left off / next action" block at the top accurate — it's the first thing the next agent reads.
  3. When you start a milestone, copy its task list from docs/roadmap.md into a section here.