Files
voice-cat/PROGRESS.md
Talon 4f89d2d32d
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
feat(deploy): Docker + Linux server deployment
Multi-stage Dockerfile (builder → export → runtime) producing a 149 MB
Ubuntu 24.04 image, verified booting end-to-end on Docker Desktop.  vcpkg
fetched via shallow git fetch at the pinned baseline, release-only overlay
triplets (x64-linux, arm64-linux) to halve intermediate disk usage, and
buildtrees deleted within the RUN layer so they never land in the image or
the BuildKit cache.  Binary cache mount (VCPKG_BINARY_SOURCES) makes
subsequent rebuilds restore pre-built packages instead of recompiling.

Also adds:
- docker-compose.yml for one-command local deploy
- .dockerignore (excludes clients/, build/, .git/)
- .github/workflows/build-linux.yml — CI cross-build for amd64 + arm64
  with downloadable artifacts (primary path for building from Windows)
- scripts/build-linux-binaries.sh — local Docker binary extraction fallback
- deploy/linux/voicecat.service — hardened systemd unit for bare-metal
- cmake/voicecat-toolchain.cmake now auto-wires VCPKG_OVERLAY_TRIPLETS

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-21 19:51:43 +02:00

45 KiB
Raw Blame History

PROGRESS — VoiceCat

Living status. Update this file in the same commit as your work so the next agent picks up instantly. Newest status at the top.

  • Date convention: ISO (YYYY-MM-DD).
  • Statuses: [ ] not started · [~] in progress · [x] done.

▶ Where we left off / next action

  • Done (2026-06-21): Docker + Linux deployment + GitHub Actions cross-build. Added the complete Linux server deployment story (the only missing platform — Windows and macOS already have native binaries):

    • Dockerfile — multi-stage (builder: ubuntu:24.04 + vcpkg + cmake --preset server-release; runtime: ubuntu:24.04, non-root voicecat user, /data volume, TCP+UDP 8384). vcpkg is fetched via the GitHub archive tarball at the exact builtin-baseline commit (d46283cf…), avoiding a full git-history clone. BuildKit cache mounts on /vcpkg/downloads, /vcpkg/buildtrees, /vcpkg/packages (scoped by TARGETARCH) keep rebuilds fast. Both voicecat-server and voicecat-admin are copied into the runtime image.
    • docker-compose.yml — single-service compose file with restart: unless-stopped, named volume voicecat-data, and port mappings for TCP+UDP 8384. command: shows how to set --name.
    • .dockerignore — excludes .git/, build/, clients/ (Swift/C# code), docs/, markdown, editor config; build context is just core/, server/, tools/, cmake/, and the three root CMake/vcpkg files.
    • deploy/linux/voicecat.service — hardened systemd unit (non-root, ProtectSystem, NoNewPrivileges, AmbientCapabilities=CAP_NET_BIND_SERVICE) for bare-metal deploys.
    • Multi-arch: docker buildx build --platform linux/amd64,linux/arm64 . works without any triplet override — cmake/voicecat-toolchain.cmake auto-detects from the host arch cmake sees inside the buildx container.
    • Quick start: docker compose up -d (or docker run -d -p 8384:8384/tcp -p 8384:8384/udp -v voicecat-data:/data voicecat). First run auto-generates identity
      • cert + DB; check logs for fingerprint + admin password.
    • GitHub Actions (.github/workflows/build-linux.yml): primary cross-platform binary build path — amd64 uses ubuntu-24.04, arm64 uses ubuntu-24.04-arm (native, not QEMU). Triggers on push to main (when C++/cmake files change) and manually via workflow_dispatch. Downloads land as 90-day artifacts. scripts/build-linux-binaries.sh is the local Docker fallback (needs ~1015 GB free disk; suits Linux dev machines, not Windows Docker Desktop).
  • Done (2026-06-21): Fix permanent voice-loss bug + harden the UDP media path (protocol v2). Field report: two iOS users lost all audio mid-call after a bad-network blip and could not recover even by restarting the apps. Root causes found in the UDP media path:

    1. Anti-replay window poisoned by unauthenticated packets (the trigger). SodiumMediaCrypto::open() advanced recv_highest_ from the plaintext header seq before verifying the AEAD tag and never rolled it back on failure. One corrupted/forged frame (a bit-flip on flaky wifi) shoved the high-water mark far ahead, after which every legitimate frame was rejected as "too old" — permanently. Fixed by reordering to replay-check → authenticate → update (RFC 3711 §3.3): the window is now touched only after a successful tag check. Regression test in test_media_aead.cpp (test_corrupted_seq_does_not_poison_window) — fails on the old code, passes now.
    2. 16-bit seq wrap with no rollover counter. The wire header carried only the low 16 bits of the nonce counter (zero-extended on receive); after 65,536 frames the reconstructed nonce diverged and all frames failed auth. Wire format widened to a full u64 seq (voice_frame.h: header 14 → 20 bytes, seq u16 → u64; crypto.cpp, client.cpp, media_relay.cpp updated; JitterBuffer::Frame::seq widened). This is a versioned wire change → VOICECAT_PROTOCOL_VERSION 1 → 2; the Hello handshake rejects on mismatch (conn_session.cpp). The voice frame is parsed only in core/+server/+tests/, so the Swift/C# clients need only a rebuild — no parser changes.
    3. Server leaked UDP state on disconnect. SessionRegistry::unregister_session() now also frees udp_endpoints_/udp_tokens_/ssrc_to_session_ (scan-and-erase by session id).
    4. Diagnostics. MediaRelay now emits rate-limited dropped-frame counters (unmapped-endpoint / no-recv-crypto / open-failed) so a wedged media path is observable.
    • Verified: cmake --build --preset dev clean; ctest --preset dev -E external_pcm 22/22 pass (incl. m2_voice e2e relay + the two new AEAD regressions). external_pcm still aborts on the pre-existing CoreAudio shutdown mutex race (confirmed identical on a clean baseline checkout under the same harness — unrelated to these changes). Docs updated: voice.md §2 (header), protocol.md (v2 + negotiation), security.md (authenticate-then-advance).
  • Done (2026-06-21): Expose all channel codec params + guest nickname in every client.

    • DRED everywhere + ABI fix. dred (Opus 1.6 Deep REDundancy) existed in the C ABI (vc_audio_config.dred) and proto but was absent from both client marshaling layers — a latent ABI mismatch: Swift AudioConfig and the C# VcAudioConfigNative blittable struct were each one int short of the native struct passed to vc_create_channel/vc_edit_channel. Added dred through Swift (Models.swift, Marshaling.swift, VoiceCatClient.toNative) and C# (Structs.cs, Models.cs, Marshaling.cs, VoiceCatClient.cs).
    • Windows: added the one missing DRED checkbox to ChannelEditDialog (all other params were already present).
    • macOS: ChannelEditSheet now exposes the previously-hidden params — application profile, sample rate, expected packet loss, complexity, and DRED (was only stereo/bitrate/frame/FEC/DTX).
    • iOS: ChannelEditView was name+topic only; rebuilt into a full create and edit form (General: name/topic/parent/password/max-users/sort-order; Audio: stereo/bitrate/sample-rate/ frame/application/packet-loss/complexity/FEC/DTX/DRED). Added SessionState.editChannel and an "Edit" swipe action (admins) in ChannelTreeView + ChannelBrowserView (iOS previously had no edit-channel UI at all). Note: the channel list doesn't carry the current audio config, so on edit the audio fields start from codec defaults — same limitation as macOS/Windows.
    • Guest nickname. Guests could not set a display name on iOS or macOS (the field was absent/disabled; only Windows had it). Added a dedicated nickname to SavedServer on both (backward-compatible Codable), a Nickname field shown in Guest mode (AddServerView / AddServerSheet), and wired the guest auth path to use it (AppState, ConnectWindowController).
    • Verified: xcodebuild Debug — macOS BUILD SUCCEEDED; iOS (sim, ARCHS=arm64) BUILD SUCCEEDED. Core ctest --preset dev 22/23 (only external_pcm aborts on a pre-existing shutdown mutex race; no C++ was changed). Windows C# not buildable on macOS — changes reviewed.
  • Done (2026-06-21): iOS iPhone-layout UX fixes. (1) Channels are now a drill-down on iPhone — new ChannelBrowserView (root list of top-level channels) → ChannelDetailView (people in the channel + sub-channels + an explicit "Join Channel" button with password prompt). The iPad 3-column NavigationSplitView is unchanged. (2) Extracted a self-contained UserRow (context menu + sheets) from UserListView so admin actions are reused in the drill-down. (3) Fixed the off-screen chat compose box: MainView now places VoiceControlsView via .safeAreaInset(edge: .bottom) instead of a floating .overlay, so it reserves layout space above the tab bar and cooperates with keyboard avoidance. (4) Collapsed Activity into Chat like macOS/Windows: ChatView renders a merged, time-sorted timeline of messages + activityLog (activity rows in gray); the separate Activity tab and ActivityLogView.swift are removed. xcodebuild Debug for generic/platform=iOS BUILD SUCCEEDED (sim slice still arm64-only → simulator run N/A). Next: on-device check of the drill-down + compose box + unified timeline.

  • Done (2026-06-21): Screen-audio sharing on macOS + iOS. macOS uses ScreenCaptureKit (ScreenAudioCapture.swift) → vc_stream_feed_pcm; iOS uses a ReplayKit Broadcast Upload Extension (VoiceCatBroadcast) that forwards captured .audioApp PCM through a shared App Group SPSC ring (BroadcastAudioRing.swift) to the host's BroadcastAudioPump, which owns the SCREEN_AUDIO stream and feeds it — single session, no creds on disk. No C++ changes (the core was already ready). macOS xcodebuild Debug BUILD SUCCEEDED; iOS app + extension build for device (the xcframework sim slice is arm64-only, so x86_64-simulator link is N/A). Next: on-device end-to-end verification (two clients hear the shared audio; iOS broadcast start/stop). NOTE: ctest --preset dev is 22/23 — external_pcm passes its assertions but aborts at shutdown (mutex lock failed), a pre-existing teardown crash unrelated to this change (no C++ was modified).

  • Done (2026-06-21): macOS per-app screen-audio selection. Before sharing, a new ScreenSharePickerSheet lets the user choose scope — share Everything / Only selected apps / All except selected apps — plus a first-class "Exclude screen reader (VoiceOver) audio" toggle. ScreenAudioCapture now takes a ScreenAudioSelection and builds the matching SCContentFilter (including: / excludingApplications:); app list comes from SCShareableContent. macOS xcodebuild Debug BUILD SUCCEEDED. iOS deliberately untouched — ReplayKit only delivers the mixed system stream, so per-app/VoiceOver filtering is impossible there (documented in voice.md §9). Still to verify on-device: which process actually carries VoiceOver speech (VoiceOver app vs. com.apple.speech.speechsynthesisd) — the exclude set covers both candidates in ScreenAudioCapture.screenReaderBundleIDs; confirm exclusion actually silences it in a real share.

  • Done (2026-06-20): macOS client UI overhaul — mirrors the Windows client's UI overhaul (commit 97fa659 + 540ec13), adapted to Mac-native conventions. Also fixed and verified the previously-uncompiled Swift changes from the external PCM feed/tap commit (615d2a8). The main window is now just toolbar + channels + users + chat; audio device settings (input mode, VAD, PTT key, device picker, level meter) moved to a modeless Settings window (⌘,). Details in M5 section below. swift test 10/10; xcodebuild Debug

    • Release BUILD SUCCEEDED with 0 Swift warnings. Next: live manual verification (toolbar toggles, unified log colors, PM windows, channel counts, volume slider, settings window); then iOS ReplayKit and macOS ScreenCaptureKit consumers of vc_stream_feed_pcm.
  • Awaiting on-device verification: iOS stereo mic kills headphone/A2DP output — REAL root cause found & fixed (2026-06-20, on Windows; verify on Mac). All prior "fixes" (the 2026-06-19 entries below) targeted the Swift IOSAudioRouter on the false premise that "miniaudio does NOT touch AVAudioSession on iOS." It does. The core opened its miniaudio devices with ma_device_init(nullptr, ...); with a NULL context, miniaudio (0.11.25) runs an iOS "hack" (miniaudio.h ~44057) that picks a session category by device type, then ma_context_init__coreaudio (~36552) calls setCategory() + setActive() on every device open — capture → AVAudioSessionCategoryRecord with zero options. That wiped the .playAndRecord category, the mode, and .allowBluetoothA2DP/.mixWithOthers/.allowAirPlay that IOSAudioRouter had just configured → headphone/A2DP (and even wired) output died. The stereo presets broke worst because they depend on the A2DP output route the wipe removed. TeamTalk never hits this: its SDK opens RemoteIO/VPIO AudioUnits directly and leaves the session entirely to the app (UtilSound.swift); miniaudio insists on managing it.

    • Fix (core, cross-platform safe): AudioEngine now owns a ma_context built by make_context_config() with coreaudio.sessionCategory = ma_ios_session_category_none + noAudioSessionActivate/noAudioSessionDeactivate = MA_TRUE, and passes it to all ma_device_init calls (playback, capture, loopback) and to enumerate_devices's context. miniaudio now never touches AVAudioSession; the Swift IOSAudioRouter is the sole owner (session is already activated on connect in AppState.swift:authResult, before any device opens, so removing miniaudio's self-activation is safe). Context is lazily inited in start(), reused across restarts, uninited in ~AudioEngine. Files: core/src/audio/audio_engine.{h,cpp}.
    • TEMP diagnostics (remove after verification): AudioSessionManager.logSessionState(_:) logs category/mode/options/route; called after ensureSessionActive, on every route change, and on .streamStarted (right after the core opens its devices). On Mac, watch the log when joining voice with the Stereo Mic preset: category must stay …PlayAndRecord with allowBluetoothA2DP and the output route must remain the headphones/A2DP device — NOT flip to …Record. If confirmed, delete the logSessionState calls + method and the prior band-aid comments in IOSAudioRouter/audio_engine.cpp can be trimmed.
    • Verified on Windows: cmake --build --preset dev clean, ctest --preset dev 23/23 (22/22 prior + test_external_pcm new binary). iOS build & on-device run still to be done by the user on the Mac.
  • Done (2026-06-20): External PCM feed/tap API (vc_stream_feed_pcm + vc_set_pcm_sink) — see detail in M5 section below. ctest --preset dev 23/23 (was 22/22 + 1 new test binary with 3 sub-tests). Next: iOS ReplayKit and macOS ScreenCaptureKit consumers of this API. A public, documented API for driving audio streams with externally-provided PCM instead of (or in addition to) miniaudio's hardware device. Motivated by four concrete use cases — all in our roadmap — that the current "miniaudio owns the device" model can't serve:

    1. ReplayKit Broadcast Upload Extension (iOS SCREEN_AUDIO) — the extension is a separate process with a ~50 MB memory cap and can't link the full AudioEngine (ma_device, capture/playback threads). It needs to feed CMSampleBuffer audio (system app audio) into the encode path without any audio hardware. The current plan in docs/voice.md §9 says the extension links "a minimal slice of the core (Opus encode + media send only)" — a public feed-PCM API is that minimal slice. The extension links Opus + the feed entry point, no ma_device needed.
    2. ScreenCaptureKit (macOS SCREEN_AUDIO)SCStream delivers CMSampleBuffer in a callback; convert to int16 and feed. No need to route through miniaudio's device layer. This is how macOS screen-audio actually gets implemented — today it does NOT work: VOICECAT_HAS_LOOPBACK is Windows-only (core/CMakeLists.txt:88-95), so on macOS AudioEngine::start_loopback_capture() hits the #else stub (audio_engine.cpp:647-649) and returns false. The macOS client's "Share Screen Audio" button (MainWindowController.swift:800-816) calls startStream(.screenAudio) which announces the stream to peers but captures zero audio — peers hear silence. The button is left in place (not touched per user request); it'll work once this API + a ScreenCaptureKit tap ship on Mac.
    3. Bots — music bot, TTS bot, radio relay, transcription bot. They create a SCREEN_AUDIO/AUX_DEVICE stream and feed synthesized or decoded PCM via the feed API. No audio hardware required — runs headless on a server. Today the only way to feed external PCM is vc_test_inject_capture (TEST-ONLY, name signals "don't ship this") or re-implementing Opus encode + AEAD + UDP framing yourself (~500 lines of duplicated crypto/codec code per consumer).
    4. Custom clients / accessibility — soundboard, DAW integration, TTS of incoming chat, recording/transcription of remote audio. Need either feed (send) or tap (receive) or both.

    What we already have (input half, gated as test-only): vc_test_inject_capture (stream_id, pcm, samples) (voicecat.h, client.cpp:1452) feeds raw int16 PCM into the encode pipeline via AudioEngine::inject_capture(kind, pcm, n). It works for any stream kind, supports multiple concurrent injection taps (one ring buffer per local kind), and goes through the full encode → AEAD → UDP path. The encode path already handles channels == 1 || 2 (proven by the WASAPI stereo loopback work, 2026-06-17 entry below). The only problems: it's marked TEST-ONLY in the header, the name signals "don't use this in production," and it hardcodes mono (no channels parameter).

    What's missing (output half): today decoded remote audio is mixed and pushed to the miniaudio playback device (on_playback). There's no way for an external consumer to intercept the decoded PCM of a specific remote stream — it all goes to the hardware device. A bot that wants to record, transcribe, or re-broadcast remote audio has no hook.

    Plan (API design — clean, append-only, no struct changes, ABI-stable):

    • vc_stream_feed_pcm — promote vc_test_inject_capture to a public, documented API and add a channels parameter:
      /* External PCM feed — replaces the hardware capture device for this stream. Caller
         provides interleaved int16 PCM at the stream's sample rate. The core frames it,
         encodes (Opus), seals (AEAD), and sends (UDP). Works for any stream kind
         (MIC/SCREEN_AUDIO/AUX_DEVICE). The stream must be started first (vc_stream_start);
         this just replaces the capture source. channels = 1 (mono) or 2 (stereo interleaved).
         Thread-safe; may be called from any thread including audio callbacks. */
      vc_result vc_stream_feed_pcm(vc_client* c, uint32_t stream_id,
                                   const int16_t* pcm, size_t samples_per_channel,
                                   uint32_t channels);
      
    • vc_set_pcm_sink — symmetric output side: receive decoded remote audio as int16 PCM instead of (or in addition to) the hardware playback device:
      /* External PCM tap — receive decoded, mixed remote audio as int16 PCM. The callback
         fires on the audio thread with the mixed output for a specific remote stream. Pass
         cb=NULL to disable (default: disabled, hardware playback only). When enabled, PCM is
         delivered to the sink AND the hardware device (dual output) so a bot can record
         without disabling local monitoring. user_id+stream_id identify the source stream.
         The callback MUST NOT block — copy what you need and return (same contract as
         vc_callbacks.on_event). */
      typedef void (*vc_pcm_sink_cb)(void* user, uint32_t user_id, uint32_t stream_id,
                                     const int16_t* pcm, size_t samples_per_channel,
                                     uint32_t channels, uint32_t sample_rate);
      vc_result vc_set_pcm_sink(vc_client* c, vc_pcm_sink_cb cb, void* user);
      
    • Core changes:
      • core/include/voicecat.h — add vc_pcm_sink_cb typedef + the two function declarations (append-only, after vc_test_inject_capture). Full doc comments on both (contract, thread-safety, lifetime, use cases).
      • core/src/voicecat.cpp — thin C trampolines → vc_client::stream_feed_pcm / set_pcm_sink.
      • core/src/core/client.{h,cpp}stream_feed_pcm: validates stream_id, looks up the LocalStream's kind, calls audio_engine_.inject_capture(kind, pcm, n) (existing path) with the channel count forwarded. set_pcm_sink: stores the callback + user pointer; on_playback (or a new fan-out in the mixer) invokes it per remote stream alongside the existing hardware write. Keep vc_test_inject_capture as a deprecated alias calling stream_feed_pcm(..., channels=1) for source compatibility.
      • core/src/audio/audio_engine.{h,cpp}inject_capture already exists per-kind; add a channels parameter to the ring-buffer write path (or a parallel stereo-aware variant). The encode path in client.cpp::on_capture_frame already handles channels==2 via the stereo encode branch — just plumb the value through. For the sink: add a pcm_sink_ member (callback + user); in on_playback after mixing, if the sink is set, copy the mixed PCM for the current stream and invoke the callback. The copy must stay off the RT-critical path — document the non-blocking contract.
    • Skeleton stub path: update client.cpp's #else (no-deps) stub section to add vc_stream_feed_pcm/vc_set_pcm_sink returning VC_ERR_NOT_IMPLEMENTED — keeps the skeleton preset green.
    • Swift VoiceCatCore: add feedPcm(streamId:pcm:samplesPerChannel:channels:) and setPcmSink(_:user:) (the Swift wrapper around vc_pcm_sink_cb — a @convention(c) closure + Unmanaged context, mirroring Callbacks.swift). Wraps both new ABI functions.
    • C# VoiceCat.Interop: add StreamFeedPcm(streamId, pcm, samples, channels) (with int16[] marshaling) and SetPcmSink (delegates via [UnmanagedCallersOnly] thunk, mirroring the event-callback pattern). Wraps both new ABI functions.
    • Tests:
      • tests/test_external_pcm.cpp (new) — test_feed_pcm_round_trip: two clients, A feeds a known mono sine wave via vc_stream_feed_pcm on a MIC stream, B receives via the normal decode path and asserts energy matches. test_feed_pcm_stereo: same with channels=2, assert L≠R end-to-end (mirrors the WASAPI loopback stereo test). test_pcm_sink: B sets a vc_pcm_sink_cb, A feeds PCM, assert the sink callback receives the decoded PCM with matching energy. All headless, no audio hardware.
      • clients/apple/Tests/VoiceCatCoreTests/ — Swift wrapper round-trip for feedPcm.
      • clients/windows/VoiceCat.Interop.Tests/ — C# wrapper round-trip.
    • Docs:
      • docs/architecture.md §4 — new subsection on external PCM feed/tap: the contract (caller provides interleaved int16 at the stream's sample rate; core frames/encodes/ seals/sends for feed; core decodes/mixes/delivers for sink; sink callback must not block), the use cases (ReplayKit, ScreenCaptureKit, bots, custom clients), and the relationship to vc_test_inject_capture (deprecated alias).
      • docs/voice.md §9 — update the iOS ReplayKit and macOS ScreenCaptureKit rows: both now consume vc_stream_feed_pcm instead of a "minimal slice of the core." Update the iOS detail bullets: the extension links Opus + vc_stream_feed_pcm (not a parallel media stack). Add a macOS ScreenCaptureKit note: convert CMSampleBuffer → int16, feed via vc_stream_feed_pcm — this is how macOS screen-audio actually ships.
      • docs/protocol.md — no protocol changes (the feed/sink are client-local; the wire format is identical whether PCM came from miniaudio or an external source). Note this explicitly.
      • docs/roadmap.md — add a milestone entry; update the iOS ReplayKit and macOS ScreenCaptureKit pending items to reference vc_stream_feed_pcm.
    • Implementation order:
      1. C ABI + core (voicecat.h, voicecat.cpp, client.{h,cpp}, audio_engine.{h,cpp}) + skeleton stub. Verify ctest --preset dev green.
      2. tests/test_external_pcm.cpp — the three behavior tests. Verify green.
      3. Swift VoiceCatCore wrapper + VoiceCatCoreTests round-trip.
      4. C# VoiceCat.Interop wrapper + VoiceCatClientSmokeTests round-trip.
      5. Docs (architecture.md, voice.md, protocol.md, roadmap.md, header comments).
      6. Then ReplayKit (iOS) and ScreenCaptureKit (macOS) become ~100-line consumers of this API instead of parallel media stacks.
    • Verification: ctest --preset dev green (3 new tests); swift test green; dotnet test green; xcodebuild (skeleton) green. The feed/sink tests are fully headless — no audio hardware, no simulator, no device — so they run in CI on every platform.
    • Files to touch:
      • Core C++: core/include/voicecat.h, core/src/voicecat.cpp, core/src/core/client.{h,cpp}, core/src/audio/audio_engine.{h,cpp}.
      • Tests: tests/test_external_pcm.cpp (new), tests/CMakeLists.txt.
      • Swift: clients/apple/Sources/VoiceCatCore/VoiceCatClient.swift, clients/apple/Sources/VoiceCatCore/Callbacks.swift, clients/apple/Tests/VoiceCatCoreTests/ExternalPcmTests.swift (new).
      • C#: clients/windows/VoiceCat.Interop/VoiceCatClient.cs, clients/windows/VoiceCat.Interop/NativeMethods.cs, clients/windows/VoiceCat.Interop.Tests/ExternalPcmTests.cs (new).
      • Docs: docs/architecture.md, docs/voice.md, docs/protocol.md, docs/roadmap.md.
    • ABI stability: append-only — two new functions + one new typedef, no existing structs/enums changed. vc_test_inject_capture stays as a deprecated alias for source compatibility. Treat as a deliberate, versioned ABI event per docs/protocol.md §8.
    • Relationship to the iOS audio routing plan: orthogonal. The iOS routing layer controls which hardware route miniaudio opens (AVAudioSession config in Swift). This plan is about bypassing miniaudio's hardware entirely (external PCM feed/tap). Both ship; they don't conflict. ReplayKit/ScreenCaptureKit consume this API; the iOS routing layer controls the mic path which still uses miniaudio's device.

Recent completed work

All items below are [x] done; ctest --preset dev 23/23 on Windows after all.

  • External PCM feed/tap API (2026-06-20): vc_stream_feed_pcm + vc_set_pcm_sink shipped. Promotes vc_test_inject_capture (mono-only, TEST-ONLY) to a public, stereo-capable API. Adds symmetric PCM sink on the playback thread. Swift wrapper (feedPcm/setPcmSink in VoiceCatClient.swift, 4 XCTest smoke tests). C# wrapper (StreamFeedPcm/SetPcmSink in VoiceCatClient.cs + NativeMethods.cs, 4 xUnit smoke tests in ExternalPcmTests.cs). Three new headless C++ ctests. Docs: architecture.md §4 new subsection, voice.md §9 updated, protocol.md §8 explicit no-protocol-change note, roadmap.md M5 entry. Files: voicecat.h, voicecat.cpp, client.{h,cpp}, audio_engine.{h,cpp}, tests/test_external_pcm.cpp, tests/CMakeLists.txt, Swift + C# wrappers.

  • iOS A2DP + stereo root cause fix (2026-06-20): miniaudio's NULL-context ma_device_init was calling AVAudioSession setCategory(Record) on every device open, wiping the session config IOSAudioRouter had set. Fixed by sharing a ma_context with sessionCategory=none + noAudioSessionActivate/Deactivate=MA_TRUE — miniaudio never touches AVAudioSession; IOSAudioRouter is the sole owner. Files: audio_engine.{h,cpp}.

  • iOS audio routing overhaul (2026-06-19): Full IOSAudioRouter singleton drives all AVAudioSession config before miniaudio opens devices. Fixed stereo mic polar-pattern setup (WWDC20 recipe: setPreferredInput + setInputDataSource + .stereo polar pattern + no setPreferredInputNumberOfChannels). Added vc_audio_restart ABI (full stop+reinit for close→reconfigure→reopen ordering). Added vc_set_capture_channels ABI (core stereo-mic support). AVAudioSession activated proactively on .authResult, not lazily on .streamStarted. Join/Leave Voice button added (parity with macOS). Channel-id sync fixed (mic button was permanently dimmed). iOS deployment target raised to 18.0.

  • iOS SwiftUI client (2026-06-19): VoiceCatiOS.xcodeproj at clients/apple/iOS/. Full feature parity with macOS/Windows: saved server list (JSON + Keychain, App Group group.cat.voice.VoiceCat), TOFU, connect flow, channel tree, user list with context menus, chat, admin sheets, voice controls, settings. xcodebuild → BUILD SUCCEEDED.

  • macOS AppKit client (2026-06-18): VoiceCatMac.xcodeproj at clients/apple/macOS/. Fixed compile errors (NSAccessibility call-site arg order, StreamSummary.id vs .streamId) and linker issues (OTHER_LDFLAGS = -lc++, ONLY_ACTIVE_ARCH = YES for Release). Debug + Release both BUILD SUCCEEDED.

  • Swift VoiceCatCore package + XCFramework (2026-06-18): Shared Swift wrapper at clients/apple/. build-xcframework.sh merges libvoicecat.a + 107 vcpkg static deps into a fat .a via libtool -static. 6/6 Swift tests green (real server, mirrors C# Interop tests). Supports macOS-arm64 + iOS-arm64 + iOS-sim slices.

  • macOS port validated (2026-06-18): 21/21 on macOS. Three cross-platform bugs fixed: missing <netdb.h> in POSIX test branch; SIGPIPE kills (added SIG_IGN); use-after-free of Asio kqueue reactor on server shutdown (fixed TcpAcceptor shutdown/connection-drain sequence).

  • CMake preset cleanup (2026-06-18): m1-devdev, devskeleton, m2-dev dropped. New release, server-release (stripped), apple-dev/apple-ios/apple-ios-sim. Cross- platform triplet auto-resolved by cmake/voicecat-toolchain.cmake.

  • Disconnect, keepalive & reaper (2026-06-18): Client sends Ping every 15 s; server reaper drops sessions after 45 s; UDP KEEPALIVE every 5 s keeps NAT alive. vc_disconnect sends graceful Disconnect proto. Stale-user LEFT broadcast on drop. PLC capped at ~2 s. Three new tests: test_disconnect_left, test_plc_cap, test_reaper_timeout.

  • Stereo screen-audio loopback (2026-06-17): WASAPI loopback opens in channel's stereo/mono mode (was hardcoded mono). Real stereo flows end-to-end through loopback → encode → decode → mixer. New test_loopback_stereo_capture.

  • Windows screen-audio UI wired (2026-06-17): btnScreenShareToggle in MainForm.cs. No core/proto/ABI changes — all the plumbing was already there. dotnet test 4/4 green.

  • Bug fixes (2026-06-16 2026-06-17):

    • AEAD nonce desync in SFU relay — relay forwarded sender's seq verbatim; recipient nonce reconstruction used the wrong counter. Fixed by rewriting the outgoing seq field to the recipient's peek_send_counter().
    • Playout clock free-ranplayout_ts advanced even during VAD/PTT silence gaps, eventually dropping all frames as too-late. Fixed with resync in on_playback via JitterBuffer::peek_front_ts().
    • Stale users after disconnectConnSession::close() didn't broadcast UserEvent::LEFT before erasing. Fixed; PLC cap added as defense-in-depth.
    • "Randomly bumped to Lobby" — server excluded the actor from its own state-change broadcasts. Fixed: UserEvent::UPDATED now goes to all clients including the actor.
    • Silent playback after joinopus_decode received hardware callback frame count as max_samples instead of the Opus frame size. Fixed with a decode ring buffer.

Milestones (see docs/roadmap.md for full detail)

  • M0 — Scaffolding ✓ complete
  • M1 — Control plane ✓ complete (2026-06-15)
  • M2 — Voice, single stream ✓ complete (2026-06-16)
  • M3 — Multi-stream & per-channel tuning ✓ complete (2026-06-16)
  • M4 — Native clients — Windows WinForms ✓ (2026-06-17); macOS AppKit ✓ (2026-06-18); iOS SwiftUI ✓ (2026-06-19)
  • [~] M5 — Moderation, polish, beyond (perms, bans, DRED; then file transfer, E2EE, …)

M0 — Scaffolding ✓

Repo layout (core/ server/ tools/ clients/ tests/), CMake + vcpkg manifest, C ABI header (voicecat.h), proto source of truth, core stubs for all six subsystems, voicecat-server + vccli skeletons, smoke CTest, .clang-format/.gitattributes/.gitignore.


M1 — Control plane ✓ (completed 2026-06-15)

Exit criterion: test_m1_integration — two clients authenticate over TLS 1.3 (guest + Argon2id), exchange channel + private text. ~1 s.

FrameCodec, TlsContext (mbedTLS 1.3, ECDSA-P256 self-signed, TOFU pins TLS leaf-cert SHA-256), WorkerPool, Database (SQLite + Argon2id), ServerIdentityManager, ConnSession state machine, SessionRegistry, vc_client full M1 C ABI, voicecat-admin CLI, dual-stack TcpAcceptor. Key bug fixed: send_frame double-framing — encode_envelope was pre-framing the protobuf; fixed by passing raw protobuf bytes.


M2 — Voice, single stream ✓ (completed 2026-06-16)

Exit criterion: test_m2_voice + test_voice_client_abi — two headless clients auth, bind UDP, 50 Opus frames relayed + re-encrypted by SFU, B receives ≥25 and decrypts. ~4 s.

14-byte UDP voice header, SodiumMediaCrypto (ChaCha20-Poly1305 + 64-bit anti-replay), OpusEncoder/OpusDecoder (FEC, PLC), UdpMediaChannel, JitterBuffer, AudioEngine (miniaudio), MediaRelay SFU. Key bug fixed: on_playback passed hardware callback frame count as opus_decode max_samples; fixed with a per-stream decode ring buffer.


M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)

Exit criterion: test_m3_multistream — client A runs two concurrent streams (MIC + SCREEN_AUDIO); B sees both; per-stream gain/mute/NR independent; effective Opus config matches channel's server-enforced settings. ~2.4 s.

Fixed server stream_id counter bug (always wrote 1). Per-channel AudioConfig populated (Lobby: mono/24kbps/VOIP + DTX; Music Room: stereo/128kbps/AUDIO). LocalStream map, pending_announce_kind_, run_talk_timer(), thread-join race in teardown_voice() fixed. New C ABI: vc_get_stream_audio_config, vc_test_inject_capture.


Post-M3 follow-up ✓ (completed 2026-06-16)

  • Device enumerationvc_list_devices/vc_set_input_device; opaque hex device ids; vc_free_device_list now frees. Works pre-connect.
  • VAD/PTT gateEnergyVadProcessor (RMS threshold ~0.025, 300 ms hang-time); vc_set_input_mode/vc_set_push_to_talk; MIC-only (SCREEN_AUDIO/AUX_DEVICE bypass).
  • True stereo playbackplayback_channels=2; stereo decoded L→L R→R in mixer; mono upmixed L=R; hardware fallback to mono on failure.
  • WASAPI loopbackloopback_device_ with ma_device_type_loopback; VOICECAT_HAS_LOOPBACK macro (Windows-only). vccli --share-screen-audio.

Known deferred (still open): AEC/NS/AGC (no working Windows/MSVC WebRTC APM build); process-specific WASAPI loopback; RT-thread rule violation in on_capture_frame (mutex lock on audio callback thread — pre-existing, needs lock-free ring-buffer refactor).


M4 — Native clients ✓ (completed 2026-06-17 2026-06-19)

Exit criterion: ctest --preset dev 21/21 green; dotnet build 0 warnings; xcodebuild BUILD SUCCEEDED (macOS + iOS); manually verified: connect, TOFU, channel tree, join, voice, text, device pickers, level meter on each platform.

New C ABI (additive): vc_list_channels/vc_list_users/vc_list_user_streams, vc_join_channel, VC_EVENT_SERVER_IDENTITY + vc_confirm_server_identity, vc_config::tofu_store_path, VC_INPUT_ALWAYS_ON, vc_set_vad_threshold, vc_audio_suspend/vc_audio_resume, vc_audio_restart, vc_set_capture_channels.

Windows (clients/windows/): VoiceCat.Interop (P/Invoke, [UnmanagedCallersOnly]), VoiceCat.App (ConnectDialog, ServerIdentityDialog, MainForm with full M5 moderation UI, PerUserTuningDialog, PttKeyCaptureDialog), VoiceCat.Interop.Tests. PTT is focus-scoped.

macOS (clients/apple/macOS/VoiceCatMac.xcodeproj): NSOutlineView channel tree, NSTableView user list, NSTextView chat, voice controls, full VoiceOver accessibility, admin menu, 17 Swift source files. build-xcframework.sh produces VoiceCatCore.xcframework.

iOS (clients/apple/iOS/VoiceCatiOS.xcodeproj): SwiftUI, NavigationSplitView/TabView, OutlineGroup channel tree, IOSAudioRouter AVAudioSession driver, 24 Swift source files, iOS 18.0 deployment target. App Group group.cat.voice.VoiceCat for Keychain sharing.


M5 — Moderation, polish, and beyond [~] (in progress 2026-06-17)

Exit criterion: four ABI-level tests green (test_m5_permissions, test_m5_kick_ban_move_mute, test_m5_admin_accounts, test_m5_channel_crud); vccli can drive all moderation/admin/channel operations against a live server.

  • Server-side moderation & permissions — per-session Permissions, kick/ban/move/ server-mute, channel CRUD, DB schema v2 (channels, bans), BLAKE2b channel passwords.
  • C ABIvc_kick_user, vc_ban_user, vc_set_permission, vc_set_server_mute, vc_move_user, vc_create_channel, vc_edit_channel, vc_delete_channel, vc_create_account, vc_reset_password, vc_delete_account, vc_list_accounts, vc_get_permissions; events VC_EVENT_GENERIC_RESULT, VC_EVENT_ACCOUNT_LIST.
  • Four M5 tests passing — ctest --preset dev 21/21.
  • vccli M5 flags: --kick, --ban, --move, --server-mute/-unmute/-deafen/ -undeafen, --set-permission, channel CRUD, account CRUD, --username/--password.
  • All three client UIs (Windows WinForms, macOS AppKit, iOS SwiftUI) expose the full M5 moderation and admin surface.
  • Docsdocs/protocol.md, docs/security.md kept in sync.
  • DRED/audio-quality polish — done (2026-06-20). bool dred added to AudioConfig proto (field 11) and vc_audio_config C ABI. Encoder: OPUS_SET_DRED_DURATION(2) when enabled (20 ms of ML redundancy per packet). Decoder: OpusDREDDecoder + per-stream OpusDRED scratch pre-allocated; JitterBuffer::try_copy_front_payload peeks at the next buffered packet on every PLC step; if DRED data is present, opus_decoder_dred_decode reconstructs the lost frame — otherwise falls back to standard PLC. New test: test_dred_toggle (ctest 22/22). Files: voicecat.proto, voicecat.h, opus_codec.{h,cpp}, audio_engine.{h,cpp}, client.cpp, session.{h,cpp}.
  • DRED toggle in client UIs — expose the dred flag in all three channel-config UIs so admins can enable it per channel. Windows: ChannelEditForm / vc_channel_info.audio.dred checkbox. macOS AppKit: channel-edit sheet. iOS SwiftUI: channel-edit form. All three UIs already have full channel CRUD wired; this is an additive checkbox on the existing audio-config section. (Core/protocol/ABI all done — this is UI-only work.)
  • macOS ScreenCaptureKit screen-audio — done 2026-06-21. ScreenAudioCapture.swift drives an SCStream (audio-only, excludesCurrentProcessAudio), converts Float32 → int16 in the channel's mono/stereo mode, and calls vc_stream_feed_pcm. Capture starts on the self .streamStarted event (when the effective config is known); wired into MainWindowController.screenAudioClicked().
  • iOS ReplayKit Broadcast Extension (VoiceCatBroadcast) — done 2026-06-21. Forward-to-host design: the extension (SampleHandler.swift) captures .audioApp, converts to 48 kHz int16 stereo, and writes a shared App Group SPSC ring (BroadcastAudioRing.swift); the host's BroadcastAudioPump owns the SCREEN_AUDIO stream and feeds via vc_stream_feed_pcm (single session, no creds on disk). UI is an RPSystemBroadcastPickerView in VoiceControlsView. (Replaced the speculative BroadcastCredentials.swift self-connecting design, now removed.)
  • External PCM feed/tap API (vc_stream_feed_pcm + vc_set_pcm_sink) — done 2026-06-20. Promotes vc_test_inject_capture (mono-only, TEST-ONLY) to a public API with stereo support. Adds a symmetric PCM sink fired on the playback thread per decoded remote stream. Full wrappers for Swift (feedPcm/setPcmSink) and C# (StreamFeedPcm/ SetPcmSink). Three new C++ ctests (test_feed_pcm_round_trip, test_feed_pcm_stereo, test_pcm_sink), 4 Swift XCTest smoke tests, 4 C# xUnit smoke tests. Docs updated (architecture.md §4 new subsection, voice.md §9 updated, protocol.md §8 explicit no-protocol-change note, roadmap.md M5 entry). ctest --preset dev 23/23.
  • macOS client UI overhaul — done 2026-06-20. Mirrors the Windows client's UI overhaul (toolbar, unified log, PM windows, channel counts, output volume, keyboard shortcuts), adapted to Mac-native conventions:
    • NSToolbar: Join Voice, Share Screen Audio, Mute, Deafen (SF Symbol toggle buttons), and Output Volume slider (NSSlider 0100, default 80). Voice actions + mute/deafen + output volume moved out of the bottom voice panel into the toolbar. Bottom panel keeps input-mode segmented control / VAD slider / PTT key / device picker / level meter.
    • Unified log: chat NSTextView + activity NSTableView collapsed into a single NSTextView — activity events in secondaryLabelColor (gray), chat in default color. Removed activityTableView and activityLog array.
    • Private messaging: scope dropdown removed; compose bar always sends to the current channel. Each PM conversation opens in its own modeless PrivateMessageWindowController (NSWindow). Incoming .textMessage with .private scope routed to the right window; outgoing PMs echoed by server arrive through the same path. "Send Private Message…" added to user context menu. "New Private Message…" (⌘⇧N) opens UserPickerSheet listing all server users.
    • Channel counts: outline view renders "Name (n)" with live user counts; refreshChannelTree() called on .userJoined/.userLeft (was missing).
    • Voice menu (⌘⇧V join/leave, ⌘⇧S share screen, ⌘⇧M mute, ⌘⇧D deafen) and Messages menu (⌘⇧N new PM) added to NSApp.mainMenu via NSMenuItem key equivalents with [.command, .shift] mask. Removed on windowWillClose. Mac-native: ⌘ not Ctrl, dispatched by the responder chain (no custom key monitor needed).
    • Output volume: setOutputVolume(_:) wrapper added to VoiceCatClient.swift (was missing — the C ABI + C# wrapper shipped in commit 97fa659 but the Swift wrapper was never added). Wired end-to-end: toolbar slider → client.setOutputVolume(gain).
    • Part A (uncompiled Swift fix): the external PCM feed/tap Swift wrapper (commit 615d2a8) was never compiled — the local xcframework predating the voicecat.h PCM additions. Fixed: rebuilt xcframework (regenerated module map), fixed UIntInt type mismatch in feedPcm (Swift imports size_t as Int not UInt), added VoiceCatPcmSinkCallback typealias (Swift-idiomatic alias for the C vc_pcm_sink_cb so consumers don't need to directly import VoiceCatC). swift test 10/10 green.
    • Audio settings moved to Settings window: the bottom voice panel (input mode, VAD slider, PTT key, device picker, level meter) was removed from the main window and moved into a new SettingsWindowController — a modeless window opened via the app menu's "Settings…" (⌘,) item. The main window is now just toolbar + channels + users + chat. Source-of-truth for audio settings (selectedInputMode, vadThresholdValue, selectedInputDeviceId, pttKeyCode) lives in MainWindowController so voice start can apply them even before the settings window has been opened; SettingsWindowController reads from and writes back to those properties and applies changes to the client immediately when voice is active. The level meter is forwarded from MainWindowController.handleLevelsettingsWindowController.updateLevel(rms:). keyCodeName helper deduplicated (was duplicated in PttKeyCaptureSheet.swift + MainWindowController.swift — now shared from MainWindowController.swift).
    • Files: MainWindowController.swift (overhauled), PrivateMessageWindowController.swift (new), UserPickerSheet.swift (new), SettingsWindowController.swift (new), VoiceCatClient.swift (setOutputVolume + VoiceCatPcmSinkCallback typealias + feedPcm type fix), ExternalPcmTests.swift (use typealias), PttKeyCaptureSheet.swift (removed duplicate keyCodeName), VoiceCatMac.xcodeproj/project.pbxproj (register 3 new files).
    • Platform-specific adaptations (vs. Windows): NSToolbar instead of ToolStrip; global menu bar + NSMenuItem key equivalents (⌘ not Ctrl, responder-chain dispatched); PM windows as modeless NSWindows; picker as Mac sheet; gray = secondaryLabelColor; SF Symbols for toolbar icons.

Decisions log

All architecture/scope decisions are settled and recorded in docs/roadmap.md §2 "Resolved decisions" and reflected across docs/. If you make a new decision, record it there and link it here.


How to update this file

  1. Check off tasks as you complete them; flip a milestone to [x] only when its exit criterion test passes.
  2. Keep the "Where we left off / next action" block at the top accurate — it's the first thing the next agent reads.
  3. When you start a milestone, copy its task list from docs/roadmap.md into a section here.