Promotes vc_test_inject_capture (mono-only, TEST-ONLY) to a public, stereo-capable production API and adds a symmetric PCM tap on the receive side. Enables ReplayKit (iOS), ScreenCaptureKit (macOS), bots, soundboards, and custom clients — all without a hardware audio device. Core C++: - voicecat.h: new vc_stream_feed_pcm, vc_pcm_sink_cb typedef, vc_set_pcm_sink; vc_test_inject_capture kept as deprecated alias - audio_engine: stereo-aware inject_capture (channels param + ring reset on channel-count change); atomic pcm_sink_ fired per decoded frame in on_playback; RemoteStream carries user_id/stream_id for RT-safe sink metadata; init_recv_stream takes user_id+stream_id - client.cpp: stream_feed_pcm / set_pcm_sink implementations; sync_remote_streams passes user_id/stream_id to init_recv_stream - voicecat.cpp: trampolines + channels=1/2 validation Tests: test_external_pcm (headless, 3 sub-tests: mono round-trip, stereo feed L≠R, sink metadata+disable). ctest 23/23. Swift: feedPcm / setPcmSink in VoiceCatClient.swift + 4 XCTest smoke tests (ExternalPcmTests.swift). C#: StreamFeedPcm / SetPcmSink in VoiceCatClient.cs + NativeMethods.cs (vc_stream_feed_pcm unsafe P/Invoke, VcPcmSinkCallback delegate, vc_set_pcm_sink via nint) + 4 xUnit smoke tests (ExternalPcmTests.cs). Docs: architecture.md §4 new subsection, voice.md §9 updated (macOS/iOS now reference vc_stream_feed_pcm), protocol.md §8 explicit no-protocol-change note, roadmap.md M5 entry. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
VoiceCat — Windows client
WinForms (.NET 10 LTS) UI over voicecat.dll (MinGW-built libvoicecat shared library).
Prerequisites
| Tool | Version | Notes |
|---|---|---|
| .NET SDK | 10.0.x | dotnet --version should report 10.0.* |
| CMake | 3.25+ | For building the C++ DLL |
| MinGW-w64 / MSYS2 UCRT64 | GCC 13+ | C:\tools\msys64\ucrt64 is the expected location |
| vcpkg | any | VCPKG_ROOT env var must point to a bootstrapped clone |
Build order
1. Build the server (for testing)
cmake --preset dev
cmake --build --preset dev --target voicecat-server
2. Build the DLL
cmake --preset windows-client
cmake --build --preset windows-client
Output: build/windows-client/bin/voicecat.dll
Verify no MinGW runtime dependencies remain:
& "C:\tools\msys64\ucrt64\bin\objdump.exe" -p build/windows-client/bin/voicecat.dll |
Select-String "DLL Name"
Expected: only Windows system DLLs (KERNEL32.dll, WS2_32.dll, BCRYPT.dll, etc.).
If libgcc_s_seh-1.dll, libstdc++-6.dll, or libwinpthread-1.dll appear, the
-static-libgcc -static-libstdc++ -static -lwinpthread link flags in core/CMakeLists.txt
are not taking effect — check the CMake log for the VOICECAT_BUILD_SHARED+WIN32 branch.
3. Build the C# solution
cd clients/windows
dotnet build VoiceCat.slnx
The app's Directory.Build.props copies voicecat.dll from ../../build/windows-client/bin/
into the output directory automatically on every build.
Running manually
# Terminal 1 — start the server
./build/dev/bin/voicecat-server.exe --name "My Server"
# Terminal 2 — launch the client
dotnet run --project clients/windows/VoiceCat.App/VoiceCat.App.csproj
On first connect to a new server:
- Enter
127.0.0.1as the host (notlocalhost— Windows resolveslocalhostto::1first, and while the server now dual-stacks,127.0.0.1is cleaner for local testing). - The server identity dialog will appear. The TLS leaf-cert SHA-256 fingerprint is shown; accept to pin it. Subsequent connects to the same server will be silent (MATCHED).
M5 — Moderation & admin UI
The WinForms client now exposes all M5 operations through the main menu and context menus:
- Admin → Server accounts… — create, reset password, and delete server accounts
(requires
can_admin_accounts). - Channel tree right-click — create, edit, and delete channels. The edit dialog exposes the full per-channel Opus configuration: mono/stereo, sample rate, bitrate, frame size, application mode, FEC, expected packet loss, DTX, and complexity.
- User list right-click — move, kick, ban, server mute/deafen, and set permissions (items are gated by your own permissions).
- Activity log shows async
GenericResultfeedback for every moderation request. - User list shows text indicators for self-mute, self-deafen, server-mute, and server-deafen states.
These operations require an admin-provisioned account with the appropriate permissions; the connect dialog already supports username/password auth.
Known limitations
- PTT is focus-scoped — the push-to-talk key only works while the VoiceCat window has
focus. A system-wide
WH_KEYBOARD_LLhook is not used in v1 (permissions + AV risk). - Receive-side noise reduction checkbox in per-user tuning is wired end-to-end but is a
passthrough no-op until a real APM/NS backend is built (no working Windows/MSVC port of
webrtc-audio-processingupstream — seedocs/tech-stack.md §1). - TOFU pins the TLS leaf cert, not the declared Ed25519 identity fingerprint. Both are
shown in the identity dialog, but the cert fingerprint is the value that is actually
verified on reconnect. See
docs/security.md §1.1.