Rationalize the preset set to match the project's actual state (past M5): - Rename dev->skeleton (no-deps stub smoke), m1-dev->dev (default dev preset) - Drop m2-dev (cache-identical to m1-dev) - Add release preset (optimized + tests on, symbols kept) - Strip server-release binaries (-s linker flag) - Add apple-dev/apple-ios/apple-ios-sim scaffolding presets for XCFramework Add cmake/voicecat-toolchain.cmake wrapper that auto-resolves the vcpkg triplet from the host platform (x64-mingw-static/x64-linux/arm64-osx) so the main presets work on Windows/Linux/macOS without per-OS variants. Update all docs (building.md, CLAUDE.md, README.md, AGENTS.md, deployment.md, tech-stack.md, client READMEs) and stale preset-name references in code comments. No C++ behavior changes — the core was already portable.
728 lines
55 KiB
Markdown
728 lines
55 KiB
Markdown
# PROGRESS — VoiceCat
|
||
|
||
Living status. **Update this file in the same commit as your work** so the next agent picks
|
||
up instantly. Newest status at the top.
|
||
|
||
- **Date convention:** ISO (YYYY-MM-DD).
|
||
- Statuses: `[ ]` not started · `[~]` in progress · `[x]` done.
|
||
|
||
---
|
||
|
||
## ▶ Where we left off / next action
|
||
|
||
- **Done:** **CMake preset cleanup + cross-platform build config** (2026-06-18). The preset
|
||
set was a mess — `dev` (never used), `m1-dev` (the one everyone used), `m2-dev`
|
||
(cache-identical to `m1-dev`, never used), no optimized+tests preset, no stripping.
|
||
Cleaned up to a sensible set + added cross-platform triplet auto-resolution + Apple
|
||
platform scaffolding:
|
||
- **Renames:** `dev`→`skeleton` (no-deps stub smoke — accurately named now); `m1-dev`→`dev`
|
||
(the default development preset — milestone-named presets were misleading since the
|
||
project is past M5); `m2-dev` **dropped** (cache-identical to `m1-dev`).
|
||
- **New presets:** `release` (optimized Release + tests on, symbols kept — run the suite
|
||
against optimized code or profile); `apple-dev` / `apple-ios` / `apple-ios-sim`
|
||
(scaffolding — static `libvoicecat.a` slices for the Swift Package / XCFramework; marked
|
||
"not yet CI-validated, build on macOS to verify").
|
||
- **`server-release`** now strips binaries (`CMAKE_EXE_LINKER_FLAGS=-s` +
|
||
`CMAKE_SHARED_LINKER_FLAGS=-s`) — smaller executables for deployment.
|
||
- **Cross-platform triplet auto-resolution:** new `cmake/voicecat-toolchain.cmake` wraps
|
||
vcpkg's toolchain and resolves `VCPKG_TARGET_TRIPLET` / `VCPKG_HOST_TRIPLET` from
|
||
`CMAKE_HOST_SYSTEM_NAME` + `CMAKE_HOST_SYSTEM_PROCESSOR` — `x64-mingw-static` on Windows,
|
||
`x64-linux` on Linux, `arm64-osx` on Apple Silicon. The main presets (`dev`, `release`,
|
||
`server-release`) now work on all three platforms without per-OS variants. Cross-compile
|
||
presets (`apple-ios`, `apple-ios-sim`) override `VCPKG_TARGET_TRIPLET` explicitly. Hidden
|
||
base preset renamed `vcpkg-base`→`vcpkg-common` (now points at the wrapper toolchain
|
||
instead of vcpkg's toolchain directly).
|
||
- **No C++ source changes** — the core was already portable (every `#ifdef _WIN32` in
|
||
`client.cpp` already had a POSIX `#else`; `voicecat.h`'s export macro already handled
|
||
GCC visibility; `transport.cpp` is pure Asio; miniaudio abstracts WASAPI/CoreAudio/ALSA).
|
||
- **Docs updated:** `docs/building.md` (full rewrite — 8-preset table, platform matrix,
|
||
preset history mapping old names to new, Apple scaffolding section), `CLAUDE.md` (build
|
||
section reframed around `dev` as default, status line updated to M5-done), `README.md`
|
||
(stale "M0 skeleton" framing replaced with current M5 reality), `AGENTS.md` (build
|
||
section updated, stale "placeholder builtin-baseline" sentence deleted),
|
||
`docs/deployment.md` (server-release now stripped + cross-platform note),
|
||
`docs/tech-stack.md` §4 (triplet auto-resolution + Apple scaffolding note),
|
||
`clients/apple/README.md` (new "Building the core for Apple platforms" section with
|
||
XCFramework workflow), `clients/windows/README.md` (m1-dev→dev).
|
||
- **Code comments updated:** `core/include/voicecat.h`, `core/src/net/transport.h`,
|
||
`core/src/voicecat.cpp`, `tests/CMakeLists.txt` — preset name references updated.
|
||
- **Historical `PROGRESS.md` entries left intact** — `m1-dev`/`m2-dev` mentions in older
|
||
entries are a true record of what was run; rewriting them would falsify history. The
|
||
preset history table in `docs/building.md` §1 maps old names to new.
|
||
- **Verified:** `cmake --list-presets` shows all 8 presets. `skeleton` + `dev` configure
|
||
and build green on Windows. Apple presets are scaffolding (won't build on Windows —
|
||
expected; they require macOS).
|
||
|
||
- **Done:** **Disconnect, timeout & keepalive system** (2026-06-18). Three reported bugs
|
||
traced to one root cause + two missing designed features, all fixed:
|
||
1. **Stale users after disconnect + eternal PLC hiss** (root cause): `ConnSession::close()`
|
||
silently erased dropped users from the registry without broadcasting `UserEvent::LEFT`.
|
||
Peers never learned the user left → their user lists stayed stale AND their audio
|
||
engines never called `remove_stream` → Opus PLC synthesized comfort noise forever
|
||
(the "soft hissing that never goes away"). **Fix:** `SessionRegistry::broadcast_left()`
|
||
helper (mirrors `kick_user`'s first half); `ConnSession::close()` now broadcasts LEFT
|
||
before erasing. Regression test `test_disconnect_left` (`tests/`).
|
||
2. **PLC cap** (defense-in-depth): `AudioEngine::on_playback` now caps pure-PLC at ~2 s
|
||
(`kPlcCapSamples`); after that it emits digital silence instead of more comfort noise,
|
||
so a stale stream can never hiss forever even if `remove_stream` is never called. Resets
|
||
automatically when fresh packets arrive. Test `test_plc_cap` (`tests/`).
|
||
3. **No timeout / no ping** (missing feature): the client never sent `Ping`, the server had
|
||
no `last_seen` / reaper, and half-open connections (NAT timeout, wifi loss, sleep) left
|
||
ghost users forever. **Fix:** client sends `Ping` every 15 s from the io thread (RTT
|
||
measured from `Pong` nonce); `ConnSession::last_seen` bumped on every inbound TCP/UDP
|
||
frame; `asio::steady_timer` reaper sweeps every 15 s and drops sessions older than 45 s
|
||
(configurable via `server::Config::reaper_timeout_ms`/`reaper_sweep_ms`). Each drop
|
||
broadcasts LEFT via fix #1. Test `test_reaper_timeout` (`tests/`, 2 s timeout for fast
|
||
turnaround).
|
||
4. **UDP KEEPALIVE** (missing feature): client sends a plaintext `KEEPALIVE` frame every
|
||
5 s (`voice_frame.h::kFrameKeepalive`); server bumps `last_seen` + echoes back. Keeps
|
||
NAT bindings alive and lets media activity defer the reaper independently of TCP.
|
||
5. **Graceful client disconnect:** `vc_disconnect()` now sends `Disconnect{code=0}` when
|
||
fully authenticated (`VC_STATE_CONNECTED`): it queues the envelope, sets a
|
||
`graceful_disconnect_pending_` flag, and joins the io thread — the io thread's
|
||
`drain_sends()` sends the Disconnect, sees the flag, sets `io_stop_`, and exits
|
||
naturally (its cleanup handles `teardown_voice()` + socket close). No main-thread
|
||
socket close → no double-close race. Server handles client-sent `Disconnect`
|
||
(`handle_client_disconnect` → `close()` → immediate LEFT broadcast, no reaper/EOF
|
||
wait). In non-authenticated states, the original force-close path runs (with TOFU
|
||
unblock).
|
||
- Docs updated: `docs/protocol.md §6/§7` (graceful disconnect, pinned N=3, reaper, last_seen),
|
||
`docs/voice.md §6` (KEEPALIVE plaintext + echo + last_seen), `docs/architecture.md §5`
|
||
(reaper timer), `server::Config` (reaper fields). `ctest --preset m2-dev` — **21/21 green**
|
||
(3 new tests: disconnect_left, plc_cap, reaper_timeout).
|
||
|
||
- **Done:** **Stereo screen-audio loopback capture on Windows** (2026-06-17). The WASAPI
|
||
loopback path (`start_loopback_capture`) used to hardcode `cfg.capture.channels = 1`,
|
||
downmixing the system's stereo mix to mono before the encoder ever saw it — so even on a
|
||
stereo channel, `SCREEN_AUDIO` was effectively mono (the encoder then upmixed L=R to
|
||
produce a *fake* stereo bitstream). Now the loopback device opens in the channel's mode:
|
||
stereo (interleaved L/R) when the channel is configured stereo, mono when mono. Real stereo
|
||
flows end-to-end through loopback → Opus encode → decode → stereo playback mixer.
|
||
- `audio_engine.h` — `CaptureCallback` gained an `int channels` parameter (the encoder
|
||
needs to know whether the PCM is real stereo or mono to avoid upmixing real stereo).
|
||
`start_loopback_capture(int kind)` → `start_loopback_capture(int kind, int channels)`.
|
||
New `loopback_channels_` member; new `feed_loopback_for_test` test hook (self-contained,
|
||
works on headless CI where the real WASAPI device can't init).
|
||
- `audio_engine.cpp` — `on_loopback` accumulator is now channel-aware (sized to
|
||
`frame_samples_*loopback_channels_`); `start_loopback_capture` sizes the accumulator off
|
||
the RT thread before `ma_device_start`, opens the device with `channels`, and falls back
|
||
to mono if the render endpoint rejects stereo (mirrors the playback path's fallback).
|
||
`on_capture`/`inject_capture`/`feed_capture_for_test` forward `channels` through the
|
||
callback (mic path always 1; loopback path 1 or 2).
|
||
- `client.cpp` — `on_capture_frame` takes `channels`; for `channels==2` (real stereo
|
||
loopback PCM) it encodes directly with no upmix; for `channels==1` on a stereo channel it
|
||
keeps the existing L=R upmix (mic stays mono in v1). `handle_stream_announce_result` reads
|
||
`effective_params.stereo` under the lock and passes `2` or `1` to `start_loopback_capture`.
|
||
- `test_vad_ptt_devices.cpp` — new `test_loopback_stereo_capture` behavior test: feeds a
|
||
loud-L / silent-R stereo signal through `feed_loopback_for_test`, encodes (as
|
||
`on_capture_frame` now does for `channels==2`), decodes, mixes, and asserts L≠R across
|
||
the frame (total_diff ~8.2M, well above the 960k threshold). A mono-downmixed-then-
|
||
upmixed bitstream would have L==R. Existing test lambda updated for the 4-arg callback.
|
||
- `docs/voice.md §8/§9` — diagram + notes updated: mic stays mono; `SCREEN_AUDIO` loopback
|
||
captures stereo when the channel is stereo.
|
||
- `cmake --build --preset m2-dev` + `ctest --preset m2-dev --parallel 1` — **18/18 green**.
|
||
The `vad_ptt_devices` VAD-gate sub-test has a pre-existing parallel-run timing flake
|
||
(passes serially and in isolation); unrelated to this change (VAD gate logic is unchanged
|
||
for `channels==1`, the only path the MIC uses).
|
||
- **Done:** **Screen-audio sharing wired into the Windows WinForms client** (2026-06-17).
|
||
The core already fully supported `SCREEN_AUDIO` capture on Windows (post-M3 WASAPI
|
||
loopback via `VOICECAT_HAS_LOOPBACK`, always on for the `windows-client` preset —
|
||
`core/CMakeLists.txt:85`; loopback start/stop at `client.cpp:1093`/`audio_engine.cpp:529`;
|
||
VAD/PTT/self-mute/server-mute correctly bypassed for non-MIC kinds at `client.cpp:862-879`)
|
||
and the C# Interop layer was already complete (`VcStreamKind.ScreenAudio`,
|
||
`StartStream`/`StopStream`/`SetRemoteStream`/`ListUserStreams` all generic). The gap was
|
||
purely UI wiring. **No core, proto, or C ABI changes were needed** — confirming the
|
||
"the core should support it already" assessment.
|
||
- `MainForm.Designer.cs` — new `btnScreenShareToggle` button in the voice panel top row
|
||
(`flpVoiceTop`), right after `btnMicToggle`, with full `AccessibleName`/
|
||
`AccessibleDescription` per the existing accessibility convention.
|
||
- `MainForm.cs` — new `_screenStreamId` field; `BtnScreenShareToggle_Click` handler
|
||
mirroring `BtnMicToggle_Click` but with no device picker / VAD / PTT / mode / mute
|
||
(screen audio bypasses all of those in the core). Independent of mic — can share
|
||
without joining voice and vice versa. `HandleDisconnected` now resets
|
||
`_screenStreamId` and disables the screen toggle. `OnFormClosed` now explicitly stops
|
||
both mic and screen streams before `Disconnect()` (clean `StreamStop` messages go out
|
||
before the control channel closes). `HandleStreamStarted` already labeled
|
||
`ScreenAudio => "screen audio"`; `PerUserTuningDialog.ApplySettings` already iterates
|
||
all of a peer's streams — both unchanged, peers can independently volume-tune a
|
||
screen-audio stream vs that user's mic.
|
||
- `VoiceCatClientSmokeTests.cs` — new `ScreenAudioStream_Starts_And_Stops` test:
|
||
connects + TOFU + guest auth, `StartStream(ScreenAudio)`, asserts
|
||
`VC_EVENT_STREAM_STARTED` arrives with matching `StreamId`, `StopStream`, asserts
|
||
`VC_EVENT_STREAM_STOPPED`. Passes headless (the `StreamAnnounce` succeeds regardless
|
||
of whether the loopback device initializes on a CI box).
|
||
- `dotnet build` — 0 warnings/errors. `dotnet test` — **4/4 tests green**
|
||
(3 existing + 1 new). Event trace confirms the full
|
||
`StreamStarted → UserUpdated → StreamStopped` path through P/Invoke against a live
|
||
`voicecat-server.exe`.
|
||
- **Not yet confirmed audible by ear** — pending manual two-instance live test (one
|
||
shares screen audio while something plays on the default render endpoint, the other
|
||
hears it). This is the observable-behavior exit criterion per `AGENTS.md`.
|
||
- **Documented caveat** (`docs/voice.md §9`, unchanged): whole-device WASAPI loopback
|
||
inherently re-captures this app's own incoming voice mix (self-echo loop) — accepted
|
||
characteristic, not a bug. Process-specific loopback (Windows 10 2004+
|
||
`AUDIOCLIENT_ACTIVATION_PARAMS`) is a future enhancement; miniaudio doesn't expose it.
|
||
- **In progress:** **M5 — moderation & admin** (2026-06-17). Server-side and C ABI are
|
||
implemented and tested: permissions, kick/ban/move/server-mute, channel CRUD, in-app account
|
||
management. Four new tests pass: `test_m5_permissions`, `test_m5_kick_ban_move_mute`,
|
||
`test_m5_admin_accounts`, `test_m5_channel_crud`. `vccli` now exposes all M5 operations via
|
||
CLI flags (`--kick`, `--ban`, `--move`, `--server-mute`/`-unmute`/`-deafen`/`-undeafen`,
|
||
`--set-permission`, `--create-channel`, `--edit-channel`, `--delete-channel`,
|
||
`--create-account`, `--reset-password`, `--delete-account`, `--list-accounts`) plus
|
||
`--username`/`--password` for account auth and `--self-mute`/`--self-deafen`. Docs updated:
|
||
`docs/protocol.md` (envelope tags for `ServerMuteRequest`/`ListAccountsResult`, `User.server_deafened`,
|
||
`GenericResult` usage), `docs/security.md` (BLAKE2b channel passwords, `bans` schema).
|
||
`ctest --preset m1-dev` — **18/18 green**. Windows WinForms UI now exposes all M5 operations:
|
||
channel CRUD with full per-channel Opus audio config, user moderation (kick/ban/move/server
|
||
mute/server deafen/set permissions), and server account management. `dotnet test` of the
|
||
Windows solution passes. Still to do: DRED/audio-quality polish.
|
||
- **Done:** **Fixed multi-user voice — relayed frames failed AEAD decryption (nonce desync)**
|
||
(2026-06-17, reported live: with 2+ people in a channel, audio was one-directional — "I can
|
||
hear them but they can't hear me" — and a 3rd joiner heard nobody). Root cause was in the SFU
|
||
relay (`server/src/media_relay.cpp`). The media AEAD nonce is an *implicit per-direction
|
||
monotonic counter*; `open()` reconstructs it from the 14-byte header's `seq` field (the AAD),
|
||
so the wire contract is `header.seq == the counter seal() used` (the client honors this at
|
||
`client.cpp:907`). The relay decrypted each inbound frame with the sender's key, then re-sealed
|
||
with the **recipient's** `send_crypto` (its own counter) but **forwarded the sender's header
|
||
verbatim** — so `header.seq` carried the sender's counter, not the recipient's. The recipient's
|
||
`open()` rebuilt the wrong nonce → every relayed frame failed auth and was silently dropped. It
|
||
only "worked" while the sender's counter coincidentally equalled the server→recipient counter
|
||
(a single first-ever sender into a fresh recipient), which is exactly why the first/sole talker
|
||
was heard but reverse/3rd-party audio was not. **Fix:** before re-sealing, the relay rewrites the
|
||
outgoing header's `seq` (bytes [8..9]) to the recipient's `peek_send_counter()`, so each
|
||
server→client direction is one contiguous monotonic counter and the nonce always matches (the
|
||
anti-replay window also stops seeing false replays from interleaved senders). Safe because the
|
||
jitter buffer orders by `timestamp`, not `seq` (`audio_engine.h`); `seq` exists only to carry the
|
||
AEAD counter. No wire-format/proto/ABI change. Regression test added in `tests/test_media_aead.cpp`
|
||
(`test_relay_interleaved_reseal`): two senders interleaved into one recipient all decrypt with the
|
||
fix, and the verbatim-seq path is asserted to fail. `ctest --preset m1-dev` — **18/18 green** (run
|
||
via PowerShell; Git Bash can't resolve the runtime DLLs. `vad_ptt_devices` is timing-flaky over
|
||
loopback — passes on re-run — unrelated to this fix). **Latent, separate:** the secondary
|
||
"3rd joiner sometimes can't see other users" report is a control-plane (TCP snapshot/UserEvent)
|
||
issue, not this AEAD bug — re-verify after live testing before investigating. Also still latent:
|
||
the 16-bit `seq` wraps after 65536 frames per direction (faster on a busy relay) with no ROC, so
|
||
the implicit counter desyncs on long continuous sessions (`crypto.cpp` open() TODO).
|
||
- **Done:** **Fixed "randomly bumped to Lobby" in the Windows client — actors were excluded
|
||
from their own state-change broadcasts** (2026-06-17, reported live: a connected client would
|
||
intermittently snap from its joined channel back to Lobby in the UI). Root cause was a design
|
||
inconsistency, not a disconnect: the server delivered self-initiated state changes (channel
|
||
join/leave, stream announce/stop) only as a private `*Result` to the actor and broadcast the
|
||
authoritative `UserEvent::UPDATED` to *everyone else*. The core never applied the join result
|
||
to its `SessionModel`, so `vc_list_users()` kept self in the old channel; the Windows
|
||
`HandleUserUpdated` rebuilds `_currentChannelId` from `vc_list_users()` on **any** user's
|
||
UPDATED event, so the next unrelated event (someone joining, announcing/stopping a stream,
|
||
being muted) surfaced the stale self-channel → "bumped to Lobby." Flaky because it depended on
|
||
other users' activity. **Fix (broad, per the actor-sees-own-changes principle):** the server
|
||
now broadcasts these `UserEvent::UPDATED`s to **all** clients including the actor
|
||
(`server/src/conn_session.cpp`, `broadcast(…, /*exclude*/ 0)`), and text fan-out now includes
|
||
the sender (`resolve_text_targets`), so every client converges via one authoritative path.
|
||
The `*Result` is now purely ack/correlation/actor-private payload; response handlers no longer
|
||
mutate the local model. Windows client drops its optimistic text echo (the relay comes back)
|
||
and renders the sender's own message via `HandleTextMessage`. Documented the
|
||
response-vs-broadcast contract in `docs/protocol.md` §6. Registry-level admin broadcasts
|
||
(move/mute/kick/channel CRUD) already used `exclude=0` and were correct. `ctest --test-dir
|
||
build/m1-dev` — **18/18 green** (PowerShell); Windows `VoiceCat.App` builds 0 warnings.
|
||
~~**Latent, not fixed:** the client never sends the `Ping` keepalive that `docs/protocol.md` §7
|
||
describes (only the server answers pings) — unrelated to this bug, noted for later.~~
|
||
**Resolved** (2026-06-18): full keepalive/timeout/disconnect system implemented — see
|
||
"disconnect, timeout & keepalive" entry below.
|
||
- **Done:** **Fixed a *second* silent-playback bug — the playout clock free-ran and drifted off
|
||
the stream** (2026-06-17, reported live: both `vccli` and the Windows client showed `talking=1/0`
|
||
correctly on VAD/PTT, mic + screen-share were recognized by peers, but nothing was audible).
|
||
Root cause: `RemoteStream::playout_ts` was only ever seeded to `0` and then advanced one Opus
|
||
frame per playback callback **via the PLC path too** (`core/src/audio/audio_engine.cpp`
|
||
`on_playback`), so it free-ran at ~1× wall-clock regardless of whether the sender was
|
||
transmitting. The sender's frame timestamps only advance while it actually sends (the VAD/PTT
|
||
gate in `core/src/core/client.cpp` returns before `ls.timestamp += samples`). Across a late join
|
||
or any VAD/PTT silence gap the two clocks diverged without bound; once past the jitter buffer's
|
||
500 ms late-drop window, every real frame was dropped-as-late (clock ahead) or never-due (clock
|
||
behind) → permanent silence, while the talk indicator (driven by `push_recv_frame`, independent
|
||
of the jitter buffer) stayed lit. The M3 E2E test missed it because clients there talked
|
||
continuously right after joining, keeping the clocks aligned. **Fix:** `JitterBuffer` gained
|
||
`peek_front_ts()` (try-lock, RT-safe); `on_playback` now seeds/re-syncs `playout_ts` to the
|
||
earliest buffered frame on the first frame and whenever it has drifted past ±200/500 ms
|
||
(`kResyncAheadSamples`/`kResyncBehindSamples`), which both seeds startup and recovers after every
|
||
silence gap. New regression test `test_playout_resync` (`tests/test_vad_ptt_devices.cpp`):
|
||
free-runs the clock ~2 s past the drop window, pushes a `ts=0` frame, asserts audible output —
|
||
verified to fail (energy=0) with the fix disabled, pass (energy≈15M) with it. `ctest --test-dir
|
||
build/m1-dev` — **14/14 green** (run via PowerShell; Git Bash exec gotcha for these binaries, see
|
||
`docs/building.md`). **Not yet confirmed audible by ear** — pending the user re-running their
|
||
live test.
|
||
- **Done:** **Fixed silent-playback bug in `AudioEngine::on_playback`** (2026-06-16, found via
|
||
live manual test: two `vccli --voice` clients, control-plane events and VAD all correct, but
|
||
zero audible output). Root cause: `opus_decode()`'s `max_samples` was being passed the
|
||
*hardware playback callback's* frame count (miniaudio's own choice, frequently smaller than
|
||
one Opus frame — e.g. ~480 samples on default low-latency WASAPI periods), instead of the
|
||
decoder's fixed frame size (960 @ 20ms/48kHz). Since the real packet almost always decodes to
|
||
more samples than that, `opus_decode` returned `OPUS_BUFFER_TOO_SMALL` on nearly every
|
||
callback — frames were correctly received/decrypted/jitter-buffered, just never decoded into
|
||
audible PCM. `mix_for_test()`'s white-box test masked this because it always called
|
||
`on_playback` with `frames == frame_samples`, the one case where the bug is invisible.
|
||
Fix: `RemoteStream` (`core/src/audio/audio_engine.h`) gained a small ring buffer
|
||
(`init_ring`/`push_ring`/`pop_ring`) that decouples decode cadence from playback-callback
|
||
cadence — `on_playback` (`core/src/audio/audio_engine.cpp`) now tops the ring up by decoding
|
||
whole Opus frames (`decoder.frame_samples()`, never the hardware `frames`) and drains exactly
|
||
`frames` samples-per-channel from it each callback, silence-padding (PLC) on underrun. Also
|
||
fixes a latent `playout_ts` bug: it now advances by the actual decoded sample count per Opus
|
||
frame, not by the hardware callback's (unrelated) frame count, which was the wrong unit for
|
||
jitter-buffer timestamp comparisons. `ctest --test-dir build/m1-dev` — 12/12 green (run via
|
||
PowerShell; Git Bash exec gotcha for these binaries, see `docs/building.md`). **Not yet
|
||
confirmed audible by ear** — pending the user re-running their live two-`vccli` test.
|
||
- **Done:** **Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI
|
||
loopback** ✓ complete (2026-06-16). Closes all three items M3 explicitly carried forward as
|
||
out of scope (see the dated section below for the full file-by-file change list).
|
||
`ctest --test-dir build/m1-dev` — **12/12 tests** green (3 consecutive full-suite runs),
|
||
including the new `test_vad_ptt_devices` (real `vc_client`s against a real server, plus a
|
||
white-box `AudioEngine` stereo-mix check — same ABI-level-coverage lesson as M2/M3).
|
||
Manually verified live: `vccli --list-devices` against real hardware, and
|
||
`vccli --voice --input-mode vad` connecting/streaming without incident.
|
||
**Still explicitly out of scope** (carried forward, not silently dropped):
|
||
- Real `webrtc-audio-processing`/AEC — no working Windows/MSVC build upstream; v1 ships a
|
||
lightweight energy/RMS VAD instead (see docs/roadmap.md §2, docs/voice.md §8/§11). There is
|
||
**no AEC, NS, or AGC implementation at all**, not just a deferred VAD.
|
||
`vc_set_remote_stream(..., noise_reduction)`'s per-stream NS toggle is unaffected by this
|
||
pass and stays exactly as inert as it was after M3 (`ApmPassthrough`, no PCM modification).
|
||
- macOS/iOS `SCREEN_AUDIO` capture (ScreenCaptureKit / ReplayKit) — this pass is
|
||
Windows-only for real loopback capture; other platforms keep `vc_test_inject_capture` as
|
||
the only way to feed `SCREEN_AUDIO`.
|
||
- Process-specific WASAPI loopback — miniaudio's loopback mode captures the whole render
|
||
endpoint (including this app's own incoming voice mix), not a single process.
|
||
- The pre-existing RT-thread rule violation in `on_capture_frame`/`AudioEngine::on_capture`
|
||
(mutex lock, heap allocation, blocking `sendto` on the miniaudio real-time callback
|
||
thread) — predates this work, documented but not fixed; fixing it needs the lock-free
|
||
ring-buffer hand-off `docs/architecture.md §3` specifies, a separate, larger refactor.
|
||
- **Done:** **M4 — Windows WinForms C# client** ✓ complete (2026-06-17). Full details in the
|
||
M4 section below. `ctest --preset m1-dev` — **14/14 tests** green. `dotnet build` — 0
|
||
warnings/errors across all three C# projects. Manually verified: saved servers, TOFU
|
||
first-connect dialog, channel tree, join, voice (VAD/PTT/always-on + sensitivity slider +
|
||
per-user gain/mute/NR), text chat (channel + private), device pickers, level meter.
|
||
- **Next:** macOS/iOS Swift client (M4 continued) and/or **M5** (admin UI, moderation, kick/
|
||
ban). The C ABI is complete and stable through M4 — both directions are unblocked. See
|
||
`docs/roadmap.md §M4–M5`.
|
||
|
||
---
|
||
|
||
## Milestones (see [docs/roadmap.md](docs/roadmap.md) for full detail)
|
||
|
||
- [x] **M0 — Scaffolding** ✓ complete
|
||
- [x] **M1 — Control plane** ✓ complete (2026-06-15)
|
||
- [x] **M2 — Voice, single stream** ✓ complete (2026-06-16)
|
||
- [x] **M3 — Multi-stream & per-channel tuning** ✓ complete (2026-06-16)
|
||
- [x] **M4 — Native clients** — Windows WinForms ✓ (2026-06-17); macOS/iOS Swift pending
|
||
- [~] **M5 — Moderation, polish, beyond** (perms, bans, DRED; then file transfer, E2EE, …)
|
||
|
||
---
|
||
|
||
## M0 — Scaffolding ✓ (completed)
|
||
|
||
- [x] Repo layout (`core/ server/ tools/ clients/ tests/`), CMake + presets, vcpkg manifest.
|
||
- [x] C ABI header `core/include/voicecat.h` (full surface, stubbed).
|
||
- [x] Protocol source-of-truth `core/proto/voicecat.proto` (matches docs/protocol.md).
|
||
- [x] Core stubs for all six subsystems (net/crypto/codec/protocol/session/audio) + `vc_client`.
|
||
- [x] `voicecat-server` (arg parsing, config, stub run) and `vccli` (drives the C ABI).
|
||
- [x] CTest **smoke test** asserting the C ABI contract (not just "it compiles").
|
||
- [x] `.gitattributes` (LF), `.gitignore`, `.clang-format`, onboarding docs.
|
||
- **Verified:** `cmake --preset dev && cmake --build --preset dev && ctest --preset dev` → green.
|
||
|
||
---
|
||
|
||
## M1 — Control plane ✓ (completed 2026-06-15)
|
||
|
||
**Exit criterion:** ✓ `test_m1_integration` — two clients authenticate over TLS 1.3 (guest
|
||
+ Argon2id password), exchange channel and private text messages. Passes in ~1 s.
|
||
|
||
- [x] vcpkg baseline + `m1-dev` preset; `find_package` for protobuf/mbedTLS/libsodium/asio/sqlite3.
|
||
- [x] `FrameCodec` feed + emit; `encode_envelope` / `decode_envelope`.
|
||
- [x] Asio TCP acceptor + `TcpServerConn` (TLS path: blocking handshake thread + `tls_read_loop`).
|
||
- [x] `TlsContext` (mbedTLS 1.3, server cert/identity, ECDSA-P256 self-signed, TOFU on client).
|
||
- [x] `WorkerPool` (3 threads, used for Argon2id).
|
||
- [x] `Database` — SQLite, Argon2id via libsodium, `create_account` / `authenticate` / bootstrap admin.
|
||
- [x] `voicecat-admin` — account add/reset/del/list against live DB file.
|
||
- [x] `ServerIdentityManager` — generate/persist Ed25519 key + cert; fingerprint display.
|
||
- [x] `ConnSession` — WaitingHello → WaitingAuth → Authenticated state machine; full protocol relay.
|
||
- [x] `SessionRegistry` — channel tree, user map, broadcast, text routing.
|
||
- [x] `vc_client` (`client.cpp`) — full M1 C ABI: connect/TLS/ClientHello/AuthRequest/text/disconnect.
|
||
- [x] `Server::run()` — io_context, acceptor, worker pool, signal handling, `on_ready` callback.
|
||
- [x] `test_m1_integration` — M1 exit criterion. Verified green 2026-06-15.
|
||
|
||
**Key bug fixed:** double-framing in `ConnSession::send_envelope` — `encode_envelope` was
|
||
pre-framing the protobuf, then `TcpServerConn::send_frame` re-framed it. Fixed by serializing
|
||
raw protobuf bytes directly and letting `send_frame` add the single `[4-byte len]` prefix.
|
||
|
||
---
|
||
|
||
---
|
||
|
||
## M2 — Voice, single stream ✓ (completed 2026-06-16)
|
||
|
||
**Exit criterion:** ✓ `test_m2_voice` — two headless clients authenticate over TLS, bind UDP,
|
||
announce a MIC stream, send 50 encrypted Opus frames; server SFU relay re-encrypts + forwards
|
||
to the second client; B receives ≥ 25 frames and all decrypt correctly. Passes in ~4 s.
|
||
|
||
- [x] `m2-dev` preset (inherits `vcpkg-base`, binaryDir `build/m2-dev`); `m1-dev` also builds all M2 code.
|
||
- [x] `core/CMakeLists.txt` — `find_package(Opus)`, `find_path(MINIAUDIO_INCLUDE_DIR)`.
|
||
- [x] `core/src/net/voice_frame.h` — 14-byte UDP header (type/flags/codec/ssrc/seq/ts), serialize/parse, `make_udp_binding_packet`.
|
||
- [x] `SodiumMediaCrypto` — ChaCha20-Poly1305 AEAD; counter-nonce; 64-bit sliding-window anti-replay; `derive_send/recv` from TLS RFC 5705 exporter.
|
||
- [x] `OpusEncoder` / `OpusDecoder` — libopus 1.6, FEC, DTX, PLC (free; nullptr → decoder extrapolates).
|
||
- [x] `UdpMediaChannel` — async UDP socket (asio); thread-safe `send_to`; async recv loop.
|
||
- [x] `JitterBuffer` — per-ssrc, EWMA jitter estimation, adaptive depth 20–200 ms, late-drop at 500 ms.
|
||
- [x] `AudioEngine` — miniaudio capture+playback; `inject_capture()` bypass for headless tests; per-ssrc RemoteStream with OpusDecoder + JitterBuffer.
|
||
- [x] `ApmProcessor` — `ApmPassthrough` stub (VAD always open); WebRTC APM deferred until M3.
|
||
- [x] `on_tls_ready` callback in `TcpChannelCallbacks` — server derives and stores media AEAD keys immediately after TLS handshake.
|
||
- [x] `ConnSession` M2 — `udp_token` generated at construction; included in `AuthResult`; `handle_udp_binding` (verifies token, TCP ack); `handle_stream_announce` (assigns SSRC via registry); `udp_media_port` in `ServerHello`.
|
||
- [x] `SessionRegistry` M2 — `register_udp_token`, `find_by_udp_token`, `register_udp_endpoint`, `find_by_udp_endpoint`, `assign_ssrc`, `find_channel_sessions`, `user_channel`.
|
||
- [x] `MediaRelay` — SFU UDP relay; `kFrameUdpBinding` → endpoint binding; `kFrameVoice` → decrypt/re-encrypt/forward to channel members.
|
||
- [x] `Server::run()` — creates and binds `MediaRelay`; passes media port to `ConnSession`; wires `on_tls_ready` to derive per-connection media AEAD keys.
|
||
- [x] `test_voice_frame` — header round-trip, big-endian layout, binding packet format.
|
||
- [x] `test_media_aead` — seal/open round-trip, anti-replay, tamper detection, multi-packet sequence.
|
||
- [x] `test_opus_codec` — encode/decode round-trip energy check (within 3 dB), PLC, frame-samples helper.
|
||
- [x] `test_m2_voice` — M2 exit criterion (raw-socket harness). Verified green 2026-06-16.
|
||
|
||
**Follow-up (same day):** the above made `test_m2_voice` pass, but `vc_client`'s public voice
|
||
methods were still stubs — the *actual* M2 exit criterion ("two vccli/early-GUI clients talk")
|
||
wasn't met. Closed the gap:
|
||
|
||
- [x] `core/src/core/client.cpp` — real `stream_start`/`stream_stop`/`set_self_mute`/
|
||
`set_remote_stream`; UDP-binding handshake (`start_udp_binding`/`handle_udp_binding_ack`/
|
||
`finish_udp_binding`); media key derivation from `tls_` (RFC 5705 exporter); `run_udp_recv`
|
||
(AEAD-open → `JitterBuffer::Frame` → `audio_engine_.push_recv_frame`); `on_capture_frame`
|
||
(encode → seal → `sendto`); `sync_remote_streams` (diffs a `User` proto's `streams` against
|
||
`remote_streams_`, wiring up `OpusDecoder`s and emitting `STREAM_STARTED`/`STOPPED`).
|
||
`set_input_device`/`set_input_mode`/`set_push_to_talk`/`list_devices` remain
|
||
`VC_ERR_NOT_IMPLEMENTED` — no device-enumeration backend yet; scoped to M3 (VAD/PTT).
|
||
- [x] `core/src/session/session.cpp/h` — `SessionModel::find_user`, `find_user_by_ssrc`,
|
||
`Stream{stream_id, ssrc, kind, label, sample_rate, frame_ms}`.
|
||
- [x] `server/src/conn_session.cpp/h` — `handle_stream_announce`/`handle_stream_stop` now
|
||
broadcast via `SessionRegistry::set_user_stream`/`clear_user_stream` → `UserEvent::UPDATED`.
|
||
- [x] `server/src/session_registry.cpp/h` — `set_user_stream`/`clear_user_stream` (mutate a
|
||
user's `StreamInfo` list, return the updated `User` proto for broadcast).
|
||
- [x] `tests/test_voice_client_abi.cpp` — drives two real `vc_client` instances through
|
||
`vc_connect`/`vc_authenticate_guest`/`vc_stream_start`/`vc_stream_stop`; asserts client B
|
||
observes client A's `STREAM_STARTED`/`STOPPED` events. Verified green 2026-06-16.
|
||
- [x] `tools/vccli/src/main.cpp` — argv parsing (`--host/--port/--nick/--channel/--voice/
|
||
--mute/--text`); `--voice` starts a MIC stream and blocks on SIGINT, printing `on_event`
|
||
callbacks live (unbuffered stdout — MinGW/MSVCRT treat `_IOLBF` as full buffering for
|
||
non-console streams). Dropped the originally-planned `--voice-loopback` and the
|
||
`tx=N rx=M lost=K jitter=J` stats line: `voicecat.h` exposes no PCM-injection hook or
|
||
jitter/loss stats getter publicly, only `on_event` + `on_level` (RMS). Manually verified:
|
||
two `vccli --voice` instances see each other's stream start in real time.
|
||
|
||
---
|
||
|
||
## M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)
|
||
|
||
**Exit criterion:** ✓ `test_m3_multistream` — a real `vc_client` (A) runs two concurrent local
|
||
streams (MIC + SCREEN_AUDIO) with distinct stream ids; a second client (B) sees both as
|
||
separate `STREAM_STARTED` events and a `VC_EVENT_TALK_STATE` talking edge for A's MIC stream;
|
||
B independently sets gain/mute/noise-reduction on each of A's streams without one call
|
||
affecting the other; A then joins "Music Room" (channel 2: stereo/128kbps/`OPUS_AUDIO`/no
|
||
DTX) and announces a fresh MIC stream there, while B stays in "Lobby" (channel 1: mono/24kbps/
|
||
`OPUS_VOIP`/DTX on) — `vc_get_stream_audio_config` shows their effective Opus config differs
|
||
exactly as the server enforces per channel. Passes in ~2.4s; verified across 8 consecutive
|
||
standalone runs + 3 consecutive full-suite runs with no flakes.
|
||
|
||
Exploration before implementing turned up several bugs/gaps where the wire format already
|
||
supported this milestone but the client/server logic didn't — these were fixed as part of M3,
|
||
not treated as pre-existing-and-out-of-scope:
|
||
|
||
- [x] **Server `stream_id` bug** — `handle_stream_announce` always wrote `stream_id=1`, so a
|
||
second stream from the same user silently overwrote the first in
|
||
`SessionRegistry::set_user_stream`'s replace-by-id logic. Fixed with a per-session counter
|
||
(`ConnSession::next_stream_id_`) + `announced_stream_ids_` (also now validated in
|
||
`handle_stream_stop`, rejecting stops for ids the session never announced).
|
||
- [x] **Per-channel `AudioConfig` was modeled but never populated/enforced.**
|
||
`SessionRegistry::init_default_channels()` now seeds Lobby (id=1: mono, 24kbps, `OPUS_VOIP`,
|
||
FEC+DTX on) and a new "Music Room" (id=2: stereo, 128kbps, `OPUS_AUDIO`, FEC+DTX off) with
|
||
real `AudioConfig`s; new `SessionRegistry::channel_audio_config(channel_id)` accessor (there
|
||
was no per-id channel getter before, only `channel_snapshot()`). `handle_stream_announce`
|
||
now treats the channel's config as authoritative (mode/frame_ms/application/fec/dtx/
|
||
complexity), clamping (not overriding) `bitrate_bps` to the channel's ceiling.
|
||
- [x] **Client silently dropped `mode`/`dtx`/`complexity`/`application` from `effective_audio`**
|
||
even for the single M2 stream — `handle_stream_announce_result` and `sync_remote_streams`
|
||
only copied `sample_rate`/`bitrate_bps`/`frame_ms`/`fec` into `OpusParams`. New shared
|
||
`opus_params_from_audio_config()` helper (`client.cpp`) fixes both the send and receive
|
||
paths.
|
||
- [x] `core/src/codec/opus_codec.h/.cpp` — new `OpusApplication` enum + `OpusParams::application`
|
||
field; `OpusEncoder::init` now honors it instead of hardcoding `OPUS_APPLICATION_VOIP`.
|
||
- [x] `core/src/session/session.h/.cpp` — `Stream` struct extended with the full `AudioConfig`
|
||
(mode/bitrate_bps/application/fec/expected_packet_loss/dtx/complexity), not just
|
||
sample_rate/frame_ms; `copy_streams()` now copies all of it.
|
||
- [x] `core/src/core/client.h/.cpp` — local-stream state is now a `std::unordered_map<int,
|
||
LocalStream>` keyed by `vc_stream_kind` (one active stream per kind — MIC/SCREEN_AUDIO/
|
||
AUX_DEVICE are each singletons for a client), replacing the M2 single-stream fields.
|
||
`StreamAnnounce`/`StreamAnnounceResult` round-trips are now correlated by `request_id`
|
||
(already round-tripped on the wire; just wasn't read) via `pending_announce_kind_`, so
|
||
multiple concurrent announces from one client resolve to the right `LocalStream`.
|
||
`on_capture_frame` takes a `kind` and `channels` parameter; for mono capture (`channels==1`)
|
||
on a stereo channel it upmixes L=R, and for real stereo capture (`channels==2`, the
|
||
`SCREEN_AUDIO` loopback path on a stereo channel) it encodes directly with no upmix. `vc_set_self_mute`'s `mic_muted` only
|
||
gates the `MIC` kind — a concurrent `SCREEN_AUDIO` share keeps playing while muted.
|
||
`set_remote_stream` now actually wires `noise_reduction` through (previously parsed and
|
||
discarded). New `run_talk_timer()` (a small dedicated thread, started alongside the UDP
|
||
media path, never the miniaudio callback thread) polls both remote talk-state edges
|
||
(`AudioEngine::poll_talk_transitions()`) and local capture-activity edges, emitting
|
||
`VC_EVENT_TALK_STATE`.
|
||
- [x] **Fixed a thread-join race in `teardown_voice()`** — it's called both from `run_io()`'s
|
||
own cleanup and from `disconnect()`, on different threads; without serialization both could
|
||
see `udp_thread_`/`talk_timer_thread_` as `joinable()` simultaneously and race to `join()`
|
||
the same `std::thread` (UB; surfaced as an intermittent `std::system_error: No such process`
|
||
under `ctest`). Added a `teardown_mu_` guard around the whole function. This pre-existed for
|
||
`udp_thread_` alone (likely the same root cause as the `test_m1_integration`/`test_m2_voice`
|
||
cleanup-path flake noted in the M2 section above) — adding `talk_timer_thread_`'s join just
|
||
made it surface more often, so it was fixed properly here rather than carried forward again.
|
||
- [x] `core/src/audio/audio_engine.h/.cpp` — `CaptureCallback` gained a `kind` parameter
|
||
(the real miniaudio capture device is always tagged `kind=0`/MIC; a second concurrent local
|
||
stream is fed via its own `inject_capture(kind, ...)` ring buffer — `inject_taps_`, keyed by
|
||
kind — since there is only one real hardware capture device in M3). Fixed a buffer-sizing
|
||
bug in `on_playback`'s per-stream decode (`opus_decode`'s `frame_size` parameter is
|
||
samples-*per-channel*, not total samples — the old code passed `frames * params_.channels`,
|
||
which would have overflowed the decode buffer for any stereo stream). Stereo decoder output
|
||
is downmixed (avg L/R) into the engine's mono mix accumulator immediately after decode.
|
||
`RemoteStream` gained `recv_ns`/`noise_reduction_enabled` (lazy `ApmProcessor` instantiation
|
||
— freed on disable, so no separate instance cap is needed per the roadmap's guidance) and
|
||
`last_voice_ms`/`talking` (talk-indicator edge state, updated in `push_recv_frame`); new
|
||
`set_stream_noise_reduction()` and `poll_talk_transitions()`. Note: until `VOICECAT_HAS_APM`
|
||
is wired to a real WebRTC APM build, the NS toggle is plumbed end-to-end but behaviorally a
|
||
passthrough no-op (`ApmPassthrough` doesn't touch PCM) — same situation send-side APM has
|
||
been in since M2; M3's job was the plumbing, not the DSP backend.
|
||
- [x] **New C ABI surface** (`core/include/voicecat.h`, additive only):
|
||
`vc_audio_config` struct + `vc_get_stream_audio_config(c, user_id, stream_id, out)` — the
|
||
effective Opus config for a stream you own or a peer's, reading from the (now richer)
|
||
`LocalStream`/`session::Stream`. `vc_test_inject_capture(c, stream_id, pcm, samples)` —
|
||
clearly-marked **test-only**, forwards to `AudioEngine::inject_capture`, so
|
||
`test_m3_multistream` can drive two concurrent synthetic-audio streams through the real ABI
|
||
without a microphone.
|
||
- [x] `tests/test_m3_multistream.cpp` — the M3 exit criterion (ABI-level, mirrors
|
||
`test_voice_client_abi.cpp`'s approach per the M2 lesson). Registered in `tests/CMakeLists.txt`.
|
||
|
||
**Explicitly out of scope for this pass** (confirmed with the user before implementing):
|
||
- `vc_set_input_device`/`vc_set_input_mode`/`vc_set_push_to_talk`/`vc_list_devices` (device
|
||
enumeration + VAD/PTT input gate) — still `VC_ERR_NOT_IMPLEMENTED`. These were mentioned as
|
||
"scoped to M3" in the M2 follow-up notes above, but docs/roadmap.md's M3 bullets never
|
||
actually listed them — deferred again, now tracked explicitly rather than implicitly.
|
||
- Real WASAPI desktop-audio loopback capture for `SCREEN_AUDIO` — the engine now supports
|
||
feeding a second concurrent local stream via `inject_capture`, but only synthetic PCM is
|
||
wired up; a real loopback capture device is a follow-up.
|
||
- True stereo *playback output* — `AudioEngine`'s mixer/output device stays mono; stereo
|
||
streams are downmixed after decode (see above). The Opus wire format itself is fully
|
||
stereo-correct.
|
||
|
||
---
|
||
|
||
## Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI loopback ✓ (completed 2026-06-16)
|
||
|
||
Closes all three items the M3 section above explicitly carried forward as out of scope.
|
||
|
||
**Exit verification:** `ctest --test-dir build/m1-dev` — **12/12 tests** green (3 consecutive
|
||
full-suite runs), including the new `test_vad_ptt_devices` (device enumeration + VAD/PTT gate
|
||
through real `vc_client`s against a real server, plus a white-box `AudioEngine` stereo-mix
|
||
check — no audio hardware needed for that last part). Also verified 5 consecutive standalone
|
||
runs of the new test alone, no flakes. Manually verified live on Windows: `vccli
|
||
--list-devices` against real hardware (3 input / 4 output devices, correct `is_default`
|
||
flags), and `vccli --voice --input-mode vad` connecting + streaming without incident.
|
||
|
||
- [x] **Device enumeration** (`vc_list_devices`) — `AudioEngine::enumerate_devices(bool
|
||
capture)` (static, works without a running engine — inits a throwaway `ma_context` via
|
||
`ma_context_get_devices`). `device_id`/`vc_device.id` is an opaque hex-encoded raw
|
||
`ma_device_id` (not the device name — names aren't guaranteed unique); documented as an
|
||
internal contract callers must round-trip, never construct by hand. `vc_client::list_devices`
|
||
works in any connection state (no `VC_STATE_CONNECTED` gate) since device pickers need to
|
||
populate pre-connect. `vc_free_device_list` is now a real free (was a no-op stub).
|
||
- [x] **Input device selection** (`vc_set_input_device`) — stores the device id on the
|
||
targeted `LocalStream` (new field); for the MIC stream, if the engine is already running,
|
||
restarts it (`stop()` + `ensure_audio_running()`) to pick up the new device. Simplified: it
|
||
restarts unconditionally rather than trying to detect whether the device id actually
|
||
changed (`AudioEngine` has no getter for "current device").
|
||
- [x] **VAD/PTT input gate** (`vc_set_input_mode`, `vc_set_push_to_talk`) — new
|
||
`EnergyVadProcessor` (`core/src/audio/apm_processor.cpp`) implementing the existing
|
||
`ApmProcessor` interface: energy/RMS threshold (default ~0.025 normalized) + hang-time
|
||
(default 300 ms, matching `kTalkHangoverMs`). New factory `ApmProcessor::create_vad()`,
|
||
kept separate from `create()` (which recv-side per-stream NS still uses, unaffected by this
|
||
pass). `vc_client` gained `current_input_mode_`/`ptt_active_`/`mic_vad_`; the gate is
|
||
inserted in `on_capture_frame`, **MIC-only** — `SCREEN_AUDIO`/`AUX_DEVICE` always bypass it
|
||
(gating a desktop-audio share on the user's own voice activity would silently drop shared
|
||
music/video audio). `last_capture_ms` (drives the talk indicator) is now updated *after* the
|
||
gate check, not before, so a VAD/PTT-closed frame never shows as "talking". `mic_vad_` is
|
||
constructed once the MIC stream's `StreamAnnounceResult` lands (on `io_thread_`), not lazily
|
||
inside the capture path.
|
||
- [x] **True stereo playback** — `AudioParams::channels` split into `capture_channels` (stays
|
||
1) and `playback_channels` (now 2, unconditionally). `AudioEngine::on_playback` no longer
|
||
downmixes decoded stereo streams to mono before mixing — stereo decode output is mixed
|
||
directly (L→L, R→R); mono decode output is upmixed (duplicated into both channels). Falls
|
||
back to a 1-channel playback device once if the 2-channel `ma_device_init` fails (unusual
|
||
hardware). New test-only `AudioEngine::mix_for_test()` exposes the mixer for white-box
|
||
testing without a real `ma_device`.
|
||
- [x] **WASAPI loopback capture for `SCREEN_AUDIO`** — new `VOICECAT_HAS_LOOPBACK` macro
|
||
(`core/CMakeLists.txt`, Windows-only). `AudioEngine` gained a separate `loopback_device_`
|
||
(own lifecycle, decoupled from the mic capture/playback devices) with
|
||
`start_loopback_capture()`/`stop_loopback_capture()`, using miniaudio's
|
||
`ma_device_type_loopback` against the default render endpoint. Its callback feeds
|
||
`capture_cb_` directly (same pattern as the real mic capture device), **not** through
|
||
`inject_capture()`'s test-only ring. Wired into `vc_client::handle_stream_announce_result`
|
||
(start, alongside `ensure_audio_running()`) and `stream_stop` (stop) for
|
||
`VC_STREAM_SCREEN_AUDIO`. Non-Windows builds keep `vc_test_inject_capture` as the only way to
|
||
feed `SCREEN_AUDIO`.
|
||
- [x] `tools/vccli/src/main.cpp` — new flags `--list-devices`, `--input-device`,
|
||
`--input-mode vad|ptt`, `--share-screen-audio`; while `--voice` is running, a background
|
||
stdin-reader thread accepts `ptt on`/`ptt off`/`mode vad`/`mode ptt` (the most portable way
|
||
to drive PTT interactively from a headless CLI — no SIGUSR1 equivalent on Windows). Also
|
||
prints `VC_EVENT_TALK_STATE`. Known minor caveat: on Windows the stdin-reader thread is
|
||
detached (not joined) on exit, since `std::getline` can't be interrupted from another thread
|
||
— a `vc_client*` use-after-free is theoretically possible if a command line arrives in the
|
||
brief window between teardown and process exit; acceptable for a headless test/dev tool.
|
||
- [x] `tests/test_vad_ptt_devices.cpp` — new test covering all four items above; registered in
|
||
`tests/CMakeLists.txt`. `tests/test_smoke.cpp`'s device-list assertion is now conditional on
|
||
`VOICECAT_HAS_AUDIO` (was a hard `VC_ERR_NOT_IMPLEMENTED` assertion) — `VC_OK` only, never
|
||
`count > 0` (a headless CI build agent may legitimately report zero audio devices).
|
||
|
||
**Still explicitly out of scope** (carried forward, not silently dropped):
|
||
- Real `webrtc-audio-processing`/AEC — no working Windows/MSVC build upstream (see
|
||
docs/roadmap.md §2's superseding decision-log entry). There is **no AEC, NS, or AGC
|
||
implementation at all**, not just a deferred VAD. The per-stream NS toggle
|
||
(`vc_set_remote_stream(..., noise_reduction)`) is unaffected by this pass and stays exactly
|
||
as inert as it was after M3 (`ApmPassthrough`, no PCM modification) — don't mistake this
|
||
pass for having fixed it.
|
||
- macOS/iOS `SCREEN_AUDIO` capture (ScreenCaptureKit / ReplayKit) — Windows-only loopback in
|
||
this pass.
|
||
- Process-specific WASAPI loopback — whole-device capture only; inherently captures this
|
||
app's own incoming voice mix along with everything else playing.
|
||
- The pre-existing RT-thread rule violation in `on_capture_frame`/`AudioEngine::on_capture`
|
||
(mutex lock, heap allocation for the stereo-upmix path, blocking `sendto`, all on the
|
||
miniaudio real-time callback thread) — predates this work (was already present in M2/M3);
|
||
documented here explicitly rather than silently carried forward again. Fixing it properly
|
||
needs the lock-free ring-buffer hand-off `docs/architecture.md §3` specifies — a separate,
|
||
larger refactor, out of scope for this pass.
|
||
|
||
---
|
||
|
||
## M4 — Windows WinForms C# client ✓ (completed 2026-06-17)
|
||
|
||
**Exit criterion:** ✓ `ctest --preset m1-dev` — **14/14 tests** green (existing 12 + 2 new
|
||
C++ tests: `test_channel_user_list_abi`, `test_tofu_flow`). `dotnet build` — 0 warnings/errors.
|
||
Manually verified: connect, TOFU first-connect dialog, channel tree, join, voice, text, device
|
||
pickers, level meter. Accessibility: explicit `AccessibleName`/`AccessibleDescription` on every
|
||
control, `&` mnemonics on every button, activity-log `ListBox` as screen-reader record.
|
||
|
||
**New C++ ABI surface** (all additive, backward-compatible):
|
||
- `vc_list_channels` / `vc_list_users` / `vc_list_user_streams` — pull-based snapshot getters
|
||
for the channel tree + user list; `session_model_mu_` added for cross-thread safety.
|
||
`SessionModel::apply_snapshot`/`apply_channel_event` fixed to populate `parent_id`,
|
||
`password_protected`, `max_users` (were permanently zeroed despite the struct declaring them).
|
||
- `vc_join_channel(channel_id, password)` — join with optional password; server replies via
|
||
new `VC_EVENT_JOIN_RESULT`.
|
||
- `VC_EVENT_SERVER_IDENTITY` + `vc_confirm_server_identity(accept)` — TOFU gate that blocks
|
||
`io_thread_` until the UI approves or rejects. Pins the TLS leaf-cert SHA-256 fingerprint
|
||
(verifiable directly at handshake), **not** the declared Ed25519 value (see docs/security.md
|
||
§1.1 for why — the TLS cert and Ed25519 key are generated independently, no binding).
|
||
`vc_get_server_identity_display` exposes the Ed25519 fingerprint for human-readable display.
|
||
- `vc_config::tofu_store_path` — optional per-user pin file path; defaults to a relative
|
||
`"./voicecat_tofu_pins.txt"` so existing tests need no change.
|
||
- `TcpAcceptor` now dual-stacks (IPv6 + IPv4 fallback) so `localhost` → `::1` on Windows
|
||
connects correctly without forcing users to type `127.0.0.1`.
|
||
- `VC_INPUT_ALWAYS_ON = 2` in `vc_input_mode` — transmit unconditionally (no VAD gate).
|
||
- `vc_set_vad_threshold(float)` — live VAD threshold update; `EnergyVadProcessor` stores it
|
||
atomically so the RT capture path reads without a lock or allocation.
|
||
|
||
**New C++ tests:**
|
||
- `test_channel_user_list_abi` — snapshot getters, `parent_id`/`password_protected`/`max_users`
|
||
regression, per-user stream list, invalid-user-id error, double-free idempotency.
|
||
- `test_tofu_flow` — first-connect blocks until confirmed; reject doesn't persist; reconnect to
|
||
same identity reports `MATCHED`; rotated identity reports `MISMATCH`; `vc_confirm_*` with
|
||
nothing pending returns an error.
|
||
|
||
**`windows-client` CMake preset** — Release, `VOICECAT_BUILD_SHARED=ON`, static MinGW runtime
|
||
(`-static-libgcc -static-libstdc++ -static -lwinpthread`), no tools/tests. Outputs
|
||
`build/windows-client/bin/voicecat.dll` with zero MinGW DLL dependencies (only Windows system
|
||
DLLs remain — verified via `objdump -p`).
|
||
|
||
**C# solution** (`clients/windows/`, .NET 10 LTS `net10.0-windows`):
|
||
- `VoiceCat.Interop` — `[LibraryImport]` P/Invoke surface, `[UnmanagedCallersOnly]` callbacks,
|
||
`System.Threading.Channels.Channel<VoiceCatEvent>` event delivery drained by 30ms WinForms
|
||
Timer; `VoiceCatClientHandle : SafeHandle` guarantees `vc_client_destroy`.
|
||
- `VoiceCat.App` — WinForms UI:
|
||
- `ConnectDialog` — saved-server `ListBox`, Add/Remove/Edit; servers persisted to
|
||
`%AppData%\VoiceCat\servers.json`; passwords DPAPI-encrypted (`ProtectedData`, opt-in).
|
||
- `ServerIdentityDialog` — shown only on `FIRST_CONNECT`/`MISMATCH` (never `MATCHED`);
|
||
mismatch text and button ordering are starkly different ("WARNING" framing, Cancel default).
|
||
- `MainForm` — `TreeView` channel tree, `ListBox` user list, `RichTextBox` chat, scope
|
||
`ComboBox` (Channel/Private), activity-log `ListBox`, voice panel with mic toggle,
|
||
mute/deafen checkboxes, VAD/PTT/Always-On radio group, VAD sensitivity `TrackBar`
|
||
(1–100, hidden for non-VAD modes), device `ComboBox` + refresh, level `ProgressBar`.
|
||
- `PerUserTuningDialog` — real-time gain `TrackBar` + mute/NR checkboxes; applied to all
|
||
of a user's streams immediately (no OK/Cancel round-trip).
|
||
- `PttKeyCaptureDialog` — focus-scoped PTT key capture.
|
||
- `VoiceCat.Interop.Tests` — xunit smoke test: connect → TOFU → guest auth → list channels
|
||
purely via P/Invoke against a live `voicecat-server.exe`.
|
||
|
||
**Explicitly out of scope for this pass:**
|
||
- macOS/iOS Swift client — pending.
|
||
- Admin/moderation UI (kick/ban/permissions/account provisioning) — server-side dispatch for
|
||
these messages is M5's job; building the UI now would require building the server side too.
|
||
- PTT hotkey is **focus-scoped only** (works while VoiceCat window has focus). A system-wide
|
||
`WH_KEYBOARD_LL` hook would require escalated permissions and risk AV flagging — documented
|
||
limitation, not silently omitted.
|
||
- Receive-side noise reduction (`vc_set_remote_stream(..., noise_reduction)`) is end-to-end
|
||
plumbed but behaviorally a passthrough no-op (`ApmPassthrough`, no PCM modification) —
|
||
same as before M4. The per-user NR checkbox in `PerUserTuningDialog` is labeled accordingly.
|
||
|
||
---
|
||
|
||
## M5 — Moderation, polish, and beyond [~] (in progress 2026-06-17)
|
||
|
||
**Exit criterion:** four ABI-level tests green (`test_m5_permissions`,
|
||
`test_m5_kick_ban_move_mute`, `test_m5_admin_accounts`, `test_m5_channel_crud`);
|
||
`vccli` can drive all moderation/admin/channel operations against a live server.
|
||
|
||
- [x] **Server-side moderation & permissions:**
|
||
- `server/src/session_registry.h/.cpp` — per-session `Permissions`, permission helpers
|
||
(`can_kick`, `can_ban`, etc.), kick/ban/move/server-mute, channel CRUD, DB-backed channel
|
||
tree load/save, in-memory channel state.
|
||
- `server/src/conn_session.cpp` — M5 dispatch handlers, permission checks, channel-password
|
||
+ `max_users` enforcement, `UserEvent::UPDATED` broadcast on join/leave.
|
||
- `core/proto/voicecat.proto` — `ServerMuteRequest`, `UserEvent.reason`, `User.server_deafened`,
|
||
`ListAccountsResult`, `AccountEntry`.
|
||
- [x] **C ABI / client-side:**
|
||
- `core/include/voicecat.h` — `vc_permissions`, `vc_channel_info`, `vc_kick_user`,
|
||
`vc_ban_user`, `vc_set_permission`, `vc_set_server_mute`, `vc_move_user`,
|
||
`vc_create_channel`, `vc_edit_channel`, `vc_delete_channel`, `vc_create_account`,
|
||
`vc_reset_password`, `vc_delete_account`, `vc_list_accounts`, `vc_get_permissions`;
|
||
new events `VC_EVENT_GENERIC_RESULT` and `VC_EVENT_ACCOUNT_LIST`.
|
||
- `core/src/voicecat.cpp`, `core/src/core/client.h/.cpp` — implementations + server-mute/deafen
|
||
gating on the client.
|
||
- [x] **Database:** `server/src/db.h/.cpp` schema v2 (`channels`, `bans`), Argon2id accounts,
|
||
BLAKE2b channel passwords, migrations.
|
||
- [x] **Tests:** four new M5 tests registered in `tests/CMakeLists.txt`:
|
||
- `test_m5_permissions` — grant/revoke permissions, verify enforcement.
|
||
- `test_m5_kick_ban_move_mute` — kick, ban, move, server-mute/deafen.
|
||
- `test_m5_admin_accounts` — create/reset/delete/list accounts.
|
||
- `test_m5_channel_crud` — create/edit/delete channels, password + max_users enforcement.
|
||
- [x] **vccli** (`tools/vccli/src/main.cpp`) — all M5 operations exposed via flags; account auth
|
||
via `--username`/`--password`; async `VC_EVENT_GENERIC_RESULT`/`VC_EVENT_ACCOUNT_LIST` handling.
|
||
- [x] **Docs** kept in sync: `docs/protocol.md`, `docs/security.md`, `PROGRESS.md`.
|
||
|
||
**Key bug fixed:** `test_m5_channel_crud` failed because `SessionRegistry::create_channel`
|
||
broadcast `ChannelEvent::CREATED` from a moved-from `entry.proto` after
|
||
`channels_[id] = std::move(entry)`. Fixed by building the event before moving into the map.
|
||
|
||
**Still to do:**
|
||
- DRED/audio-quality polish.
|
||
- macOS/iOS Swift client (carried from M4).
|
||
|
||
---
|
||
|
||
## Decisions log
|
||
|
||
All architecture/scope decisions are settled and recorded in
|
||
[docs/roadmap.md §2 "Resolved decisions"](docs/roadmap.md) and reflected across `docs/`.
|
||
If you make a *new* decision, record it there and link it here.
|
||
|
||
---
|
||
|
||
## How to update this file
|
||
|
||
1. Check off tasks as you complete them; flip a milestone to `[x]` only when its **exit
|
||
criterion test** passes.
|
||
2. Keep the **"Where we left off / next action"** block at the top accurate — it's the first
|
||
thing the next agent reads.
|
||
3. When you start a milestone, copy its task list from `docs/roadmap.md` into a section here.
|