Files
voice-cat/PROGRESS.md
Talon 5f6c223526 feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as
out of scope:

- Device enumeration (vc_list_devices) + input device selection
  (vc_set_input_device), backed by AudioEngine::enumerate_devices() via
  miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded
  ma_device_id strings.
- VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk).
  webrtc-audio-processing (the originally-planned APM) has no working
  Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard
  abseil-cpp dependency), so VAD is a new lightweight, dependency-free
  energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor
  interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it.
- True stereo playback: AudioEngine's mixer and output device now carry
  stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded
  stereo streams to mono before mixing.
- Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via
  miniaudio's loopback device type), replacing test-only injection as the
  production capture path.

Also: vccli gains --list-devices, --input-device, --input-mode, and
--share-screen-audio flags, plus a stdin command loop (ptt on/off, mode
vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all
four items (ABI-level + a white-box AudioEngine stereo-mix check).

Docs updated to match: voice.md, roadmap.md (decision-log entry superseding
the original webrtc-audio-processing choice), tech-stack.md, README.md,
architecture.md, CLAUDE.md, PROGRESS.md.

Still explicitly out of scope, documented not silently dropped: real
webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS
SCREEN_AUDIO capture, process-specific loopback, and a pre-existing
RT-thread rule violation in the capture path that predates this work.

Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev
and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive
standalone runs; manually verified live (vccli --list-devices against real
hardware, vccli --voice --input-mode vad streaming without incident).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 16:11:52 +02:00

347 lines
25 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# PROGRESS — VoiceCat
Living status. **Update this file in the same commit as your work** so the next agent picks
up instantly. Newest status at the top.
- **Date convention:** ISO (YYYY-MM-DD).
- Statuses: `[ ]` not started · `[~]` in progress · `[x]` done.
---
## ▶ Where we left off / next action
- **Done:** **Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI
loopback** ✓ complete (2026-06-16). Closes all three items M3 explicitly carried forward as
out of scope (see the dated section below for the full file-by-file change list).
`ctest --test-dir build/m1-dev`**12/12 tests** green (3 consecutive full-suite runs),
including the new `test_vad_ptt_devices` (real `vc_client`s against a real server, plus a
white-box `AudioEngine` stereo-mix check — same ABI-level-coverage lesson as M2/M3).
Manually verified live: `vccli --list-devices` against real hardware, and
`vccli --voice --input-mode vad` connecting/streaming without incident.
**Still explicitly out of scope** (carried forward, not silently dropped):
- Real `webrtc-audio-processing`/AEC — no working Windows/MSVC build upstream; v1 ships a
lightweight energy/RMS VAD instead (see docs/roadmap.md §2, docs/voice.md §8/§11). There is
**no AEC, NS, or AGC implementation at all**, not just a deferred VAD.
`vc_set_remote_stream(..., noise_reduction)`'s per-stream NS toggle is unaffected by this
pass and stays exactly as inert as it was after M3 (`ApmPassthrough`, no PCM modification).
- macOS/iOS `SCREEN_AUDIO` capture (ScreenCaptureKit / ReplayKit) — this pass is
Windows-only for real loopback capture; other platforms keep `vc_test_inject_capture` as
the only way to feed `SCREEN_AUDIO`.
- Process-specific WASAPI loopback — miniaudio's loopback mode captures the whole render
endpoint (including this app's own incoming voice mix), not a single process.
- The pre-existing RT-thread rule violation in `on_capture_frame`/`AudioEngine::on_capture`
(mutex lock, heap allocation, blocking `sendto` on the miniaudio real-time callback
thread) — predates this work, documented but not fixed; fixing it needs the lock-free
ring-buffer hand-off `docs/architecture.md §3` specifies, a separate, larger refactor.
- **Next:** **M4 — native clients** (Windows C#, macOS/iOS Swift). See `docs/roadmap.md §M4`.
---
## Milestones (see [docs/roadmap.md](docs/roadmap.md) for full detail)
- [x] **M0 — Scaffolding** ✓ complete
- [x] **M1 — Control plane** ✓ complete (2026-06-15)
- [x] **M2 — Voice, single stream** ✓ complete (2026-06-16)
- [x] **M3 — Multi-stream & per-channel tuning** ✓ complete (2026-06-16)
- [ ] **M4 — Native clients** (Windows C#, macOS/iOS Swift) ← next
- [ ] **M5 — Moderation, polish, beyond** (perms, bans, DRED; then file transfer, E2EE, …)
---
## M0 — Scaffolding ✓ (completed)
- [x] Repo layout (`core/ server/ tools/ clients/ tests/`), CMake + presets, vcpkg manifest.
- [x] C ABI header `core/include/voicecat.h` (full surface, stubbed).
- [x] Protocol source-of-truth `core/proto/voicecat.proto` (matches docs/protocol.md).
- [x] Core stubs for all six subsystems (net/crypto/codec/protocol/session/audio) + `vc_client`.
- [x] `voicecat-server` (arg parsing, config, stub run) and `vccli` (drives the C ABI).
- [x] CTest **smoke test** asserting the C ABI contract (not just "it compiles").
- [x] `.gitattributes` (LF), `.gitignore`, `.clang-format`, onboarding docs.
- **Verified:** `cmake --preset dev && cmake --build --preset dev && ctest --preset dev` → green.
---
## M1 — Control plane ✓ (completed 2026-06-15)
**Exit criterion:**`test_m1_integration` — two clients authenticate over TLS 1.3 (guest
+ Argon2id password), exchange channel and private text messages. Passes in ~1 s.
- [x] vcpkg baseline + `m1-dev` preset; `find_package` for protobuf/mbedTLS/libsodium/asio/sqlite3.
- [x] `FrameCodec` feed + emit; `encode_envelope` / `decode_envelope`.
- [x] Asio TCP acceptor + `TcpServerConn` (TLS path: blocking handshake thread + `tls_read_loop`).
- [x] `TlsContext` (mbedTLS 1.3, server cert/identity, ECDSA-P256 self-signed, TOFU on client).
- [x] `WorkerPool` (3 threads, used for Argon2id).
- [x] `Database` — SQLite, Argon2id via libsodium, `create_account` / `authenticate` / bootstrap admin.
- [x] `voicecat-admin` — account add/reset/del/list against live DB file.
- [x] `ServerIdentityManager` — generate/persist Ed25519 key + cert; fingerprint display.
- [x] `ConnSession` — WaitingHello → WaitingAuth → Authenticated state machine; full protocol relay.
- [x] `SessionRegistry` — channel tree, user map, broadcast, text routing.
- [x] `vc_client` (`client.cpp`) — full M1 C ABI: connect/TLS/ClientHello/AuthRequest/text/disconnect.
- [x] `Server::run()` — io_context, acceptor, worker pool, signal handling, `on_ready` callback.
- [x] `test_m1_integration` — M1 exit criterion. Verified green 2026-06-15.
**Key bug fixed:** double-framing in `ConnSession::send_envelope``encode_envelope` was
pre-framing the protobuf, then `TcpServerConn::send_frame` re-framed it. Fixed by serializing
raw protobuf bytes directly and letting `send_frame` add the single `[4-byte len]` prefix.
---
---
## M2 — Voice, single stream ✓ (completed 2026-06-16)
**Exit criterion:**`test_m2_voice` — two headless clients authenticate over TLS, bind UDP,
announce a MIC stream, send 50 encrypted Opus frames; server SFU relay re-encrypts + forwards
to the second client; B receives ≥ 25 frames and all decrypt correctly. Passes in ~4 s.
- [x] `m2-dev` preset (inherits `vcpkg-base`, binaryDir `build/m2-dev`); `m1-dev` also builds all M2 code.
- [x] `core/CMakeLists.txt``find_package(Opus)`, `find_path(MINIAUDIO_INCLUDE_DIR)`.
- [x] `core/src/net/voice_frame.h` — 14-byte UDP header (type/flags/codec/ssrc/seq/ts), serialize/parse, `make_udp_binding_packet`.
- [x] `SodiumMediaCrypto` — ChaCha20-Poly1305 AEAD; counter-nonce; 64-bit sliding-window anti-replay; `derive_send/recv` from TLS RFC 5705 exporter.
- [x] `OpusEncoder` / `OpusDecoder` — libopus 1.6, FEC, DTX, PLC (free; nullptr → decoder extrapolates).
- [x] `UdpMediaChannel` — async UDP socket (asio); thread-safe `send_to`; async recv loop.
- [x] `JitterBuffer` — per-ssrc, EWMA jitter estimation, adaptive depth 20200 ms, late-drop at 500 ms.
- [x] `AudioEngine` — miniaudio capture+playback; `inject_capture()` bypass for headless tests; per-ssrc RemoteStream with OpusDecoder + JitterBuffer.
- [x] `ApmProcessor``ApmPassthrough` stub (VAD always open); WebRTC APM deferred until M3.
- [x] `on_tls_ready` callback in `TcpChannelCallbacks` — server derives and stores media AEAD keys immediately after TLS handshake.
- [x] `ConnSession` M2 — `udp_token` generated at construction; included in `AuthResult`; `handle_udp_binding` (verifies token, TCP ack); `handle_stream_announce` (assigns SSRC via registry); `udp_media_port` in `ServerHello`.
- [x] `SessionRegistry` M2 — `register_udp_token`, `find_by_udp_token`, `register_udp_endpoint`, `find_by_udp_endpoint`, `assign_ssrc`, `find_channel_sessions`, `user_channel`.
- [x] `MediaRelay` — SFU UDP relay; `kFrameUdpBinding` → endpoint binding; `kFrameVoice` → decrypt/re-encrypt/forward to channel members.
- [x] `Server::run()` — creates and binds `MediaRelay`; passes media port to `ConnSession`; wires `on_tls_ready` to derive per-connection media AEAD keys.
- [x] `test_voice_frame` — header round-trip, big-endian layout, binding packet format.
- [x] `test_media_aead` — seal/open round-trip, anti-replay, tamper detection, multi-packet sequence.
- [x] `test_opus_codec` — encode/decode round-trip energy check (within 3 dB), PLC, frame-samples helper.
- [x] `test_m2_voice` — M2 exit criterion (raw-socket harness). Verified green 2026-06-16.
**Follow-up (same day):** the above made `test_m2_voice` pass, but `vc_client`'s public voice
methods were still stubs — the *actual* M2 exit criterion ("two vccli/early-GUI clients talk")
wasn't met. Closed the gap:
- [x] `core/src/core/client.cpp` — real `stream_start`/`stream_stop`/`set_self_mute`/
`set_remote_stream`; UDP-binding handshake (`start_udp_binding`/`handle_udp_binding_ack`/
`finish_udp_binding`); media key derivation from `tls_` (RFC 5705 exporter); `run_udp_recv`
(AEAD-open → `JitterBuffer::Frame``audio_engine_.push_recv_frame`); `on_capture_frame`
(encode → seal → `sendto`); `sync_remote_streams` (diffs a `User` proto's `streams` against
`remote_streams_`, wiring up `OpusDecoder`s and emitting `STREAM_STARTED`/`STOPPED`).
`set_input_device`/`set_input_mode`/`set_push_to_talk`/`list_devices` remain
`VC_ERR_NOT_IMPLEMENTED` — no device-enumeration backend yet; scoped to M3 (VAD/PTT).
- [x] `core/src/session/session.cpp/h``SessionModel::find_user`, `find_user_by_ssrc`,
`Stream{stream_id, ssrc, kind, label, sample_rate, frame_ms}`.
- [x] `server/src/conn_session.cpp/h``handle_stream_announce`/`handle_stream_stop` now
broadcast via `SessionRegistry::set_user_stream`/`clear_user_stream``UserEvent::UPDATED`.
- [x] `server/src/session_registry.cpp/h``set_user_stream`/`clear_user_stream` (mutate a
user's `StreamInfo` list, return the updated `User` proto for broadcast).
- [x] `tests/test_voice_client_abi.cpp` — drives two real `vc_client` instances through
`vc_connect`/`vc_authenticate_guest`/`vc_stream_start`/`vc_stream_stop`; asserts client B
observes client A's `STREAM_STARTED`/`STOPPED` events. Verified green 2026-06-16.
- [x] `tools/vccli/src/main.cpp` — argv parsing (`--host/--port/--nick/--channel/--voice/
--mute/--text`); `--voice` starts a MIC stream and blocks on SIGINT, printing `on_event`
callbacks live (unbuffered stdout — MinGW/MSVCRT treat `_IOLBF` as full buffering for
non-console streams). Dropped the originally-planned `--voice-loopback` and the
`tx=N rx=M lost=K jitter=J` stats line: `voicecat.h` exposes no PCM-injection hook or
jitter/loss stats getter publicly, only `on_event` + `on_level` (RMS). Manually verified:
two `vccli --voice` instances see each other's stream start in real time.
---
## M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)
**Exit criterion:** ✓ `test_m3_multistream` — a real `vc_client` (A) runs two concurrent local
streams (MIC + SCREEN_AUDIO) with distinct stream ids; a second client (B) sees both as
separate `STREAM_STARTED` events and a `VC_EVENT_TALK_STATE` talking edge for A's MIC stream;
B independently sets gain/mute/noise-reduction on each of A's streams without one call
affecting the other; A then joins "Music Room" (channel 2: stereo/128kbps/`OPUS_AUDIO`/no
DTX) and announces a fresh MIC stream there, while B stays in "Lobby" (channel 1: mono/24kbps/
`OPUS_VOIP`/DTX on) — `vc_get_stream_audio_config` shows their effective Opus config differs
exactly as the server enforces per channel. Passes in ~2.4s; verified across 8 consecutive
standalone runs + 3 consecutive full-suite runs with no flakes.
Exploration before implementing turned up several bugs/gaps where the wire format already
supported this milestone but the client/server logic didn't — these were fixed as part of M3,
not treated as pre-existing-and-out-of-scope:
- [x] **Server `stream_id` bug** — `handle_stream_announce` always wrote `stream_id=1`, so a
second stream from the same user silently overwrote the first in
`SessionRegistry::set_user_stream`'s replace-by-id logic. Fixed with a per-session counter
(`ConnSession::next_stream_id_`) + `announced_stream_ids_` (also now validated in
`handle_stream_stop`, rejecting stops for ids the session never announced).
- [x] **Per-channel `AudioConfig` was modeled but never populated/enforced.**
`SessionRegistry::init_default_channels()` now seeds Lobby (id=1: mono, 24kbps, `OPUS_VOIP`,
FEC+DTX on) and a new "Music Room" (id=2: stereo, 128kbps, `OPUS_AUDIO`, FEC+DTX off) with
real `AudioConfig`s; new `SessionRegistry::channel_audio_config(channel_id)` accessor (there
was no per-id channel getter before, only `channel_snapshot()`). `handle_stream_announce`
now treats the channel's config as authoritative (mode/frame_ms/application/fec/dtx/
complexity), clamping (not overriding) `bitrate_bps` to the channel's ceiling.
- [x] **Client silently dropped `mode`/`dtx`/`complexity`/`application` from `effective_audio`**
even for the single M2 stream — `handle_stream_announce_result` and `sync_remote_streams`
only copied `sample_rate`/`bitrate_bps`/`frame_ms`/`fec` into `OpusParams`. New shared
`opus_params_from_audio_config()` helper (`client.cpp`) fixes both the send and receive
paths.
- [x] `core/src/codec/opus_codec.h/.cpp` — new `OpusApplication` enum + `OpusParams::application`
field; `OpusEncoder::init` now honors it instead of hardcoding `OPUS_APPLICATION_VOIP`.
- [x] `core/src/session/session.h/.cpp` — `Stream` struct extended with the full `AudioConfig`
(mode/bitrate_bps/application/fec/expected_packet_loss/dtx/complexity), not just
sample_rate/frame_ms; `copy_streams()` now copies all of it.
- [x] `core/src/core/client.h/.cpp` — local-stream state is now a `std::unordered_map<int,
LocalStream>` keyed by `vc_stream_kind` (one active stream per kind — MIC/SCREEN_AUDIO/
AUX_DEVICE are each singletons for a client), replacing the M2 single-stream fields.
`StreamAnnounce`/`StreamAnnounceResult` round-trips are now correlated by `request_id`
(already round-tripped on the wire; just wasn't read) via `pending_announce_kind_`, so
multiple concurrent announces from one client resolve to the right `LocalStream`.
`on_capture_frame` takes a `kind` parameter and upmixes mono capture to stereo (duplicate
L=R) when a stream's channel config calls for it. `vc_set_self_mute`'s `mic_muted` only
gates the `MIC` kind — a concurrent `SCREEN_AUDIO` share keeps playing while muted.
`set_remote_stream` now actually wires `noise_reduction` through (previously parsed and
discarded). New `run_talk_timer()` (a small dedicated thread, started alongside the UDP
media path, never the miniaudio callback thread) polls both remote talk-state edges
(`AudioEngine::poll_talk_transitions()`) and local capture-activity edges, emitting
`VC_EVENT_TALK_STATE`.
- [x] **Fixed a thread-join race in `teardown_voice()`** — it's called both from `run_io()`'s
own cleanup and from `disconnect()`, on different threads; without serialization both could
see `udp_thread_`/`talk_timer_thread_` as `joinable()` simultaneously and race to `join()`
the same `std::thread` (UB; surfaced as an intermittent `std::system_error: No such process`
under `ctest`). Added a `teardown_mu_` guard around the whole function. This pre-existed for
`udp_thread_` alone (likely the same root cause as the `test_m1_integration`/`test_m2_voice`
cleanup-path flake noted in the M2 section above) — adding `talk_timer_thread_`'s join just
made it surface more often, so it was fixed properly here rather than carried forward again.
- [x] `core/src/audio/audio_engine.h/.cpp` — `CaptureCallback` gained a `kind` parameter
(the real miniaudio capture device is always tagged `kind=0`/MIC; a second concurrent local
stream is fed via its own `inject_capture(kind, ...)` ring buffer — `inject_taps_`, keyed by
kind — since there is only one real hardware capture device in M3). Fixed a buffer-sizing
bug in `on_playback`'s per-stream decode (`opus_decode`'s `frame_size` parameter is
samples-*per-channel*, not total samples — the old code passed `frames * params_.channels`,
which would have overflowed the decode buffer for any stereo stream). Stereo decoder output
is downmixed (avg L/R) into the engine's mono mix accumulator immediately after decode.
`RemoteStream` gained `recv_ns`/`noise_reduction_enabled` (lazy `ApmProcessor` instantiation
— freed on disable, so no separate instance cap is needed per the roadmap's guidance) and
`last_voice_ms`/`talking` (talk-indicator edge state, updated in `push_recv_frame`); new
`set_stream_noise_reduction()` and `poll_talk_transitions()`. Note: until `VOICECAT_HAS_APM`
is wired to a real WebRTC APM build, the NS toggle is plumbed end-to-end but behaviorally a
passthrough no-op (`ApmPassthrough` doesn't touch PCM) — same situation send-side APM has
been in since M2; M3's job was the plumbing, not the DSP backend.
- [x] **New C ABI surface** (`core/include/voicecat.h`, additive only):
`vc_audio_config` struct + `vc_get_stream_audio_config(c, user_id, stream_id, out)` — the
effective Opus config for a stream you own or a peer's, reading from the (now richer)
`LocalStream`/`session::Stream`. `vc_test_inject_capture(c, stream_id, pcm, samples)` —
clearly-marked **test-only**, forwards to `AudioEngine::inject_capture`, so
`test_m3_multistream` can drive two concurrent synthetic-audio streams through the real ABI
without a microphone.
- [x] `tests/test_m3_multistream.cpp` — the M3 exit criterion (ABI-level, mirrors
`test_voice_client_abi.cpp`'s approach per the M2 lesson). Registered in `tests/CMakeLists.txt`.
**Explicitly out of scope for this pass** (confirmed with the user before implementing):
- `vc_set_input_device`/`vc_set_input_mode`/`vc_set_push_to_talk`/`vc_list_devices` (device
enumeration + VAD/PTT input gate) — still `VC_ERR_NOT_IMPLEMENTED`. These were mentioned as
"scoped to M3" in the M2 follow-up notes above, but docs/roadmap.md's M3 bullets never
actually listed them — deferred again, now tracked explicitly rather than implicitly.
- Real WASAPI desktop-audio loopback capture for `SCREEN_AUDIO` — the engine now supports
feeding a second concurrent local stream via `inject_capture`, but only synthetic PCM is
wired up; a real loopback capture device is a follow-up.
- True stereo *playback output* — `AudioEngine`'s mixer/output device stays mono; stereo
streams are downmixed after decode (see above). The Opus wire format itself is fully
stereo-correct.
---
## Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI loopback ✓ (completed 2026-06-16)
Closes all three items the M3 section above explicitly carried forward as out of scope.
**Exit verification:** `ctest --test-dir build/m1-dev` — **12/12 tests** green (3 consecutive
full-suite runs), including the new `test_vad_ptt_devices` (device enumeration + VAD/PTT gate
through real `vc_client`s against a real server, plus a white-box `AudioEngine` stereo-mix
check — no audio hardware needed for that last part). Also verified 5 consecutive standalone
runs of the new test alone, no flakes. Manually verified live on Windows: `vccli
--list-devices` against real hardware (3 input / 4 output devices, correct `is_default`
flags), and `vccli --voice --input-mode vad` connecting + streaming without incident.
- [x] **Device enumeration** (`vc_list_devices`) — `AudioEngine::enumerate_devices(bool
capture)` (static, works without a running engine — inits a throwaway `ma_context` via
`ma_context_get_devices`). `device_id`/`vc_device.id` is an opaque hex-encoded raw
`ma_device_id` (not the device name — names aren't guaranteed unique); documented as an
internal contract callers must round-trip, never construct by hand. `vc_client::list_devices`
works in any connection state (no `VC_STATE_CONNECTED` gate) since device pickers need to
populate pre-connect. `vc_free_device_list` is now a real free (was a no-op stub).
- [x] **Input device selection** (`vc_set_input_device`) — stores the device id on the
targeted `LocalStream` (new field); for the MIC stream, if the engine is already running,
restarts it (`stop()` + `ensure_audio_running()`) to pick up the new device. Simplified: it
restarts unconditionally rather than trying to detect whether the device id actually
changed (`AudioEngine` has no getter for "current device").
- [x] **VAD/PTT input gate** (`vc_set_input_mode`, `vc_set_push_to_talk`) — new
`EnergyVadProcessor` (`core/src/audio/apm_processor.cpp`) implementing the existing
`ApmProcessor` interface: energy/RMS threshold (default ~0.025 normalized) + hang-time
(default 300 ms, matching `kTalkHangoverMs`). New factory `ApmProcessor::create_vad()`,
kept separate from `create()` (which recv-side per-stream NS still uses, unaffected by this
pass). `vc_client` gained `current_input_mode_`/`ptt_active_`/`mic_vad_`; the gate is
inserted in `on_capture_frame`, **MIC-only** — `SCREEN_AUDIO`/`AUX_DEVICE` always bypass it
(gating a desktop-audio share on the user's own voice activity would silently drop shared
music/video audio). `last_capture_ms` (drives the talk indicator) is now updated *after* the
gate check, not before, so a VAD/PTT-closed frame never shows as "talking". `mic_vad_` is
constructed once the MIC stream's `StreamAnnounceResult` lands (on `io_thread_`), not lazily
inside the capture path.
- [x] **True stereo playback** — `AudioParams::channels` split into `capture_channels` (stays
1) and `playback_channels` (now 2, unconditionally). `AudioEngine::on_playback` no longer
downmixes decoded stereo streams to mono before mixing — stereo decode output is mixed
directly (L→L, R→R); mono decode output is upmixed (duplicated into both channels). Falls
back to a 1-channel playback device once if the 2-channel `ma_device_init` fails (unusual
hardware). New test-only `AudioEngine::mix_for_test()` exposes the mixer for white-box
testing without a real `ma_device`.
- [x] **WASAPI loopback capture for `SCREEN_AUDIO`** — new `VOICECAT_HAS_LOOPBACK` macro
(`core/CMakeLists.txt`, Windows-only). `AudioEngine` gained a separate `loopback_device_`
(own lifecycle, decoupled from the mic capture/playback devices) with
`start_loopback_capture()`/`stop_loopback_capture()`, using miniaudio's
`ma_device_type_loopback` against the default render endpoint. Its callback feeds
`capture_cb_` directly (same pattern as the real mic capture device), **not** through
`inject_capture()`'s test-only ring. Wired into `vc_client::handle_stream_announce_result`
(start, alongside `ensure_audio_running()`) and `stream_stop` (stop) for
`VC_STREAM_SCREEN_AUDIO`. Non-Windows builds keep `vc_test_inject_capture` as the only way to
feed `SCREEN_AUDIO`.
- [x] `tools/vccli/src/main.cpp` — new flags `--list-devices`, `--input-device`,
`--input-mode vad|ptt`, `--share-screen-audio`; while `--voice` is running, a background
stdin-reader thread accepts `ptt on`/`ptt off`/`mode vad`/`mode ptt` (the most portable way
to drive PTT interactively from a headless CLI — no SIGUSR1 equivalent on Windows). Also
prints `VC_EVENT_TALK_STATE`. Known minor caveat: on Windows the stdin-reader thread is
detached (not joined) on exit, since `std::getline` can't be interrupted from another thread
— a `vc_client*` use-after-free is theoretically possible if a command line arrives in the
brief window between teardown and process exit; acceptable for a headless test/dev tool.
- [x] `tests/test_vad_ptt_devices.cpp` — new test covering all four items above; registered in
`tests/CMakeLists.txt`. `tests/test_smoke.cpp`'s device-list assertion is now conditional on
`VOICECAT_HAS_AUDIO` (was a hard `VC_ERR_NOT_IMPLEMENTED` assertion) — `VC_OK` only, never
`count > 0` (a headless CI build agent may legitimately report zero audio devices).
**Still explicitly out of scope** (carried forward, not silently dropped):
- Real `webrtc-audio-processing`/AEC — no working Windows/MSVC build upstream (see
docs/roadmap.md §2's superseding decision-log entry). There is **no AEC, NS, or AGC
implementation at all**, not just a deferred VAD. The per-stream NS toggle
(`vc_set_remote_stream(..., noise_reduction)`) is unaffected by this pass and stays exactly
as inert as it was after M3 (`ApmPassthrough`, no PCM modification) — don't mistake this
pass for having fixed it.
- macOS/iOS `SCREEN_AUDIO` capture (ScreenCaptureKit / ReplayKit) — Windows-only loopback in
this pass.
- Process-specific WASAPI loopback — whole-device capture only; inherently captures this
app's own incoming voice mix along with everything else playing.
- The pre-existing RT-thread rule violation in `on_capture_frame`/`AudioEngine::on_capture`
(mutex lock, heap allocation for the stereo-upmix path, blocking `sendto`, all on the
miniaudio real-time callback thread) — predates this work (was already present in M2/M3);
documented here explicitly rather than silently carried forward again. Fixing it properly
needs the lock-free ring-buffer hand-off `docs/architecture.md §3` specifies — a separate,
larger refactor, out of scope for this pass.
---
## Decisions log
All architecture/scope decisions are settled and recorded in
[docs/roadmap.md §2 "Resolved decisions"](docs/roadmap.md) and reflected across `docs/`.
If you make a *new* decision, record it there and link it here.
---
## How to update this file
1. Check off tasks as you complete them; flip a milestone to `[x]` only when its **exit
criterion test** passes.
2. Keep the **"Where we left off / next action"** block at the top accurate — it's the first
thing the next agent reads.
3. When you start a milestone, copy its task list from `docs/roadmap.md` into a section here.