Files
voice-cat/PROGRESS.md
Talon 5f6c223526 feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as
out of scope:

- Device enumeration (vc_list_devices) + input device selection
  (vc_set_input_device), backed by AudioEngine::enumerate_devices() via
  miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded
  ma_device_id strings.
- VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk).
  webrtc-audio-processing (the originally-planned APM) has no working
  Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard
  abseil-cpp dependency), so VAD is a new lightweight, dependency-free
  energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor
  interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it.
- True stereo playback: AudioEngine's mixer and output device now carry
  stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded
  stereo streams to mono before mixing.
- Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via
  miniaudio's loopback device type), replacing test-only injection as the
  production capture path.

Also: vccli gains --list-devices, --input-device, --input-mode, and
--share-screen-audio flags, plus a stdin command loop (ptt on/off, mode
vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all
four items (ABI-level + a white-box AudioEngine stereo-mix check).

Docs updated to match: voice.md, roadmap.md (decision-log entry superseding
the original webrtc-audio-processing choice), tech-stack.md, README.md,
architecture.md, CLAUDE.md, PROGRESS.md.

Still explicitly out of scope, documented not silently dropped: real
webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS
SCREEN_AUDIO capture, process-specific loopback, and a pre-existing
RT-thread rule violation in the capture path that predates this work.

Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev
and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive
standalone runs; manually verified live (vccli --list-devices against real
hardware, vccli --voice --input-mode vad streaming without incident).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 16:11:52 +02:00

25 KiB
Raw Blame History

PROGRESS — VoiceCat

Living status. Update this file in the same commit as your work so the next agent picks up instantly. Newest status at the top.

  • Date convention: ISO (YYYY-MM-DD).
  • Statuses: [ ] not started · [~] in progress · [x] done.

▶ Where we left off / next action

  • Done: Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI loopback ✓ complete (2026-06-16). Closes all three items M3 explicitly carried forward as out of scope (see the dated section below for the full file-by-file change list). ctest --test-dir build/m1-dev12/12 tests green (3 consecutive full-suite runs), including the new test_vad_ptt_devices (real vc_clients against a real server, plus a white-box AudioEngine stereo-mix check — same ABI-level-coverage lesson as M2/M3). Manually verified live: vccli --list-devices against real hardware, and vccli --voice --input-mode vad connecting/streaming without incident. Still explicitly out of scope (carried forward, not silently dropped):
    • Real webrtc-audio-processing/AEC — no working Windows/MSVC build upstream; v1 ships a lightweight energy/RMS VAD instead (see docs/roadmap.md §2, docs/voice.md §8/§11). There is no AEC, NS, or AGC implementation at all, not just a deferred VAD. vc_set_remote_stream(..., noise_reduction)'s per-stream NS toggle is unaffected by this pass and stays exactly as inert as it was after M3 (ApmPassthrough, no PCM modification).
    • macOS/iOS SCREEN_AUDIO capture (ScreenCaptureKit / ReplayKit) — this pass is Windows-only for real loopback capture; other platforms keep vc_test_inject_capture as the only way to feed SCREEN_AUDIO.
    • Process-specific WASAPI loopback — miniaudio's loopback mode captures the whole render endpoint (including this app's own incoming voice mix), not a single process.
    • The pre-existing RT-thread rule violation in on_capture_frame/AudioEngine::on_capture (mutex lock, heap allocation, blocking sendto on the miniaudio real-time callback thread) — predates this work, documented but not fixed; fixing it needs the lock-free ring-buffer hand-off docs/architecture.md §3 specifies, a separate, larger refactor.
  • Next: M4 — native clients (Windows C#, macOS/iOS Swift). See docs/roadmap.md §M4.

Milestones (see docs/roadmap.md for full detail)

  • M0 — Scaffolding ✓ complete
  • M1 — Control plane ✓ complete (2026-06-15)
  • M2 — Voice, single stream ✓ complete (2026-06-16)
  • M3 — Multi-stream & per-channel tuning ✓ complete (2026-06-16)
  • M4 — Native clients (Windows C#, macOS/iOS Swift) ← next
  • M5 — Moderation, polish, beyond (perms, bans, DRED; then file transfer, E2EE, …)

M0 — Scaffolding ✓ (completed)

  • Repo layout (core/ server/ tools/ clients/ tests/), CMake + presets, vcpkg manifest.
  • C ABI header core/include/voicecat.h (full surface, stubbed).
  • Protocol source-of-truth core/proto/voicecat.proto (matches docs/protocol.md).
  • Core stubs for all six subsystems (net/crypto/codec/protocol/session/audio) + vc_client.
  • voicecat-server (arg parsing, config, stub run) and vccli (drives the C ABI).
  • CTest smoke test asserting the C ABI contract (not just "it compiles").
  • .gitattributes (LF), .gitignore, .clang-format, onboarding docs.
  • Verified: cmake --preset dev && cmake --build --preset dev && ctest --preset dev → green.

M1 — Control plane ✓ (completed 2026-06-15)

Exit criterion:test_m1_integration — two clients authenticate over TLS 1.3 (guest

  • Argon2id password), exchange channel and private text messages. Passes in ~1 s.
  • vcpkg baseline + m1-dev preset; find_package for protobuf/mbedTLS/libsodium/asio/sqlite3.
  • FrameCodec feed + emit; encode_envelope / decode_envelope.
  • Asio TCP acceptor + TcpServerConn (TLS path: blocking handshake thread + tls_read_loop).
  • TlsContext (mbedTLS 1.3, server cert/identity, ECDSA-P256 self-signed, TOFU on client).
  • WorkerPool (3 threads, used for Argon2id).
  • Database — SQLite, Argon2id via libsodium, create_account / authenticate / bootstrap admin.
  • voicecat-admin — account add/reset/del/list against live DB file.
  • ServerIdentityManager — generate/persist Ed25519 key + cert; fingerprint display.
  • ConnSession — WaitingHello → WaitingAuth → Authenticated state machine; full protocol relay.
  • SessionRegistry — channel tree, user map, broadcast, text routing.
  • vc_client (client.cpp) — full M1 C ABI: connect/TLS/ClientHello/AuthRequest/text/disconnect.
  • Server::run() — io_context, acceptor, worker pool, signal handling, on_ready callback.
  • test_m1_integration — M1 exit criterion. Verified green 2026-06-15.

Key bug fixed: double-framing in ConnSession::send_envelopeencode_envelope was pre-framing the protobuf, then TcpServerConn::send_frame re-framed it. Fixed by serializing raw protobuf bytes directly and letting send_frame add the single [4-byte len] prefix.



M2 — Voice, single stream ✓ (completed 2026-06-16)

Exit criterion:test_m2_voice — two headless clients authenticate over TLS, bind UDP, announce a MIC stream, send 50 encrypted Opus frames; server SFU relay re-encrypts + forwards to the second client; B receives ≥ 25 frames and all decrypt correctly. Passes in ~4 s.

  • m2-dev preset (inherits vcpkg-base, binaryDir build/m2-dev); m1-dev also builds all M2 code.
  • core/CMakeLists.txtfind_package(Opus), find_path(MINIAUDIO_INCLUDE_DIR).
  • core/src/net/voice_frame.h — 14-byte UDP header (type/flags/codec/ssrc/seq/ts), serialize/parse, make_udp_binding_packet.
  • SodiumMediaCrypto — ChaCha20-Poly1305 AEAD; counter-nonce; 64-bit sliding-window anti-replay; derive_send/recv from TLS RFC 5705 exporter.
  • OpusEncoder / OpusDecoder — libopus 1.6, FEC, DTX, PLC (free; nullptr → decoder extrapolates).
  • UdpMediaChannel — async UDP socket (asio); thread-safe send_to; async recv loop.
  • JitterBuffer — per-ssrc, EWMA jitter estimation, adaptive depth 20200 ms, late-drop at 500 ms.
  • AudioEngine — miniaudio capture+playback; inject_capture() bypass for headless tests; per-ssrc RemoteStream with OpusDecoder + JitterBuffer.
  • ApmProcessorApmPassthrough stub (VAD always open); WebRTC APM deferred until M3.
  • on_tls_ready callback in TcpChannelCallbacks — server derives and stores media AEAD keys immediately after TLS handshake.
  • ConnSession M2 — udp_token generated at construction; included in AuthResult; handle_udp_binding (verifies token, TCP ack); handle_stream_announce (assigns SSRC via registry); udp_media_port in ServerHello.
  • SessionRegistry M2 — register_udp_token, find_by_udp_token, register_udp_endpoint, find_by_udp_endpoint, assign_ssrc, find_channel_sessions, user_channel.
  • MediaRelay — SFU UDP relay; kFrameUdpBinding → endpoint binding; kFrameVoice → decrypt/re-encrypt/forward to channel members.
  • Server::run() — creates and binds MediaRelay; passes media port to ConnSession; wires on_tls_ready to derive per-connection media AEAD keys.
  • test_voice_frame — header round-trip, big-endian layout, binding packet format.
  • test_media_aead — seal/open round-trip, anti-replay, tamper detection, multi-packet sequence.
  • test_opus_codec — encode/decode round-trip energy check (within 3 dB), PLC, frame-samples helper.
  • test_m2_voice — M2 exit criterion (raw-socket harness). Verified green 2026-06-16.

Follow-up (same day): the above made test_m2_voice pass, but vc_client's public voice methods were still stubs — the actual M2 exit criterion ("two vccli/early-GUI clients talk") wasn't met. Closed the gap:

  • core/src/core/client.cpp — real stream_start/stream_stop/set_self_mute/ set_remote_stream; UDP-binding handshake (start_udp_binding/handle_udp_binding_ack/ finish_udp_binding); media key derivation from tls_ (RFC 5705 exporter); run_udp_recv (AEAD-open → JitterBuffer::Frameaudio_engine_.push_recv_frame); on_capture_frame (encode → seal → sendto); sync_remote_streams (diffs a User proto's streams against remote_streams_, wiring up OpusDecoders and emitting STREAM_STARTED/STOPPED). set_input_device/set_input_mode/set_push_to_talk/list_devices remain VC_ERR_NOT_IMPLEMENTED — no device-enumeration backend yet; scoped to M3 (VAD/PTT).
  • core/src/session/session.cpp/hSessionModel::find_user, find_user_by_ssrc, Stream{stream_id, ssrc, kind, label, sample_rate, frame_ms}.
  • server/src/conn_session.cpp/hhandle_stream_announce/handle_stream_stop now broadcast via SessionRegistry::set_user_stream/clear_user_streamUserEvent::UPDATED.
  • server/src/session_registry.cpp/hset_user_stream/clear_user_stream (mutate a user's StreamInfo list, return the updated User proto for broadcast).
  • tests/test_voice_client_abi.cpp — drives two real vc_client instances through vc_connect/vc_authenticate_guest/vc_stream_start/vc_stream_stop; asserts client B observes client A's STREAM_STARTED/STOPPED events. Verified green 2026-06-16.
  • tools/vccli/src/main.cpp — argv parsing (--host/--port/--nick/--channel/--voice/ --mute/--text); --voice starts a MIC stream and blocks on SIGINT, printing on_event callbacks live (unbuffered stdout — MinGW/MSVCRT treat _IOLBF as full buffering for non-console streams). Dropped the originally-planned --voice-loopback and the tx=N rx=M lost=K jitter=J stats line: voicecat.h exposes no PCM-injection hook or jitter/loss stats getter publicly, only on_event + on_level (RMS). Manually verified: two vccli --voice instances see each other's stream start in real time.

M3 — Multi-stream & per-channel tuning ✓ (completed 2026-06-16)

Exit criterion:test_m3_multistream — a real vc_client (A) runs two concurrent local streams (MIC + SCREEN_AUDIO) with distinct stream ids; a second client (B) sees both as separate STREAM_STARTED events and a VC_EVENT_TALK_STATE talking edge for A's MIC stream; B independently sets gain/mute/noise-reduction on each of A's streams without one call affecting the other; A then joins "Music Room" (channel 2: stereo/128kbps/OPUS_AUDIO/no DTX) and announces a fresh MIC stream there, while B stays in "Lobby" (channel 1: mono/24kbps/ OPUS_VOIP/DTX on) — vc_get_stream_audio_config shows their effective Opus config differs exactly as the server enforces per channel. Passes in ~2.4s; verified across 8 consecutive standalone runs + 3 consecutive full-suite runs with no flakes.

Exploration before implementing turned up several bugs/gaps where the wire format already supported this milestone but the client/server logic didn't — these were fixed as part of M3, not treated as pre-existing-and-out-of-scope:

  • Server stream_id bughandle_stream_announce always wrote stream_id=1, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream's replace-by-id logic. Fixed with a per-session counter (ConnSession::next_stream_id_) + announced_stream_ids_ (also now validated in handle_stream_stop, rejecting stops for ids the session never announced).
  • Per-channel AudioConfig was modeled but never populated/enforced. SessionRegistry::init_default_channels() now seeds Lobby (id=1: mono, 24kbps, OPUS_VOIP, FEC+DTX on) and a new "Music Room" (id=2: stereo, 128kbps, OPUS_AUDIO, FEC+DTX off) with real AudioConfigs; new SessionRegistry::channel_audio_config(channel_id) accessor (there was no per-id channel getter before, only channel_snapshot()). handle_stream_announce now treats the channel's config as authoritative (mode/frame_ms/application/fec/dtx/ complexity), clamping (not overriding) bitrate_bps to the channel's ceiling.
  • Client silently dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream — handle_stream_announce_result and sync_remote_streams only copied sample_rate/bitrate_bps/frame_ms/fec into OpusParams. New shared opus_params_from_audio_config() helper (client.cpp) fixes both the send and receive paths.
  • core/src/codec/opus_codec.h/.cpp — new OpusApplication enum + OpusParams::application field; OpusEncoder::init now honors it instead of hardcoding OPUS_APPLICATION_VOIP.
  • core/src/session/session.h/.cppStream struct extended with the full AudioConfig (mode/bitrate_bps/application/fec/expected_packet_loss/dtx/complexity), not just sample_rate/frame_ms; copy_streams() now copies all of it.
  • core/src/core/client.h/.cpp — local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind (one active stream per kind — MIC/SCREEN_AUDIO/ AUX_DEVICE are each singletons for a client), replacing the M2 single-stream fields. StreamAnnounce/StreamAnnounceResult round-trips are now correlated by request_id (already round-tripped on the wire; just wasn't read) via pending_announce_kind_, so multiple concurrent announces from one client resolve to the right LocalStream. on_capture_frame takes a kind parameter and upmixes mono capture to stereo (duplicate L=R) when a stream's channel config calls for it. vc_set_self_mute's mic_muted only gates the MIC kind — a concurrent SCREEN_AUDIO share keeps playing while muted. set_remote_stream now actually wires noise_reduction through (previously parsed and discarded). New run_talk_timer() (a small dedicated thread, started alongside the UDP media path, never the miniaudio callback thread) polls both remote talk-state edges (AudioEngine::poll_talk_transitions()) and local capture-activity edges, emitting VC_EVENT_TALK_STATE.
  • Fixed a thread-join race in teardown_voice() — it's called both from run_io()'s own cleanup and from disconnect(), on different threads; without serialization both could see udp_thread_/talk_timer_thread_ as joinable() simultaneously and race to join() the same std::thread (UB; surfaced as an intermittent std::system_error: No such process under ctest). Added a teardown_mu_ guard around the whole function. This pre-existed for udp_thread_ alone (likely the same root cause as the test_m1_integration/test_m2_voice cleanup-path flake noted in the M2 section above) — adding talk_timer_thread_'s join just made it surface more often, so it was fixed properly here rather than carried forward again.
  • core/src/audio/audio_engine.h/.cppCaptureCallback gained a kind parameter (the real miniaudio capture device is always tagged kind=0/MIC; a second concurrent local stream is fed via its own inject_capture(kind, ...) ring buffer — inject_taps_, keyed by kind — since there is only one real hardware capture device in M3). Fixed a buffer-sizing bug in on_playback's per-stream decode (opus_decode's frame_size parameter is samples-per-channel, not total samples — the old code passed frames * params_.channels, which would have overflowed the decode buffer for any stereo stream). Stereo decoder output is downmixed (avg L/R) into the engine's mono mix accumulator immediately after decode. RemoteStream gained recv_ns/noise_reduction_enabled (lazy ApmProcessor instantiation — freed on disable, so no separate instance cap is needed per the roadmap's guidance) and last_voice_ms/talking (talk-indicator edge state, updated in push_recv_frame); new set_stream_noise_reduction() and poll_talk_transitions(). Note: until VOICECAT_HAS_APM is wired to a real WebRTC APM build, the NS toggle is plumbed end-to-end but behaviorally a passthrough no-op (ApmPassthrough doesn't touch PCM) — same situation send-side APM has been in since M2; M3's job was the plumbing, not the DSP backend.
  • New C ABI surface (core/include/voicecat.h, additive only): vc_audio_config struct + vc_get_stream_audio_config(c, user_id, stream_id, out) — the effective Opus config for a stream you own or a peer's, reading from the (now richer) LocalStream/session::Stream. vc_test_inject_capture(c, stream_id, pcm, samples) — clearly-marked test-only, forwards to AudioEngine::inject_capture, so test_m3_multistream can drive two concurrent synthetic-audio streams through the real ABI without a microphone.
  • tests/test_m3_multistream.cpp — the M3 exit criterion (ABI-level, mirrors test_voice_client_abi.cpp's approach per the M2 lesson). Registered in tests/CMakeLists.txt.

Explicitly out of scope for this pass (confirmed with the user before implementing):

  • vc_set_input_device/vc_set_input_mode/vc_set_push_to_talk/vc_list_devices (device enumeration + VAD/PTT input gate) — still VC_ERR_NOT_IMPLEMENTED. These were mentioned as "scoped to M3" in the M2 follow-up notes above, but docs/roadmap.md's M3 bullets never actually listed them — deferred again, now tracked explicitly rather than implicitly.
  • Real WASAPI desktop-audio loopback capture for SCREEN_AUDIO — the engine now supports feeding a second concurrent local stream via inject_capture, but only synthetic PCM is wired up; a real loopback capture device is a follow-up.
  • True stereo playback outputAudioEngine's mixer/output device stays mono; stereo streams are downmixed after decode (see above). The Opus wire format itself is fully stereo-correct.

Post-M3 follow-up — device enumeration, VAD/PTT gate, stereo playback, WASAPI loopback ✓ (completed 2026-06-16)

Closes all three items the M3 section above explicitly carried forward as out of scope.

Exit verification: ctest --test-dir build/m1-dev12/12 tests green (3 consecutive full-suite runs), including the new test_vad_ptt_devices (device enumeration + VAD/PTT gate through real vc_clients against a real server, plus a white-box AudioEngine stereo-mix check — no audio hardware needed for that last part). Also verified 5 consecutive standalone runs of the new test alone, no flakes. Manually verified live on Windows: vccli --list-devices against real hardware (3 input / 4 output devices, correct is_default flags), and vccli --voice --input-mode vad connecting + streaming without incident.

  • Device enumeration (vc_list_devices) — AudioEngine::enumerate_devices(bool capture) (static, works without a running engine — inits a throwaway ma_context via ma_context_get_devices). device_id/vc_device.id is an opaque hex-encoded raw ma_device_id (not the device name — names aren't guaranteed unique); documented as an internal contract callers must round-trip, never construct by hand. vc_client::list_devices works in any connection state (no VC_STATE_CONNECTED gate) since device pickers need to populate pre-connect. vc_free_device_list is now a real free (was a no-op stub).
  • Input device selection (vc_set_input_device) — stores the device id on the targeted LocalStream (new field); for the MIC stream, if the engine is already running, restarts it (stop() + ensure_audio_running()) to pick up the new device. Simplified: it restarts unconditionally rather than trying to detect whether the device id actually changed (AudioEngine has no getter for "current device").
  • VAD/PTT input gate (vc_set_input_mode, vc_set_push_to_talk) — new EnergyVadProcessor (core/src/audio/apm_processor.cpp) implementing the existing ApmProcessor interface: energy/RMS threshold (default ~0.025 normalized) + hang-time (default 300 ms, matching kTalkHangoverMs). New factory ApmProcessor::create_vad(), kept separate from create() (which recv-side per-stream NS still uses, unaffected by this pass). vc_client gained current_input_mode_/ptt_active_/mic_vad_; the gate is inserted in on_capture_frame, MIC-onlySCREEN_AUDIO/AUX_DEVICE always bypass it (gating a desktop-audio share on the user's own voice activity would silently drop shared music/video audio). last_capture_ms (drives the talk indicator) is now updated after the gate check, not before, so a VAD/PTT-closed frame never shows as "talking". mic_vad_ is constructed once the MIC stream's StreamAnnounceResult lands (on io_thread_), not lazily inside the capture path.
  • True stereo playbackAudioParams::channels split into capture_channels (stays
    1. and playback_channels (now 2, unconditionally). AudioEngine::on_playback no longer downmixes decoded stereo streams to mono before mixing — stereo decode output is mixed directly (L→L, R→R); mono decode output is upmixed (duplicated into both channels). Falls back to a 1-channel playback device once if the 2-channel ma_device_init fails (unusual hardware). New test-only AudioEngine::mix_for_test() exposes the mixer for white-box testing without a real ma_device.
  • WASAPI loopback capture for SCREEN_AUDIO — new VOICECAT_HAS_LOOPBACK macro (core/CMakeLists.txt, Windows-only). AudioEngine gained a separate loopback_device_ (own lifecycle, decoupled from the mic capture/playback devices) with start_loopback_capture()/stop_loopback_capture(), using miniaudio's ma_device_type_loopback against the default render endpoint. Its callback feeds capture_cb_ directly (same pattern as the real mic capture device), not through inject_capture()'s test-only ring. Wired into vc_client::handle_stream_announce_result (start, alongside ensure_audio_running()) and stream_stop (stop) for VC_STREAM_SCREEN_AUDIO. Non-Windows builds keep vc_test_inject_capture as the only way to feed SCREEN_AUDIO.
  • tools/vccli/src/main.cpp — new flags --list-devices, --input-device, --input-mode vad|ptt, --share-screen-audio; while --voice is running, a background stdin-reader thread accepts ptt on/ptt off/mode vad/mode ptt (the most portable way to drive PTT interactively from a headless CLI — no SIGUSR1 equivalent on Windows). Also prints VC_EVENT_TALK_STATE. Known minor caveat: on Windows the stdin-reader thread is detached (not joined) on exit, since std::getline can't be interrupted from another thread — a vc_client* use-after-free is theoretically possible if a command line arrives in the brief window between teardown and process exit; acceptable for a headless test/dev tool.
  • tests/test_vad_ptt_devices.cpp — new test covering all four items above; registered in tests/CMakeLists.txt. tests/test_smoke.cpp's device-list assertion is now conditional on VOICECAT_HAS_AUDIO (was a hard VC_ERR_NOT_IMPLEMENTED assertion) — VC_OK only, never count > 0 (a headless CI build agent may legitimately report zero audio devices).

Still explicitly out of scope (carried forward, not silently dropped):

  • Real webrtc-audio-processing/AEC — no working Windows/MSVC build upstream (see docs/roadmap.md §2's superseding decision-log entry). There is no AEC, NS, or AGC implementation at all, not just a deferred VAD. The per-stream NS toggle (vc_set_remote_stream(..., noise_reduction)) is unaffected by this pass and stays exactly as inert as it was after M3 (ApmPassthrough, no PCM modification) — don't mistake this pass for having fixed it.
  • macOS/iOS SCREEN_AUDIO capture (ScreenCaptureKit / ReplayKit) — Windows-only loopback in this pass.
  • Process-specific WASAPI loopback — whole-device capture only; inherently captures this app's own incoming voice mix along with everything else playing.
  • The pre-existing RT-thread rule violation in on_capture_frame/AudioEngine::on_capture (mutex lock, heap allocation for the stereo-upmix path, blocking sendto, all on the miniaudio real-time callback thread) — predates this work (was already present in M2/M3); documented here explicitly rather than silently carried forward again. Fixing it properly needs the lock-free ring-buffer hand-off docs/architecture.md §3 specifies — a separate, larger refactor, out of scope for this pass.

Decisions log

All architecture/scope decisions are settled and recorded in docs/roadmap.md §2 "Resolved decisions" and reflected across docs/. If you make a new decision, record it there and link it here.


How to update this file

  1. Check off tasks as you complete them; flip a milestone to [x] only when its exit criterion test passes.
  2. Keep the "Where we left off / next action" block at the top accurate — it's the first thing the next agent reads.
  3. When you start a milestone, copy its task list from docs/roadmap.md into a section here.