Files
voice-cat/docs/roadmap.md
Talon 5f6c223526 feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as
out of scope:

- Device enumeration (vc_list_devices) + input device selection
  (vc_set_input_device), backed by AudioEngine::enumerate_devices() via
  miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded
  ma_device_id strings.
- VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk).
  webrtc-audio-processing (the originally-planned APM) has no working
  Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard
  abseil-cpp dependency), so VAD is a new lightweight, dependency-free
  energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor
  interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it.
- True stereo playback: AudioEngine's mixer and output device now carry
  stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded
  stereo streams to mono before mixing.
- Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via
  miniaudio's loopback device type), replacing test-only injection as the
  production capture path.

Also: vccli gains --list-devices, --input-device, --input-mode, and
--share-screen-audio flags, plus a stdin command loop (ptt on/off, mode
vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all
four items (ABI-level + a white-box AudioEngine stereo-mix check).

Docs updated to match: voice.md, roadmap.md (decision-log entry superseding
the original webrtc-audio-processing choice), tech-stack.md, README.md,
architecture.md, CLAUDE.md, PROGRESS.md.

Still explicitly out of scope, documented not silently dropped: real
webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS
SCREEN_AUDIO capture, process-specific loopback, and a pre-existing
RT-thread rule violation in the capture path that predates this work.

Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev
and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive
standalone runs; manually verified live (vccli --list-devices against real
hardware, vccli --voice --input-mode vad streaming without incident).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 16:11:52 +02:00

6.2 KiB

Roadmap & Open Questions

1. Milestones

Each milestone is shippable/testable on its own. The headless C++ test client (vccli) exists from M1 so the protocol can be exercised long before any GUI.

M0 — Scaffolding

  • Repo layout (see architecture.md §6), CMake + vcpkg manifest, CI matrix.
  • core/proto/ skeleton; protoc codegen wired for C++ (C#/Swift later).
  • Empty libvoicecat with the C ABI header and stub implementations.
  • Exit: core + server + vccli compile and link on Linux/macOS/Windows.

M1 — Control plane (TCP/TLS, no audio yet)

  • TLS 1.3 transport; framing; Envelope; ClientHello/ServerHello negotiation.
  • Auth: guest + admin-provisioned local account (Argon2id, SQLite); server identity (TOFU/Ed25519); voicecat-admin account add/reset/del/list.
  • Channel tree: snapshot + deltas; join/leave; create/edit/delete (perm-checked).
  • Text chat (ephemeral): channel + private messages, acks; live relay, no history store.
  • vccli can connect, auth, browse channels, and chat.
  • Exit: two vccli instances chat through a real server over TLS.

M2 — Voice, single stream

  • UDP transport with exported-key + ChaCha20-Poly1305 AEAD (mandatory, no plaintext path); UDP token binding; anti-replay.
  • miniaudio capture/playback; libopus encode/decode; one MIC stream per user.
  • Send-side webrtc APM (AEC + NS + AGC + VAD) and a VAD/PTT input gate (both modes, client-configurable) — AEC is in from the start, not deferred.
  • Per-ssrc adaptive jitter buffer; mixer; FEC/PLC/DTX.
  • Per-channel AudioConfig enforced (incl. server max_bitrate_bps ceiling); SFU relay.
  • Exit: talk between two vccli/early-GUI clients in a channel; loss resilience visible.

M3 — Multi-stream & per-channel tuning

  • Multiple concurrent streams per user (MIC, SCREEN_AUDIO, AUX_DEVICE).
  • Per-stream receiver gain/mute; listener-side per-user noise reduction (APM NS on the receive path, per ssrc, local-only); talk indicators.
  • Full per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX).
  • Exit: a user shares mic + desktop audio; listeners control each independently.

M4 — Native clients

  • Windows (C#/WinUI): connect, saved-server list, channel tree, voice, text, device pickers, meters, VAD/PTT + per-user NR controls.
  • macOS (Swift/SwiftUI): same.
  • iOS (Swift): AVAudioSession integration, mic permission, foreground voice; ReplayKit broadcast extension for SCREEN_AUDIO.
  • In-app admin interface (account provisioning, bans) for admin users.
  • Exit: non-technical user installs a client, saves a server, and joins.

M5 — Moderation, polish, and beyond

  • Permissions/roles, kick/ban/server-mute, channel passwords UI.
  • DRED toggle, audio-quality polish. (AEC and VAD/PTT already shipped in M2.)
  • Then (post-v1, protocol already reserves space): file transfer, E2EE option, CallKit/PushKit background voice, key-based identity, server-side text history, multi-node server.

2. Resolved decisions

Settled and reflected throughout the docs:

  • Media crypto: exported-key + ChaCha20-Poly1305 AEAD from day one, mandatory — no DTLS, no plaintext mode. (security.md §2)
  • TLS library / licensing: mbedTLS (Apache-2.0) + libsodium (ISC). No GPL/LGPL anywhere; code is redistributable closed-source. wolfSSL is rejected. (tech-stack.md §5)
  • iOS screen/system audio: supported via a ReplayKit Broadcast Upload Extension (.audioApp); audio-only stays within the extension memory cap. (voice.md §9)
  • Self-host UX: zero-config, encrypted-by-default; Docker / single binary / source build. (deployment.md)
  • DSP engine: webrtc-audio-processing (APM) — AEC in from the start, plus NS/AGC/VAD. (voice.md §8)
  • Input activation: VAD and PTT, both modes client-configurable. (voice.md §11)
  • Noise reduction is two-sided: sender can denoise its mic, and each listener can apply NS to a specific other user, locally, with no protocol traffic. (voice.md §10)
  • Accounts: admin-provisioned (no self-serve registration) via voicecat-admin or the in-app admin interface. (security.md §4, protocol.md §3, deployment.md §3a)
  • Text: ephemeral — live relay, no server-side history in v1. (protocol.md §5)
  • Connect UX: pure direct-connect with a client-side saved-server list (no central directory). (deployment.md §3)
  • Bitrate ceiling: server-config opus.limits.max_bitrate_bps. (deployment.md §2)
  • Name: "VoiceCat" stays as the internal placeholder.
  • DSP engine, superseded (2026-06-16): the "webrtc-audio-processing (APM)" decision above (AEC + NS/AGC/VAD in one module) could not be carried out — it has no working Windows/MSVC build upstream (GCC-only Meson build, MinGW support unfinished, hard abseil-cpp dependency, Linux-tested only). v1 ships a lightweight, dependency-free energy/RMS VAD instead, behind the same ApmProcessor interface; there is no AEC/NS/AGC implementation at all yet. Real webrtc-audio-processing stays a tracked future swap (e.g. if/when a Linux build target exists). (voice.md §8, §11)

3. Open questions

All initial open questions are resolved (§2). Two second-order considerations to keep in mind during implementation — not blockers:

  • APM in constrained contexts. webrtc-audio-processing is a heavier build; confirm it static-links cleanly for the single-binary goal, and note the iOS broadcast extension only does Opus encode + send (no APM), so it stays under the ~50 MB cap. Listener-side per-user NS runs only in the full host app.
  • Receive-side NR cost at scale. A per-ssrc APM NS instance per flagged user adds CPU on busy channels; instantiate lazily (only for flagged streams) and cap concurrent instances.

4. What's intentionally deferred

To keep v1 focused (voice + text), these are designed-for but not built: file transfer, end-to-end encryption, key-based identities, multi-node/federated servers, mobile background VoIP push, and any server-side audio mixing/transcoding. The protocol's versioning + feature negotiation + reserved tag ranges (protocol.md §8) ensure each can be added without breaking deployed clients.