Files
voice-cat/docs/README.md
Talon 5f6c223526 feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as
out of scope:

- Device enumeration (vc_list_devices) + input device selection
  (vc_set_input_device), backed by AudioEngine::enumerate_devices() via
  miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded
  ma_device_id strings.
- VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk).
  webrtc-audio-processing (the originally-planned APM) has no working
  Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard
  abseil-cpp dependency), so VAD is a new lightweight, dependency-free
  energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor
  interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it.
- True stereo playback: AudioEngine's mixer and output device now carry
  stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded
  stereo streams to mono before mixing.
- Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via
  miniaudio's loopback device type), replacing test-only injection as the
  production capture path.

Also: vccli gains --list-devices, --input-device, --input-mode, and
--share-screen-audio flags, plus a stdin command loop (ptt on/off, mode
vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all
four items (ABI-level + a white-box AudioEngine stereo-mix check).

Docs updated to match: voice.md, roadmap.md (decision-log entry superseding
the original webrtc-audio-processing choice), tech-stack.md, README.md,
architecture.md, CLAUDE.md, PROGRESS.md.

Still explicitly out of scope, documented not silently dropped: real
webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS
SCREEN_AUDIO capture, process-specific loopback, and a pre-existing
RT-thread rule violation in the capture path that predates this work.

Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev
and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive
standalone runs; manually verified live (vccli --list-devices against real
hardware, vccli --voice --input-mode vad streaming without incident).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 16:11:52 +02:00

69 lines
5.3 KiB
Markdown

# VoiceCat — Design Documentation
VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of
classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a
single server you own and run. It deliberately avoids WebRTC. The media path is plain
**UDP**, the control path is plain **TCP**, and both are encrypted.
This folder is the design spec. No code yet — these documents define the architecture,
the wire protocol, the audio pipeline, the security model, and the dependency list, so
that implementation can start from a shared, agreed plan.
## Decisions locked so far
| Area | Decision |
|------|----------|
| Code architecture | **Shared C++ core** (`libvoicecat`) consumed by native UIs over a **C ABI**. Server reuses the same core. |
| Native clients | macOS/iOS in **Swift** (SwiftUI; Swift↔C++ interop), Windows in **C#** (`LibraryImport` P/Invoke). |
| Control transport | **TCP + TLS 1.3** (mbedTLS) |
| Media transport | **UDP** secured by **TLS-exported keys + ChaCha20-Poly1305 AEAD** — mandatory, no plaintext mode (see [security.md](security.md)) |
| Crypto libraries | **mbedTLS** (TLS 1.3) + **libsodium** (AEAD, Argon2id, Ed25519) — both permissive, **no GPL/LGPL anywhere** |
| Voice codec | **Opus** (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX |
| Audio DSP | **webrtc-audio-processing (APM)** was the plan for AEC/NS/AGC/VAD, but has no working Windows/MSVC build upstream — v1 ships a lightweight energy/RMS VAD instead, no AEC/NS/AGC yet (see [tech-stack.md](tech-stack.md), [roadmap.md](roadmap.md) §2). NR is **two-sided**: sender can denoise, and each listener can denoise a *specific* other user locally — this plumbing exists but is currently inert pending a real DSP backend. Input gate supports **VAD and PTT**, client-configurable. |
| Identity | **Guests + admin-provisioned local accounts** (Argon2id, SQLite). No self-serve registration; guests toggleable per server. |
| Text | **Ephemeral** — live relay, no server-side history in v1. |
| Serialization | **Protocol Buffers** for the control plane; **custom binary** for voice frames |
| Deployment | **One `docker run`, one static binary, or `cmake --build`** — zero-config, secured by default (see [deployment.md](deployment.md)) |
## Document index
1. [architecture.md](architecture.md) — System layers, the shared core, threading model, the C ABI, server design.
2. [protocol.md](protocol.md) — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility.
3. [voice.md](voice.md) — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling.
4. [security.md](security.md) — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model.
5. [tech-stack.md](tech-stack.md) — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes.
6. [deployment.md](deployment.md) — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults.
7. [roadmap.md](roadmap.md) — Milestones, what ships when, and the list of open questions still to resolve.
## Design principles
- **One core, many faces.** Protocol, crypto, Opus, networking, jitter buffering, and
mixing live once in C++. UIs are thin. This keeps behavior identical across platforms
and the security-sensitive code reviewed in a single place.
- **Boringly simple transport.** TCP for control, UDP for media. No ICE, no SDP, no
TURN. A user opens a port (or port-forwards) and runs a server.
- **Extensible from day one.** Every message rides in a versioned envelope; capabilities
are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio,
and moderation slot in without breaking older clients.
- **Real-time correctness.** The audio thread never blocks, never allocates, never takes a
lock. Network and audio communicate through lock-free ring buffers.
- **Encrypted, always.** There is no unencrypted mode to misconfigure. The server has no
plaintext listener; encryption is on because it can't be turned off. And it's free to the
operator — the server self-provisions its key/cert on first run.
- **Stupid-easy to self-host.** The target reaction is "oh, I (or my agent) can stand this up
in a few minutes." One `docker run`, or one static binary, or a plain `cmake --build` — no
certificate wrangling, no external database, sane defaults out of the box. Permissive
licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source.
## Glossary
- **Core** — `libvoicecat`, the shared C++ library.
- **Control channel** — the TCP/TLS connection carrying protobuf messages.
- **Media channel** — the UDP connection carrying voice frames.
- **Stream** — one audio source from one user (e.g. mic, screen audio, second device). A
user may publish several streams at once; each is independently controllable.
- **Channel** — a room in the channel tree. Voice is scoped to a channel.
- **Session** — an authenticated connection; ties a TCP control channel to a UDP 5-tuple.
- **SFU relay** — the server forwards Opus packets between channel members without decoding
them (selective forwarding, no transcoding).