Closes the three items PROGRESS.md's M3 section explicitly carried forward as out of scope: - Device enumeration (vc_list_devices) + input device selection (vc_set_input_device), backed by AudioEngine::enumerate_devices() via miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded ma_device_id strings. - VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk). webrtc-audio-processing (the originally-planned APM) has no working Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard abseil-cpp dependency), so VAD is a new lightweight, dependency-free energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it. - True stereo playback: AudioEngine's mixer and output device now carry stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded stereo streams to mono before mixing. - Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via miniaudio's loopback device type), replacing test-only injection as the production capture path. Also: vccli gains --list-devices, --input-device, --input-mode, and --share-screen-audio flags, plus a stdin command loop (ptt on/off, mode vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all four items (ABI-level + a white-box AudioEngine stereo-mix check). Docs updated to match: voice.md, roadmap.md (decision-log entry superseding the original webrtc-audio-processing choice), tech-stack.md, README.md, architecture.md, CLAUDE.md, PROGRESS.md. Still explicitly out of scope, documented not silently dropped: real webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS SCREEN_AUDIO capture, process-specific loopback, and a pre-existing RT-thread rule violation in the capture path that predates this work. Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive standalone runs; manually verified live (vccli --list-devices against real hardware, vccli --voice --input-mode vad streaming without incident). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
69 lines
5.3 KiB
Markdown
69 lines
5.3 KiB
Markdown
# VoiceCat — Design Documentation
|
|
|
|
VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of
|
|
classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a
|
|
single server you own and run. It deliberately avoids WebRTC. The media path is plain
|
|
**UDP**, the control path is plain **TCP**, and both are encrypted.
|
|
|
|
This folder is the design spec. No code yet — these documents define the architecture,
|
|
the wire protocol, the audio pipeline, the security model, and the dependency list, so
|
|
that implementation can start from a shared, agreed plan.
|
|
|
|
## Decisions locked so far
|
|
|
|
| Area | Decision |
|
|
|------|----------|
|
|
| Code architecture | **Shared C++ core** (`libvoicecat`) consumed by native UIs over a **C ABI**. Server reuses the same core. |
|
|
| Native clients | macOS/iOS in **Swift** (SwiftUI; Swift↔C++ interop), Windows in **C#** (`LibraryImport` P/Invoke). |
|
|
| Control transport | **TCP + TLS 1.3** (mbedTLS) |
|
|
| Media transport | **UDP** secured by **TLS-exported keys + ChaCha20-Poly1305 AEAD** — mandatory, no plaintext mode (see [security.md](security.md)) |
|
|
| Crypto libraries | **mbedTLS** (TLS 1.3) + **libsodium** (AEAD, Argon2id, Ed25519) — both permissive, **no GPL/LGPL anywhere** |
|
|
| Voice codec | **Opus** (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX |
|
|
| Audio DSP | **webrtc-audio-processing (APM)** was the plan for AEC/NS/AGC/VAD, but has no working Windows/MSVC build upstream — v1 ships a lightweight energy/RMS VAD instead, no AEC/NS/AGC yet (see [tech-stack.md](tech-stack.md), [roadmap.md](roadmap.md) §2). NR is **two-sided**: sender can denoise, and each listener can denoise a *specific* other user locally — this plumbing exists but is currently inert pending a real DSP backend. Input gate supports **VAD and PTT**, client-configurable. |
|
|
| Identity | **Guests + admin-provisioned local accounts** (Argon2id, SQLite). No self-serve registration; guests toggleable per server. |
|
|
| Text | **Ephemeral** — live relay, no server-side history in v1. |
|
|
| Serialization | **Protocol Buffers** for the control plane; **custom binary** for voice frames |
|
|
| Deployment | **One `docker run`, one static binary, or `cmake --build`** — zero-config, secured by default (see [deployment.md](deployment.md)) |
|
|
|
|
## Document index
|
|
|
|
1. [architecture.md](architecture.md) — System layers, the shared core, threading model, the C ABI, server design.
|
|
2. [protocol.md](protocol.md) — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility.
|
|
3. [voice.md](voice.md) — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling.
|
|
4. [security.md](security.md) — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model.
|
|
5. [tech-stack.md](tech-stack.md) — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes.
|
|
6. [deployment.md](deployment.md) — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults.
|
|
7. [roadmap.md](roadmap.md) — Milestones, what ships when, and the list of open questions still to resolve.
|
|
|
|
## Design principles
|
|
|
|
- **One core, many faces.** Protocol, crypto, Opus, networking, jitter buffering, and
|
|
mixing live once in C++. UIs are thin. This keeps behavior identical across platforms
|
|
and the security-sensitive code reviewed in a single place.
|
|
- **Boringly simple transport.** TCP for control, UDP for media. No ICE, no SDP, no
|
|
TURN. A user opens a port (or port-forwards) and runs a server.
|
|
- **Extensible from day one.** Every message rides in a versioned envelope; capabilities
|
|
are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio,
|
|
and moderation slot in without breaking older clients.
|
|
- **Real-time correctness.** The audio thread never blocks, never allocates, never takes a
|
|
lock. Network and audio communicate through lock-free ring buffers.
|
|
- **Encrypted, always.** There is no unencrypted mode to misconfigure. The server has no
|
|
plaintext listener; encryption is on because it can't be turned off. And it's free to the
|
|
operator — the server self-provisions its key/cert on first run.
|
|
- **Stupid-easy to self-host.** The target reaction is "oh, I (or my agent) can stand this up
|
|
in a few minutes." One `docker run`, or one static binary, or a plain `cmake --build` — no
|
|
certificate wrangling, no external database, sane defaults out of the box. Permissive
|
|
licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source.
|
|
|
|
## Glossary
|
|
|
|
- **Core** — `libvoicecat`, the shared C++ library.
|
|
- **Control channel** — the TCP/TLS connection carrying protobuf messages.
|
|
- **Media channel** — the UDP connection carrying voice frames.
|
|
- **Stream** — one audio source from one user (e.g. mic, screen audio, second device). A
|
|
user may publish several streams at once; each is independently controllable.
|
|
- **Channel** — a room in the channel tree. Voice is scoped to a channel.
|
|
- **Session** — an authenticated connection; ties a TCP control channel to a UDP 5-tuple.
|
|
- **SFU relay** — the server forwards Opus packets between channel members without decoding
|
|
them (selective forwarding, no transcoding).
|