# VoiceCat — Design Documentation VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a single server you own and run. It deliberately avoids WebRTC. The media path is plain **UDP**, the control path is plain **TCP**, and both are encrypted. This folder is the design spec. No code yet — these documents define the architecture, the wire protocol, the audio pipeline, the security model, and the dependency list, so that implementation can start from a shared, agreed plan. ## Decisions locked so far | Area | Decision | |------|----------| | Code architecture | **Shared C++ core** (`libvoicecat`) consumed by native UIs over a **C ABI**. Server reuses the same core. | | Native clients | macOS/iOS in **Swift** (SwiftUI; Swift↔C++ interop), Windows in **C#** (`LibraryImport` P/Invoke). | | Control transport | **TCP + TLS 1.3** (mbedTLS) | | Media transport | **UDP** secured by **TLS-exported keys + ChaCha20-Poly1305 AEAD** — mandatory, no plaintext mode (see [security.md](security.md)) | | Crypto libraries | **mbedTLS** (TLS 1.3) + **libsodium** (AEAD, Argon2id, Ed25519) — both permissive, **no GPL/LGPL anywhere** | | Voice codec | **Opus** (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX | | Audio DSP | **webrtc-audio-processing (APM)** was the plan for AEC/NS/AGC/VAD, but has no working Windows/MSVC build upstream — v1 ships a lightweight energy/RMS VAD instead, no AEC/NS/AGC yet (see [tech-stack.md](tech-stack.md), [roadmap.md](roadmap.md) §2). NR is **two-sided**: sender can denoise, and each listener can denoise a *specific* other user locally — this plumbing exists but is currently inert pending a real DSP backend. Input gate supports **VAD and PTT**, client-configurable. | | Identity | **Guests + admin-provisioned local accounts** (Argon2id, SQLite). No self-serve registration; guests toggleable per server. | | Text | **Ephemeral** — live relay, no server-side history in v1. | | Serialization | **Protocol Buffers** for the control plane; **custom binary** for voice frames | | Deployment | **One `docker run`, one static binary, or `cmake --build`** — zero-config, secured by default (see [deployment.md](deployment.md)) | ## Document index 1. [architecture.md](architecture.md) — System layers, the shared core, threading model, the C ABI, server design. 2. [protocol.md](protocol.md) — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility. 3. [voice.md](voice.md) — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling. 4. [security.md](security.md) — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model. 5. [tech-stack.md](tech-stack.md) — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes. 6. [deployment.md](deployment.md) — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults. 7. [roadmap.md](roadmap.md) — Milestones, what ships when, and the list of open questions still to resolve. ## Design principles - **One core, many faces.** Protocol, crypto, Opus, networking, jitter buffering, and mixing live once in C++. UIs are thin. This keeps behavior identical across platforms and the security-sensitive code reviewed in a single place. - **Boringly simple transport.** TCP for control, UDP for media. No ICE, no SDP, no TURN. A user opens a port (or port-forwards) and runs a server. - **Extensible from day one.** Every message rides in a versioned envelope; capabilities are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio, and moderation slot in without breaking older clients. - **Real-time correctness.** The audio thread never blocks, never allocates, never takes a lock. Network and audio communicate through lock-free ring buffers. - **Encrypted, always.** There is no unencrypted mode to misconfigure. The server has no plaintext listener; encryption is on because it can't be turned off. And it's free to the operator — the server self-provisions its key/cert on first run. - **Stupid-easy to self-host.** The target reaction is "oh, I (or my agent) can stand this up in a few minutes." One `docker run`, or one static binary, or a plain `cmake --build` — no certificate wrangling, no external database, sane defaults out of the box. Permissive licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source. ## Glossary - **Core** — `libvoicecat`, the shared C++ library. - **Control channel** — the TCP/TLS connection carrying protobuf messages. - **Media channel** — the UDP connection carrying voice frames. - **Stream** — one audio source from one user (e.g. mic, screen audio, second device). A user may publish several streams at once; each is independently controllable. - **Channel** — a room in the channel tree. Voice is scoped to a channel. - **Session** — an authenticated connection; ties a TCP control channel to a UDP 5-tuple. - **SFU relay** — the server forwards Opus packets between channel members without decoding them (selective forwarding, no transcoding).