Establish the design spec in docs/ before implementation: - README: overview, locked decisions, principles, glossary - architecture: shared C++ core + C ABI, native UIs (Swift/C#), threading model, server design (SFU relay) - protocol: TCP/TLS control plane, protobuf Envelope + message catalog, connection lifecycle, extensibility rules - voice: UDP media frame format, per-channel Opus config, multi-stream model, two-sided noise reduction, VAD/PTT, jitter buffer, iOS ReplayKit screen-audio - security: mandatory encryption (TLS 1.3 + exported-key AEAD), TOFU server identity, admin-provisioned accounts, anti-replay - tech-stack: permissive-only deps (mbedTLS, libsodium, opus, miniaudio, webrtc-apm, ...), build tooling, no GPL/LGPL - deployment: zero-config self-host (Docker / binary / source) - roadmap: M0-M5 milestones, resolved decisions Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
69 lines
5.1 KiB
Markdown
69 lines
5.1 KiB
Markdown
# VoiceCat — Design Documentation
|
|
|
|
VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of
|
|
classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a
|
|
single server you own and run. It deliberately avoids WebRTC. The media path is plain
|
|
**UDP**, the control path is plain **TCP**, and both are encrypted.
|
|
|
|
This folder is the design spec. No code yet — these documents define the architecture,
|
|
the wire protocol, the audio pipeline, the security model, and the dependency list, so
|
|
that implementation can start from a shared, agreed plan.
|
|
|
|
## Decisions locked so far
|
|
|
|
| Area | Decision |
|
|
|------|----------|
|
|
| Code architecture | **Shared C++ core** (`libvoicecat`) consumed by native UIs over a **C ABI**. Server reuses the same core. |
|
|
| Native clients | macOS/iOS in **Swift** (SwiftUI; Swift↔C++ interop), Windows in **C#** (`LibraryImport` P/Invoke). |
|
|
| Control transport | **TCP + TLS 1.3** (mbedTLS) |
|
|
| Media transport | **UDP** secured by **TLS-exported keys + ChaCha20-Poly1305 AEAD** — mandatory, no plaintext mode (see [security.md](security.md)) |
|
|
| Crypto libraries | **mbedTLS** (TLS 1.3) + **libsodium** (AEAD, Argon2id, Ed25519) — both permissive, **no GPL/LGPL anywhere** |
|
|
| Voice codec | **Opus** (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX |
|
|
| Audio DSP | **webrtc-audio-processing (APM)** — AEC/NS/AGC/VAD. NR is **two-sided**: sender can denoise, and each listener can denoise a *specific* other user locally. Input gate supports **VAD and PTT**, client-configurable. |
|
|
| Identity | **Guests + admin-provisioned local accounts** (Argon2id, SQLite). No self-serve registration; guests toggleable per server. |
|
|
| Text | **Ephemeral** — live relay, no server-side history in v1. |
|
|
| Serialization | **Protocol Buffers** for the control plane; **custom binary** for voice frames |
|
|
| Deployment | **One `docker run`, one static binary, or `cmake --build`** — zero-config, secured by default (see [deployment.md](deployment.md)) |
|
|
|
|
## Document index
|
|
|
|
1. [architecture.md](architecture.md) — System layers, the shared core, threading model, the C ABI, server design.
|
|
2. [protocol.md](protocol.md) — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility.
|
|
3. [voice.md](voice.md) — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling.
|
|
4. [security.md](security.md) — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model.
|
|
5. [tech-stack.md](tech-stack.md) — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes.
|
|
6. [deployment.md](deployment.md) — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults.
|
|
7. [roadmap.md](roadmap.md) — Milestones, what ships when, and the list of open questions still to resolve.
|
|
|
|
## Design principles
|
|
|
|
- **One core, many faces.** Protocol, crypto, Opus, networking, jitter buffering, and
|
|
mixing live once in C++. UIs are thin. This keeps behavior identical across platforms
|
|
and the security-sensitive code reviewed in a single place.
|
|
- **Boringly simple transport.** TCP for control, UDP for media. No ICE, no SDP, no
|
|
TURN. A user opens a port (or port-forwards) and runs a server.
|
|
- **Extensible from day one.** Every message rides in a versioned envelope; capabilities
|
|
are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio,
|
|
and moderation slot in without breaking older clients.
|
|
- **Real-time correctness.** The audio thread never blocks, never allocates, never takes a
|
|
lock. Network and audio communicate through lock-free ring buffers.
|
|
- **Encrypted, always.** There is no unencrypted mode to misconfigure. The server has no
|
|
plaintext listener; encryption is on because it can't be turned off. And it's free to the
|
|
operator — the server self-provisions its key/cert on first run.
|
|
- **Stupid-easy to self-host.** The target reaction is "oh, I (or my agent) can stand this up
|
|
in a few minutes." One `docker run`, or one static binary, or a plain `cmake --build` — no
|
|
certificate wrangling, no external database, sane defaults out of the box. Permissive
|
|
licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source.
|
|
|
|
## Glossary
|
|
|
|
- **Core** — `libvoicecat`, the shared C++ library.
|
|
- **Control channel** — the TCP/TLS connection carrying protobuf messages.
|
|
- **Media channel** — the UDP connection carrying voice frames.
|
|
- **Stream** — one audio source from one user (e.g. mic, screen audio, second device). A
|
|
user may publish several streams at once; each is independently controllable.
|
|
- **Channel** — a room in the channel tree. Voice is scoped to a channel.
|
|
- **Session** — an authenticated connection; ties a TCP control channel to a UDP 5-tuple.
|
|
- **SFU relay** — the server forwards Opus packets between channel members without decoding
|
|
them (selective forwarding, no transcoding).
|