2026-06-15 20:47:09 +02:00
# VoiceCat — Design Documentation
VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of
classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a
single server you own and run. It deliberately avoids WebRTC. The media path is plain
**UDP** , the control path is plain **TCP** , and both are encrypted.
This folder is the design spec. No code yet — these documents define the architecture,
the wire protocol, the audio pipeline, the security model, and the dependency list, so
that implementation can start from a shared, agreed plan.
## Decisions locked so far
| Area | Decision |
|------|----------|
| Code architecture | **Shared C++ core** (`libvoicecat` ) consumed by native UIs over a **C ABI** . Server reuses the same core. |
2026-09-19 00:39:04 +02:00
| Native clients | Managed **C#** WinForms and AppKit replacements are implemented; Swift macOS remains the migration/release oracle pending manual cutover gates, and iOS remains Swift. |
2026-06-15 20:47:09 +02:00
| Control transport | **TCP + TLS 1.3** (mbedTLS) |
| Media transport | **UDP** secured by **TLS-exported keys + ChaCha20-Poly1305 AEAD** — mandatory, no plaintext mode (see [security.md ](security.md )) |
| Crypto libraries | **mbedTLS** (TLS 1.3) + **libsodium** (AEAD, Argon2id, Ed25519) — both permissive, **no GPL/LGPL anywhere** |
| Voice codec | **Opus** (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX |
2026-06-16 16:11:52 +02:00
| Audio DSP | **webrtc-audio-processing (APM)** was the plan for AEC/NS/AGC/VAD, but has no working Windows/MSVC build upstream — v1 ships a lightweight energy/RMS VAD instead, no AEC/NS/AGC yet (see [tech-stack.md ](tech-stack.md ), [roadmap.md ](roadmap.md ) §2). NR is **two-sided** : sender can denoise, and each listener can denoise a *specific* other user locally — this plumbing exists but is currently inert pending a real DSP backend. Input gate supports **VAD and PTT** , client-configurable. |
2026-06-15 20:47:09 +02:00
| Identity | **Guests + admin-provisioned local accounts** (Argon2id, SQLite). No self-serve registration; guests toggleable per server. |
| Text | **Ephemeral** — live relay, no server-side history in v1. |
| Serialization | **Protocol Buffers** for the control plane; **custom binary** for voice frames |
| Deployment | **One `docker run`, one static binary, or `cmake --build`** — zero-config, secured by default (see [deployment.md ](deployment.md )) |
## Document index
1. [architecture.md ](architecture.md ) — System layers, the shared core, threading model, the C ABI, server design.
2. [protocol.md ](protocol.md ) — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility.
3. [voice.md ](voice.md ) — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling.
4. [security.md ](security.md ) — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model.
5. [tech-stack.md ](tech-stack.md ) — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes.
6. [deployment.md ](deployment.md ) — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults.
7. [roadmap.md ](roadmap.md ) — Milestones, what ships when, and the list of open questions still to resolve.
2026-09-19 00:39:04 +02:00
8. [porting-to-dotnet.md ](porting-to-dotnet.md ) — **Active migration.** Step-by-step plan and implementation checkpoints for replacing the C++ core, C++ server, and replaceable Swift clients with .NET 10 / C#. Dependency map, TLS exporter, real-time-audio design, and phased cutover gates.
2026-06-15 20:47:09 +02:00
## Design principles
- **One core, many faces.** Protocol, crypto, Opus, networking, jitter buffering, and
mixing live once in C++. UIs are thin. This keeps behavior identical across platforms
and the security-sensitive code reviewed in a single place.
- **Boringly simple transport.** TCP for control, UDP for media. No ICE, no SDP, no
TURN. A user opens a port (or port-forwards) and runs a server.
- **Extensible from day one.** Every message rides in a versioned envelope; capabilities
are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio,
and moderation slot in without breaking older clients.
- **Real-time correctness.** The audio thread never blocks, never allocates, never takes a
lock. Network and audio communicate through lock-free ring buffers.
- **Encrypted, always.** There is no unencrypted mode to misconfigure. The server has no
plaintext listener; encryption is on because it can't be turned off. And it's free to the
operator — the server self-provisions its key/cert on first run.
- **Stupid-easy to self-host.** The target reaction is "oh, I (or my agent) can stand this up
in a few minutes." One `docker run` , or one static binary, or a plain `cmake --build` — no
certificate wrangling, no external database, sane defaults out of the box. Permissive
licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source.
## Glossary
- **Core** — `libvoicecat` , the shared C++ library.
- **Control channel** — the TCP/TLS connection carrying protobuf messages.
- **Media channel** — the UDP connection carrying voice frames.
- **Stream** — one audio source from one user (e.g. mic, screen audio, second device). A
user may publish several streams at once; each is independently controllable.
- **Channel** — a room in the channel tree. Voice is scoped to a channel.
- **Session** — an authenticated connection; ties a TCP control channel to a UDP 5-tuple.
- **SFU relay** — the server forwards Opus packets between channel members without decoding
them (selective forwarding, no transcoding).