docs: initial design baseline for VoiceCat voice/text chat
Establish the design spec in docs/ before implementation:
- README: overview, locked decisions, principles, glossary
- architecture: shared C++ core + C ABI, native UIs (Swift/C#),
threading model, server design (SFU relay)
- protocol: TCP/TLS control plane, protobuf Envelope + message
catalog, connection lifecycle, extensibility rules
- voice: UDP media frame format, per-channel Opus config,
multi-stream model, two-sided noise reduction, VAD/PTT,
jitter buffer, iOS ReplayKit screen-audio
- security: mandatory encryption (TLS 1.3 + exported-key AEAD),
TOFU server identity, admin-provisioned accounts, anti-replay
- tech-stack: permissive-only deps (mbedTLS, libsodium, opus,
miniaudio, webrtc-apm, ...), build tooling, no GPL/LGPL
- deployment: zero-config self-host (Docker / binary / source)
- roadmap: M0-M5 milestones, resolved decisions
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:47:09 +02:00
# VoiceCat — Design Documentation
VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of
classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a
single server you own and run. It deliberately avoids WebRTC. The media path is plain
**UDP**, the control path is plain **TCP** , and both are encrypted.
This folder is the design spec. No code yet — these documents define the architecture,
the wire protocol, the audio pipeline, the security model, and the dependency list, so
that implementation can start from a shared, agreed plan.
## Decisions locked so far
| Area | Decision |
|------|----------|
| Code architecture | **Shared C++ core** (`libvoicecat` ) consumed by native UIs over a **C ABI** . Server reuses the same core. |
| Native clients | macOS/iOS in **Swift** (SwiftUI; Swift↔C++ interop), Windows in **C#** (`LibraryImport` P/Invoke). |
| Control transport | **TCP + TLS 1.3** (mbedTLS) |
| Media transport | **UDP** secured by **TLS-exported keys + ChaCha20-Poly1305 AEAD** — mandatory, no plaintext mode (see [security.md ](security.md )) |
| Crypto libraries | **mbedTLS** (TLS 1.3) + **libsodium** (AEAD, Argon2id, Ed25519) — both permissive, **no GPL/LGPL anywhere** |
| Voice codec | **Opus** (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX |
feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as
out of scope:
- Device enumeration (vc_list_devices) + input device selection
(vc_set_input_device), backed by AudioEngine::enumerate_devices() via
miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded
ma_device_id strings.
- VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk).
webrtc-audio-processing (the originally-planned APM) has no working
Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard
abseil-cpp dependency), so VAD is a new lightweight, dependency-free
energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor
interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it.
- True stereo playback: AudioEngine's mixer and output device now carry
stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded
stereo streams to mono before mixing.
- Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via
miniaudio's loopback device type), replacing test-only injection as the
production capture path.
Also: vccli gains --list-devices, --input-device, --input-mode, and
--share-screen-audio flags, plus a stdin command loop (ptt on/off, mode
vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all
four items (ABI-level + a white-box AudioEngine stereo-mix check).
Docs updated to match: voice.md, roadmap.md (decision-log entry superseding
the original webrtc-audio-processing choice), tech-stack.md, README.md,
architecture.md, CLAUDE.md, PROGRESS.md.
Still explicitly out of scope, documented not silently dropped: real
webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS
SCREEN_AUDIO capture, process-specific loopback, and a pre-existing
RT-thread rule violation in the capture path that predates this work.
Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev
and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive
standalone runs; manually verified live (vccli --list-devices against real
hardware, vccli --voice --input-mode vad streaming without incident).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 16:11:52 +02:00
| Audio DSP | **webrtc-audio-processing (APM)** was the plan for AEC/NS/AGC/VAD, but has no working Windows/MSVC build upstream — v1 ships a lightweight energy/RMS VAD instead, no AEC/NS/AGC yet (see [tech-stack.md ](tech-stack.md ), [roadmap.md ](roadmap.md ) §2). NR is **two-sided** : sender can denoise, and each listener can denoise a *specific* other user locally — this plumbing exists but is currently inert pending a real DSP backend. Input gate supports **VAD and PTT** , client-configurable. |
docs: initial design baseline for VoiceCat voice/text chat
Establish the design spec in docs/ before implementation:
- README: overview, locked decisions, principles, glossary
- architecture: shared C++ core + C ABI, native UIs (Swift/C#),
threading model, server design (SFU relay)
- protocol: TCP/TLS control plane, protobuf Envelope + message
catalog, connection lifecycle, extensibility rules
- voice: UDP media frame format, per-channel Opus config,
multi-stream model, two-sided noise reduction, VAD/PTT,
jitter buffer, iOS ReplayKit screen-audio
- security: mandatory encryption (TLS 1.3 + exported-key AEAD),
TOFU server identity, admin-provisioned accounts, anti-replay
- tech-stack: permissive-only deps (mbedTLS, libsodium, opus,
miniaudio, webrtc-apm, ...), build tooling, no GPL/LGPL
- deployment: zero-config self-host (Docker / binary / source)
- roadmap: M0-M5 milestones, resolved decisions
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:47:09 +02:00
| Identity | **Guests + admin-provisioned local accounts** (Argon2id, SQLite). No self-serve registration; guests toggleable per server. |
| Text | **Ephemeral** — live relay, no server-side history in v1. |
| Serialization | **Protocol Buffers** for the control plane; **custom binary** for voice frames |
| Deployment | **One `docker run`, one static binary, or `cmake --build`** — zero-config, secured by default (see [deployment.md ](deployment.md )) |
## Document index
1. [architecture.md ](architecture.md ) — System layers, the shared core, threading model, the C ABI, server design.
2. [protocol.md ](protocol.md ) — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility.
3. [voice.md ](voice.md ) — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling.
4. [security.md ](security.md ) — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model.
5. [tech-stack.md ](tech-stack.md ) — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes.
6. [deployment.md ](deployment.md ) — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults.
7. [roadmap.md ](roadmap.md ) — Milestones, what ships when, and the list of open questions still to resolve.
docs: add the .NET/C# porting plan
Proposal for replacing the C++ core, C++ server, and Swift macOS/iOS
clients with a single .NET 10 / C# codebase.
Covers the dependency map (8 vcpkg deps + 1 vendored -> 3 native libs),
the real-time-audio design, per-client strategy, a test-porting plan for
all 29 ctest cases, doc-sync work, an 11-phase migration, and a risk
register.
Two findings drive the shape of the plan:
- SslStream has no RFC 5705 keying-material exporter, which the media
AEAD key derivation depends on (docs/security.md 2). The API is an
unapproved proposal and SChannel structurally cannot export secrets.
Recommends BouncyCastle's managed TLS 1.3 stack, which does implement
the exporter and keeps the wire format byte-compatible with the C++
implementation -- so the existing tree stays usable as a conformance
oracle throughout the port. Protocol-v3 in-band media keys documented
as the fallback.
- The iOS ReplayKit broadcast upload extension stays in Swift: 50 MB
jetsam cap plus an unsupported extension type in .NET for iOS, and it
already doesn't link the core. Leaves one Swift file plus the shared
App Group ring.
Indexed in docs/README.md. Nothing here is implemented yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 03:19:59 +02:00
8. [porting-to-dotnet.md ](porting-to-dotnet.md ) — **Proposal.** Step-by-step plan to replace the C++ core, C++ server, and Swift clients with a single .NET 10 / C# codebase. Dependency map, the TLS-exporter blocker, real-time-audio design, phased migration.
docs: initial design baseline for VoiceCat voice/text chat
Establish the design spec in docs/ before implementation:
- README: overview, locked decisions, principles, glossary
- architecture: shared C++ core + C ABI, native UIs (Swift/C#),
threading model, server design (SFU relay)
- protocol: TCP/TLS control plane, protobuf Envelope + message
catalog, connection lifecycle, extensibility rules
- voice: UDP media frame format, per-channel Opus config,
multi-stream model, two-sided noise reduction, VAD/PTT,
jitter buffer, iOS ReplayKit screen-audio
- security: mandatory encryption (TLS 1.3 + exported-key AEAD),
TOFU server identity, admin-provisioned accounts, anti-replay
- tech-stack: permissive-only deps (mbedTLS, libsodium, opus,
miniaudio, webrtc-apm, ...), build tooling, no GPL/LGPL
- deployment: zero-config self-host (Docker / binary / source)
- roadmap: M0-M5 milestones, resolved decisions
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:47:09 +02:00
## Design principles
- **One core, many faces.** Protocol, crypto, Opus, networking, jitter buffering, and
mixing live once in C++. UIs are thin. This keeps behavior identical across platforms
and the security-sensitive code reviewed in a single place.
- **Boringly simple transport.** TCP for control, UDP for media. No ICE, no SDP, no
TURN. A user opens a port (or port-forwards) and runs a server.
- **Extensible from day one.** Every message rides in a versioned envelope; capabilities
are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio,
and moderation slot in without breaking older clients.
- **Real-time correctness.** The audio thread never blocks, never allocates, never takes a
lock. Network and audio communicate through lock-free ring buffers.
- **Encrypted, always.** There is no unencrypted mode to misconfigure. The server has no
plaintext listener; encryption is on because it can't be turned off. And it's free to the
operator — the server self-provisions its key/cert on first run.
- **Stupid-easy to self-host.** The target reaction is "oh, I (or my agent) can stand this up
in a few minutes." One `docker run` , or one static binary, or a plain `cmake --build` — no
certificate wrangling, no external database, sane defaults out of the box. Permissive
licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source.
## Glossary
- **Core** — `libvoicecat` , the shared C++ library.
- **Control channel** — the TCP/TLS connection carrying protobuf messages.
- **Media channel** — the UDP connection carrying voice frames.
- **Stream** — one audio source from one user (e.g. mic, screen audio, second device). A
user may publish several streams at once; each is independently controllable.
- **Channel** — a room in the channel tree. Voice is scoped to a channel.
- **Session** — an authenticated connection; ties a TCP control channel to a UDP 5-tuple.
- **SFU relay** — the server forwards Opus packets between channel members without decoding
them (selective forwarding, no transcoding).