Files
voice-cat/docs
Talon bad9c7533a
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
feat(audio): real noise suppression via vendored RNNoise (send + receive)
The two-sided NR plumbing (RemoteStream::recv_ns + the per-listener
vc_set_remote_stream noise_reduction toggle) was wired but inert:
ApmProcessor::create() returned a no-op passthrough, because the
originally-planned webrtc-audio-processing has no working Windows/macOS
build. Drop in RNNoise as the real backend behind the same ApmProcessor
interface, lighting up both NR paths.

- Vendor RNNoise (BSD-3 + CC0) at third_party/rnnoise/ — the vcpkg port
  is !windows !arm, so it can't cover our primary targets. Shrunk int8
  model (78MB -> 11.7MB via upstream scripts/shrink_model.sh), built as a
  standalone C static lib with no RTCD (portable scalar path on x86,
  auto-NEON on arm64) under -DDISABLE_DEBUG_FLOAT. Model is baked in
  (rnnoise_create(NULL)); no runtime file.
- New RnnoiseProcessor (core/src/audio/apm_processor.cpp) selected by
  ApmProcessor::create() when VOICECAT_HAS_NS. Mono/48kHz/480-sample;
  our clock is fixed 48kHz and Opus frame sizes are multiples of 480, so
  no resampling. RT-safe: allocates at construction, lock-free in the
  capture/playback callbacks.
- Receive-side: lit up via the factory; gated to mono streams (a stereo
  stream is a screen-audio share, not voice).
- Send-side (new): vc_set_input_noise_reduction(client, enable) ABI +
  vc_client::mic_ns_, run before input gain/VAD in on_capture_frame. A
  stereo mic is downmixed to mono ONLY when NR is on — with NR off a
  stereo mic keeps full stereo (never collapse mic quality unasked).
- Enable C as a project language for the vendored lib.
- New noise_suppression test: white noise through ApmProcessor::create()
  drops ~99.9% RMS. ctest --preset dev green, 28/28. windows-client DLL
  builds clean with vc_set_input_noise_reduction exported, system-only deps.
- Docs synced: voice.md §10, tech-stack.md §1/§5, third_party/README.md,
  vcpkg.json note, PROGRESS.md, CLAUDE.md.

Client on/off UI toggles (Windows/macOS/iOS) are the remaining follow-up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 13:30:54 +02:00
..

VoiceCat — Design Documentation

VoiceCat is a self-hosted, server-based voice and text chat system in the spirit of classic TeamSpeak / Mumble: channel-based voice, channel and private text chat, and a single server you own and run. It deliberately avoids WebRTC. The media path is plain UDP, the control path is plain TCP, and both are encrypted.

This folder is the design spec. No code yet — these documents define the architecture, the wire protocol, the audio pipeline, the security model, and the dependency list, so that implementation can start from a shared, agreed plan.

Decisions locked so far

Area Decision
Code architecture Shared C++ core (libvoicecat) consumed by native UIs over a C ABI. Server reuses the same core.
Native clients macOS/iOS in Swift (SwiftUI; Swift↔C++ interop), Windows in C# (LibraryImport P/Invoke).
Control transport TCP + TLS 1.3 (mbedTLS)
Media transport UDP secured by TLS-exported keys + ChaCha20-Poly1305 AEAD — mandatory, no plaintext mode (see security.md)
Crypto libraries mbedTLS (TLS 1.3) + libsodium (AEAD, Argon2id, Ed25519) — both permissive, no GPL/LGPL anywhere
Voice codec Opus (libopus 1.6), per-channel configurable mono/stereo, bitrate, frame size, FEC/DTX
Audio DSP webrtc-audio-processing (APM) was the plan for AEC/NS/AGC/VAD, but has no working Windows/MSVC build upstream — v1 ships a lightweight energy/RMS VAD instead, no AEC/NS/AGC yet (see tech-stack.md, roadmap.md §2). NR is two-sided: sender can denoise, and each listener can denoise a specific other user locally — this plumbing exists but is currently inert pending a real DSP backend. Input gate supports VAD and PTT, client-configurable.
Identity Guests + admin-provisioned local accounts (Argon2id, SQLite). No self-serve registration; guests toggleable per server.
Text Ephemeral — live relay, no server-side history in v1.
Serialization Protocol Buffers for the control plane; custom binary for voice frames
Deployment One docker run, one static binary, or cmake --build — zero-config, secured by default (see deployment.md)

Document index

  1. architecture.md — System layers, the shared core, threading model, the C ABI, server design.
  2. protocol.md — The control protocol: framing, connection lifecycle, the full message catalog, encoding, versioning/extensibility.
  3. voice.md — The UDP media protocol: voice frame format, Opus configuration, multi-stream model, jitter buffer, packet-loss handling.
  4. security.md — Mandatory encryption (TLS 1.3 + exported-key media AEAD), server identity (TOFU), authentication, accounts, anti-replay, threat model.
  5. tech-stack.md — Concrete libraries with versions and rationale, the permissive-license rule, build tooling, per-platform notes.
  6. deployment.md — The "set it up in a few minutes" story: Docker, single binary, source build, zero-config defaults.
  7. roadmap.md — Milestones, what ships when, and the list of open questions still to resolve.

Design principles

  • One core, many faces. Protocol, crypto, Opus, networking, jitter buffering, and mixing live once in C++. UIs are thin. This keeps behavior identical across platforms and the security-sensitive code reviewed in a single place.
  • Boringly simple transport. TCP for control, UDP for media. No ICE, no SDP, no TURN. A user opens a port (or port-forwards) and runs a server.
  • Extensible from day one. Every message rides in a versioned envelope; capabilities are negotiated at connect time; unknown fields are ignored. File transfer, screen-audio, and moderation slot in without breaking older clients.
  • Real-time correctness. The audio thread never blocks, never allocates, never takes a lock. Network and audio communicate through lock-free ring buffers.
  • Encrypted, always. There is no unencrypted mode to misconfigure. The server has no plaintext listener; encryption is on because it can't be turned off. And it's free to the operator — the server self-provisions its key/cert on first run.
  • Stupid-easy to self-host. The target reaction is "oh, I (or my agent) can stand this up in a few minutes." One docker run, or one static binary, or a plain cmake --build — no certificate wrangling, no external database, sane defaults out of the box. Permissive licenses only (no GPL/LGPL) so it can be redistributed freely, including closed-source.

Glossary

  • Corelibvoicecat, the shared C++ library.
  • Control channel — the TCP/TLS connection carrying protobuf messages.
  • Media channel — the UDP connection carrying voice frames.
  • Stream — one audio source from one user (e.g. mic, screen audio, second device). A user may publish several streams at once; each is independently controllable.
  • Channel — a room in the channel tree. Voice is scoped to a channel.
  • Session — an authenticated connection; ties a TCP control channel to a UDP 5-tuple.
  • SFU relay — the server forwards Opus packets between channel members without decoding them (selective forwarding, no transcoding).