Three reported bugs traced to one root cause plus two missing designed features:
1. Stale users + eternal PLC hiss (root cause): ConnSession::close() silently
erased dropped users without broadcasting UserEvent::LEFT, so peers never
learned the user left and their audio engines never called remove_stream —
Opus PLC synthesized comfort noise forever. Fix: broadcast_left() helper
+ close() broadcasts LEFT before erasing.
2. PLC cap (defense-in-depth): on_playback now caps pure PLC at ~2s, then
emits digital silence so a stale stream can never hiss forever even if
remove_stream is skipped. Resets automatically on fresh packets.
3. No timeout / no ping: client never sent Ping, server had no last_seen /
reaper, so half-open connections (NAT timeout, wifi loss, sleep) left
ghost users forever. Fix: client Ping every 15s with RTT measurement,
ConnSession::last_seen bumped on every inbound TCP/UDP frame, steady_timer
reaper sweeps every 15s and drops sessions older than 45s (configurable
via server::Config).
4. UDP KEEPALIVE: client sends plaintext kFrameKeepalive every 5s; server
bumps last_seen + echoes back. Keeps NAT bindings alive and lets media
activity defer the reaper independently of TCP.
5. Graceful client disconnect: vc_disconnect() sends Disconnect{code=0} via
a flag-based io-thread exit (no double-close race); server handles
client-sent Disconnect with immediate close() + LEFT broadcast.
3 new tests: disconnect_left, plc_cap, reaper_timeout. 21/21 ctest green.
Docs: protocol.md §6/§7, voice.md §6, architecture.md §5, PROGRESS.md.
14 KiB
Architecture
1. The shared-core model
All non-UI logic lives in one C++ library, libvoicecat. The same library is linked
into every client and into the server. Platform UIs are thin and call the core through a
stable C ABI (voicecat.h).
┌───────────────────────────────────────────┐
macOS / iOS (Swift) │ │ Windows (C#)
┌──────────────────┐ │ libvoicecat (C++) │ ┌──────────────────┐
│ SwiftUI views │ │ ┌─────────────────────────────────────┐ │ │ WinForms (.NET 10│
│ AVAudioSession │──┼─▶│ C ABI (voicecat.h) │◀─┼──│ LibraryImport │
│ Swift↔C++ interop│ │ ├─────────────────────────────────────┤ │ │ P/Invoke │
└──────────────────┘ │ │ Session / Protocol state machine │ │ └──────────────────┘
│ │ Text + voice signaling │ │
Linux/macOS/Windows │ │ Audio engine: capture→encode→send, │ │
server │ │ recv→jitter→decode→mix→playback │ │
┌──────────────────┐ │ │ Codec layer (Opus 1.6) │ │
│ voicecat-server │──┼─▶│ Crypto + transport (TLS 1.3 + AEAD) │ │
│ (reuses core) │ │ │ Net I/O (Asio: TCP + UDP + timers) │ │
└──────────────────┘ │ └─────────────────────────────────────┘ │
└───────────────────────────────────────────┘
Why this shape:
- Swift (5.9+) can import C++ directly, but we still ship a C ABI because it is the
lowest-friction, most stable boundary and it is what C# needs (
LibraryImport/ P/Invoke). One ABI serves both. - The server is not a separate codebase. It links the same protocol, crypto, and Opus code as the client, so framing/encryption can never drift between the two ends.
2. Layered design inside the core
From the OS up:
| Layer | Responsibility | Key deps |
|---|---|---|
| Platform I/O | Sockets, timers; audio device capture/playback | Asio, miniaudio |
| Transport | TLS 1.3 (TCP), exported-key ChaCha20-Poly1305 AEAD (UDP), framing, anti-replay | mbedTLS, libsodium |
| Codec & DSP | Opus encode/decode; APM (AEC/NS/AGC/VAD) send-side + per-user NR receive-side; resample; jitter buffer; mixer | libopus, webrtc-audio-processing, speexdsp |
| Protocol | Message (de)serialization, request/response correlation, state machine | protobuf |
| Session/domain | Channels, users, streams, permissions, text routing | — |
| C ABI façade | Handle-based API + event callbacks exposed to UIs | — |
A UI never sees a socket, an Opus packet, or a protobuf message. It sees: "connect", "join channel", "start a stream from this device", "send this text", and a stream of events ("user joined", "user is talking", "message received", "level meter = 0.4").
3. Threading model
Three classes of thread, with strict rules.
┌──────────────┐ lock-free ┌──────────────┐ lock-free ┌──────────────┐
│ Audio capture│ ──ring buffer─▶│ Net thread │ ──ring buffer─▶│Audio playback│
│ (RT, miniaudio│ │ (Asio loop) │ │ (RT, miniaudio│
│ callback) │◀───ring buffer─│ │◀───ring buffer─│ callback) │
│ capture→Opus │ │ TLS + AEAD, │ │ jitter→Opus │
│ encode │ │ route, relay │ │ decode→mix │
└──────────────┘ └──────────────┘ └──────────────┘
│
┌─────▼──────┐
│ Worker pool│ DB, Argon2id, file I/O,
│ (blocking) │ TLS handshakes, codec setup
└────────────┘
Rules:
- Audio (real-time) threads are driven by the OS audio callback. They must not allocate, lock, log, or do syscalls beyond the ring-buffer hand-off. Opus encode/decode runs here (it is allocation-free after init).
- Net thread(s) run the Asio event loop: TLS records, AEAD seal/open, protobuf parse, channel routing, jitter-buffer feed. On the server, this is where the SFU relay copies packets to subscribers.
- Worker pool absorbs anything that can block: SQLite, Argon2id verification, DNS, TLS handshake CPU, codec (re)configuration.
- Communication between audio and net is single-producer/single-consumer lock-free ring buffers (one per direction per stream). Control-plane events to the UI go through a thread-safe queue drained on the UI's terms.
4. The C ABI (voicecat.h) — shape
Handle-based, opaque pointers, C-linkage. Illustrative (final names in implementation):
typedef struct vc_client vc_client;
typedef struct {
void (*on_event)(void* user, const vc_event* ev); // state changes, messages
void (*on_level)(void* user, uint32_t stream_id, float rms); // meters (throttled)
void* user;
} vc_callbacks;
vc_client* vc_client_create(const vc_config* cfg, vc_callbacks cb);
void vc_client_destroy(vc_client*);
int vc_connect(vc_client*, const char* host, uint16_t port); // async; result via event
int vc_authenticate_guest(vc_client*, const char* nickname);
int vc_authenticate_user(vc_client*, const char* user, const char* password);
int vc_join_channel(vc_client*, uint32_t channel_id, const char* password /*nullable*/);
int vc_leave_channel(vc_client*);
// Streams (mic / screen audio / aux device)
int vc_stream_start(vc_client*, const vc_stream_desc* desc, uint32_t* out_stream_id);
int vc_stream_stop(vc_client*, uint32_t stream_id);
int vc_set_input_device(vc_client*, uint32_t stream_id, const char* device_id);
int vc_set_self_mute(vc_client*, bool mic_muted, bool deafened);
// Text
int vc_send_text(vc_client*, vc_text_scope scope, uint32_t target_id, const char* utf8);
// Enumeration helpers for UI device pickers
int vc_list_devices(vc_client*, vc_device_kind kind, vc_device_list* out);
Design notes:
- Async, event-driven. Calls return immediately; results and state changes arrive via
on_event. This maps cleanly onto SwiftUI/asyncand C#event/Taskpatterns. - The core owns audio. Capture, encode, decode, mixing, and playback happen inside the core via miniaudio. The UI only selects devices, starts/stops streams, and renders meters/state. This keeps the real-time path identical on every OS. (iOS is the one exception that needs UI-side cooperation — see below.)
- Device enumeration works pre-connect.
vc_list_devicesneeds no live session — device pickers can populate beforevc_connect.vc_device.idis an opaque, internally-encoded handle (currently a hex-encodedma_device_id) — always round-trip an id that came fromvc_list_devices/vc_get_stream_audio_config; never construct one by hand. Tolerate an empty list (a machine can legitimately have zero input or output devices). - Strings are UTF-8
const char*; ownership is explicit. Output buffers are caller-allocated or returned with a pairedvc_free.
Per-platform binding notes
- Swift / Apple. Import the C ABI via a module map; SwiftUI on top. On iOS the app
must still own
AVAudioSession(category.playAndRecord,.voiceChatmode), request mic permission, and handle interruptions/route changes — the core exposes hooks (vc_audio_suspend/vc_audio_resume) the Swift layer calls fromAVAudioSessionnotifications. Background voice and VoIP push (CallKit/PushKit) are a later milestone. - iOS screen / system-audio sharing is supported via a ReplayKit Broadcast Upload
Extension (the same mechanism Discord uses; triggered from Control Center's screen-record
button via
RPSystemBroadcastPickerView). The extension receivesRPSampleBufferType.audioApp(system/app audio) and.audioMic. We capture.audioAppfor theSCREEN_AUDIO"listen together" stream and ignore video. The extension runs in a separate process with a ~50 MB memory cap — that cap is a problem only for video frames, so audio-only stays well within budget. It links a minimal slice of the core (Opus encode + media send), shares the session/credentials with the host app through an App Group, and re-derives its own media keys. This is detailed in voice.md §9. - C# / Windows.
[LibraryImport](source-generated P/Invoke, .NET 7+) over the C ABI.[UnmanagedCallersOnly]static methods foron_event/on_levelto avoid delegate-lifetime pitfalls. UI in WinForms (.NET 10) — chosen over WinUI 3/Avalonia for its mature, predictable screen-reader (NVDA/JAWS/Narrator) UIA support (see roadmap.md §2). Events are delivered viaSystem.Threading.Channels.Channel<VoiceCatEvent>, drained by a 30msSystem.Windows.Forms.Timeron the UI thread — simpler than a message-only HWND +PostMessagewith no meaningful latency cost.VoiceCatClientHandle : SafeHandlewraps thevc_client*and guaranteesvc_client_destroyruns on GC/Dispose.
5. Server architecture
voicecat-server is a headless process linking the core.
TCP/TLS 1.3 UDP + media AEAD
│ │
┌────────▼─────────┐ ┌─────────▼──────────┐
│ Connection mgr │ │ UDP demux │
│ (accept, TLS, │ │ 5-tuple → session │
│ per-conn state) │ │ anti-replay window │
└────────┬─────────┘ └─────────┬──────────┘
│ │
┌────────▼───────────────────────────────────▼──────────┐
│ Session registry (session_id ↔ TCP conn ↔ UDP tuple) │
└────────┬───────────────────────────────┬───────────────┘
│ │
┌────────▼─────────┐ ┌──────────────┐ ┌▼─────────────────┐
│ Channel manager │ │ Text router │ │ Voice router/SFU │
│ tree, configs, │ │ channel + PM │ │ relay Opus to │
│ membership, perms│ │ │ │ channel members │
└────────┬─────────┘ └──────────────┘ └──────────────────┘
│
┌────────▼─────────┐
│ Persistence │ accounts (Argon2id), channels, bans, config
│ SQLite │
└──────────────────┘
- Voice router is a relay, not a mixer. For each incoming voice frame it looks up the sender's channel and forwards the unmodified Opus payload (restamped with the sender's user id) to every other subscribed member. No server-side decode/transcode → low CPU, low latency, and end-to-content is just Opus. Per-channel Opus params are enforced so all members are mutually decodable.
- Subscriptions. Clients implicitly subscribe to their current channel's voice; text and presence can be subscribed more broadly. This keeps fan-out bounded on big servers.
- Stateless-ish media. UDP carries no auth per packet beyond the media-AEAD session; the 5-tuple→session binding is established once via a token (see protocol.md §4).
- Keepalive reaper. An
asio::steady_timersweeps every 15 s and drops any session whoselast_seen(bumped on every inbound TCP or UDP frame) is older than 45 s. Each drop broadcastsUserEvent::LEFTso peers clean up immediately. This catches half-open connections that never produce a TCP EOF. Configurable viaserver::Config. - Single process, scalable later. v1 is one process, one machine. The session registry and router are written behind interfaces so a future build can sit them behind a shared bus for multi-node, but that is explicitly out of scope for now.
6. Repository layout (proposed)
voice-cat/
├── docs/ # this folder
├── core/ # libvoicecat (C++)
│ ├── include/voicecat.h # the C ABI
│ ├── src/{net,crypto,codec,protocol,session,audio}/
│ └── proto/ # .proto definitions (shared source of truth)
├── server/ # voicecat-server (C++, links core)
├── clients/
│ ├── apple/ # Swift package + Xcode project (macOS + iOS)
│ └── windows/ # .NET solution (C#)
├── tools/
│ └── vccli/ # headless test client (C++), for protocol bring-up
├── third_party/ # vendored / vcpkg manifest
└── CMakeLists.txt
Build is CMake with vcpkg (manifest mode) for C/C++ deps; the Apple and Windows UI projects consume the built core as a binary + headers. See tech-stack.md.