Files
voice-cat/docs/architecture.md
Talon fdcc84fb42 fix(ios): stereo mic + A2DP output, add vc_audio_restart ABI
Diagnosed by comparing against TeamTalk5 (Client/iTeamTalk), which
achieves stereo mic + A2DP output. Five fixes:

1. configureStereoCapture now calls setPreferredInput +
   setInputDataSource (mirroring TeamTalk5's SoundDevicesModel).
   Previously omitted based on incorrect diagnosis that
   setPreferredInput collapsed A2DP — the real culprit was
   setPreferredInputNumberOfChannels(2), which neither project uses.

2. New C ABI: vc_audio_restart (full stop + re-init, unlike
   suspend/resume which only stop/start). Swift wrapper added.
   The withAudioSuspend wrapper that used it was removed after
   on-device testing showed it killed all audio (including
   VoiceOver) when switching presets — the core's
   set_capture_channels handles engine restart internally.

3. Bluetooth options: Voice Chat preset now includes BOTH
   .allowBluetoothHFP AND .allowBluetoothA2DP (matching TeamTalk5's
   UtilSound.swift:228). Previously HFP-only blocked A2DP headphones.

4. Capture channels now reset when switching stereo→mono via
   selectCaptureChannels/applyPreset. AudioSessionManager tracks
   activeMicStreamId (set by SessionState on join/leave voice).

5. Docs synced: voice.md, tech-stack.md, architecture.md,
   PROGRESS.md. Removed stale setPreferredInputNumberOfChannels(2)
   references.

Verified: ctest --preset dev 21/21 green, iOS client builds.
Stereo mic + A2DP output still needs on-device debugging — the
core recipe is correct but iOS 26 route behavior requires
hands-on testing with a debugger.
2026-06-19 16:58:21 +02:00

16 KiB

Architecture

1. The shared-core model

All non-UI logic lives in one C++ library, libvoicecat. The same library is linked into every client and into the server. Platform UIs are thin and call the core through a stable C ABI (voicecat.h).

                         ┌───────────────────────────────────────────┐
   macOS / iOS (Swift)   │                                           │   Windows (C#)
   ┌──────────────────┐  │            libvoicecat (C++)              │  ┌──────────────────┐
   │ SwiftUI views    │  │  ┌─────────────────────────────────────┐ │  │ WinForms (.NET 10│
   │ AVAudioSession   │──┼─▶│ C ABI  (voicecat.h)                 │◀─┼──│ LibraryImport    │
   │ Swift↔C++ interop│  │  ├─────────────────────────────────────┤ │  │ P/Invoke         │
   └──────────────────┘  │  │ Session / Protocol state machine    │ │  └──────────────────┘
                         │  │ Text + voice signaling              │ │
   Linux/macOS/Windows   │  │ Audio engine: capture→encode→send,  │ │
   server                │  │   recv→jitter→decode→mix→playback   │ │
   ┌──────────────────┐  │  │ Codec layer (Opus 1.6)              │ │
   │ voicecat-server  │──┼─▶│ Crypto + transport (TLS 1.3 + AEAD) │ │
   │ (reuses core)    │  │  │ Net I/O (Asio: TCP + UDP + timers)  │ │
   └──────────────────┘  │  └─────────────────────────────────────┘ │
                         └───────────────────────────────────────────┘

Why this shape:

  • Swift (5.9+) can import C++ directly, but we still ship a C ABI because it is the lowest-friction, most stable boundary and it is what C# needs (LibraryImport/ P/Invoke). One ABI serves both.
  • The server is not a separate codebase. It links the same protocol, crypto, and Opus code as the client, so framing/encryption can never drift between the two ends.

2. Layered design inside the core

From the OS up:

Layer Responsibility Key deps
Platform I/O Sockets, timers; audio device capture/playback Asio, miniaudio
Transport TLS 1.3 (TCP), exported-key ChaCha20-Poly1305 AEAD (UDP), framing, anti-replay mbedTLS, libsodium
Codec & DSP Opus encode/decode; APM (AEC/NS/AGC/VAD) send-side + per-user NR receive-side; resample; jitter buffer; mixer libopus, webrtc-audio-processing, speexdsp
Protocol Message (de)serialization, request/response correlation, state machine protobuf
Session/domain Channels, users, streams, permissions, text routing
C ABI façade Handle-based API + event callbacks exposed to UIs

A UI never sees a socket, an Opus packet, or a protobuf message. It sees: "connect", "join channel", "start a stream from this device", "send this text", and a stream of events ("user joined", "user is talking", "message received", "level meter = 0.4").

3. Threading model

Three classes of thread, with strict rules.

 ┌──────────────┐   lock-free    ┌──────────────┐   lock-free    ┌──────────────┐
 │ Audio capture│ ──ring buffer─▶│  Net thread  │ ──ring buffer─▶│Audio playback│
 │ (RT, miniaudio│                │  (Asio loop) │                │ (RT, miniaudio│
 │  callback)   │◀───ring buffer─│              │◀───ring buffer─│  callback)   │
 │ capture→Opus │                │ TLS + AEAD,  │                │ jitter→Opus  │
 │  encode      │                │ route, relay │                │ decode→mix   │
 └──────────────┘                └──────────────┘                └──────────────┘
                                        │
                                  ┌─────▼──────┐
                                  │ Worker pool│  DB, Argon2id, file I/O,
                                  │ (blocking) │  TLS handshakes, codec setup
                                  └────────────┘

Rules:

  • Audio (real-time) threads are driven by the OS audio callback. They must not allocate, lock, log, or do syscalls beyond the ring-buffer hand-off. Opus encode/decode runs here (it is allocation-free after init).
  • Net thread(s) run the Asio event loop: TLS records, AEAD seal/open, protobuf parse, channel routing, jitter-buffer feed. On the server, this is where the SFU relay copies packets to subscribers.
  • Worker pool absorbs anything that can block: SQLite, Argon2id verification, DNS, TLS handshake CPU, codec (re)configuration.
  • Communication between audio and net is single-producer/single-consumer lock-free ring buffers (one per direction per stream). Control-plane events to the UI go through a thread-safe queue drained on the UI's terms.

4. The C ABI (voicecat.h) — shape

Handle-based, opaque pointers, C-linkage. Illustrative (final names in implementation):

typedef struct vc_client vc_client;

typedef struct {
    void (*on_event)(void* user, const vc_event* ev);   // state changes, messages
    void (*on_level)(void* user, uint32_t stream_id, float rms);  // meters (throttled)
    void* user;
} vc_callbacks;

vc_client*  vc_client_create(const vc_config* cfg, vc_callbacks cb);
void        vc_client_destroy(vc_client*);

int  vc_connect(vc_client*, const char* host, uint16_t port);     // async; result via event
int  vc_authenticate_guest(vc_client*, const char* nickname);
int  vc_authenticate_user(vc_client*, const char* user, const char* password);

int  vc_join_channel(vc_client*, uint32_t channel_id, const char* password /*nullable*/);
int  vc_leave_channel(vc_client*);

// Streams (mic / screen audio / aux device)
int  vc_stream_start(vc_client*, const vc_stream_desc* desc, uint32_t* out_stream_id);
int  vc_stream_stop(vc_client*, uint32_t stream_id);
int  vc_set_input_device(vc_client*, uint32_t stream_id, const char* device_id);
int  vc_set_self_mute(vc_client*, bool mic_muted, bool deafened);

// Text
int  vc_send_text(vc_client*, vc_text_scope scope, uint32_t target_id, const char* utf8);

// Enumeration helpers for UI device pickers
int  vc_list_devices(vc_client*, vc_device_kind kind, vc_device_list* out);

Design notes:

  • Async, event-driven. Calls return immediately; results and state changes arrive via on_event. This maps cleanly onto SwiftUI/async and C# event/Task patterns.
  • The core owns audio. Capture, encode, decode, mixing, and playback happen inside the core via miniaudio. The UI only selects devices, starts/stops streams, and renders meters/state. This keeps the real-time path identical on every OS. (iOS is the one exception that needs UI-side cooperation — see below.)
  • Device enumeration works pre-connect. vc_list_devices needs no live session — device pickers can populate before vc_connect. vc_device.id is an opaque, internally-encoded handle (currently a hex-encoded ma_device_id) — always round-trip an id that came from vc_list_devices/vc_get_stream_audio_config; never construct one by hand. Tolerate an empty list (a machine can legitimately have zero input or output devices).
  • Strings are UTF-8 const char*; ownership is explicit. Output buffers are caller-allocated or returned with a paired vc_free.

Per-platform binding notes

  • Swift / Apple. Import the C ABI via a module map (module VoiceCatC { header "voicecat.h" }) staged into the XCFramework headers by clients/apple/scripts/build-xcframework.sh — Swift gets a clean import VoiceCatC with all C enums/structs/functions available directly (no manual redeclaration, unlike the C# P/Invoke layer). A Swift wrapper (VoiceCatCore package at clients/apple/) provides Swift-idiomatic types (VoiceCatResult, VoiceCatEvent, Channel, User, etc.) on top, mirroring the C# VoiceCat.Interop layer. Callbacks use @convention(c) closures (plain C function pointers, not ARC-managed closures) + Unmanaged.passUnretained(self) as the user context (the Swift analog of C#'s [UnmanagedCallersOnly] + GCHandle). Events are delivered on @MainActor via a coalesced DispatchQueue.main drain (one async block scheduled at a time) — the Swift analog of C#'s Channel<VoiceCatEvent> + 30ms WinForms Timer pump. deinit calls vc_client_destroy (joins all threads) then frees native CString config storage (the core stores raw pointers, doesn't copy). macOS UI: AppKit (chosen over SwiftUI for the most mature VoiceOver accessibility story — same rationale as the Windows client's WinForms choice); iOS UI: SwiftUI (narrower control surface, sufficient VoiceOver support). On iOS the app owns AVAudioSession (category .playAndRecord), requests mic permission, and handles interruptions/route changes — the core exposes hooks (vc_audio_suspend/vc_audio_resume/vc_audio_restart, implemented) the Swift layer calls from AVAudioSession notifications and IOSAudioRouter setting changes. All iOS audio routing (input port selection, mic orientation/polar patterns, HFP vs A2DP, measurement/raw mode, stereo capture) is driven from the Swift IOSAudioRouter singleton via AVAudioSession before the core (miniaudio) opens its device — miniaudio does NOT touch AVAudioSession on iOS. The core is told the capture channel count via vc_set_capture_channels (append-only ABI). vc_audio_restart does a full stop + re-init (unlike suspend/resume which only stop/start) so devices reopen against a new route after AVAudioSession reconfiguration. iOS 18.0 deployment target. Background voice and VoIP push (CallKit/PushKit) are a later milestone. The XCFramework carries a fat static library (libvoicecat-fat.a) bundling libvoicecat.a + all vcpkg static deps so the Swift Package links a single self-contained .a per slice.
  • iOS screen / system-audio sharing is supported via a ReplayKit Broadcast Upload Extension (the same mechanism Discord uses; triggered from Control Center's screen-record button via RPSystemBroadcastPickerView). The extension receives RPSampleBufferType.audioApp (system/app audio) and .audioMic. We capture .audioApp for the SCREEN_AUDIO "listen together" stream and ignore video. The extension runs in a separate process with a ~50 MB memory cap — that cap is a problem only for video frames, so audio-only stays well within budget. It links a minimal slice of the core (Opus encode + media send), shares the session/credentials with the host app through an App Group, and re-derives its own media keys. This is detailed in voice.md §9.
  • C# / Windows. [LibraryImport] (source-generated P/Invoke, .NET 7+) over the C ABI. [UnmanagedCallersOnly] static methods for on_event/on_level to avoid delegate-lifetime pitfalls. UI in WinForms (.NET 10) — chosen over WinUI 3/Avalonia for its mature, predictable screen-reader (NVDA/JAWS/Narrator) UIA support (see roadmap.md §2). Events are delivered via System.Threading.Channels.Channel<VoiceCatEvent>, drained by a 30ms System.Windows.Forms.Timer on the UI thread — simpler than a message-only HWND + PostMessage with no meaningful latency cost. VoiceCatClientHandle : SafeHandle wraps the vc_client* and guarantees vc_client_destroy runs on GC/Dispose.

5. Server architecture

voicecat-server is a headless process linking the core.

        TCP/TLS 1.3                       UDP + media AEAD
            │                                   │
   ┌────────▼─────────┐               ┌─────────▼──────────┐
   │ Connection mgr   │               │ UDP demux          │
   │ (accept, TLS,    │               │ 5-tuple → session  │
   │  per-conn state) │               │ anti-replay window │
   └────────┬─────────┘               └─────────┬──────────┘
            │                                   │
   ┌────────▼───────────────────────────────────▼──────────┐
   │ Session registry  (session_id ↔ TCP conn ↔ UDP tuple)  │
   └────────┬───────────────────────────────┬───────────────┘
            │                                │
   ┌────────▼─────────┐   ┌──────────────┐  ┌▼─────────────────┐
   │ Channel manager  │   │ Text router  │  │ Voice router/SFU │
   │ tree, configs,   │   │ channel + PM │  │ relay Opus to    │
   │ membership, perms│   │              │  │ channel members  │
   └────────┬─────────┘   └──────────────┘  └──────────────────┘
            │
   ┌────────▼─────────┐
   │ Persistence      │  accounts (Argon2id), channels, bans, config
   │ SQLite           │
   └──────────────────┘
  • Voice router is a relay, not a mixer. For each incoming voice frame it looks up the sender's channel and forwards the unmodified Opus payload (restamped with the sender's user id) to every other subscribed member. No server-side decode/transcode → low CPU, low latency, and end-to-content is just Opus. Per-channel Opus params are enforced so all members are mutually decodable.
  • Subscriptions. Clients implicitly subscribe to their current channel's voice; text and presence can be subscribed more broadly. This keeps fan-out bounded on big servers.
  • Stateless-ish media. UDP carries no auth per packet beyond the media-AEAD session; the 5-tuple→session binding is established once via a token (see protocol.md §4).
  • Keepalive reaper. An asio::steady_timer sweeps every 15 s and drops any session whose last_seen (bumped on every inbound TCP or UDP frame) is older than 45 s. Each drop broadcasts UserEvent::LEFT so peers clean up immediately. This catches half-open connections that never produce a TCP EOF. Configurable via server::Config.
  • Single process, scalable later. v1 is one process, one machine. The session registry and router are written behind interfaces so a future build can sit them behind a shared bus for multi-node, but that is explicitly out of scope for now.

6. Repository layout (proposed)

voice-cat/
├── docs/                  # this folder
├── core/                  # libvoicecat (C++)
│   ├── include/voicecat.h # the C ABI
│   ├── src/{net,crypto,codec,protocol,session,audio}/
│   └── proto/             # .proto definitions (shared source of truth)
├── server/                # voicecat-server (C++, links core)
├── clients/
│   ├── apple/             # Swift package + Xcode project (macOS + iOS)
│   └── windows/           # .NET solution (C#)
├── tools/
│   └── vccli/             # headless test client (C++), for protocol bring-up
├── third_party/           # vendored / vcpkg manifest
└── CMakeLists.txt

Build is CMake with vcpkg (manifest mode) for C/C++ deps; the Apple and Windows UI projects consume the built core as a binary + headers. See tech-stack.md.