Lays the groundwork for the macOS (AppKit) and iOS (SwiftUI) clients with a shared Swift core wrapping the C ABI, mirroring the proven Windows VoiceCat.Interop layer. Architecture decision: macOS UI = AppKit (not SwiftUI) for the most mature VoiceOver accessibility story — same rationale as the Windows client's WinForms-over-WinUI-3 decision. iOS stays SwiftUI. Recorded in docs/roadmap.md §2. Build infrastructure (Phase 0): - clients/apple/scripts/build-xcframework.sh: runs cmake --preset apple-dev, merges libvoicecat.a + 107 vcpkg static deps into a single ~30 MB fat static library (libvoicecat-fat.a) via libtool -static (SPM binary targets link one .a per slice), stages voicecat.h + a generated module.modulemap (module VoiceCatC) into the headers, runs xcodebuild -create-xcframework -> clients/apple/VoiceCatCore.xcframework. VoiceCatCore Swift Package (Phase 1): - Package.swift: binary target (VoiceCatCoreXCF) + library (VoiceCatCore) + test target. - Sources/VoiceCatCore/: 7 files mirroring the C# VoiceCat.Interop patterns adapted to Swift native C interop — Enums (9 Swift mirrors of C enums, UInt32-backed), Config, Event (copies ev.text to String inside the callback — the #1 lifetime rule), Models (10 Swift value types), Marshaling (C arrays -> Swift + immediate vc_free_*), Callbacks (@convention(c) + Unmanaged.passUnretained, the Swift analog of C#'s [UnmanagedCallersOnly] + GCHandle), VoiceCatClient (owns vc_client* as OpaquePointer, all 38 C ABI functions, deinit -> vc_client_destroy then frees config CStrings, event delivery on main queue via coalesced DispatchQueue.main drain). Tests — 6/6 green (swift test against a real voicecat-server): - testConnectTofuAuthListChannelsRoundTrips, testAdminChannelCrudAccountCrudRoundTrips, testScreenAudioStreamStartsAndStops, testPerStreamRecvControlsRoundTrip, plus two static smoke tests. Catches Swift-specific interop bugs (@convention(c) callback lifetime, Unmanaged pointer resolution, CString memory management, enum raw-value bridging, struct layout) that C++ ctest cannot. C++ suite still 21/21 green. Docs updated (house rule): tech-stack.md §2, architecture.md §4, roadmap.md M4 + §2, clients/apple/README.md (full rewrite), PROGRESS.md, .gitignore.
9.7 KiB
9.7 KiB
Tech Stack & Dependencies
Concrete library choices with versions and rationale. Everything in the core is C++ (C++20). UIs are Swift and C#. Build is CMake + vcpkg.
1. Core library (libvoicecat, C++20)
| Concern | Choice | Version (as of 2026-06) | Why / notes |
|---|---|---|---|
| Sockets, timers, async | Standalone Asio | 1.30.x | Header-only, no Boost dependency, cross-platform TCP+UDP+timers, one reactor for client and server. (Boost.Asio is interchangeable if we already pull Boost.) |
| TLS 1.3 (control) | mbedTLS 3.6 LTS | 3.6.x (LTS ≥ Mar 2027) | Apache-2.0 (permissive — clean for eventual closed-source distribution). TLS 1.3 client+server, plus mbedtls_ssl_export_keying_material() to seed the media AEAD. Static-links cleanly → single self-host binary. OpenSSL 3.x (Apache-2.0) is an interchangeable alternative. No DTLS/wolfSSL (GPL) — see security.md §2. |
| Crypto primitives + password hashing + media AEAD | libsodium | 1.0.20 | ISC. Argon2id (crypto_pwhash), ChaCha20-Poly1305 (per-frame media encryption), Ed25519 server identity, X25519, CSPRNG. Audited, hard to misuse. |
| Audio codec | libopus | 1.6 (2025-12) | Per-channel mono/stereo, bitrate, frame size; in-band FEC, DTX, PLC, and optional DRED deep redundancy; Opus HD/96 kHz available. The whole reason the design is codec-flexible. |
| Audio capture/playback | miniaudio | 0.11.x | Single-header, public-domain, backends for WASAPI / CoreAudio / ALSA / PulseAudio. One real-time abstraction across all desktop targets; keeps the RT path identical. |
| Audio DSP — AEC/NS/AGC/VAD | webrtc-audio-processing (APM) — planned, not built | 1.x (standalone APM) | BSD-3, but has no working Windows/MSVC build upstream (GCC-only Meson, MinGW support unfinished, hard abseil-cpp dep — see roadmap.md §2). v1 ships a lightweight, dependency-free energy/RMS VAD instead (core/src/audio/apm_processor.cpp); there is no AEC, NS, or AGC implementation at all yet. Real APM stays a tracked future swap behind the same ApmProcessor interface. |
| Resampling + jitter ref | speexdsp | 1.2.x | BSD. Resampler for non-48 kHz devices; lightweight jitter-buffer reference. (No longer the NS/AGC/VAD source — APM replaces it.) |
| Control serialization | Protocol Buffers (protobuf-lite) | 5.x (proto3) | Codegen for C++/C#/Swift; additive, forward/backward compatible; oneof envelopes. nanopb is a fallback if footprint matters. |
| Server persistence | SQLite | 3.4x | Accounts, channels, bans, config. Zero-admin, single file, ships everywhere. |
| Logging | spdlog | 1.14.x | Fast, async-capable; off the RT path. |
Resampling note: Opus runs internally at 48 kHz; miniaudio can deliver 48 kHz directly, so explicit resampling (speexdsp/libsamplerate) is only needed when a device can't do 48 kHz.
2. Clients
macOS / iOS — Swift
| Concern | Choice | Notes |
|---|---|---|
| Language | Swift 5.9+ | Direct Swift↔C interop — the C ABI (voicecat.h) is imported as a Clang module (import VoiceCatC) via a module map in the XCFramework headers; no manual struct/function redeclaration (unlike the C# P/Invoke layer). A Swift wrapper (VoiceCatCore package) provides Swift-idiomatic types on top. |
| UI — macOS | AppKit | Chosen over SwiftUI for the most mature, granular VoiceOver accessibility story (per-control accessibilityLabel/accessibilityHelp/accessibilityRole, NSAccessibility.post(.announcement) for live announcements) — the same rationale that drove the Windows client to WinForms over WinUI 3 for screen-reader (NVDA/JAWS/Narrator) UIA support (resolved decision in docs/roadmap.md). macOS 14 (Sonoma) deployment target. |
| UI — iOS | SwiftUI | iOS has a narrower control surface (no channel-tree moderation, etc.) and SwiftUI's VoiceOver support is sufficient; revisit if gaps emerge. |
| Shared core | VoiceCatCore Swift Package | One Swift library wrapping the C ABI, consumed by both the macOS AppKit app and the iOS SwiftUI app. Mirrors the C# VoiceCat.Interop layer. Events delivered on @MainActor via a coalesced DispatchQueue.main drain (the Swift analog of C#'s Channel<VoiceCatEvent> + 30ms WinForms Timer pump). |
| Audio session (iOS) | AVAudioSession | App owns category .playAndRecord + .voiceChat mode, mic permission, interruption/route-change handling; calls vc_audio_suspend/resume on the core. macOS uses CoreAudio via the core directly. (vc_audio_suspend/vc_audio_resume ABI hooks are deferred until the iOS client milestone — keep ABI stable.) |
| Packaging | Swift Package + Xcode project | Core shipped as an XCFramework binary target — a fat static library (libvoicecat-fat.a) bundling libvoicecat.a + all vcpkg static deps (protobuf/mbedtls/sodium/opus/sqlite3/spdlog/asio), so the Swift Package links a single self-contained .a per slice. macOS slice validated; iOS device + sim slices are scaffolding. |
| Future | CallKit / PushKit | For background VoIP + incoming-call UX on iOS. Post-v1. |
Windows — C# (shipped in M4, 2026-06-17)
| Concern | Choice | Notes |
|---|---|---|
| Runtime | .NET 10 LTS (net10.0-windows) |
In-service until 2028. |
| Interop | [LibraryImport] (source-gen P/Invoke) over the C ABI |
[UnmanagedCallersOnly] static methods for on_event/on_level; VoiceCatClientHandle : SafeHandle owns the vc_client* lifetime. |
| Event delivery | System.Threading.Channels.Channel<VoiceCatEvent> |
Single-writer/reader, unbounded; drained by a 30ms System.Windows.Forms.Timer on the UI thread. Simpler than a message-only HWND with no meaningful latency cost. |
| UI | WinForms | Chosen over WinUI 3 / Avalonia for mature, predictable NVDA/JAWS/Narrator UIA support. Win32 HWND controls have the most complete accessibility story on .NET 10 today. See roadmap.md §2. |
| Persistence | System.Text.Json (servers.json), ProtectedData (DPAPI) |
Saved-server list in %AppData%\VoiceCat\; passwords DPAPI-encrypted at rest, opt-in, CurrentUser scope. |
| Audio | Handled by the core (miniaudio/WASAPI) | C# only drives device selection + meters. |
3. Server (voicecat-server)
- Pure C++ linking the core; no GUI. Runs on Linux (primary), macOS, Windows.
- Config via a
server.toml(allow_guests, ports, channel defaults, Opus policy, TLS cert paths or auto-self-signed + Ed25519 identity, Argon2id cost params, rate limits). - SQLite for state. Single process for v1; interfaces drawn so a multi-node build is possible later but explicitly out of scope.
- Packaging: static-ish binary per OS; systemd unit + Docker image for Linux.
4. Build & tooling
| Tool | Use |
|---|---|
| CMake (3.25+) | One build graph for core + server + test CLI; UI projects consume the built core. |
| vcpkg (manifest mode) | Pin C/C++ deps (opus, libsodium, mbedtls, protobuf, sqlite3, spdlog, asio, miniaudio — see vcpkg.json). webrtc-audio-processing/speexdsp are not in the manifest: no working vcpkg port / no working Windows/MSVC build exists upstream for the former; the latter was never actually wired up (the lightweight VAD needs no resampler). Reproducible across OSes. Triplet auto-resolved from the host platform by cmake/voicecat-toolchain.cmake — x64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon. Apple platform scaffolding presets (apple-dev/apple-ios/apple-ios-sim) produce static libvoicecat.a slices for XCFramework consumption. |
| protoc | Generate C++/C#/Swift from core/proto/*.proto (single source of truth). |
| clang-format / clang-tidy | Style + static analysis on the core. |
| CTest + a fuzz target | Unit/integration tests; fuzz the frame parser and protobuf boundary (security-sensitive). |
| GitHub Actions (or similar) | Matrix CI: Linux/macOS/Windows core+server; Xcode build for Apple; dotnet build for Windows. |
5. Licensing — permissive only (hard rule)
The code will eventually be distributed in closed-source form, so no GPL/LGPL dependencies are permitted. Every dependency below is BSD / MIT / ISC / Apache-2.0 / public-domain:
- mbedTLS — Apache-2.0 ✅ · libsodium — ISC ✅ · libopus — BSD ✅ · miniaudio — public domain / MIT-0 ✅ · protobuf — BSD ✅ · SQLite — public domain ✅ · Asio (standalone) — Boost ✅ · spdlog — MIT ✅. webrtc-audio-processing would be BSD-3 ✅ if/when it's actually built in (see §1) — not a live dependency today, so not part of the resolved vcpkg graph the license scanner below checks.
- Explicitly rejected: wolfSSL (GPLv2/commercial) and any DTLS stack that would drag in copyleft. The exported-keys + AEAD media design (security.md §2) removes the need for one entirely.
- CI runs a license scanner over the resolved vcpkg graph and fails the build on any GPL/LGPL transitive dependency, so this rule can't silently regress.
6. Why not the obvious alternatives
- WebRTC — explicitly rejected: ICE/SDP/TURN complexity, huge dependency, opaque. We want plain TCP+UDP we fully control.
- QUIC — capable (reliable streams + datagrams + TLS 1.3 in one), but heavier and drifts toward the complexity we're avoiding. Revisit only if NAT traversal/multiplexing pain appears.
- gRPC for control — pulls HTTP/2 and a lot of surface for what is a simple framed message stream over TLS. Plain protobuf-over-framed-TLS is enough.
- A Rust core — viable and memory-safe, but the user prefers C++ and the Swift/C# binding story is marginally simpler from C++ (Swift can even consume C++ directly).