Files
voice-cat/docs/roadmap.md
Talon b4766d2f24 feat(apple): VoiceCatCore Swift package + XCFramework build for macOS/iOS clients
Lays the groundwork for the macOS (AppKit) and iOS (SwiftUI) clients with a shared
Swift core wrapping the C ABI, mirroring the proven Windows VoiceCat.Interop layer.

Architecture decision: macOS UI = AppKit (not SwiftUI) for the most mature VoiceOver
accessibility story — same rationale as the Windows client's WinForms-over-WinUI-3
decision. iOS stays SwiftUI. Recorded in docs/roadmap.md §2.

Build infrastructure (Phase 0):
- clients/apple/scripts/build-xcframework.sh: runs cmake --preset apple-dev, merges
  libvoicecat.a + 107 vcpkg static deps into a single ~30 MB fat static library
  (libvoicecat-fat.a) via libtool -static (SPM binary targets link one .a per slice),
  stages voicecat.h + a generated module.modulemap (module VoiceCatC) into the headers,
  runs xcodebuild -create-xcframework -> clients/apple/VoiceCatCore.xcframework.

VoiceCatCore Swift Package (Phase 1):
- Package.swift: binary target (VoiceCatCoreXCF) + library (VoiceCatCore) + test target.
- Sources/VoiceCatCore/: 7 files mirroring the C# VoiceCat.Interop patterns adapted to
  Swift native C interop — Enums (9 Swift mirrors of C enums, UInt32-backed), Config,
  Event (copies ev.text to String inside the callback — the #1 lifetime rule), Models
  (10 Swift value types), Marshaling (C arrays -> Swift + immediate vc_free_*),
  Callbacks (@convention(c) + Unmanaged.passUnretained, the Swift analog of C#'s
  [UnmanagedCallersOnly] + GCHandle), VoiceCatClient (owns vc_client* as OpaquePointer,
  all 38 C ABI functions, deinit -> vc_client_destroy then frees config CStrings, event
  delivery on main queue via coalesced DispatchQueue.main drain).

Tests — 6/6 green (swift test against a real voicecat-server):
- testConnectTofuAuthListChannelsRoundTrips, testAdminChannelCrudAccountCrudRoundTrips,
  testScreenAudioStreamStartsAndStops, testPerStreamRecvControlsRoundTrip, plus two
  static smoke tests. Catches Swift-specific interop bugs (@convention(c) callback
  lifetime, Unmanaged pointer resolution, CString memory management, enum raw-value
  bridging, struct layout) that C++ ctest cannot. C++ suite still 21/21 green.

Docs updated (house rule): tech-stack.md §2, architecture.md §4, roadmap.md M4 + §2,
clients/apple/README.md (full rewrite), PROGRESS.md, .gitignore.
2026-06-18 14:20:38 +02:00

10 KiB

Roadmap & Open Questions

1. Milestones

Each milestone is shippable/testable on its own. The headless C++ test client (vccli) exists from M1 so the protocol can be exercised long before any GUI.

M0 — Scaffolding

  • Repo layout (see architecture.md §6), CMake + vcpkg manifest, CI matrix.
  • core/proto/ skeleton; protoc codegen wired for C++ (C#/Swift later).
  • Empty libvoicecat with the C ABI header and stub implementations.
  • Exit: core + server + vccli compile and link on Linux/macOS/Windows.

M1 — Control plane (TCP/TLS, no audio yet)

  • TLS 1.3 transport; framing; Envelope; ClientHello/ServerHello negotiation.
  • Auth: guest + admin-provisioned local account (Argon2id, SQLite); server identity (TOFU/Ed25519); voicecat-admin account add/reset/del/list.
  • Channel tree: snapshot + deltas; join/leave; create/edit/delete (perm-checked).
  • Text chat (ephemeral): channel + private messages, acks; live relay, no history store.
  • vccli can connect, auth, browse channels, and chat.
  • Exit: two vccli instances chat through a real server over TLS.

M2 — Voice, single stream

  • UDP transport with exported-key + ChaCha20-Poly1305 AEAD (mandatory, no plaintext path); UDP token binding; anti-replay.
  • miniaudio capture/playback; libopus encode/decode; one MIC stream per user.
  • Send-side webrtc APM (AEC + NS + AGC + VAD) and a VAD/PTT input gate (both modes, client-configurable) — AEC is in from the start, not deferred.
  • Per-ssrc adaptive jitter buffer; mixer; FEC/PLC/DTX.
  • Per-channel AudioConfig enforced (incl. server max_bitrate_bps ceiling); SFU relay.
  • Exit: talk between two vccli/early-GUI clients in a channel; loss resilience visible.

M3 — Multi-stream & per-channel tuning

  • Multiple concurrent streams per user (MIC, SCREEN_AUDIO, AUX_DEVICE).
  • Per-stream receiver gain/mute; listener-side per-user noise reduction (APM NS on the receive path, per ssrc, local-only); talk indicators.
  • Full per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX).
  • Exit: a user shares mic + desktop audio; listeners control each independently.

M4 — Native clients

Windows (C#/WinForms, .NET 10) ✓ complete 2026-06-17:

  • Connect, saved-server list (JSON, DPAPI-encrypted passwords), TOFU identity dialog.
  • Channel tree (TreeView), user list, join (incl. password-protected channels).
  • Voice: mic start/stop, mute/deafen, VAD/PTT/always-on mode, VAD sensitivity slider (live threshold update via vc_set_vad_threshold), device picker, level meter.
  • Per-user gain/mute/NR tuning (PerUserTuningDialog).
  • Channel + private text chat. Activity log (screen-reader primary path).
  • Explicit AccessibleName/AccessibleDescription on every control; & mnemonics; AutomationNotification curated live announcements.
  • Focus-scoped PTT (documented limitation — no system-wide hook in v1).
  • Admin/moderation UI out of scope — needs server-side dispatch first (M5).

macOS (Swift/AppKit) — in progress:

  • Same feature set as Windows over the same C ABI (now stable and complete).
  • Shared Swift core (VoiceCatCore package) ✓ complete 2026-06-18 — wraps all 38 C ABI functions; 6/6 XCTest smoke tests pass against a real server (connect → TOFU → auth → channels → moderation → admin CRUD → per-stream recv controls). See clients/apple/README.md.
  • UI: AppKit (not SwiftUI) — chosen for the most mature VoiceOver accessibility story (per-control accessibilityLabel/accessibilityHelp/accessibilityRole, NSAccessibility.post(.announcement) for live announcements). Same rationale as the Windows client's WinForms-over-WinUI-3 decision (see §2 below). macOS 14 (Sonoma) deployment target.
  • AVAudioSession not needed on macOS (CoreAudio via the core directly).

iOS (Swift/SwiftUI) — pending:

  • SwiftUI app consuming the same VoiceCatCore package.
  • AVAudioSession, mic permission, foreground voice.
  • ReplayKit broadcast extension for SCREEN_AUDIO.
  • vc_audio_suspend/vc_audio_resume ABI hooks (deferred until this milestone).

Exit: non-technical user installs a client, saves a server, and joins.

M5 — Moderation, polish, and beyond

  • Permissions/roles, kick/ban/server-mute, channel passwords UI. Windows WinForms UI complete (channel CRUD with full Opus config, user moderation, server account management); macOS/iOS Swift UI pending.
  • DRED toggle, audio-quality polish. (AEC and VAD/PTT already shipped in M2.)
  • Then (post-v1, protocol already reserves space): file transfer, E2EE option, CallKit/PushKit background voice, key-based identity, server-side text history, multi-node server.

2. Resolved decisions

Settled and reflected throughout the docs:

  • Media crypto: exported-key + ChaCha20-Poly1305 AEAD from day one, mandatory — no DTLS, no plaintext mode. (security.md §2)
  • TLS library / licensing: mbedTLS (Apache-2.0) + libsodium (ISC). No GPL/LGPL anywhere; code is redistributable closed-source. wolfSSL is rejected. (tech-stack.md §5)
  • iOS screen/system audio: supported via a ReplayKit Broadcast Upload Extension (.audioApp); audio-only stays within the extension memory cap. (voice.md §9)
  • Self-host UX: zero-config, encrypted-by-default; Docker / single binary / source build. (deployment.md)
  • DSP engine: webrtc-audio-processing (APM) — AEC in from the start, plus NS/AGC/VAD. (voice.md §8)
  • Input activation: VAD and PTT, both modes client-configurable. (voice.md §11)
  • Noise reduction is two-sided: sender can denoise its mic, and each listener can apply NS to a specific other user, locally, with no protocol traffic. (voice.md §10)
  • Accounts: admin-provisioned (no self-serve registration) via voicecat-admin or the in-app admin interface. (security.md §4, protocol.md §3, deployment.md §3a)
  • Text: ephemeral — live relay, no server-side history in v1. (protocol.md §5)
  • Connect UX: pure direct-connect with a client-side saved-server list (no central directory). (deployment.md §3)
  • Bitrate ceiling: server-config opus.limits.max_bitrate_bps. (deployment.md §2)
  • Name: "VoiceCat" stays as the internal placeholder.
  • DSP engine, superseded (2026-06-16): the "webrtc-audio-processing (APM)" decision above (AEC + NS/AGC/VAD in one module) could not be carried out — it has no working Windows/MSVC build upstream (GCC-only Meson build, MinGW support unfinished, hard abseil-cpp dependency, Linux-tested only). v1 ships a lightweight, dependency-free energy/RMS VAD instead, behind the same ApmProcessor interface; there is no AEC/NS/AGC implementation at all yet. Real webrtc-audio-processing stays a tracked future swap (e.g. if/when a Linux build target exists). (voice.md §8, §11)
  • Windows client UI framework (2026-06-17): WinForms (.NET 10), not WinUI 3 or Avalonia. Reason: Win32 HWND controls have the most mature, predictable screen-reader (NVDA/JAWS/Narrator) support of any current .NET UI stack. WinUI 3's accessibility UIA tree has known rough edges on .NET 10; Avalonia's accessibility story is thinner still. This overrides the earlier WinUI/Avalonia mention in docs/tech-stack.md §2 and docs/architecture.md §4. (clients/windows/)
  • TOFU pins TLS leaf cert, not Ed25519 (2026-06-17): the original design said to pin the server's declared Ed25519 identity fingerprint from ServerHello. This is circular — the Ed25519 key and the TLS cert are generated independently with no cryptographic binding, so accepting/rejecting based on a value sent inside the channel being trust-decided is meaningless. Decision: pin the TLS leaf certificate's own SHA-256 fingerprint, which is verifiable directly from the TLS handshake before any application data is trusted. The Ed25519 value is still shown in the identity dialog for human-readable display only (informational). See docs/security.md §1.1. (core/src/crypto/tofu_store.*, vc_confirm_server_identity)
  • PTT is focus-scoped in v1 (2026-06-17): PTT hotkey capture uses Form.KeyDown/KeyUp (works only while the VoiceCat window has focus), not a system-wide WH_KEYBOARD_LL hook. Reason: a low-level keyboard hook requires escalated permissions, risks AV flagging, and is disproportionate complexity for a v1 client. Documented in the UI as a known limitation. Can be revisited for v2 if users request it. (clients/windows/VoiceCat.App/Forms/MainForm.cs)
  • macOS client UI framework (2026-06-18): AppKit, not SwiftUI. Reason: AppKit has the most mature, granular VoiceOver accessibility story on macOS — per-control accessibilityLabel/accessibilityHelp/accessibilityRole, NSAccessibility.post(.announcement) for curated live announcements, and decades of real-world screen-reader usage. This is the same rationale that drove the Windows client to WinForms over WinUI 3 (screen-reader support is the deciding factor). SwiftUI's VoiceOver support has improved but still has gaps in complex AppKit-bridged control surfaces (outline views, data-table column headers, live-region announcements). iOS stays SwiftUI — its control surface is narrower and SwiftUI's VoiceOver support is sufficient there. This overrides the earlier "SwiftUI for both macOS + iOS" mention in docs/tech-stack.md §2 and docs/architecture.md §4. macOS 14 (Sonoma) deployment target. (clients/apple/)

3. Open questions

All initial open questions are resolved (§2). Two second-order considerations to keep in mind during implementation — not blockers:

  • APM in constrained contexts. webrtc-audio-processing is a heavier build; confirm it static-links cleanly for the single-binary goal, and note the iOS broadcast extension only does Opus encode + send (no APM), so it stays under the ~50 MB cap. Listener-side per-user NS runs only in the full host app.
  • Receive-side NR cost at scale. A per-ssrc APM NS instance per flagged user adds CPU on busy channels; instantiate lazily (only for flagged streams) and cap concurrent instances.

4. What's intentionally deferred

To keep v1 focused (voice + text), these are designed-for but not built: file transfer, end-to-end encryption, key-based identities, multi-node/federated servers, mobile background VoIP push, and any server-side audio mixing/transcoding. The protocol's versioning + feature negotiation + reserved tag ranges (protocol.md §8) ensure each can be added without breaking deployed clients.