46 Commits

Author SHA1 Message Date
c6c003b8a7 docs: add the .NET/C# porting plan
Proposal for replacing the C++ core, C++ server, and Swift macOS/iOS
clients with a single .NET 10 / C# codebase.

Covers the dependency map (8 vcpkg deps + 1 vendored -> 3 native libs),
the real-time-audio design, per-client strategy, a test-porting plan for
all 29 ctest cases, doc-sync work, an 11-phase migration, and a risk
register.

Two findings drive the shape of the plan:

- SslStream has no RFC 5705 keying-material exporter, which the media
  AEAD key derivation depends on (docs/security.md 2). The API is an
  unapproved proposal and SChannel structurally cannot export secrets.
  Recommends BouncyCastle's managed TLS 1.3 stack, which does implement
  the exporter and keeps the wire format byte-compatible with the C++
  implementation -- so the existing tree stays usable as a conformance
  oracle throughout the port. Protocol-v3 in-band media keys documented
  as the fallback.

- The iOS ReplayKit broadcast upload extension stays in Swift: 50 MB
  jetsam cap plus an unsupported extension type in .NET for iOS, and it
  already doesn't link the core. Leaves one Swift file plus the shared
  App Group ring.

Indexed in docs/README.md. Nothing here is implemented yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 03:19:59 +02:00
4f71b784fe docs: condense implementation comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-23 13:37:05 +02:00
575e2907d0 Cleanup root cmakelists
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:59:12 +01:00
f687370e47 Cleanup voicecat.cpp impl comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:56:14 +01:00
552ffceb66 Cleanup session/
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:55:17 +01:00
a4b3838125 Cleanup protocol/ comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:52:46 +01:00
eafa5eb90c Cleanup net/ comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:49:32 +01:00
158a2df062 Cleanup crypto comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:37:39 +01:00
826b3bfb86 Cleanup client and worker pool comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 15:29:09 +01:00
cc81d19d02 Cleanup audio engine and codec comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 14:59:26 +01:00
8844325efa Cleanup audio engine comments
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 13:13:35 +01:00
bd844e4710 Cleanup core header and proto code
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
2026-07-03 12:46:51 +01:00
5c03e5f261 build: bundle vcpkg as a git submodule, pinned to the manifest baseline
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Adds vcpkg as a submodule at vcpkg/, pinned to the exact commit vcpkg.json
already declares as builtin-baseline, so the bundled checkout and the
manifest's resolved port versions can never drift apart.

cmake/voicecat-toolchain.cmake, scripts/common.sh, and
clients/apple/scripts/build-xcframework.sh now resolve vcpkg as:
VCPKG_ROOT env var (external checkout) > bundled submodule. Docs updated
to describe the new one-time setup.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-03 10:42:42 +01:00
bba605401d chore: comment cleanup pass ahead of open-sourcing
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Removes leftover debug scaffolding (stray Console.WriteLine/NSLog traces,
dead nick_buf_ptr, a no-op --print-config flag now implemented for real),
fixes stale/misleading comments (channel passwords are no longer a "future
M5+" feature, a wrong cross-reference, a stale TlsContext::close() mention,
an incomplete BanRecord::subject_type doc, and a smoke test pointing at a
build/m1-dev preset that no longer exists), strips internal M1-M5 milestone
jargon from comments now that the roadmap is done, trims comments that just
restated the following line, and consolidates a few "why" explanations that
were duplicated 2-3 times in the same file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-03 10:20:18 +01:00
bda37ec27b fix(ios): add AppIcon asset catalog and fix launch screen for TestFlight
Adds Assets.xcassets with a placeholder 1024x1024 AppIcon so Xcode
compiles CFBundleIconName and the required icon sizes into the bundle.
Replaces UILaunchStoryboardName (missing storyboard) with UILaunchScreen
dict, valid for iOS 14+ (min target is iOS 18). Fixes all four App Store
Connect validation errors blocking TestFlight upload.
2026-06-30 13:22:14 +01:00
9612b0af89 chore(ios): update bundle ID to me.iamtalon.voicecat for App Store
Changes main app bundle ID from cat.voice.VoiceCatiOS, broadcast extension
from cat.voice.VoiceCatiOS.broadcast, and App Group from group.cat.voice.VoiceCat
to match the registered App Store identifier.
2026-06-30 11:48:02 +01:00
2c8178fa02 chore: remove skeleton build mode and stub #ifdef scaffolding
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Drop the M0 no-deps skeleton preset and all VOICECAT_HAS_NET/AUDIO/OPUS/NS
guards that it required. Every subsystem is fully implemented; the stub
#else paths were dead code that added noise to every header and source file.

- CMakePresets.json: remove skeleton configure/build/test entries
- CMakeLists.txt (root/core/tests): remove VOICECAT_USE_VCPKG_DEPS option
  and guards; all targets now build unconditionally
- 17 C++ source files: unwrap HAS_* guards, delete stub #else blocks
- apm_processor.cpp: delete ApmPassthrough no-op class; create() always
  returns RnnoiseProcessor
- 18 test files: remove HAS_* guards and stub int main() skip bodies
- docs/building.md: remove skeleton from preset table and prose

VOICECAT_HAS_LOOPBACK (Windows WASAPI loopback platform gate) unchanged.
29/29 ctest green.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 11:32:22 +01:00
9a20953c08 fix(ios): break AirPods-disconnect reinitialize loop on A2DP presets
Recovery from the 99937c9 audio-device-change commit broadened the
route-change recovery set to "everything except categoryChange /
routeConfigurationChange", which added .override. But .override is
fired by our own applyA2dpSpeakerFallback() -> overrideOutputAudioPort,
which recoverAudio() calls on every recovery. On an A2DP preset
(Stereo Mic / Mono Mic), disconnecting AirPods ping-ponged:

  oldDeviceUnavailable -> recoverAudio -> applyA2dpSpeakerFallback
  -> overrideOutputAudioPort(.speaker) -> .override routeChange
  -> recoverAudio -> applyConfiguration (setCategory resets override)
  -> applyA2dpSpeakerFallback -> overrideOutputAudioPort -> .override -> ...

Each iteration also rebuilt the AVAudioEngine via reconfigure() ->
rebuild() -- the audible reinitialize loop + CPU spin. Voice Chat and
Built-in Mic + Speaker were unaffected (applyA2dpSpeakerFallback
early-returns for non-A2DP modes, so no overrideOutputAudioPort call).

Two-part fix (pure Swift iOS-app target; no C ABI / proto / docs changes):
1. AudioSessionManager.handleRouteChange: added .override to the skip
   list alongside .categoryChange / .routeConfigurationChange. .override
   is only ever fired by our own overrideOutputAudioPort call, so
   treating it as a recovery reason is the loop by definition. The
   AVAudioEngineConfigurationChange observer in IOSVoiceProcessingEngine
   remains as the backstop if an override ever actually stops the engine.
2. IOSAudioRouter.applyA2dpSpeakerFallback: made idempotent via a
   lastAppliedOutputOverride tracker that skips the redundant
   overrideOutputAudioPort call when the desired state (.none for
   external output present, .speaker otherwise) already matches. Reset
   to nil at the top of applyConfiguration() (setCategory can reset the
   override) and on a failed call. Defense-in-depth on top of fix 1.

Build: xcodebuild -project clients/apple/iOS/VoiceCatiOS.xcodeproj
-scheme VoiceCatiOS -destination 'generic/platform=iOS' build green
(Xcode 26.5 / iOS 18.0).
2026-06-25 16:18:21 +02:00
99937c9446 feat(ios): auto-reconnect + audio-device-change recovery
Network drops (e.g. Wi-Fi -> cellular) and audio-device plug/unplug (wired
headphones, AirPods) used to leave the iOS client in a dead/zombie state:
the engine went silent, no reconnect was attempted, and a live-session
disconnect waited 30-60 s for the C core's TCP keepalive/reaper timeout.

Reconnect (AppState.swift, SessionState.swift):
- Two-layer reconcile. Once SessionState overwrites client.onEvent at auth
  success, AppState.handleConnectEvent no longer sees live-session events.
  Added a weak SessionState.appState; SessionState.handleEvent .disconnected
  calls appState.onLiveSessionDisconnected after the cue -- the single path
  AppState learns a live session dropped. Shared teardownLiveSessionAndReconnect
  snapshots lastSession, stops audio, releases session/VoiceCatClient (io-
  thread join via vc_client_destroy), resets the backoff, and arms
  scheduleReconnect (exponential 1s -> 30s cap, indefinite, restored on auth
  success via existing TOFU_MATCHED auto-confirm + idempotent join_channel).
- NWPathMonitor now runs WHILE CONNECTED (not only mid-reconnect). On a Wi-Fi
  <-> cellular interface change or path .unsatisfied it calls
  proactiveReconnect: tearing the session down BEFORE the C core notices the
  dead socket collapses the 30-60 s reaper wait into ~1 s + first backoff
  tick. Same-interface refreshes (BSSID roams) are ignored via pathSignature.
  While mid-reconnect a .satisfied path resets the backoff for a fast retry.
- User-initiated disconnect()/cancelConnect() set userInitiatedDisconnect
  and cancel all reconnect state (task + monitor + lastSession + connectedServer).

Audio recovery (AudioSessionManager.swift, IOSVoiceProcessingEngine.swift):
- Intent-gated recoverAudio() replaces the narrow .oldDeviceUnavailable/
  .newDeviceAvailable route-change guard; fires on every externally-initiated
  route change reason except the ones we cause ourselves (.categoryChange/
  .routeConfigurationChange) to avoid a notification loop. Interruption-end
  now always recovers instead of only when .shouldResume is set.
- Added AVAudioEngineConfigurationChange observer on the engine so a system
  self-stop after our route-change handler wins the race is caught.
- IOSAudioEngine.rebuild() does a one-shot reactivation-retry on
  engine.start() failure (iOS sometimes refuses until the session is
  re-reactivated -- the silent-death case).

No C ABI / voicecat.h / proto / core changes. Swift-only. iOS sim build green
via scripts/build-ios-client.sh --no-configure (Xcode 26.5 / iOS 18.0 sim).
2026-06-25 14:57:13 +02:00
44a336cc89 fix(ios): cast audio complexity UInt32 to Int in ChannelEditView 2026-06-24 16:49:37 +02:00
6fe7bf0158 feat: fix voice join/leave, channel edit defaults, channel-update stream restart
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Three bugs fixed across the full stack (proto/server/core/ABI/Win/macOS/iOS):

1. Join/Leave Voice now truly subscribes/unsubscribes from the voice plane.
   Previously the button only toggled the local mic — receiving was always on
   (gated by channel membership alone). Added a protocol-level voice subscription
   concept: new SubscribeVoiceRequest/UnsubscribeVoiceRequest/VoiceSubscriptionResult
   proto messages, User.voice_subscribed field, vc_join_voice/vc_leave_voice C ABI
   functions, VC_EVENT_VOICE_STATE event, server-side voice_subscribed flag checked
   by the SFU relay recipient filter, and core-client gating of remote-stream
   decoder setup. All three clients rewired to subscribe+mic on Join / unsubscribe
   on Leave. Text chat works regardless of voice subscription.

2. Channel edit dialog now shows the channel's actual current settings. The read
   struct vc_channel was missing sort_order and audio fields — only the write
   struct vc_channel_info had them. Extended vc_channel with both (additive, no
   ABI break), updated the session model and list_channels marshaling to populate
   them, and updated all three clients' edit callers to use actual channel info
   instead of hardcoded defaults.

3. Channel parameter updates now automatically restart everyone's streams.
   Previously editing a channel's audio config persisted and broadcast a
   ChannelEvent::UPDATED, but no layer restarted streams — encoders/decoders are
   frozen at announce time. handle_channel_event now detects audio-config changes
   on the user's current channel and stop->starts each active local stream. The
   server reads the updated config on re-announce; peers wire up fresh decoders
   at the new ssrc.

All 29 CTest tests pass; Windows DLL + C# client build clean. Apple clients not
yet compile-verified (Windows environment).
2026-06-24 14:29:39 +02:00
2baefddbe4 feat(windows): system-wide push-to-talk via Raw Input
PTT was focus-scoped (WinForms KeyDown/KeyUp) so it died the moment the
window lost focus. Add an optional system-wide path using the Raw Input
API (RegisterRawInputDevices + WM_INPUT with RIDEV_INPUTSINK) instead of
a WH_KEYBOARD_LL low-level hook -- the latter is the keylogger pattern AV
heuristics flag, which is worse for our unsigned MinGW binary. Raw Input
involves no DLL injection or global hook and passes keys through.

- New VoiceCat.App/Native/RawInput.cs: P/Invoke + structs; register the
  keyboard sink, decode WM_INPUT to vkey/up-down, GetAsyncKeyState helper.
- MainForm overrides OnHandleCreated/OnHandleDestroyed/WndProc to manage
  the sink and route WM_INPUT to PTT; gates the focus-scoped KeyDown/KeyUp
  off when system-wide is on; makes the Deactivate force-release
  conditional; adds a GetAsyncKeyState watchdog on the pump timer so a
  missed key-up (RDP/lock-screen) cannot leave PTT stuck.
- VoiceSettings.SystemWidePtt (default ON) + system-wide checkbox in the
  Audio settings PTT section.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 13:14:01 +02:00
65736df464 fix(windows): polish hotkeys, user list, titlebar, and PM window
- Suppress the system ding on global hotkeys/PTT and the user-list Enter
  key by setting SuppressKeyPress (Handled alone leaves WM_CHAR to beep).
- Preserve the user-list keyboard selection across talking/mute refreshes
  instead of resetting it on every Items.Clear().
- Include the connected server name in the main window titlebar.
- Close the private-message window on Escape.
- Show the PM window without an owner so focus is no longer trapped to it
  and the main window can be worked in while a PM is open.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 12:30:48 +02:00
8bb2ba933c feat(clients): unify volume boost cap to 4x across all clients
Mic input gain (was 3x) and per-user receive gain (was 2x) had
asymmetric boost ceilings. Raise both, plus the desktop aux input
gain (was 3x), to a uniform 4x (400%) on macOS, iOS, and Windows.

The master Output volume slider is unchanged (still 1x). No core
changes needed: the C ABI only clamps negatives, so the ceilings
live entirely in the client UI sliders.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 12:18:10 +02:00
7ef560ba8a fix(windows): statically link MinGW runtime into all binaries
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
The static-MinGW recipe (-static-libgcc -static-libstdc++ -static) was
only applied inside the VOICECAT_BUILD_SHARED branch in core/CMakeLists,
so it covered voicecat.dll but not voicecat-server.exe / vccli. With the
x64-mingw-static triplet only vcpkg's own deps are static; the GCC/MinGW
runtime stays dynamic, so the server failed to start on clean Windows
machines with missing libgcc_s_seh-1.dll / libwinpthread-1.dll /
libstdc++-6.dll.

Move the recipe to a global add_link_options (WIN32 AND MINGW) so every
produced binary embeds the runtime. Verified with objdump -p: only
Windows system DLLs remain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 01:28:29 +02:00
47124b15a2 fix(windows): list all user-session apps in the app-audio picker
The picker gated its list on visible-window processes plus audio sessions
on only the default render device. That both showed non-audio apps (any
window) and missed real ones (windowless or routed to a secondary device).
Process loopback targets a PID and its child tree regardless of whether the
app is currently playing, so the gate fought the capture layer.

Now enumerate every process in the user's interactive session (windowed or
not), deduped by executable with the windowed tree-root as the capture PID;
scan all active render endpoints to flag currently-playing apps with a > and
sort them first. Adds a filter box and persists checks across filtering.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 00:24:20 +02:00
04bdb70d47 fix(codec): scale DRED duration to frame size instead of hardcoding 2
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
OPUS_SET_DRED_DURATION was hardcoded to 2 (20 ms), meaning DRED only
covered 1/3 of a lost 60 ms frame and was useless above 20 ms channels.
Now computed as max(2, ceil(frame_ms/10)) so DRED always embeds enough
redundancy to reconstruct one full previous frame regardless of frame size.
The floor of 2 preserves two-frame burst-loss coverage at 10 ms channels.
Decoder side and server are unaffected (server relays payloads verbatim).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-23 23:44:00 +02:00
b44a200b95 fix(audio): apply receive-side NR to stereo mic streams
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
The per-listener noise-reduction toggle (vc_set_remote_stream) did nothing
on Windows/macOS/iOS. The decode loop gated the RNNoise pass on
dec_channels == 1 as a proxy for "this stream is voice" (assuming
stereo => screen-share). The stereo-mic capture commit broke that: a stereo
mic with send-side NR off transmits stereo Opus, so the receiver decoded
two channels and skipped NR entirely. gain/mute have no channel guard, which
is why only NR appeared broken.

Thread the stream kind through init_recv_stream into RemoteStream::is_voice
(set from si.kind() == STREAM_MIC), gate receive NR on is_voice instead of
channel count, and fold a stereo voice frame to mono -> denoise -> duplicate
back across both channels in place (symmetric with the send-side downmix;
RNNoise is mono-only). Screen-audio shares are never denoised.

New test test_recv_noise_reduction drives AudioEngine and asserts a stereo
voice stream's noise floor collapses with NR on (RMS 1046 -> 0.1) while a
screen-audio share stays unchanged. ctest --preset dev green 29/29.
Docs: voice.md section 10. Shared-core fix; clients need only a rebuild.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 21:11:03 +02:00
f72219ddf3 feat(audio): stereo mic capture on Windows & macOS desktop clients
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Both desktop mics were hard-mono: the core defaults capture_channels=1 and
neither client ever called vc_set_capture_channels (only iOS did). Add a
persisted "Stereo microphone" toggle to each client's Audio settings, applied
when the mic stream starts and live via vc_set_capture_channels + vc_audio_restart.
Expose both ABI calls in the Windows interop; the macOS wrapper already had them.

Core fix: encode_and_send_frame now folds a stereo mic frame to mono on a mono
channel - previously the channels==2 branch encoded interleaved L/R directly even
on a mono channel, feeding a mono opus_encode 2x its samples (wrong pitch/garbage).
Real stereo still only reaches the wire on a stereo channel; on a mono channel the
mic is cleanly downmixed.

Test: test_stereo_mic_mono_channel. ctest --preset dev green (28/28). Docs: voice.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 20:48:26 +02:00
b14cf2a4e8 fix(ios): stop core opening a second mic device — dual capture/crackle
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
On iOS the remote end heard the mic twice and crackly (BT headset + internal
mic in Voice Chat; mono + stereo copies of the internal mic in Stereo Mic).

Root cause is an io-thread ordering race. `ensure_audio_running()` only set
`external_capture` when a MIC stream already existed, but it also runs from
`sync_remote_streams` on the post-auth `ServerStateSnapshot` — before the user
joins voice. With `external_playback_` still false and no MIC stream,
`AudioEngine::start()` opened a real miniaudio capture device that stayed open
all session (later calls early-return on running()), racing the AVAudioEngine
input tap fed via `vc_stream_feed_pcm`. `on_capture_frame` then encoded+sent
both paths — the mic transmitted twice, the two unsynchronized capture clocks
producing the crackle.

- core (ensure_audio_running): force `external_capture = true` whenever
  `external_playback_` is set, so iOS unified mode never opens a hardware
  capture device. No-op on desktop.
- ios (AppState): move `setExternalPlayback(true)` before `connect()`, so the
  flag is set before the io thread processes any message — closing the race.

Verified on device: remote end hears the iOS mic once and clean in both Voice
Chat (+ BT) and Stereo Mic.
2026-06-23 17:47:49 +02:00
19c2fb6ec9 fix(ios): pace mic feed with a prebuffer cushion to stop flutter/crackle
The iOS mic was unusable — a consistent ~40-60ms flutter + volume fade
('slow fan') on every preset. The core sends each captured frame
synchronously (no send pacer), so packet cadence == capture cadence, and
the receiver's playout keeps near-zero buffering by design (its jitter
estimate keys off the regular sender timestamp, so it's blind to arrival
jitter). That's smooth only for a steady sender (desktop miniaudio =
steady 20ms); the iOS AVAudioEngine tap delivers ~2 frames per ~40ms
callback -> bursty -> receiver underruns -> PLC fade.

Fix (iOS-only): the mic tap writes converted 48kHz int16 to an SPSC ring;
a 20ms feed pump drains it and calls feedPcm at a steady cadence. The pump
primes a small prebuffer cushion (3 frames ~60ms, self-healing up to
~120ms on underrun) before releasing, so the tap's bursts can't drain it
to empty. Never reads a partial frame (read consumes what it returns ->
partials were the crackle), and rebuilds with the current channel count
each rebuild() (a frozen count fed mono-as-stereo = octave-up on a
Stereo->Voice Chat switch).

Trade-off: ~60-120ms added mic-send latency, unavoidable when de-bursting
for a near-zero-buffer receiver. PROGRESS.md notes the proper follow-up:
make the jitter buffer measure real RFC-3550 arrival jitter so the
receiver absorbs bursts itself.

Verified: xcodebuild Debug BUILD SUCCEEDED (iOS Simulator, arm64).
2026-06-23 15:40:54 +02:00
cd9c08a47a fix(apple): merge vendored librnnoise.a into xcframework fat static lib
Both VoiceCatMac and VoiceCatiOS failed to link with 'Undefined symbols
for architecture arm64: _rnnoise_create/_rnnoise_destroy/_rnnoise_process_frame'.
Root cause: build-xcframework.sh merged vcpkg deps into the fat static lib but
NOT the locally-built vendored librnnoise.a (a CMake target from
third_party/rnnoise/ linked privately into voicecat via VOICECAT_HAS_NS — not a
vcpkg dep). The xcframework had been rebuilt after the RNNoise commit but still
omitted the symbols, so every slice's libvoicecat-fat.a referenced _rnnoise_*
with no defining object. The iOS slices were also stale (pre-rnnoise) and
absent from the xcframework entirely.

Fix: build-xcframework.sh now also collects .a files from build/<preset>/lib/
(excluding libvoicecat*) so vendored CMake-target static libs like librnnoise.a
get merged in. Future-proof: any new vendored static-lib target landing in
build/<preset>/lib/ is picked up automatically. README 'Fat static library'
section updated.

Verify: rebuilt VoiceCatCore.xcframework --all → all 3 slices (macos-arm64,
ios-arm64, ios-arm64-simulator) now carry the 10 _rnnoise_* symbols; fat lib
~30 MB → ~33 MB. xcodebuild Debug BUILD SUCCEEDED for VoiceCatMac, VoiceCatiOS
(iphonesimulator arm64), and VoiceCatiOS (iphoneos arm64). No core/ABI/proto
changes — xcframework artifact + build script only.
2026-06-23 14:25:30 +02:00
2e0e0caccb feat(clients): wire RNNoise mic noise reduction into Windows, macOS, and iOS
Expose the existing send-side vc_set_input_noise_reduction C ABI (MIC-only,
mono, LOCAL — denoises captured mic PCM before input gain and VAD/PTT gate)
as a persisted global toggle in each client's audio settings, applied live
and re-applied on Join Voice. Mirrors the existing mic-gain wiring pattern.

- Shared Swift (VoiceCatCore): add setInputNoiseReduction(_:) wrapper
- Windows: P/Invoke + SetInputNoiseReduction wrapper, MicNoiseReduction in
  VoiceSettings, new checkbox in AudioSettingsForm (layout shifted +28px),
  apply on Join Voice; also fix stale 'planned - currently passthrough'
  label on the receive-side per-user NR checkbox (RNNoise now backs it)
- macOS: inputNoiseReduction state + UserDefaults in MainWindowController,
  NR checkbox + nrChanged action in SettingsWindowController
- iOS: inputNoiseReduction in VoiceState + setter + restore in SessionState,
  NR Toggle in SettingsView Voice section

Aux/screen are out of scope by design (core's NR guards kind == MIC). Apple
builds require a rebuilt VoiceCatCore.xcframework with VOICECAT_HAS_NS.
2026-06-23 14:11:18 +02:00
bad9c7533a feat(audio): real noise suppression via vendored RNNoise (send + receive)
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
The two-sided NR plumbing (RemoteStream::recv_ns + the per-listener
vc_set_remote_stream noise_reduction toggle) was wired but inert:
ApmProcessor::create() returned a no-op passthrough, because the
originally-planned webrtc-audio-processing has no working Windows/macOS
build. Drop in RNNoise as the real backend behind the same ApmProcessor
interface, lighting up both NR paths.

- Vendor RNNoise (BSD-3 + CC0) at third_party/rnnoise/ — the vcpkg port
  is !windows !arm, so it can't cover our primary targets. Shrunk int8
  model (78MB -> 11.7MB via upstream scripts/shrink_model.sh), built as a
  standalone C static lib with no RTCD (portable scalar path on x86,
  auto-NEON on arm64) under -DDISABLE_DEBUG_FLOAT. Model is baked in
  (rnnoise_create(NULL)); no runtime file.
- New RnnoiseProcessor (core/src/audio/apm_processor.cpp) selected by
  ApmProcessor::create() when VOICECAT_HAS_NS. Mono/48kHz/480-sample;
  our clock is fixed 48kHz and Opus frame sizes are multiples of 480, so
  no resampling. RT-safe: allocates at construction, lock-free in the
  capture/playback callbacks.
- Receive-side: lit up via the factory; gated to mono streams (a stereo
  stream is a screen-audio share, not voice).
- Send-side (new): vc_set_input_noise_reduction(client, enable) ABI +
  vc_client::mic_ns_, run before input gain/VAD in on_capture_frame. A
  stereo mic is downmixed to mono ONLY when NR is on — with NR off a
  stereo mic keeps full stereo (never collapse mic quality unasked).
- Enable C as a project language for the vendored lib.
- New noise_suppression test: white noise through ApmProcessor::create()
  drops ~99.9% RMS. ctest --preset dev green, 28/28. windows-client DLL
  builds clean with vc_set_input_noise_reduction exported, system-only deps.
- Docs synced: voice.md §10, tech-stack.md §1/§5, third_party/README.md,
  vcpkg.json note, PROGRESS.md, CLAUDE.md.

Client on/off UI toggles (Windows/macOS/iOS) are the remaining follow-up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 13:30:54 +02:00
7249a8fd30 feat(clients): add aux outgoing stream (mic + second input device) on Windows + macOS
Lets a user transmit a second hardware input device (e.g. line-in / aux)
alongside the mic, with its own device picker and volume, from Audio Settings.

No core/ABI/proto changes: the aux is a VC_STREAM_AUX_DEVICE stream started
with external_feed=1 and fed via vc_stream_feed_pcm (the same external-feed
pipeline screen-audio uses). Per-kind local_streams_ already allows mic +
screen + one aux to coexist; volume is a client-side gain multiply (the core's
vc_set_input_gain is mic-only/global). Aux is always-on (core never gates
AUX_DEVICE on VAD/PTT) and is tied to the voice session.

Windows: new Audio/InputDeviceCapture.cs (WASAPI shared-mode capture from a
real input endpoint + capture-endpoint enumeration); aux section in
AudioSettingsForm.cs; lifecycle in MainForm.cs; persistence in VoiceSettings.cs.

macOS: new Audio/InputDeviceCapture.swift (AVAudioEngine input-node tap pinned
to the chosen Core Audio device + device enumeration by stable UID); aux section
in SettingsWindowController.swift; lifecycle + UserDefaults persistence in
MainWindowController.swift; file registered in project.pbxproj.

Windows verified (C# solution builds clean; aux confirmed working). macOS build
+ E2E pending a Mac.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 12:56:09 +02:00
a48b47d4ca feat(windows): move audio settings to dedicated dialog, fix device ComboBox accessibility
Audio input settings (device picker, VAD/PTT mode, VAD sensitivity, mic gain,
PTT key) move from the always-visible bottom panel into Settings > Audio...,
matching the macOS/iOS pattern.

The device ComboBox now uses DataSource + DisplayMember="Name" instead of
Items.Add() with no DisplayMember — this fixes both the display bug (was
showing the full DeviceInfo record ToString()) and the NVDA silence on
dropdown open (DataSource binding exposes proper MSAA text per item).

Changes apply live for immediate feedback; Cancel reverts. VoiceSettings
gains InputDeviceId to persist the chosen device across sessions.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-23 12:15:18 +02:00
95f1fb70b0 feat(clients): persist input settings, add mic input gain, fix iOS chat + VoiceOver
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Input mode (VAD/PTT/Always-On), VAD threshold, and the new mic gain were
applied to the core + UI but never saved, so every relaunch reset to VAD
defaults. Each client now persists them and re-applies on connect:
  - iOS: UserDefaults (SessionState.loadAndApplyVoiceSettings + setter writes)
  - macOS: UserDefaults via MainWindowController didSet + loadPersistedAudioSettings
    (settings window also restores the VAD slider from the stored threshold)
  - Windows: new Models/VoiceSettings.cs (JSON at %AppData%\VoiceCat\voice.json,
    mirrors FeedbackSettings) loaded/applied in MainForm

Add global send-side mic gain API vc_set_input_gain (applied to MIC PCM in
on_capture_frame before the VAD gate, clamped to int16) + Swift/C# bindings,
and a 0-300% (default 100%) mic-volume slider on all three clients.

Fix iOS chat: ChatView called sendText(scope:.channel) with no targetId (0),
so channel messages went nowhere; now passes session.currentChannelId.

Fix iOS per-user tuning for VoiceOver: the tuning sheet was long-press
.contextMenu only (invisible to VoiceOver); UserRow now also exposes the same
buttons via .accessibilityActions (no visual change).

Verified: core builds clean; ctest 24/27 (3 pre-existing teardown crashes,
reproduced with changes stashed); VoiceCatMac + VoiceCatiOS (arm64 sim) build
SUCCEEDED; VoiceCat.Interop dotnet build succeeded. Windows App not built
(WinForms can't build on macOS) — follows existing patterns.
2026-06-23 03:35:26 +02:00
d30c4ee2f5 fix(ios-audio): unify iOS audio onto one always-external AVAudioEngine
The iOS audio path was a hybrid: Voice-Chat-class presets ran a native
VPIO AVAudioEngine (core external) while Stereo/Studio/A2DP presets ran
the core's miniaudio devices. Nearly every "no input / no output / both"
bug lived in the seam between the two paths — the lingering miniaudio
capture unit fighting VPIO, the audioRestart ordering dance, the
route-change "glitching" loop, stereo<->mono stickiness, and
"can't hear anyone". Switching presets/routes mid-call routinely dropped
a direction.

Drive ALL iOS audio through one AVAudioEngine with the core fully
external at all times: setExternalPlayback(1) once at connect, every MIC
stream external_feed=1, mic via vc_stream_feed_pcm, playback via
vc_set_mixed_output_sink (drained by an always-on AVAudioSourceNode so
remote audio plays before joining voice). VPIO + AGC toggle per preset.
Every preset/route/interruption change funnels through one deterministic
Swift-only reconfigure (stop -> apply session config -> rebuild -> start)
— no second path to hand off to, so a change can't drop a direction.

- IOSVoiceProcessingEngine.swift -> IOSAudioEngine: always-on source-node
  playback, conditional mic tap, VPIO/AGC; one rebuild() backing
  startListening/stop/startMic/stopMic/reconfigure/setCaptureChannels.
- IOSAudioRouter: 7 presets -> 4 (Voice Chat / Stereo Mic / Mono Mic /
  Advanced); persisted voiceProcessingEnabled + agcEnabled; setters call
  IOSAudioEngine.reconfigure() instead of audioRestart/reconcileVoicePath.
- AudioSessionManager slimmed; SessionState mic lifecycle collapsed;
  AppState wires external playback + listening at connect, stop at
  disconnect; SettingsView shows 4 presets + Advanced VPIO/AGC toggles.

No core/ABI/test changes — relies on the already-shipped external API
(test_external_pcm, test_external_playback). xcodebuild iOS device Debug
BUILD SUCCEEDED. Updates docs/voice.md §8 and PROGRESS.md.
2026-06-23 02:45:53 +02:00
7547b8e140 fix(audio): wire in-band FEC into the decoder loss path
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
The encoder set OPUS_SET_INBAND_FEC, but the decoder never invoked FEC --
the loss path went DRED -> PLC, so FEC redundancy was emitted (and paid for
in bitrate) yet never consumed.

Wire FEC recovery into AudioEngine::on_playback between DRED and PLC: copy
the next buffered packet once, try DRED, else (if the stream negotiated FEC)
decode(next_pkt, ..., fec=true), else PLC. Recovery priority is now
DRED -> FEC -> PLC. Add per-stream RemoteStream::fec_enabled_, captured from
OpusParams in init_recv_stream. Docs (voice.md) updated to match.

ctest --preset dev: 27/27 green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 20:18:29 +02:00
ce2035f271 fix(audio): bound playout depth to stop voice latency ratcheting up
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
Latency between speakers grew to multiple seconds and only reset on
rejoining voice. Root cause was the receiver playout logic, not the
codec settings: the playout clock free-ran in real time while the
sender omitted silence from its timestamps (and set no header flags),
and the only correction snapped the clock to the *oldest* buffered
frame — which could only ever add standing latency. target_depth_ms_
was computed but never enforced, so latency could only grow or reset.

Fix: bound playout against the stream's leading edge (newest frame).
(Re)seed to the leading edge on start/marker/starve (no prebuffer, so
latency stays low), and frame-skip catch-up trims any backlog beyond
target+hysteresis — the missing downward force.

Hardening: sender now stamps kFlagMarker (talkspurt start) and kFlagDtx,
consumed on recv for clean resync; adaptive late-drop window; EWMA
outlier rejection so silence gaps/stragglers don't poison the estimate;
duplicate counting and ring-underrun diagnostics.

New test_jitter_depth asserts depth stays bounded (<200ms) while
arrivals outrun playback. ctest --preset dev green (27/27).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 20:02:20 +02:00
e155e342f4 feat(audio): channel sample_rate caps Opus bandwidth (narrowband/wideband)
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
The per-channel sample_rate field was inert after pinning the codec to
48 kHz. Make it meaningful without changing the 48 kHz clock: carry it as
OpusParams::max_bandwidth_hz and apply OPUS_SET_MAX_BANDWIDTH in
OpusEncoder::init (8000->narrowband, 16000->wideband, 24000->super-wideband,
48000->full). A low-bitrate room can now shed out-of-band content while
every endpoint keeps a single 48 kHz clock.

Make sample_rate channel-authoritative on the server: conn_session no
longer overrides effective sample_rate with the client's always-48000
request (it now behaves like frame_ms/mode). vc_get_stream_audio_config
reports the channel's configured rate for own streams too.

New ctest channel_samplerate: a 7 kHz tone is attenuated ~1000x on an
8 kHz (narrowband) channel vs a 48 kHz (full-band) channel, proving the cap
is in effect. ctest --preset dev 26/26.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 16:56:49 +02:00
a460009a2f fix(audio): reframe send path to channel frame_ms; pin codec to 48 kHz
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
The AudioEngine capture clock is fixed at 48 kHz / 20 ms (960-sample
frames), but a channel may set any Opus frame_ms (2.5..60 ms, voice.md
§3) and the server enforces it unclamped. on_capture_frame handed the
engine's 960-sample frame straight to an encoder configured for the
channel's window: frame_ms > 20 was silently ignored, and frame_ms < 20
broke entirely (receiver sized its decode buffer too small ->
OPUS_BUFFER_TOO_SMALL -> dead audio). Affected the hardware mic and
vc_stream_feed_pcm alike.

Reframe each captured/fed block to ls.frame_samples via a per-LocalStream
accumulator (pre-sized at announce, no RT-thread alloc) before
encode_and_send_frame; the 20 ms case stays a zero-copy fast path. Also
pin the codec to 48 kHz in opus_params_from_audio_config — it was honoring
a non-48k effective sample_rate against a 48 kHz PCM clock.

New ctest frame_ms_reframe covers 40 ms (accumulate) and 10 ms (split)
feed->encode->relay->decode->sink round trips. ctest --preset dev 25/25.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 16:45:02 +02:00
50416c33a2 feat(clients): event sound effects + optional text-to-speech
Add audible cues and optional spoken announcements for session events
(join/leave, channel + PM sent/recv, login, logout/connection-lost,
mic on/off, voice-activity, PTT) across all three clients, driven off
the shared C ABI vc_event stream so the mapping stays consistent.

TTS is off by default; when enabled it announces events and reads
message/PM bodies aloud. Master toggles + a sound-volume slider; the
per-utterance voice-activity and PTT cues default off. WAVs ship from
assets/sounds/.

Windows (built + verified): new VoiceCat.App/Notifications/ layer
(FeedbackSettings -> %AppData%\VoiceCat\feedback.json, SoundPlayerPool
via System.Media.SoundPlayer, SpeechAnnouncer via Prismatoid 0.3.0,
EventFeedback dispatcher); MainForm hooks; NotificationSettingsForm
under Settings > Notifications; csproj adds the Prismatoid PackageRef
and copies the WAVs into sounds\.

macOS + iOS (written, not yet built -- needs a Mac): shared
VoiceCatCore/Feedback/ (SoundEvent, EventFeedback = AVAudioPlayer pool
+ native AVSpeechSynthesizer, FeedbackSettings over UserDefaults); WAVs
bundled via Package.swift resources (.process). Hooks in SessionState/
AppState (iOS) and MainWindowController (macOS); settings UI in
SettingsView (iOS) and SettingsWindowController (macOS).

No core/server code touched; ctest --preset dev unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 15:20:31 +02:00
725bd8e925 fix(windows): screen-reader accessibility for context menus and channel tree
- ContextMenuStrip: select first item on Opened so keyboard-invoked menus
  (Shift+F10 / Apps key) raise the UIA focus event immediately instead of
  staying silent until the first arrow key.
- Name the SplitContainer/SplitterPanel containers so screen readers announce
  orientation instead of a stack of anonymous pane nodes.
- Take the resize splitters out of the Tab cycle (TabStop = false) so focus
  moves control-to-control.
- RefreshChannelTree: restore keyboard focus and re-announce the current node
  after a Nodes.Clear()/rebuild, so the tree no longer loses focus when
  channels/users change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 13:34:26 +02:00
483f889910 fix(server): bind UDP media to the TCP port so self-host needs one forward rule
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
media_port defaulted to 0 (OS-assigned) and --port only set the TCP bind_port,
so the UDP relay bound a random high port and advertised it to clients in HELLO.
Self-hosters forwarding only 8384/udp saw connect-OK-but-no-voice, contradicting
docs/deployment.md (control and media share one port). Media now follows
bind_port when media_port is unset; 0=OS-assigned survives when bind_port is also
0 so ephemeral-port tests are unaffected. Banner now reads TCP :8384 UDP :8384.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 12:48:13 +02:00
5e18dfa1c9 fix(windows): real native exclude + self-echo removal for screen audio
"All apps except selected" previously captured the complement of a frozen
app snapshot in INCLUDE mode (missed late-launched apps and system sounds,
wasted captures on silent windows). It now opens a single ProcessLoopbackCapture
in EXCLUDE mode (AUDIOCLIENT_PROCESS_LOOPBACK_MODE_EXCLUDE_TARGET_PROCESS_TREE)
of the one chosen app — true system-mix-minus-one, dynamic so apps launched
after sharing starts are included. The picker enforces single-selection in
exclude mode (the activation params take one target PID).

Adds an "Exclude VoiceCat's own audio (prevents echo)" checkbox (default on,
entire-desktop only) that routes the desktop capture through the same EXCLUDE
path targeting our own process id, killing the whole-device self-echo loop.

No C++/ABI changes. Updates voice.md and PROGRESS.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 12:28:53 +02:00
213 changed files with 389800 additions and 2600 deletions

5
.gitignore vendored
View File

@@ -14,9 +14,8 @@
*.exe *.exe
*.pdb *.pdb
# vcpkg # vcpkg (bundled as a submodule at /vcpkg — see docs/building.md §2)
/vcpkg_installed/ /vcpkg_installed/
/vcpkg/
# Generated protobuf # Generated protobuf
*.pb.cc *.pb.cc
@@ -36,7 +35,7 @@
.DS_Store .DS_Store
Thumbs.db Thumbs.db
# Apple / Windows client build artifacts (added in M4) # Apple / Windows client build artifacts
clients/apple/**/build/ clients/apple/**/build/
clients/apple/**/*.xcodeproj/xcuserdata/ clients/apple/**/*.xcodeproj/xcuserdata/
clients/apple/**/*.xcodeproj/project.xcworkspace/ clients/apple/**/*.xcodeproj/project.xcworkspace/

3
.gitmodules vendored Normal file
View File

@@ -0,0 +1,3 @@
[submodule "vcpkg"]
path = vcpkg
url = https://github.com/microsoft/vcpkg.git

View File

@@ -38,15 +38,20 @@ making progress.
## Build ## Build
Default development preset (real deps via vcpkg — works on Windows/Linux/macOS): Default development preset (real deps via vcpkg — works on Windows/Linux/macOS). vcpkg is
bundled as a git submodule at `vcpkg/`, pinned to `vcpkg.json`'s `builtin-baseline`:
```bash ```bash
export VCPKG_ROOT=/path/to/vcpkg # bootstrap vcpkg first; cross-platform git submodule update --init vcpkg # one-time, after cloning
./vcpkg/bootstrap-vcpkg.sh # .bat on Windows
cmake --preset dev cmake --preset dev
cmake --build --preset dev cmake --build --preset dev
ctest --preset dev ctest --preset dev
``` ```
To use an external vcpkg checkout instead, `export VCPKG_ROOT=/path/to/vcpkg` — it always
takes priority over the bundled submodule.
Skeleton (no third-party deps — works immediately, no vcpkg needed): Skeleton (no third-party deps — works immediately, no vcpkg needed):
```bash ```bash
@@ -68,7 +73,8 @@ for the full matrix.
`vcpkg.json` pins all deps to a fixed vcpkg baseline — `cmake --preset dev` resolves them `vcpkg.json` pins all deps to a fixed vcpkg baseline — `cmake --preset dev` resolves them
automatically on first configure. The vcpkg triplet is auto-resolved from the host platform automatically on first configure. The vcpkg triplet is auto-resolved from the host platform
by [`cmake/voicecat-toolchain.cmake`](cmake/voicecat-toolchain.cmake). by [`cmake/voicecat-toolchain.cmake`](cmake/voicecat-toolchain.cmake), which also resolves
`VCPKG_ROOT` (env var override, else the bundled `vcpkg/` submodule).
## Where each subsystem lives (and its doc) ## Where each subsystem lives (and its doc)

View File

@@ -8,10 +8,14 @@ and what's next* read [`PROGRESS.md`](PROGRESS.md); for *design* read [`docs/`](
> server-mute, channel CRUD, in-app account management, disconnect/keepalive/reaper. Windows > server-mute, channel CRUD, in-app account management, disconnect/keepalive/reaper. Windows
> WinForms C# client shipped (M4). **macOS AppKit client shipped** — `VoiceCatMac.xcodeproj` > WinForms C# client shipped (M4). **macOS AppKit client shipped** — `VoiceCatMac.xcodeproj`
> at `clients/apple/macOS/`. **iOS SwiftUI client shipped** — `VoiceCatiOS.xcodeproj` at > at `clients/apple/macOS/`. **iOS SwiftUI client shipped** — `VoiceCatiOS.xcodeproj` at
> `clients/apple/iOS/`. `ctest --preset dev` green — 24/24 tests. > `clients/apple/iOS/`. `ctest --preset dev` green — 29/29 tests.
> External PCM feed/tap API (`vc_stream_feed_pcm` + `vc_set_pcm_sink`) shipped. > External PCM feed/tap API (`vc_stream_feed_pcm` + `vc_set_pcm_sink`) shipped.
> **Screen-audio sharing shipped on macOS (ScreenCaptureKit) and iOS (ReplayKit Broadcast > **Screen-audio sharing shipped on macOS (ScreenCaptureKit) and iOS (ReplayKit Broadcast
> Upload Extension → host App Group ring → `vc_stream_feed_pcm`).** See [`PROGRESS.md`](PROGRESS.md). > Upload Extension → host App Group ring → `vc_stream_feed_pcm`).**
> **Noise suppression shipped (RNNoise, vendored at `third_party/rnnoise/`)** — both send-side
> mic NR (`vc_set_input_noise_reduction`) and per-listener receive NR; client on/off toggles ship
on all three clients (receive NR now denoises stereo mic streams too — fixed 2026-06-23).
> See [`PROGRESS.md`](PROGRESS.md).
VoiceCat = self-hosted native voice & text chat (TeamSpeak/Mumble-style). Plain TCP (control) VoiceCat = self-hosted native voice & text chat (TeamSpeak/Mumble-style). Plain TCP (control)
+ UDP (media), no WebRTC, encrypted by default. A shared C++ core (`libvoicecat`) drives + UDP (media), no WebRTC, encrypted by default. A shared C++ core (`libvoicecat`) drives
@@ -64,10 +68,18 @@ Vcpkg triplet is auto-resolved from the host platform by
Windows, `x64-linux` on Linux, `arm64-osx` on Apple Silicon. See docs/building.md §1 Windows, `x64-linux` on Linux, `arm64-osx` on Apple Silicon. See docs/building.md §1
"Platform matrix" for details. "Platform matrix" for details.
One-time vcpkg setup: vcpkg is bundled as a git submodule at `vcpkg/`, pinned to the commit in `vcpkg.json`'s
`builtin-baseline`. One-time setup after cloning:
```bash
git submodule update --init vcpkg
./vcpkg/bootstrap-vcpkg.sh # .bat on Windows
```
To use an external vcpkg checkout instead (e.g. one shared across projects), set
`VCPKG_ROOT` — it always takes priority over the bundled submodule:
```bash ```bash
# one-time: git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh (.bat on Windows)
export VCPKG_ROOT=/path/to/vcpkg # works on Linux / macOS / Windows export VCPKG_ROOT=/path/to/vcpkg # works on Linux / macOS / Windows
``` ```

View File

@@ -3,7 +3,8 @@ cmake_minimum_required(VERSION 3.25)
project(voicecat project(voicecat
VERSION 0.0.1 VERSION 0.0.1
DESCRIPTION "Self-hosted native voice & text chat (see docs/)" DESCRIPTION "Self-hosted native voice & text chat (see docs/)"
LANGUAGES CXX) # C is needed for the vendored RNNoise noise-suppression lib (third_party/rnnoise).
LANGUAGES CXX C)
# On iOS, audio_engine.cpp includes miniaudio.h which pulls in AVFoundation Objective-C # On iOS, audio_engine.cpp includes miniaudio.h which pulls in AVFoundation Objective-C
# headers. Those cannot be compiled as C++; we set audio_engine.cpp's LANGUAGE to OBJCXX # headers. Those cannot be compiled as C++; we set audio_engine.cpp's LANGUAGE to OBJCXX
@@ -13,11 +14,6 @@ if(CMAKE_SYSTEM_NAME STREQUAL "iOS")
endif() endif()
# ── Options ─────────────────────────────────────────────────────────────────── # ── Options ───────────────────────────────────────────────────────────────────
# The M0 skeleton compiles with NO third-party dependencies: every subsystem is a
# stub that returns VC_ERR_NOT_IMPLEMENTED. As each subsystem is built out, flip
# VOICECAT_USE_VCPKG_DEPS=ON so CMake pulls the real libraries (mbedTLS, libsodium,
# opus, protobuf, ...) via the vcpkg toolchain (see vcpkg.json / docs/tech-stack.md).
option(VOICECAT_USE_VCPKG_DEPS "Link real third-party deps via vcpkg" OFF)
option(VOICECAT_BUILD_SERVER "Build voicecat-server" ON) option(VOICECAT_BUILD_SERVER "Build voicecat-server" ON)
option(VOICECAT_BUILD_TOOLS "Build the vccli headless test client" ON) option(VOICECAT_BUILD_TOOLS "Build the vccli headless test client" ON)
option(VOICECAT_BUILD_TESTS "Build tests" ON) option(VOICECAT_BUILD_TESTS "Build tests" ON)
@@ -36,6 +32,20 @@ set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/bin)
set(CMAKE_LIBRARY_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/lib) set(CMAKE_LIBRARY_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/lib)
set(CMAKE_ARCHIVE_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/lib) set(CMAKE_ARCHIVE_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/lib)
# ── Static MinGW runtime (Windows) ────────────────────────────────────────────
# x64-mingw-static only statically links vcpkg's OWN deps (protobuf, sodium, ...);
# the GCC/MinGW runtime stays dynamic by default, so every produced .exe/.dll
# otherwise depends on libgcc_s_seh-1.dll / libwinpthread-1.dll / libstdc++-6.dll
# at runtime — DLLs that exist on a dev box (MSYS2/UCRT64) but not on a clean
# Windows machine, where the server then fails to start with "… was not found".
# Apply the fully-static-MinGW recipe to ALL targets so binaries are portable.
# (The SHARED voicecat.dll already sets these in core/CMakeLists.txt; the duplicate
# is harmless. Verify with: objdump -p build/<preset>/bin/voicecat-server.exe |
# grep "DLL Name" — only Windows system DLLs should remain.)
if(WIN32 AND MINGW)
add_link_options(-static-libgcc -static-libstdc++ -static -lwinpthread)
endif()
# ── Targets ─────────────────────────────────────────────────────────────────── # ── Targets ───────────────────────────────────────────────────────────────────
add_subdirectory(core) add_subdirectory(core)
@@ -45,9 +55,7 @@ endif()
if(VOICECAT_BUILD_TOOLS) if(VOICECAT_BUILD_TOOLS)
add_subdirectory(tools/vccli) add_subdirectory(tools/vccli)
if(VOICECAT_USE_VCPKG_DEPS) add_subdirectory(tools/voicecat-admin)
add_subdirectory(tools/voicecat-admin)
endif()
endif() endif()
if(VOICECAT_BUILD_TESTS) if(VOICECAT_BUILD_TESTS)
@@ -56,5 +64,4 @@ if(VOICECAT_BUILD_TESTS)
endif() endif()
message(STATUS "VoiceCat ${PROJECT_VERSION} configured " message(STATUS "VoiceCat ${PROJECT_VERSION} configured "
"(vcpkg deps: ${VOICECAT_USE_VCPKG_DEPS}, " "(server: ${VOICECAT_BUILD_SERVER}, tools: ${VOICECAT_BUILD_TOOLS})")
"server: ${VOICECAT_BUILD_SERVER}, tools: ${VOICECAT_BUILD_TOOLS})")

View File

@@ -5,28 +5,17 @@
{ {
"name": "vcpkg-common", "name": "vcpkg-common",
"hidden": true, "hidden": true,
"description": "Shared base for all presets that link real deps via vcpkg. Uses cmake/voicecat-toolchain.cmake, which auto-resolves VCPKG_TARGET_TRIPLET / VCPKG_HOST_TRIPLET from the host platform (x64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon). Cross-compile presets override VCPKG_TARGET_TRIPLET in their cacheVariables. Requires VCPKG_ROOT in the environment.", "description": "Shared base for all presets that link real deps via vcpkg. Uses cmake/voicecat-toolchain.cmake, which auto-resolves VCPKG_TARGET_TRIPLET / VCPKG_HOST_TRIPLET from the host platform (x64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon). Cross-compile presets override VCPKG_TARGET_TRIPLET in their cacheVariables. Resolves vcpkg from the bundled git submodule (vcpkg/) unless VCPKG_ROOT points at an external checkout.",
"generator": "Ninja", "generator": "Ninja",
"toolchainFile": "${sourceDir}/cmake/voicecat-toolchain.cmake", "toolchainFile": "${sourceDir}/cmake/voicecat-toolchain.cmake",
"cacheVariables": { "cacheVariables": {
"VOICECAT_USE_VCPKG_DEPS": "ON" "VOICECAT_USE_VCPKG_DEPS": "ON"
} }
}, },
{
"name": "skeleton",
"displayName": "Skeleton (no third-party deps)",
"description": "Builds the stub skeleton with just a C++20 compiler — no vcpkg needed. Subsystems return VC_ERR_NOT_IMPLEMENTED. Good for 'does the repo even build' smoke checks. Runs 2 tests (smoke + frame_codec).",
"generator": "Ninja",
"binaryDir": "${sourceDir}/build/skeleton",
"cacheVariables": {
"CMAKE_BUILD_TYPE": "Debug",
"VOICECAT_USE_VCPKG_DEPS": "OFF"
}
},
{ {
"name": "dev", "name": "dev",
"displayName": "Dev (full real-deps build, vcpkg)", "displayName": "Dev (full real-deps build, vcpkg)",
"description": "Day-to-day development preset. Real protocol, crypto, voice, server — everything from M1 onward. Builds server + tools + tests (21 tests). Auto-triplet: x64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon. Requires VCPKG_ROOT.", "description": "Day-to-day development preset. Real protocol, crypto, voice, server. Builds server + tools + tests. Auto-triplet: x64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon. Requires vcpkg bootstrapped (bundled submodule or VCPKG_ROOT).",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/dev", "binaryDir": "${sourceDir}/build/dev",
"cacheVariables": { "cacheVariables": {
@@ -38,7 +27,7 @@
{ {
"name": "release", "name": "release",
"displayName": "Release (optimized, tests on, symbols kept)", "displayName": "Release (optimized, tests on, symbols kept)",
"description": "Optimized build with the full test suite enabled. Use to run tests against optimized code, profile, or catch optimizer-sensitive bugs. Symbols are kept (not stripped) so stack traces and profiling remain useful. Auto-triplet. Requires VCPKG_ROOT.", "description": "Optimized build with the full test suite enabled. Use to run tests against optimized code, profile, or catch optimizer-sensitive bugs. Symbols are kept (not stripped) so stack traces and profiling remain useful. Auto-triplet. Requires vcpkg bootstrapped (bundled submodule or VCPKG_ROOT).",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/release", "binaryDir": "${sourceDir}/build/release",
"cacheVariables": { "cacheVariables": {
@@ -50,7 +39,7 @@
{ {
"name": "server-release", "name": "server-release",
"displayName": "Server Release (optimized, stripped, no tests)", "displayName": "Server Release (optimized, stripped, no tests)",
"description": "Production-shaped build for deployment. Optimized (Release) with stripped binaries (-s linker flag), no tests. This is what you'd ship/run — see docs/deployment.md. Auto-triplet. Requires VCPKG_ROOT.", "description": "Production-shaped build for deployment. Optimized (Release) with stripped binaries (-s linker flag), no tests. This is what you'd ship/run — see docs/deployment.md. Auto-triplet. Requires vcpkg bootstrapped (bundled submodule or VCPKG_ROOT).",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/server-release", "binaryDir": "${sourceDir}/build/server-release",
"cacheVariables": { "cacheVariables": {
@@ -63,7 +52,7 @@
}, },
{ {
"name": "windows-client", "name": "windows-client",
"displayName": "Windows client (voicecat.dll for C# WinForms, M4)", "displayName": "Windows client (voicecat.dll for C# WinForms)",
"description": "Produces a redistributable Release voicecat.dll with no MinGW runtime DLL dependencies (see core/CMakeLists.txt's static-runtime link flags and clients/windows/README.md). Server/tools/tests are off — this preset exists only to build the DLL. Windows only.", "description": "Produces a redistributable Release voicecat.dll with no MinGW runtime DLL dependencies (see core/CMakeLists.txt's static-runtime link flags and clients/windows/README.md). Server/tools/tests are off — this preset exists only to build the DLL. Windows only.",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/windows-client", "binaryDir": "${sourceDir}/build/windows-client",
@@ -78,7 +67,7 @@
{ {
"name": "apple-dev", "name": "apple-dev",
"displayName": "Apple macOS (libvoicecat.a for Swift Package, scaffolding)", "displayName": "Apple macOS (libvoicecat.a for Swift Package, scaffolding)",
"description": "SCAFFOLDING — not yet CI-validated; build on macOS to verify. Produces a static libvoicecat.a for macOS (arm64-osx on Apple Silicon, x64-osx on Intel) for consumption by the Swift Package / XCFramework. Server/tools/tests off. Requires VCPKG_ROOT.", "description": "SCAFFOLDING — not yet CI-validated; build on macOS to verify. Produces a static libvoicecat.a for macOS (arm64-osx on Apple Silicon, x64-osx on Intel) for consumption by the Swift Package / XCFramework. Server/tools/tests off. Requires vcpkg bootstrapped (bundled submodule or VCPKG_ROOT).",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/apple-dev", "binaryDir": "${sourceDir}/build/apple-dev",
"cacheVariables": { "cacheVariables": {
@@ -91,7 +80,7 @@
{ {
"name": "apple-ios", "name": "apple-ios",
"displayName": "Apple iOS device (XCFramework slice)", "displayName": "Apple iOS device (XCFramework slice)",
"description": "Cross-compiles a static libvoicecat.a for iOS device (arm64). One slice of the XCFramework. Server/tools/tests off. Requires VCPKG_ROOT and a macOS host with iOS SDK. Uses cmake/vcpkg-overlays/triplets/arm64-ios.cmake (release-only, correct autoconf host triple).", "description": "Cross-compiles a static libvoicecat.a for iOS device (arm64). One slice of the XCFramework. Server/tools/tests off. Requires vcpkg bootstrapped (bundled submodule or VCPKG_ROOT) and a macOS host with iOS SDK. Uses cmake/vcpkg-overlays/triplets/arm64-ios.cmake (release-only, correct autoconf host triple).",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/apple-ios", "binaryDir": "${sourceDir}/build/apple-ios",
"cacheVariables": { "cacheVariables": {
@@ -111,7 +100,7 @@
{ {
"name": "apple-ios-sim", "name": "apple-ios-sim",
"displayName": "Apple iOS simulator (XCFramework slice)", "displayName": "Apple iOS simulator (XCFramework slice)",
"description": "Cross-compiles a static libvoicecat.a for iOS simulator (arm64-ios-simulator). One slice of the XCFramework. Server/tools/tests off. Requires VCPKG_ROOT and a macOS host with iOS simulator SDK. Uses cmake/vcpkg-overlays/triplets/arm64-ios-simulator.cmake (release-only, correct autoconf host triple).", "description": "Cross-compiles a static libvoicecat.a for iOS simulator (arm64-ios-simulator). One slice of the XCFramework. Server/tools/tests off. Requires vcpkg bootstrapped (bundled submodule or VCPKG_ROOT) and a macOS host with iOS simulator SDK. Uses cmake/vcpkg-overlays/triplets/arm64-ios-simulator.cmake (release-only, correct autoconf host triple).",
"inherits": "vcpkg-common", "inherits": "vcpkg-common",
"binaryDir": "${sourceDir}/build/apple-ios-sim", "binaryDir": "${sourceDir}/build/apple-ios-sim",
"cacheVariables": { "cacheVariables": {
@@ -130,7 +119,6 @@
} }
], ],
"buildPresets": [ "buildPresets": [
{ "name": "skeleton", "configurePreset": "skeleton" },
{ "name": "dev", "configurePreset": "dev" }, { "name": "dev", "configurePreset": "dev" },
{ "name": "release", "configurePreset": "release" }, { "name": "release", "configurePreset": "release" },
{ "name": "server-release", "configurePreset": "server-release" }, { "name": "server-release", "configurePreset": "server-release" },
@@ -140,7 +128,6 @@
{ "name": "apple-ios-sim", "configurePreset": "apple-ios-sim" } { "name": "apple-ios-sim", "configurePreset": "apple-ios-sim" }
], ],
"testPresets": [ "testPresets": [
{ "name": "skeleton", "configurePreset": "skeleton", "output": { "outputOnFailure": true } },
{ "name": "dev", "configurePreset": "dev", "output": { "outputOnFailure": true } }, { "name": "dev", "configurePreset": "dev", "output": { "outputOnFailure": true } },
{ "name": "release", "configurePreset": "release", "output": { "outputOnFailure": true } } { "name": "release", "configurePreset": "release", "output": { "outputOnFailure": true } }
] ]

View File

@@ -10,8 +10,490 @@ up instantly. Newest status at the top.
## ▶ Where we left off / next action ## ▶ Where we left off / next action
- **Awaiting on-device verification (2026-06-22):** **iOS real echo cancellation / noise suppression - **Done (2026-07-23):** **First comment-density cleanup across core, server, and native
via native VPIO.** Root cause of "voice chat doesn't sound like a call" (echo + no NR): real iOS clients.** Condensed comments in the highest-noise audio, reconnect, registry, and binding
files; removed implementation history and narration; retained ABI ownership, threading,
real-time, ordering, and OS-API invariants. Moved the durable iOS audio-routing/pacing rules
to `docs/voice.md` and client reconnection policy to `docs/protocol.md`. No behavior, wire
format, or C ABI changes. **Verification:** `cmake --build --preset dev` green;
`ctest --preset dev` 29/29 green. `dotnet build VoiceCat.slnx` restores dependencies and
builds `VoiceCat.Interop` + `VoiceCat.App`, then fails in the unchanged test project because
`ExternalPcmTests.cs:49` references internal `NativeMethods` (`CS0122`).
- **Done (2026-06-25):** **Fixed iOS AirPods-disconnect reinitialize loop on A2DP presets
(Stereo Mic / Mono Mic).** Regression from the 2026-06-25 audio-device-change recovery
commit below, which broadened the route-change recovery set from
`{oldDeviceUnavailable, newDeviceAvailable}` to "everything except
categoryChange/routeConfigurationChange". That added `.override` to the recovery set, and
`.override` is fired by our own `applyA2dpSpeakerFallback()`
`overrideOutputAudioPort(.speaker)` — which `recoverAudio()` calls on every recovery. On an
A2DP preset with AirPods connected, disconnecting them ran:
`oldDeviceUnavailable``recoverAudio()``applyA2dpSpeakerFallback()` (no external
output now) → `overrideOutputAudioPort(.speaker)``.override` routeChange →
`recoverAudio()``applyConfiguration()` (setCategory resets the override) →
`applyA2dpSpeakerFallback()``overrideOutputAudioPort(.speaker)``.override` → …
Each iteration also called `IOSAudioEngine.reconfigure()``rebuild()` (a full
stop/restart of `AVAudioEngine`), which is the audible reinitialize loop + CPU spin the
user reported. Voice Chat (`.btHfpVoice`) and Built-in Mic + Speaker were unaffected
because `applyA2dpSpeakerFallback` early-returns for non-A2DP modes (no
`overrideOutputAudioPort` call, no `.override` notification).
Two-part fix (no C ABI / proto / docs changes — pure Swift iOS-app target):
1. **`AudioSessionManager.handleRouteChange`** (`AudioSessionManager.swift:182`): added
`.override` to the skip list alongside `.categoryChange`/`.routeConfigurationChange`.
`.override` is only ever fired by our own `overrideOutputAudioPort` call, so treating
it as a recovery reason is the loop by definition. The
`AVAudioEngineConfigurationChange` observer in `IOSVoiceProcessingEngine` remains as
the backstop for the case where an override actually stops the engine.
2. **`IOSAudioRouter.applyA2dpSpeakerFallback`** (`IOSAudioRouter.swift`): made idempotent
via a `lastAppliedOutputOverride` tracker. Skips the `overrideOutputAudioPort` call
when the desired override (`.none` for external output present, `.speaker` otherwise)
already matches the last successfully applied value — so even if some other path
re-enters, the redundant override (and its `.override` notification) isn't fired. The
tracker is reset to `nil` at the top of `applyConfiguration()` (setCategory can reset
the override) and on a failed call. Defense-in-depth on top of fix 1.
**Build:** `xcodebuild -project clients/apple/iOS/VoiceCatiOS.xcodeproj -scheme VoiceCatiOS
-destination 'generic/platform=iOS' build` green (Xcode 26.5 / iOS 18.0). The standalone
`swift test` in `clients/apple/` fails with `no such module 'VoiceCatC'` — pre-existing
(confirmed by stashing the changes: fails identically without them); the `VoiceCatC` C ABI
XCFramework isn't on SwiftPM's resolver path in this workspace. Not caused by this change
(the edit is in the iOS app target, not the `VoiceCatCore` SwiftPM package).
**Next (manual, on-device):** connect on the Stereo Mic preset, join voice, disconnect
AirPods — expect ONE `oldDeviceUnavailable` → one `recoverAudio` → one `engine started`
one `override` routeChange (skipped, no further `recoverAudio`) and steady audio through
the loudspeaker. Also sanity-check AirPods reconnect and wired headphone plug/unplug
recover exactly once.
- **Done (2026-06-25):** **iOS robustness — auto-reconnect after a network change + audio
recovery when audio devices plug/unplug.** Two layers of bugs the iOS client had:
(a) a `VC_EVENT_DISCONNECTED` from the C core on a Wi-Fi→cellular flip / DNS outage /
server restart used to leave the session dead with no retry; (b) unplugging wired
headphones or AirPods left the engine stopped forever — mic stopped transmitting and
remote audio stayed silent (the server connection itself survived, but the audio graph
did not recover).
The first attempt wired reconnect into `AppState.handleConnectEvent`, but that handler
never runs for a live-session disconnect: once `SessionState.init` overwrites
`client.onEvent` (`SessionState.swift:87`), the `.disconnected` event is delivered to
`SessionState.handleEvent`, which used to play a cue and do nothing else. So the live
session would sit as a zombie for ~30-60 s (the C core's TCP keepalive/reaper timeout)
and then play the "connection lost" sound with no reconnect armed — exactly what the
user saw. The fix below has two parts addressing both the missing reconnect AND the
long wait.
1. **Event-driven reconnect** (`AppState.swift`, `SessionState.swift`): added a
`weak var appState: AppState?` to `SessionState`, set by AppState on auth success.
`SessionState.handleEvent` `.disconnected` now plays the cue and calls
`appState?.onLiveSessionDisconnected()` — the SINGLE path by which AppState learns a
live session dropped (since its own `handleConnectEvent` is bypassed for live-session
events). `onLiveSessionDisconnected` calls a shared `teardownLiveSessionAndReconnect`
that snapshots the live session into `LastSession`, stops the audio engine,
deactivates the AVAudioSession, nil's `session` (which releases `VoiceCatClient`
`vc_client_destroy` joins the io thread), resets the backoff counter, and arms
`scheduleReconnect`.
2. **Path-driven proactive reconnect** (`AppState.swift`): an `NWPathMonitor`
(`Network.framework`) now runs the whole time we're CONNECTED (started on auth
success, not only when armed for reconnect) and stays armed across reconnects. Its
`pathUpdateHandler` (dispatched to @MainActor) does two things:
- While connected: a primary-interface change (Wi-Fi↔cellular) OR the path becoming
`.unsatisfied` triggers `proactiveReconnect()` — tearing the live session down
BEFORE the C core notices the dead TCP read. This is what collapses the 30-60 s
reaper wait into ~1 s + the first backoff tick. Same-interface refreshes (Wi-Fi
BSSID roams, signal-strength changes) are intentionally ignored (signature
comparison via `pathSignature`); those usually don't break the TCP connection.
- While mid-reconnect (no session): a path becoming `.satisfied` resets the backoff
counter and arms `scheduleReconnect` for a fast-fresh retry.
`userInitiatedDisconnect` distinguishes manual `disconnect()`/`cancelConnect()` (which
set it true → cancel all reconnect state) from a network drop (which leaves it false).
On a successful reconnect, `reconnectAttempt` resets and `lastSession` clears; the
path monitor keeps watching for the next change. On user-initiated disconnect, all
reconnect state (task + path monitor + `lastSession` + `connectedServer`) is cancelled.
3. **Backoff + restore**: exponential backoff 1s → 2s → 4s → 8s → 16s → 30s cap,
indefinite. TOFU pins match on the second connect (`VC_TOFU_MATCHED`) so the identity
gate auto-confirms; on auth success `SessionState.requestRestore` issues a
`joinChannel` and re-arms voice + restores the local mute/deafen state on the
resulting `.joinResult`.
4. **Audio recovery** (`AudioSessionManager.swift`, `IOSVoiceProcessingEngine.swift`):
replaced the route-change handler's narrow `.oldDeviceUnavailable`/
`.newDeviceAvailable` guard with a single intent-gated `recoverAudio()` path that
re-activates the AVAudioSession, re-applies the route config, and rebuilds the
engine; it runs on every externally-initiated route change reason except
`.categoryChange`/`.routeConfigurationChange` (those we cause ourselves and would
loop). Interruption-end now always calls `recoverAudio()` instead of only when
`.shouldResume` is set (which left the session permanently dead after Siri). Added
an `AVAudioEngineConfigurationChange` observer on the engine in `IOSAudioEngine`
that catches the case where iOS stops the engine itself AFTER our route-change
handler already rebuilt it (the previous rebuilds raced the engine's own self-stop
and lost). And `IOSAudioEngine.rebuild()` now does a one-shot reactivation-retry on
`engine.start()` failure — iOS sometimes refuses to start until the AVAudioSession is
re-activated, which is the silent-death case.
**Build:** `scripts/build-ios-client.sh --no-configure` green (Xcode 26.5 / iOS 18.0 sim
SDK, Swift 5 mode). No C ABI / `voicecat.h` / `voicecat.proto` / C core changes; the
existing TOFU auto-confirm (`VC_TOFU_MATCHED`) and idempotent `vc_join_channel` make
reconnect+restore possible without new C ABI. macOS and Windows clients unchanged.
**Next (manual, on-device):** verify unplugging AirPods/wired headphones mid-call keeps
audio going through the loudspeaker; verify Wi-Fi→cellular flip mid-call now triggers a
FAST reconnect (within a couple seconds, not 30-60 s) and lands in the same channel with
voice re-armed; verify tapping Disconnect mid-reconnect-abort cancels cleanly.
- **Done (2026-06-24):** **Three bug fixes — voice join/leave, channel edit defaults, channel-update stream restart.**
1. **Join/Leave Voice now truly subscribes/unsubscribes from the voice plane.** Previously
"Join Voice" only started the local mic — receiving was always on (gated by channel
membership alone). Added a protocol-level voice subscription concept: new
`SubscribeVoiceRequest`/`UnsubscribeVoiceRequest`/`VoiceSubscriptionResult` proto messages
(`core/proto/voicecat.proto`), `User.voice_subscribed` field, `vc_join_voice`/`vc_leave_voice`
C ABI functions (`core/include/voicecat.h`), `VC_EVENT_VOICE_STATE` event, server-side
`voice_subscribed_` flag on `ConnSession` checked by the SFU relay's recipient filter
(`SessionRegistry::find_channel_sessions` excludes non-subscribers; `MediaRelay::on_udp_frame`
also skips non-subscribed senders). The core client gates `sync_remote_streams` on
`voice_subscribed_`, tears down all remote decoders + stops local streams on leave, and
re-syncs from the session model on join. All three clients (Windows/macOS/iOS) rewired
their Join/Leave Voice button to call `joinVoice`+start mic / `leaveVoice`+core stops mic.
The configured input mode (PTT/VAD/AlwaysOn) takes effect on join — no extra mic button.
Text chat works regardless of voice subscription. **Apple clients not yet compile-verified
(Windows environment).**
2. **Channel edit dialog now shows the channel's actual current settings.** The read struct
`vc_channel` (`voicecat.h`) was missing `sort_order` and `audio` fields — only the write
struct `vc_channel_info` had them. Extended `vc_channel` with both (additive, no ABI break),
updated the session model (`session::Channel`) and `apply_snapshot`/`apply_channel_event`
to populate them, and updated `vc_list_channels` marshaling. All three clients now build
the edit descriptor from the actual channel info instead of hardcoded defaults.
3. **Channel parameter updates now automatically restart everyone's streams.** Previously
editing a channel's audio config (codec/bitrate/sample-rate/FEC/DTX/etc.) persisted and
broadcast a `ChannelEvent::UPDATED`, but no layer restarted streams — encoders/decoders
are frozen at announce time. `handle_channel_event` (`core/src/core/client.cpp`) now
detects audio-config changes on the user's current channel and calls
`restart_active_streams_for_channel`, which stop→starts each active local stream. The
server reads the updated channel config on re-announce, and peers' `sync_remote_streams`
wire up fresh decoders at the new ssrc. The `LocalStream` struct now retains the stream
label across restarts. No server or protocol change needed.
- **[ ] Soon — jitter buffer should measure REAL arrival jitter (RFC 3550), not sender
timestamps.** `JitterBuffer::push` (`core/src/audio/audio_engine.cpp:84-108`) estimates
jitter from `gap = ts - last_push_ts_`, where `ts` is the **sender's timestamp** — which is
perfectly regular (`ls.timestamp += samples` every frame, independent of when the packet is
actually sent). So `diff` is always ~0, `jitter_est_` stays 0, and `target_depth_ms_` is
pinned at its ~20 ms floor. The buffer is therefore **blind to real network/arrival jitter
and to bursty senders** — it never deepens. Combined with the playout deliberately seeding
to near-zero depth (`on_playback`, ~line 715), the receiver tolerates only a *steady*
sender. This is exactly why the iOS mic needed a send-side pacing cushion (below) and why
genuine network jitter would also cause underruns. **Fix:** measure inter-arrival jitter
the RFC 3550 way — `D = (arrival_j - arrival_i) - (ts_j - ts_i)` using a wall-clock arrival
stamp captured in `push()` — and drive `target_depth_ms_` off that EWMA (keep the existing
marker/silence-gap outlier rejection). Then the receiver absorbs bursts itself and the iOS
send cushion could be reduced or removed. Shared-core change → add a test and re-verify
desktop↔desktop stays low-latency (steady sender ⇒ ~0 arrival jitter ⇒ no regression).
- **Done (2026-06-24):** **Windows PTT can now work system-wide (in the background).** Previously
the PTT key was focus-scoped (WinForms `KeyDown`/`KeyUp`, dead the moment the window lost
focus). Added an AV-safe global path using the **Raw Input API** (`RegisterRawInputDevices` +
`WM_INPUT` with `RIDEV_INPUTSINK`) — *not* a `WH_KEYBOARD_LL` low-level hook, which is the
keylogger pattern AV heuristics flag (worse for our unsigned MinGW binary). New
`clients/windows/VoiceCat.App/Native/RawInput.cs` (P/Invoke + structs); `MainForm` overrides
`OnHandleCreated`/`OnHandleDestroyed`/`WndProc` to register the keyboard sink and handle
`WM_INPUT`, gates the focus-scoped `KeyDown`/`KeyUp` handlers off when system-wide is on, makes
the `Deactivate` force-release conditional, and adds a `GetAsyncKeyState` watchdog on the pump
timer so a missed key-up (RDP/lock-screen focus switch) can't leave PTT stuck. New
`VoiceSettings.SystemWidePtt` (default ON) with a "Works in the background (system-wide)"
checkbox in the Audio settings PTT section. Build green (`dotnet build`, 0 warnings). **Next
(manual):** verify background PTT against a live server, and confirm the binary trips no AV
keyboard-hook detection.
- **Done (2026-06-23):** **Fixed: receive-side noise reduction silently skipped on stereo mic
streams (regression from stereo-mic capture below).** The per-listener NR toggle
(`vc_set_remote_stream(... noise_reduction)`) did nothing on Windows/macOS/iOS — the UI and
the whole C-ABI→core path were correctly wired, but the decode loop gated the RNNoise pass on
`dec_channels == 1` (`core/src/audio/audio_engine.cpp`), an old proxy for "this stream is
voice" that assumed *stereo ⇒ screen-share*. The stereo-mic commit broke it: a stereo mic with
**send-side NR off** transmits stereo Opus, so the receiver decoded `dec_channels == 2` and
skipped NR entirely (gain/mute have no channel guard, which is why only NR looked broken).
**Fix:** thread the stream *kind* through `init_recv_stream` into `RemoteStream::is_voice`
(set from `si.kind() == STREAM_MIC` in `client.cpp`), gate receive NR on `is_voice` instead of
channel count, and fold a stereo voice frame to mono → denoise → duplicate back across both
channels in place (symmetric with the send-side downmix; RNNoise is mono-only). A stereo voice
stream now plays mono while NR is on; a screen-audio share is never touched. New test
`tests/test_recv_noise_reduction.cpp` drives `AudioEngine` and asserts a stereo voice stream's
noise floor collapses with NR on (RMS 1046 → 0.1) while a screen-audio share stays unchanged
(RMS ≈ 1015). Full `ctest --preset dev` green — **29/29**. Docs: voice.md §10. Clients need no
change (shared-core fix). Not yet re-verified two-client E2E on real hardware.
- **Done (2026-06-23):** **Stereo mic capture on Windows & macOS desktop clients.** Both
desktop mics were hard-mono: `ensure_audio_running()` defaults `capture_channels = 1` and
neither client ever called `vc_set_capture_channels` (only iOS did). Added a **"Stereo
microphone" toggle** to each client's Audio settings (off by default, persisted —
`VoiceSettings.StereoMic` on Windows, `MainWindowController.stereoMic` /
`voice.stereoMic` UserDefaults on macOS). It's applied to the core when the mic stream
starts (stored on the stream before the announce round-trip, so the first device open picks
it up) and live in settings via `vc_set_capture_channels` + `vc_audio_restart`. Exposed both
ABI calls in the Windows interop (`NativeMethods`/`VoiceCatClient`); the macOS wrapper already
had them. **Core fix:** `encode_and_send_frame` (`core/src/core/client.cpp`) now folds a
stereo mic frame to mono when the channel is mono — previously the `channels == 2` branch
encoded interleaved L/R directly even on a mono channel, feeding a mono `opus_encode` 2× its
samples (wrong pitch / garbage). Real stereo still only reaches the wire on a **stereo
channel** (encoder channel count = channel's Opus mode); on a mono channel the mic is cleanly
downmixed. Test: `test_stereo_mic_mono_channel` in `tests/test_vad_ptt_devices.cpp`. Full
`ctest --preset dev` green — 28/28. macOS Xcode build not compiled here (Windows host); the
Swift changes follow existing `nrChanged`/`setInputDevice` patterns. Docs: voice.md §8.
- **Done (2026-06-23):** **Fixed iOS dual-stream / crackly mic — core opened a second
(miniaudio) capture device alongside the AVAudioEngine tap.** Symptom: with two clients in
a channel, the remote end heard the iOS mic **twice** and crackly. With Voice Chat + a BT
headset, both the BT mic and the internal mic were captured; with Stereo Mic, both a mono
and a stereo copy of the internal mic were sent simultaneously. Root cause is a timing gap
in `vc_client::ensure_audio_running()` (`core/src/core/client.cpp`): `external_capture` was
only set when a MIC stream already existed, but `ensure_audio_running` is also called from
`sync_remote_streams` (triggered by the post-auth `ServerStateSnapshot`) **before** the user
joins voice — so with no MIC stream, `external_capture` stayed `false` and
`AudioEngine::start()` opened a real miniaudio capture device. Later the user joined voice →
`IOSAudioEngine.startMic` installed the AVAudioEngine input tap → `feedPcm`
`inject_capture``on_capture_frame`. The miniaudio device was still open (the engine was
already `running()`, so the later `ensure_audio_running` early-returned and never applied
`external_feed`), and `on_capture_frame` encodes+sends every frame with **no deduplication**
→ the mic was sent twice. The two unsynchronized capture clocks interleaving in the encoder
is the crackle; the mono miniaudio device + stereo AVAudioEngine tap is the "mono and stereo
at the same time" on Stereo Mic.
- **Fix 1 (core, `core/src/core/client.cpp:ensure_audio_running`):** force
`p.external_capture = true` whenever `external_playback_` is set. In iOS unified mode the
core must never open a hardware capture device — the AVAudioEngine owns the only mic path.
No-op on desktop (`external_playback_` is never set there).
- **Fix 2 (iOS, `clients/apple/iOS/VoiceCatiOS/AppState.swift`):** move
`client.setExternalPlayback(true)` from the `authResult` handler to **before**
`client.connect(...)`. The server sends `AuthResult` immediately followed by
`ServerStateSnapshot`; `handle_server_state` runs `ensure_audio_running` on the io thread
before the main thread drains `authResult`, so setting the flag post-auth raced. Setting it
pre-connect guarantees `external_playback_` is true before any message is processed —
eliminating the playback-device race too (the mixer timer + AVAudioEngine playback path +
VPIO AEC reference are correct from the first frame).
- **Verify:** `cmake --build --preset dev` clean; `ctest --preset dev` = 24/28 — the 4
failures (`vad_ptt_devices`, `external_pcm`, `frame_ms_reframe`, `channel_samplerate`) are
a **pre-existing** teardown `mutex lock failed` race, reproduced identically with the
changes stashed. `external_playback` (the one test exercising this code path) **passes**.
No xcframework rebuild needed (no new symbols). **Next (manual, on device):** two clients
in a channel — Voice Chat + BT, and Stereo Mic — confirm the remote end hears the iOS mic
once, clean (no duplicate, no crackle); confirm the iOS user hears the remote user cleanly
with AEC working in Voice Chat.
- **Done (2026-06-23):** **Fixed iOS mic flutter / crackle / octave-up.** The iOS mic was
unusable: a consistent ~4060 ms flutter with volume fade ("talking through a slow fan") on
every preset. Root cause: the core sends each captured frame **synchronously**
(`on_capture_frame``encode_and_send_frame`, no send pacer), so packet cadence == capture
cadence; and the receiver's playout keeps **near-zero buffering** by design and its jitter
estimate is blind to arrival timing (see RFC-3550 item above). That's smooth only for a
*steady* sender (desktop miniaudio = steady 20 ms), but the iOS `AVAudioEngine` input tap
delivers ~2 frames per ~40 ms callback (more under VPIO) → bursty → receiver underruns → PLC
fade.
- **Fix (iOS-only, `clients/apple/iOS/VoiceCatiOS/IOSVoiceProcessingEngine.swift`):** the
mic tap converts to 48 kHz int16 and writes a lock-free SPSC ring; a 20 ms feed pump
drains it and calls `feedPcm` at a **steady** cadence so packets leave the core every
20 ms (what the receiver expects). The pump **primes a small prebuffer cushion**
(`PumpState.targetFrames`, 3 frames ≈ 60 ms, self-healing up to ~120 ms on underrun)
before releasing, so the tap's bursts can't drain it to empty. Two correctness rules
(each had bit us): never read a partial frame (`read` consumes what it returns →
discarding partials caused crackle), and rebuild the pump with the current channel count
every `rebuild()` (a frozen channel count fed mono-as-stereo = octave-up on a Stereo→Voice
Chat switch). Trade-off: ~60120 ms added mic-send latency — unavoidable when de-bursting
for a near-zero-buffer receiver; the RFC-3550 fix above would let us shrink it.
- **Verify:** `xcodebuild` Debug **BUILD SUCCEEDED** (iOS Simulator, arm64). Audible test
requires a real device (simulator has no real mic route): mic should be smooth on Voice
Chat / Mono Mic / Stereo Mic, including switching presets while live (no octave).
- **Done (2026-06-23):** **Fixed Apple client link failure (stale xcframework missing
RNNoise).** Both `VoiceCatMac` and `VoiceCatiOS` failed to link with `Undefined symbols for
architecture arm64: _rnnoise_create / _rnnoise_destroy / _rnnoise_process_frame`. Root
cause: `clients/apple/scripts/build-xcframework.sh` merged vcpkg deps into the fat static
lib but NOT the locally-built vendored `librnnoise.a` (a CMake target from
`third_party/rnnoise/`, linked privately into `voicecat` via `VOICECAT_HAS_NS` — not a
vcpkg dep). The xcframework had been rebuilt at 14:17 after the RNNoise commit but still
omitted the symbols, so every slice's `libvoicecat-fat.a` referenced `_rnnoise_*` with no
defining object. The iOS slices were also stale (pre-rnnoise) and absent from the
xcframework entirely.
- **Fix:** `build-xcframework.sh` now collects `.a` files from `build/<preset>/lib/`
(excluding `libvoicecat*`) in addition to `vcpkg_installed/<triplet>/lib/`, so vendored
CMake-target static libs like `librnnoise.a` are merged into the fat lib. Future-proof:
any new vendored static-lib target landing in `build/<preset>/lib/` is picked up
automatically. README "Fat static library" section updated.
- **Verify:** rebuilt `VoiceCatCore.xcframework --all` → all 3 slices (macos-arm64,
ios-arm64, ios-arm64-simulator) now carry 10 `_rnnoise_*` symbols each; fat lib
~30 MB → ~33 MB. `xcodebuild` Debug **BUILD SUCCEEDED** for `VoiceCatMac`,
`VoiceCatiOS` (iphonesimulator arm64), and `VoiceCatiOS` (iphoneos arm64,
`CODE_SIGNING_ALLOWED=NO`). No core/ABI/proto changes — xcframework artifact only.
- **Done (2026-06-23):** **Remote-stream noise suppression — real backend (RNNoise).** The
two-sided NR plumbing (`RemoteStream::recv_ns` + `vc_set_remote_stream(... noise_reduction)`)
was wired but **inert**`ApmProcessor::create()` returned a no-op passthrough, because the
originally-planned `webrtc-audio-processing` won't build on Windows/macOS. Replaced with
**RNNoise** (BSD-3 + CC0), vendored at `third_party/rnnoise/` (the vcpkg port is `!windows
!arm`), built as a standalone C static lib + `VOICECAT_HAS_NS`. One `RnnoiseProcessor`
(`core/src/audio/apm_processor.cpp`) now backs **both** NR paths:
- **Receive-side** (per-listener, per-`ssrc`): lit up automatically via the factory; gated to
mono streams (`audio_engine.cpp` ~L791).
- **Send-side** (mic, new): `vc_set_input_noise_reduction(client, enable)` ABI +
`vc_client::mic_ns_`, run before input gain/VAD in `on_capture_frame`. A stereo mic is
downmixed to mono **only when NR is on**; with NR off a stereo mic keeps full stereo.
- RNNoise is mono/48 kHz/480-sample; our clock is fixed 48 kHz and Opus frame sizes are all
multiples of 480, so no resampling. RT-safe: alloc at construction, lock-free in the callback.
- **Verify status:** `ctest --preset dev` green — **28/28** (new `noise_suppression` test:
feeds white noise through `ApmProcessor::create()`, measures **99.9%** RMS reduction). Build
clean on the `dev` MinGW preset. **Next (manual):** add the on/off toggles to the client UIs
(Windows Audio Settings dialog, macOS/iOS settings) calling the two ABIs; build `windows-client`
+ `apple-dev` presets to confirm RNNoise compiles under MinGW-DLL and arm64; two-client E2E.
- **Done (2026-06-23):** **Aux outgoing stream (mic + a second input device) — Windows + macOS.**
Users can now transmit a second hardware input device (e.g. line-in / aux) alongside the mic, with
its own device picker and volume, from Audio Settings. **No core/ABI/proto changes** — the aux is
a `VC_STREAM_AUX_DEVICE` stream started with `external_feed=1`, captured client-side, and fed via
`vc_stream_feed_pcm` (the same external-feed pipeline screen-audio uses). Per-kind `local_streams_`
already allows mic + screen + one aux to coexist; volume is a client-side gain multiply (the core's
`vc_set_input_gain` is mic-only/global). The aux is always-on (the core never gates `AUX_DEVICE` on
VAD/PTT) and is tied to the voice session (started on Join Voice when enabled, stopped on Leave).
- **Windows:** new `Audio/InputDeviceCapture.cs` (WASAPI shared-mode capture from a real input
endpoint via `IMMDevice.Activate(IAudioClient)`, 48 kHz/s16, 20 ms frames) + `InputDeviceEnumerator`
(WASAPI capture-endpoint list — separate from the core's miniaudio ids). Aux section in
`AudioSettingsForm.cs` (enable checkbox, device combo, refresh, volume slider, accessible names,
live-apply + Cancel revert via callbacks). Lifecycle in `MainForm.cs` (`_auxStreamId` +
`InputDeviceCapture`). Persisted in `VoiceSettings.cs` (`AuxEnabled/AuxDeviceId/AuxGain`).
- **macOS:** new `Audio/InputDeviceCapture.swift` (AVAudioEngine input-node tap pinned to the chosen
Core Audio device via `kAudioOutputUnitProperty_CurrentDevice`; AVAudioConverter → 48 kHz int16;
20 ms framing modelled on `ScreenAudioCapture`) + `InputDeviceEnumerator` (Core Audio device list
by stable UID). Aux section in `SettingsWindowController.swift`; lifecycle + UserDefaults
persistence (`voice.aux*`) in `MainWindowController.swift`. New file added to `project.pbxproj`.
- **Verify status:** Windows C# solution builds clean (0 warn/0 err); `ctest` core suite unchanged
(no core edits). **Next (manual):** on a Mac, build `VoiceCatMac.xcodeproj`; then two-client E2E —
enable aux on a second input device, confirm two distinct streams for the sender and that the aux
volume slider moves the aux level independently of the mic; confirm persistence across relaunch.
- **Done (2026-06-23):** **Input-settings persistence, mic input gain, + two iOS bugs (all 3
clients).** Four fixes:
1. **Input settings now persist.** Transmission mode (VAD/PTT/Always-On), VAD threshold, and the
new mic gain were applied to the core + UI but never saved, so every relaunch reset to VAD
defaults. Each client now persists them and re-applies on connect: iOS via `UserDefaults`
(`SessionState.loadAndApplyVoiceSettings` + setter writes, keys `voice.*`); macOS via
`UserDefaults` (`MainWindowController` `didSet` + `loadPersistedAudioSettings`, also restores
the VAD slider from the stored threshold); Windows via new
`VoiceCat.App/Models/VoiceSettings.cs` (JSON at `%AppData%\VoiceCat\voice.json`, mirrors
`FeedbackSettings`) loaded/applied in `MainForm`.
2. **Microphone input gain.** New global send-side API `vc_set_input_gain` (voicecat.h →
`client.cpp::on_capture_frame`, applied to MIC PCM before the VAD gate, clamped to int16) plus
Swift (`setInputGain`) and C# (`SetInputGain`) bindings. Mic-volume slider (0300 %, default
100 %) added to all three clients' input settings, persisted with the rest.
3. **iOS chat send fixed.** `ChatView` called `sendText(scope:.channel)` with no `targetId` (→ 0),
so channel messages went nowhere; now passes `session.currentChannelId`.
4. **iOS per-user tuning reachable via VoiceOver.** The tuning sheet was long-press
`.contextMenu` only (invisible to VoiceOver); `UserRow` now also exposes the same buttons as
`.accessibilityActions` (no visual change), so the actions rotor reaches tuning + admin actions.
- **Verified:** core `cmake --build --preset dev` clean; `ctest --preset dev` = 24/27 (the 3
failures — `external_pcm`, `frame_ms_reframe`, `channel_samplerate` — are a pre-existing
teardown crash on this machine, reproduced identically with the changes stashed). xcframework
rebuilt (`--all`); **VoiceCatMac** and **VoiceCatiOS** (arm64 sim) → BUILD SUCCEEDED;
`VoiceCat.Interop` (`dotnet build`) succeeded. **Windows App not built** (WinForms
net10.0-windows can't build on macOS) — changes follow existing patterns; needs a Windows
build + manual check.
- **Next (manual):** on each client, set PTT + non-default VAD/mic-gain, relaunch → settings
restored; boost a quiet mic and confirm others hear it louder; iOS send a channel message;
iOS VoiceOver → focus a user → actions rotor opens tuning.
- **Done (2026-06-22):** **Fixed growing voice latency (jitter-buffer depth ratchet).** Symptom:
end-to-end latency grew to multiple seconds and "drifted backward," reset only by leaving/
rejoining voice (DTX/FEC/DRED on, 10% loss). Root cause was **not** the codec settings (10% loss
is just an `OPUS_SET_PACKET_LOSS_PERC` encoder hint; FEC/DRED add no standing latency) but the
receiver playout logic in `core/src/audio/audio_engine.cpp`: the playout clock free-ran in real
time while the sender omitted silence from its timestamps and set **no header flags at all**, and
the only correction snapped the clock to the *oldest* buffered frame (could only *add* latency) —
with `target_depth_ms_` computed but never enforced, so latency could only grow or be reset.
**Fix:** bounded-depth playout — (re)seed to the *leading edge* (newest frame) on start/marker/
starve, and **frame-skip catch-up** that trims a backlog beyond `target + hysteresis` (the missing
downward force). Plus hardening: adaptive late-drop window, talkspurt `kFlagMarker`/`kFlagDtx`
now actually stamped by the sender (`client.cpp` send path) and consumed on recv, EWMA outlier
rejection (silence gaps/stragglers no longer poison the estimate), duplicate counting, ring-
underrun diagnostics (`stream_underruns`/`stream_duplicates`). New regression test
`tests/test_jitter_depth.cpp` asserts depth stays bounded (<200 ms) while arrivals outrun playout
for ~4 s. `ctest --preset dev` green **27/27**. Docs: `docs/voice.md` §5 rewritten.
- **Next (manual E2E):** two clients in a channel, DTX/FEC/DRED on talk in alternating bursts
for several minutes and confirm latency stays low/stable (no backward drift, no rejoin needed).
- **Windows done / Apple awaiting Mac build (2026-06-22):** **Event sound effects + optional
text-to-speech for all clients.** Clients now play a cue per session event and can optionally
speak it (TTS off by default; when on it announces joins/leaves and reads message/PM bodies).
One canonical eventsound mapping (defined off the shared C ABI `vc_event` stream) is mirrored
across all three clients; `self` vs others is `user_id == self_user_id`, and outgoing messages
echo back as events so sent/recv cues need no separate send-path hook. Conservative defaults
(join/leave, channel/PM sent+recv, login, logout, connection-lost, mic on/off ON; per-utterance
self voice-activity `va_start/va_stop` and the PTT cue OFF). WAVs ship from `assets/sounds/`.
- **Windows (built + verified):** new `VoiceCat.App/Notifications/` (`FeedbackSettings`
`%AppData%\VoiceCat\feedback.json`, `SoundPlayerPool` via `System.Media.SoundPlayer`,
`SpeechAnnouncer` via the **Prismatoid** NuGet 0.3.0, `EventFeedback` dispatcher); hooks in
`Forms/MainForm.cs`; `Forms/NotificationSettingsForm.cs` under a new **Settings ▸ Notifications**
menu. `.csproj` adds the Prismatoid PackageRef and copies the WAVs into `sounds\`. `dotnet build`
clean; WAVs + `Prismatoid.dll` confirmed in output. Note: `SoundPlayer` has no gain control, so
volume is honoured as a mute gate (0 = silent) swap to NAudio if finer/overlap control is needed.
- **macOS + iOS (written, NOT yet built needs a Mac):** shared `Sources/VoiceCatCore/Feedback/`
(`SoundEvent`, `EventFeedback` = `AVAudioPlayer` pool + native `AVSpeechSynthesizer`,
`FeedbackSettings` over `UserDefaults`); WAVs copied into `Sources/VoiceCatCore/Sounds/` and
bundled via `Package.swift` `resources: [.process("Sounds")]` (`Bundle.module`). Hooks: iOS
`SessionState.handleEvent` (+ split `userJoined`/`userLeft`, added a `.disconnected` cue case),
`AppState` auth-success login cue, PTT cue in `setPushToTalk`; macOS `MainWindowController`
handlers + NSEvent PTT monitor. Settings UI: iOS `SettingsView` Notifications section
(`@AppStorage`), macOS `SettingsWindowController` checkboxes + volume slider. No `.pbxproj`
edits needed (shared files are SPM-managed; app files already in the projects).
- **Next:** on a Mac, `clients/apple/scripts/build-xcframework.sh --all` then build
VoiceCatMac/VoiceCatiOS; fix compile fallout. **Watch the iOS audio session:** cues/TTS play over
the live VPIO `playAndRecord` session verify they mix and don't duck/interrupt the call or get
silenced by the mute switch (most likely bug site). Then run `ctest --preset dev` (unchanged
no core/server code touched).
- **Done (2026-06-22):** **UDP media now shares the TCP port (self-host port-forward fix).** Symptom: a
remote self-hosted server (`iamtalon.me:8384`, TCP+UDP 8384 forwarded) accepted TCP connections but
passed no voice. Root cause: `Config::media_port` defaulted to `0` = OS-assigned, and `main.cpp`'s
`--port` only set `bind_port` (TCP) so the UDP relay bound a *random high port*, advertised it to
clients in HELLO (`udp_port`), and clients sent voice there. With only `8384/udp` forwarded those
packets were dropped connect OK, no audio. This contradicted `docs/deployment.md` ("Control and media
share one port number on TCP+UDP"). **Fix (`server/src/server.cpp`):** media follows bind_port when
`media_port == 0` `media_want = cfg_.media_port != 0 ? cfg_.media_port : cfg_.bind_port`. The
`0 = OS-assigned` escape hatch survives when `bind_port` is also 0, so tests that bind ephemeral ports
are unaffected (kept the logic in server.cpp rather than hardcoding 8384 as the default, which would
collide parallel tests on UDP 8384). Banner now reads `TCP :8384 UDP :8384`. Build + `ctest --preset
dev` green (24/24); live-verified banner with `--port 8390` `UDP :8390`. **Action for self-hosters:**
redeploy and confirm the startup banner shows matching TCP/UDP ports; the existing single forward rule
is now correct. If voice still fails, watch the server's rate-limited `[media] dropped frames —
unmapped-endpoint=…` line (NAT source-port rewrite would be the next suspect).
- **Done (2026-06-23, Swift-only no core/ABI change; awaiting on-device verification):** **iOS audio
stack unified one always-external `AVAudioEngine`, miniaudio dropped on iOS.** The iOS audio path was
a fragile hybrid: Voice-Chat-class presets ran a native VPIO `AVAudioEngine` (core external) while
Stereo/Studio/A2DP presets ran the core's miniaudio devices. Nearly every bug lived in the seam
(lingering miniaudio capture unit fighting VPIO, the `audioRestart` ordering dance, the route-change
"glitching" loop, stereomono stickiness, "can't hear anyone"), and switching presets/routes mid-call
routinely dropped input, output, or both. **Fix: drive *all* iOS audio through one `AVAudioEngine` with
the core fully external at all times** `vc_set_external_playback(1)` once at connect, every MIC stream
`external_feed=1`, mic via `vc_stream_feed_pcm`, playback via `vc_set_mixed_output_sink`.
- `IOSVoiceProcessingEngine.swift` **`IOSAudioEngine`** (same file): always-on `AVAudioSourceNode`
playback (runs whenever connected, so remote audio plays before you join voice); conditional mic tap;
VPIO + AGC toggled per config. One private `rebuild()` (stop set VPIO install tap start) backs
`startListening`/`stop`/`startMic`/`stopMic`/`reconfigure`/`setCaptureChannels`. Kept the `PCMRing`
and ring-stats diagnostics.
- `IOSAudioRouter`: presets cut from seven to **four** Voice Chat (VPIO mono, system output),
Stereo Mic / Mono Mic (internal built-in mic regardless of output, A2DP-capable, no VPIO), Advanced
(manual). New persisted `voiceProcessingEnabled` (master AEC+NS) + `agcEnabled`; setters now call
`IOSAudioEngine.reconfigure()` instead of `client.audioRestart()` + `reconcileVoicePath`. Kept the
proven AVAudioSession recipes (category/mode/options, stereo capsule, `applyA2dpSpeakerFallback`).
- `AudioSessionManager` slimmed (drops `client`/`activeMicStreamId`/`reconcileVoicePath`; adds
`isActive`); interruption-end & device-change now `reconfigure()` the engine. `SessionState`
`doStartMicStream`/`stopMicStream` collapsed to start-stream + `startMic`/`stopMic` (no
`setExternalPlayback`/`audioRestart` toggling); `reconcileVoicePath` deleted. `AppState` sets external
playback + `startListening` at connect, `stop()` at disconnect. `SettingsView` four presets +
Advanced VPIO/AGC toggles.
- **No core/ABI/test change** relies on the already-shipped `vc_set_external_playback` /
`external_feed` / `vc_set_mixed_output_sink` / `vc_stream_feed_pcm` path (`test_external_pcm`,
`test_external_playback`). `xcodebuild` iOS device Debug **BUILD SUCCEEDED**. **Rebuild the
xcframework is NOT required** (no new symbols).
- **Next (user, on device):** two iPhones in a channel verify BOTH directions survive every
transition and are never silent unless intended: Voice Chat (no echo, NR), listen-only before joining,
joinleave repeatedly, switch Voice ChatStereoMonoAdvanced *while in voice*, A2DP connect/unplug,
wired connect/unplug, phone-call interruption + resume, screen-audio share.
- **Superseded by the 2026-06-23 unification above (2026-06-22):** **iOS real echo cancellation / noise
suppression via native VPIO.** Root cause of "voice chat doesn't sound like a call" (echo + no NR): real iOS
AEC/NS/AGC come only from Apple's Voice-Processing I/O unit (VPIO), but the core uses miniaudio's AEC/NS/AGC come only from Apple's Voice-Processing I/O unit (VPIO), but the core uses miniaudio's
plain RemoteIO units so `.voiceChat` mode alone never engaged AEC. Fix moves both mic capture and plain RemoteIO units so `.voiceChat` mode alone never engaged AEC. Fix moves both mic capture and
playback to a native Swift `AVAudioEngine` (`setVoiceProcessingEnabled`) on the AEC presets, with the playback to a native Swift `AVAudioEngine` (`setVoiceProcessingEnabled`) on the AEC presets, with the
@@ -148,6 +630,21 @@ up instantly. Newest status at the top.
SUCCEEDED (sim slice still arm64-only → simulator run N/A). Next: on-device check of the SUCCEEDED (sim slice still arm64-only → simulator run N/A). Next: on-device check of the
drill-down + compose box + unified timeline. drill-down + compose box + unified timeline.
- **Done (2026-06-22):** **Windows exclude mode is now a real native exclude + self-echo
removal.** The "All apps except selected" mode previously captured the *complement of a frozen
app snapshot* in INCLUDE mode (missed late-launched apps, system sounds; wasted captures on
silent windows). It now opens a **single `ProcessLoopbackCapture` in EXCLUDE mode**
(`AUDIOCLIENT_PROCESS_LOOPBACK_MODE_EXCLUDE_TARGET_PROCESS_TREE`) of the one chosen app — true
system-mix-minus-one, dynamic. `AppAudioPickerDialog` enforces single-selection in exclude
mode (the API takes one target PID). Added an **"Exclude VoiceCat's own audio (prevents echo)"**
checkbox (default on, entire-desktop only) that routes the desktop capture through the same
EXCLUDE path targeting `Environment.ProcessId`, killing the whole-device self-echo loop.
Touched `ProcessAudioMixer.cs` (`ResolveCaptures`), `AppAudioPickerDialog.cs`, `MainForm.cs`,
`AudioSessionEnumerator.cs` (`EntireDesktop(bool ExcludeSelf)`); docs in voice.md §9. No C++ /
ABI changes. `dotnet build` clean. **Still to verify on-device:** exclude actually silences
the chosen app while the rest plays, late-launched apps appear without restart, and the
self-exclude checkbox removes the echo.
- **Done (2026-06-21):** **Screen-audio sharing on macOS + iOS.** macOS uses ScreenCaptureKit - **Done (2026-06-21):** **Screen-audio sharing on macOS + iOS.** macOS uses ScreenCaptureKit
(`ScreenAudioCapture.swift`) → `vc_stream_feed_pcm`; iOS uses a ReplayKit Broadcast Upload (`ScreenAudioCapture.swift`) → `vc_stream_feed_pcm`; iOS uses a ReplayKit Broadcast Upload
Extension (`VoiceCatBroadcast`) that forwards captured `.audioApp` PCM through a shared App Extension (`VoiceCatBroadcast`) that forwards captured `.audioApp` PCM through a shared App
@@ -381,7 +878,33 @@ up instantly. Newest status at the top.
## Recent completed work ## Recent completed work
All items below are `[x]` done; `ctest --preset dev` 23/23 on Windows after all. All items below are `[x]` done; `ctest --preset dev` 26/26 on Windows after all.
- **Per-channel sample_rate as a bandwidth cap** (2026-06-22): the channel `sample_rate` field
was inert (the codec is pinned to 48 kHz). Made it meaningful without changing the 48 kHz
clock: it's carried as `OpusParams::max_bandwidth_hz` and applied via `OPUS_SET_MAX_BANDWIDTH`
in `OpusEncoder::init` (8000→narrowband … 48000→full). Made it **channel-authoritative** on
the server (`conn_session.cpp` no longer overrides effective `sample_rate` with the client's
always-48000 request — like `frame_ms`/`mode`). `vc_get_stream_audio_config` now reports the
channel's configured rate for own streams too. New ctest `channel_samplerate`: a 7 kHz tone is
attenuated ~1000× on an 8 kHz channel vs a 48 kHz channel. Files: `opus_codec.{h,cpp}`,
`client.cpp`, `server/src/conn_session.cpp`, `docs/voice.md`, `tests/test_channel_samplerate.cpp`,
`tests/CMakeLists.txt`. (Future: a true non-48k stack is possible but unnecessary — 48 kHz is
what nearly all hard/software runs at; the bandwidth cap covers the narrowband use case.)
- **Non-20ms channel frame_ms fix** (2026-06-22): the AudioEngine capture clock is fixed at
48 kHz / 20 ms (960-sample frames), but a channel may set any Opus `frame_ms` (2.5…60 ms,
docs/voice.md §3) and the server enforces it unclamped. The send path handed the engine's
960-sample frame straight to an encoder configured for the channel's window — silently
ignoring `frame_ms > 20` and **breaking `frame_ms < 20` entirely** (receiver sized its decode
buffer too small → `OPUS_BUFFER_TOO_SMALL` → dead audio). Affected the hardware mic AND
`vc_stream_feed_pcm`. Fix: `vc_client::on_capture_frame` now reframes each captured/fed block
to `ls.frame_samples` via a per-`LocalStream` accumulator (pre-sized at announce, no RT-thread
alloc) before `encode_and_send_frame`; the 20 ms case stays a zero-copy fast path. Also pinned
the codec to 48 kHz internally in `opus_params_from_audio_config` (was honoring a non-48k
effective sample_rate against a 48k PCM clock). New ctest `frame_ms_reframe` (40 ms accumulate
+ 10 ms split round trips). Files: `client.{h,cpp}`, `voicecat.h` (feed doc), `docs/voice.md`,
`tests/test_frame_ms_reframe.cpp`, `tests/CMakeLists.txt`.
- **External PCM feed/tap API** (2026-06-20): `vc_stream_feed_pcm` + `vc_set_pcm_sink` shipped. - **External PCM feed/tap API** (2026-06-20): `vc_stream_feed_pcm` + `vc_set_pcm_sink` shipped.
Promotes `vc_test_inject_capture` (mono-only, TEST-ONLY) to a public, stereo-capable API. Promotes `vc_test_inject_capture` (mono-only, TEST-ONLY) to a public, stereo-capable API.
@@ -547,7 +1070,8 @@ text, device pickers, level meter on each platform.
**Windows** (`clients/windows/`): `VoiceCat.Interop` (P/Invoke, `[UnmanagedCallersOnly]`), **Windows** (`clients/windows/`): `VoiceCat.Interop` (P/Invoke, `[UnmanagedCallersOnly]`),
`VoiceCat.App` (ConnectDialog, ServerIdentityDialog, MainForm with full M5 moderation UI, `VoiceCat.App` (ConnectDialog, ServerIdentityDialog, MainForm with full M5 moderation UI,
PerUserTuningDialog, PttKeyCaptureDialog), `VoiceCat.Interop.Tests`. PTT is focus-scoped. PerUserTuningDialog, PttKeyCaptureDialog), `VoiceCat.Interop.Tests`. PTT can be system-wide
(Raw Input / WM_INPUT) or focus-scoped, toggled in Audio settings (default system-wide).
**macOS** (`clients/apple/macOS/VoiceCatMac.xcodeproj`): NSOutlineView channel tree, **macOS** (`clients/apple/macOS/VoiceCatMac.xcodeproj`): NSOutlineView channel tree,
NSTableView user list, NSTextView chat, voice controls, full VoiceOver accessibility, admin NSTableView user list, NSTextView chat, voice controls, full VoiceOver accessibility, admin
@@ -585,6 +1109,14 @@ iOS 18.0 deployment target. App Group `group.cat.voice.VoiceCat` for Keychain sh
reconstructs the lost frame — otherwise falls back to standard PLC. New test: reconstructs the lost frame — otherwise falls back to standard PLC. New test:
`test_dred_toggle` (ctest 22/22). Files: `voicecat.proto`, `voicecat.h`, `test_dred_toggle` (ctest 22/22). Files: `voicecat.proto`, `voicecat.h`,
`opus_codec.{h,cpp}`, `audio_engine.{h,cpp}`, `client.cpp`, `session.{h,cpp}`. `opus_codec.{h,cpp}`, `audio_engine.{h,cpp}`, `client.cpp`, `session.{h,cpp}`.
- [x] **In-band FEC decoder wiring** — done (2026-06-22). The encoder set `OPUS_SET_INBAND_FEC`
all along, but the decoder never invoked it — the loss path went DRED → PLC, so FEC redundancy
was emitted (and paid for in bitrate) but never consumed. Wired the FEC recovery into
`AudioEngine::on_playback`'s loss branch between DRED and PLC: copy the next buffered packet
once, try DRED, else (if the stream negotiated FEC) `decode(next_pkt, …, fec=true)`, else PLC.
Added per-stream `RemoteStream::fec_enabled_`, captured from `OpusParams` in
`init_recv_stream`. Recovery priority is now **DRED → FEC → PLC**. ctest 27/27 green. Files:
`audio_engine.{h,cpp}`, `docs/voice.md`.
- [ ] **DRED toggle in client UIs** — expose the `dred` flag in all three channel-config UIs - [ ] **DRED toggle in client UIs** — expose the `dred` flag in all three channel-config UIs
so admins can enable it per channel. Windows: `ChannelEditForm` / `vc_channel_info.audio.dred` so admins can enable it per channel. Windows: `ChannelEditForm` / `vc_channel_info.audio.dred`
checkbox. macOS AppKit: channel-edit sheet. iOS SwiftUI: channel-edit form. All three UIs checkbox. macOS AppKit: channel-edit sheet. iOS SwiftUI: channel-edit form. All three UIs

View File

@@ -23,9 +23,9 @@ The default development preset is **`dev`** — it builds everything (server + t
with real vcpkg deps. It works on Windows, Linux, and macOS (vcpkg triplet auto-resolved). with real vcpkg deps. It works on Windows, Linux, and macOS (vcpkg triplet auto-resolved).
```bash ```bash
# one-time vcpkg setup: # one-time vcpkg setup (bundled as a submodule, pinned to vcpkg.json's builtin-baseline):
git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh # .bat on Windows git submodule update --init vcpkg
export VCPKG_ROOT=/path/to/vcpkg # Linux/macOS; or $env:VCPKG_ROOT on PowerShell ./vcpkg/bootstrap-vcpkg.sh # .bat on Windows
# configure + build + test: # configure + build + test:
cmake --preset dev cmake --preset dev
@@ -33,6 +33,9 @@ cmake --build --preset dev
ctest --preset dev # 21 behavior tests ctest --preset dev # 21 behavior tests
``` ```
To use an external vcpkg checkout instead, set `VCPKG_ROOT=/path/to/vcpkg` (or
`$env:VCPKG_ROOT` on PowerShell) — it always takes priority over the bundled submodule.
Artifacts land in `build/dev/bin/` (`voicecat-server`, `vccli`, `voicecat-admin`). Artifacts land in `build/dev/bin/` (`voicecat-server`, `vccli`, `voicecat-admin`).
The `skeleton` preset (no vcpkg deps, stubs only) is a fast smoke check that needs no The `skeleton` preset (no vcpkg deps, stubs only) is a fast smoke check that needs no

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

BIN
assets/sounds/login.wav Normal file

Binary file not shown.

BIN
assets/sounds/logout.wav Normal file

Binary file not shown.

BIN
assets/sounds/pm_recv.wav Normal file

Binary file not shown.

BIN
assets/sounds/pm_sent.wav Normal file

Binary file not shown.

BIN
assets/sounds/ptt.wav Normal file

Binary file not shown.

BIN
assets/sounds/va_start.wav Normal file

Binary file not shown.

BIN
assets/sounds/va_stop.wav Normal file

Binary file not shown.

BIN
assets/sounds/voice_off.wav Normal file

Binary file not shown.

BIN
assets/sounds/voice_on.wav Normal file

Binary file not shown.

View File

@@ -40,7 +40,11 @@ let package = Package(
.target( .target(
name: "VoiceCatCore", name: "VoiceCatCore",
dependencies: ["VoiceCatCoreXCF"], dependencies: ["VoiceCatCoreXCF"],
path: "Sources/VoiceCatCore" path: "Sources/VoiceCatCore",
// Event-cue WAVs (shared with the Windows client) bundled into the package's
// resource bundle; EventFeedback loads them via Bundle.module. Copied from
// assets/sounds/ into Sources/VoiceCatCore/Sounds/.
resources: [.process("Sounds")]
), ),
// Smoke tests against a real voicecat-server mirrors clients/windows/ // Smoke tests against a real voicecat-server mirrors clients/windows/
// VoiceCat.Interop.Tests/VoiceCatClientSmokeTests.cs. Requires the `dev` CMake preset // VoiceCat.Interop.Tests/VoiceCatClientSmokeTests.cs. Requires the `dev` CMake preset

View File

@@ -106,12 +106,16 @@ scripts/build-xcframework.sh --all
The `apple-dev` CMake preset produces a 1.9 MB `libvoicecat.a` containing only voicecat's The `apple-dev` CMake preset produces a 1.9 MB `libvoicecat.a` containing only voicecat's
own object files — vcpkg's static dependencies (protobuf, mbedtls, libsodium, opus, sqlite3, own object files — vcpkg's static dependencies (protobuf, mbedtls, libsodium, opus, sqlite3,
spdlog, asio, abseil, …) are 107 separate `.a` files under `vcpkg_installed/arm64-osx/lib/`. spdlog, asio, abseil, …) are 107 separate `.a` files under `vcpkg_installed/arm64-osx/lib/`,
A Swift Package binary target can only link ONE `.a` per XCFramework slice, so and the vendored RNNoise noise-suppression lib (`third_party/rnnoise/`, built as a CMake
`build-xcframework.sh` merges them all into a single self-contained `libvoicecat-fat.a` target → `build/<preset>/lib/librnnoise.a`) is another. A Swift Package binary target can
(~30 MB) using `libtool -static`. This is the Apple equivalent of how the Windows client only link ONE `.a` per XCFramework slice, so `build-xcframework.sh` merges them all — vcpkg
deps plus the locally-built vendored libs — into a single self-contained `libvoicecat-fat.a`
(~33 MB) using `libtool -static`. This is the Apple equivalent of how the Windows client
ships a single `voicecat.dll` with all deps statically linked (via MinGW's `-static` flags ships a single `voicecat.dll` with all deps statically linked (via MinGW's `-static` flags
in [`core/CMakeLists.txt`](../../core/CMakeLists.txt)). in [`core/CMakeLists.txt`](../../core/CMakeLists.txt)). If you add another vendored (non-vcpkg)
static-lib target to the core, it's picked up automatically as long as it lands in
`build/<preset>/lib/` and isn't named `libvoicecat*`.
### Swift Package ### Swift Package

View File

@@ -1,19 +1,5 @@
// Callbacks the C function pointers passed to `vc_callbacks`. These are the Swift // C callbacks use an unretained `user` context. Client destruction joins callback threads,
// equivalent of the C# client's `[UnmanagedCallersOnly]` static methods (NativeCallbacks.cs). // and transient event pointers are copied before the callback returns.
//
// The critical patterns (carried over from the proven C# implementation):
// 1. `@convention(c)` closures plain C function pointers, NOT GC/ARC-managed closures.
// A @convention(c) closure cannot capture context, which is why the `user` pointer is
// used to resolve back to the VoiceCatClient instance (the C# version uses GCHandle for
// the same thing; Swift uses Unmanaged).
// 2. `Unmanaged.passUnretained(self).toOpaque()` as the `user` context a stable raw
// pointer to the Swift object WITHOUT incrementing the retain count. This is safe
// because `deinit` calls `vc_client_destroy` (which synchronously joins every internal
// thread) BEFORE the object's memory is freed so no callback can fire after the object
// is gone. (The C# equivalent: GCHandle.Alloc + GCHandle.Free in Dispose.)
// 3. Copy `ev.text` to a Swift `String` INSIDE `onEvent` (via `VoiceCatEvent.from(_:)`)
// before returning the raw pointer is dangling after the callback returns. This is
// the #1 lifetime rule from voicecat.h's vc_event doc comment.
import VoiceCatC import VoiceCatC
import Foundation import Foundation

View File

@@ -61,7 +61,7 @@ public enum VoiceCatConnectionState: UInt32, Sendable, Equatable {
case tlsHandshake = 2 case tlsHandshake = 2
case authenticating = 3 case authenticating = 3
case connected = 4 case connected = 4
/// M4: handshake succeeded, waiting on `confirmServerIdentity()`. /// Handshake succeeded, waiting on `confirmServerIdentity()`.
case verifyingIdentity = 5 case verifyingIdentity = 5
public init(_ cValue: vc_connection_state) { public init(_ cValue: vc_connection_state) {
@@ -125,14 +125,16 @@ public enum VoiceCatEventType: UInt32, Sendable, Equatable {
case talkState = 9 case talkState = 9
case error = 10 case error = 10
case disconnected = 11 case disconnected = 11
/// M4: reply to `joinChannel()` see `VoiceCatEvent.result` / `.channelId`. /// Reply to `joinChannel()` see `VoiceCatEvent.result` / `.channelId`.
case joinResult = 12 case joinResult = 12
/// M4: the TOFU server-identity gate see `VoiceCatEvent.tofuStatus` / `.text`. /// The TOFU server-identity gate see `VoiceCatEvent.tofuStatus` / `.text`.
case serverIdentity = 13 case serverIdentity = 13
/// M5: async result for moderation/admin/channel operations. /// Async result for moderation/admin/channel operations.
case genericResult = 14 case genericResult = 14
/// M5: reply to `requestAccountList()` call `listAccounts()` to read. /// Reply to `requestAccountList()` call `listAccounts()` to read.
case accountList = 15 case accountList = 15
/// Voice-plane subscription state. `u32a` = 1 (subscribed) or 0 (unsubscribed).
case voiceState = 16
public init(_ cValue: vc_event_type) { public init(_ cValue: vc_event_type) {
self = VoiceCatEventType(rawValue: cValue.rawValue) ?? .error self = VoiceCatEventType(rawValue: cValue.rawValue) ?? .error

View File

@@ -0,0 +1,105 @@
// EventFeedback shared audible + spoken feedback for session events, used by both the macOS
// (AppKit) and iOS (SwiftUI) clients. Mirrors the Windows client's EventFeedback policy
// (clients/windows/.../Notifications/EventFeedback.cs): the platform event handlers decide WHAT
// to play (they own the model/nickname/channel context); this type owns the "should I, and how"
// policy plus the AVFoundation playback/synthesis.
//
// Sound effects use AVAudioPlayer; spoken announcements use the OS-native AVSpeechSynthesizer.
//
// NOTE (iOS): on the voice path the app runs a play-and-record AVAudioSession (VPIO). Playing
// these cues / speech over that session can interact with the live call (ducking, route, or the
// mute switch). The session category should allow mixing verify on device. This is the most
// likely place for platform bugs to surface.
import Foundation
import AVFoundation
/// User preferences for event sounds and spoken feedback, backed by UserDefaults so the macOS
/// and iOS settings screens and this player share one source of truth.
public struct FeedbackSettings: Sendable {
public var sounds: Bool
public var speech: Bool
public var volume: Float
public var selfTalkSounds: Bool
public var pttSound: Bool
static let keySounds = "feedback.sounds"
static let keySpeech = "feedback.speech"
static let keyVolume = "feedback.volume"
static let keySelfTalk = "feedback.selfTalk"
static let keyPtt = "feedback.ptt"
/// Default values registered with UserDefaults (so "unset" reads as the intended default
/// rather than false/0).
static let defaults: [String: Any] = [
keySounds: true,
keySpeech: false,
keyVolume: 1.0,
keySelfTalk: false,
keyPtt: false,
]
/// The current settings, read live from UserDefaults.
public static var current: FeedbackSettings {
let d = UserDefaults.standard
d.register(defaults: defaults) // idempotent ensures "unset" reads as the intended default
return FeedbackSettings(
sounds: d.bool(forKey: keySounds),
speech: d.bool(forKey: keySpeech),
volume: d.float(forKey: keyVolume),
selfTalkSounds: d.bool(forKey: keySelfTalk),
pttSound: d.bool(forKey: keyPtt))
}
}
@MainActor
public final class EventFeedback {
public static let shared = EventFeedback()
private var players: [SoundEvent: AVAudioPlayer] = [:]
private let synthesizer = AVSpeechSynthesizer()
private init() {
UserDefaults.standard.register(defaults: FeedbackSettings.defaults)
}
// MARK: - Sounds
/// Play an event cue, honouring the user's settings. The two opt-in categories (your own
/// voice-activity, and the PTT cue) are gated by their own flags.
public func play(_ event: SoundEvent) {
let s = FeedbackSettings.current
guard s.sounds, s.volume > 0 else { return }
if (event == .vaStart || event == .vaStop), !s.selfTalkSounds { return }
if event == .ptt, !s.pttSound { return }
guard let player = player(for: event) else { return }
player.volume = s.volume
player.currentTime = 0
player.play()
}
/// Lazily load and cache an AVAudioPlayer for the event's bundled WAV. Returns nil (silent)
/// if the resource is missing or fails to load.
private func player(for event: SoundEvent) -> AVAudioPlayer? {
if let cached = players[event] { return cached }
guard let url = Bundle.module.url(forResource: event.resourceName, withExtension: "wav"),
let player = try? AVAudioPlayer(contentsOf: url) else {
return nil
}
player.prepareToPlay()
players[event] = player
return player
}
// MARK: - Speech
/// Speak `text` when spoken feedback is enabled. Utterances queue (do not interrupt prior
/// speech) so a burst of events is read in order.
public func speak(_ text: String) {
guard FeedbackSettings.current.speech else { return }
let trimmed = text.trimmingCharacters(in: .whitespacesAndNewlines)
guard !trimmed.isEmpty else { return }
synthesizer.speak(AVSpeechUtterance(string: trimmed))
}
}

View File

@@ -0,0 +1,43 @@
// SoundEvent the cross-platform set of audible event cues. Each maps to a WAV bundled as an
// SPM resource (Sources/VoiceCatCore/Sounds/, copied from assets/sounds/). The same logical set
// is mirrored in the Windows client (clients/windows/.../Notifications/SoundEvent.cs) so feedback
// stays consistent across platforms.
import Foundation
public enum SoundEvent: CaseIterable, Sendable {
case channelJoin // another user joined my channel
case channelLeave // another user left my channel
case channelRecv // channel text message from someone else
case channelSent // channel text message I sent
case pmRecv // private message received
case pmSent // private message I sent
case login // connected / authenticated
case logout // clean disconnect
case connectionLost // unexpected disconnect
case voiceOn // my microphone stream started
case voiceOff // my microphone stream stopped
case vaStart // my voice-activity began (off by default)
case vaStop // my voice-activity ended (off by default)
case ptt // push-to-talk engaged (off by default)
/// Resource name (without extension) as bundled in Sources/VoiceCatCore/Sounds/.
var resourceName: String {
switch self {
case .channelJoin: return "channel_join"
case .channelLeave: return "channel_leave"
case .channelRecv: return "channel_recv"
case .channelSent: return "channel_sent"
case .pmRecv: return "pm_recv"
case .pmSent: return "pm_sent"
case .login: return "login"
case .logout: return "logout"
case .connectionLost: return "connection_lost"
case .voiceOn: return "voice_on"
case .voiceOff: return "voice_off"
case .vaStart: return "va_start"
case .vaStop: return "va_stop"
case .ptt: return "ptt"
}
}
}

View File

@@ -38,7 +38,8 @@ internal enum Marshaling {
let c = items.advanced(by: i).pointee let c = items.advanced(by: i).pointee
result.append(Channel(id: c.id, parentId: c.parent_id, name: string(c.name), result.append(Channel(id: c.id, parentId: c.parent_id, name: string(c.name),
topic: string(c.topic), passwordProtected: c.password_protected != 0, topic: string(c.topic), passwordProtected: c.password_protected != 0,
maxUsers: c.max_users)) maxUsers: c.max_users, sortOrder: c.sort_order,
audio: audioConfig(c.audio)))
} }
vc_free_channel_list(&list) vc_free_channel_list(&list)
return result return result
@@ -53,7 +54,8 @@ internal enum Marshaling {
result.append(User(id: u.id, nickname: string(u.nickname), isGuest: u.is_guest != 0, result.append(User(id: u.id, nickname: string(u.nickname), isGuest: u.is_guest != 0,
channelId: u.channel_id, selfMicMuted: u.self_mic_muted != 0, channelId: u.channel_id, selfMicMuted: u.self_mic_muted != 0,
selfDeafened: u.self_deafened != 0, serverMuted: u.server_muted != 0, selfDeafened: u.self_deafened != 0, serverMuted: u.server_muted != 0,
serverDeafened: u.server_deafened != 0)) serverDeafened: u.server_deafened != 0,
voiceSubscribed: u.voice_subscribed != 0))
} }
vc_free_user_list(&list) vc_free_user_list(&list)
return result return result

View File

@@ -16,11 +16,17 @@ public struct Channel: Sendable, Equatable, Identifiable {
public let passwordProtected: Bool public let passwordProtected: Bool
/// 0 = unlimited. /// 0 = unlimited.
public let maxUsers: UInt32 public let maxUsers: UInt32
public let sortOrder: UInt32
/// Authoritative channel Opus params (docs/voice.md §3). Populated from the Channel proto
/// so the edit dialog can read back the current config.
public let audio: AudioConfig
public init(id: UInt32, parentId: UInt32, name: String, topic: String, public init(id: UInt32, parentId: UInt32, name: String, topic: String,
passwordProtected: Bool, maxUsers: UInt32) { passwordProtected: Bool, maxUsers: UInt32, sortOrder: UInt32,
audio: AudioConfig) {
self.id = id; self.parentId = parentId; self.name = name; self.topic = topic self.id = id; self.parentId = parentId; self.name = name; self.topic = topic
self.passwordProtected = passwordProtected; self.maxUsers = maxUsers self.passwordProtected = passwordProtected; self.maxUsers = maxUsers
self.sortOrder = sortOrder; self.audio = audio
} }
} }
@@ -57,17 +63,19 @@ public struct User: Sendable, Equatable, Identifiable {
public let selfDeafened: Bool public let selfDeafened: Bool
public let serverMuted: Bool public let serverMuted: Bool
public let serverDeafened: Bool public let serverDeafened: Bool
public let voiceSubscribed: Bool
public init(id: UInt32, nickname: String, isGuest: Bool, channelId: UInt32, public init(id: UInt32, nickname: String, isGuest: Bool, channelId: UInt32,
selfMicMuted: Bool, selfDeafened: Bool, serverMuted: Bool, selfMicMuted: Bool, selfDeafened: Bool, serverMuted: Bool,
serverDeafened: Bool) { serverDeafened: Bool, voiceSubscribed: Bool) {
self.id = id; self.nickname = nickname; self.isGuest = isGuest; self.channelId = channelId self.id = id; self.nickname = nickname; self.isGuest = isGuest; self.channelId = channelId
self.selfMicMuted = selfMicMuted; self.selfDeafened = selfDeafened self.selfMicMuted = selfMicMuted; self.selfDeafened = selfDeafened
self.serverMuted = serverMuted; self.serverDeafened = serverDeafened self.serverMuted = serverMuted; self.serverDeafened = serverDeafened
self.voiceSubscribed = voiceSubscribed
} }
} }
/// Permission bitset mirrors `vc_permissions` (M5). /// Permission bitset mirrors `vc_permissions`.
public struct Permissions: Sendable, Equatable { public struct Permissions: Sendable, Equatable {
public let canCreateTempChannel: Bool public let canCreateTempChannel: Bool
public let canKick: Bool public let canKick: Bool
@@ -83,7 +91,7 @@ public struct Permissions: Sendable, Equatable {
} }
} }
/// Account entry mirrors `vc_account` (M5, reply to `listAccounts()`). /// Account entry mirrors `vc_account` (reply to `listAccounts()`).
public struct Account: Sendable, Equatable { public struct Account: Sendable, Equatable {
public let username: String public let username: String
public let isAdmin: Bool public let isAdmin: Bool

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

View File

@@ -1,36 +1,6 @@
// VoiceCatClient the public, Swift-idiomatic surface over libvoicecat. This is the Swift // Swift binding invariants: native config strings outlive the handle, destroy joins callback
// analog of the C# client's `VoiceCatClient.cs` (clients/windows/VoiceCat.Interop). // threads before deallocation, and callback payloads are copied before main-queue delivery.
// // See docs/architecture.md §4 for the complete binding contract.
// Key patterns carried over from the proven C# implementation (see docs/architecture.md §4
// per-platform binding notes):
//
// 1. HANDLE OWNERSHIP: the class owns `vc_client*`; `deinit` calls `vc_client_destroy`
// (which synchronously joins every internal thread, so nothing can still be reading the
// config-string pointers or firing callbacks by the time it returns).
//
// 2. CONFIG STRING LIFETIMES: the core stores raw pointers from `vc_config` by value it
// does NOT copy the string data. `client_name`/`client_version`/`tofu_store_path` are
// read later, whenever `connect()` actually runs on the io_thread_. So the native CString
// storage (`_clientNamePtr` etc.) must outlive the WHOLE client, not just `init`. It's
// freed in `deinit`, AFTER `vc_client_destroy` has returned. (C#: Marshal.StringToCoTask
// MemUTF8 in ctor, FreeCoTaskMem in Dispose after destroy.)
//
// 3. EVENT DELIVERY THREAD HANDOFF: `on_event` fires on the core's event thread. Events are
// buffered in a lock-protected array and drained on `DispatchQueue.main` this is the
// boundary where the core's thread hands off to the UI thread. The C# analog is
// `Channel<VoiceCatEvent>` drained by a 30ms WinForms Timer; the Swift analog is a
// coalesced main-queue drain (only one async block scheduled at a time). `on_event`'s
// `text` is copied to a Swift `String` inside the callback (Callbacks.swift) before
// enqueueing the raw pointer is dangling by the time the main thread drains.
//
// 4. LEVEL METER COALESCING: `on_level` fires far more often than `on_event` and
// intermediate values are visually irrelevant coalesced to "latest sample per
// stream_id" in a lock-protected dictionary, drained on main alongside events.
// (C#: ConcurrentDictionary<uint,float> cleared in PumpEvents.)
//
// 5. IMMEDIATE vc_free_* ON LIST READS: `listChannels()`/`listUsers()`/etc. walk the native
// array, convert to Swift value types, and call `vc_free_*_list` INSIDE the function
// callers never manage native list lifetime. (C#: Marshaling.ToManaged does the same.)
import VoiceCatC import VoiceCatC
import Foundation import Foundation
@@ -56,26 +26,14 @@ public final class VoiceCatClient {
// MARK: - Stored properties // MARK: - Stored properties
/// The opaque C handle (`vc_client*` Swift imports the incomplete C struct as
/// `OpaquePointer`). Set in `init`, passed to every C function, destroyed in `deinit`.
private var handle: OpaquePointer? private var handle: OpaquePointer?
/// Unmanaged pointer to `self` passed as `vc_callbacks.user` so the C function-pointer /// Unretained callback context; destroying the handle joins callback threads first.
/// callbacks can resolve back to this instance. `passUnretained` (not `passRetained`)
/// because we want normal ARC to control the object's lifetime `deinit` calls
/// `vc_client_destroy` (joins all threads) before the object's memory is freed, so no
/// callback can fire with a dangling `user` pointer. See Callbacks.swift.
///
/// Computed (not stored) to break a circular init dependency: it needs `self`, but
/// stored properties must be initialized before `self` is available. `Unmanaged.passUn
/// retained(self).toOpaque()` always returns the same address for a given instance, so
/// computing it on demand is safe and consistent.
private var selfPointer: UnsafeMutableRawPointer { private var selfPointer: UnsafeMutableRawPointer {
Unmanaged.passUnretained(self).toOpaque() Unmanaged.passUnretained(self).toOpaque()
} }
/// Native CString storage backing `vc_config` must outlive the whole client (the core /// The core retains these pointers for the handle's lifetime.
/// stores raw pointers, doesn't copy). Freed in `deinit` after `vc_client_destroy`.
private var clientNamePtr: UnsafeMutablePointer<CChar>? private var clientNamePtr: UnsafeMutablePointer<CChar>?
private var clientVersionPtr: UnsafeMutablePointer<CChar>? private var clientVersionPtr: UnsafeMutablePointer<CChar>?
private var tofuStorePathPtr: UnsafeMutablePointer<CChar>? private var tofuStorePathPtr: UnsafeMutablePointer<CChar>?
@@ -90,7 +48,6 @@ public final class VoiceCatClient {
/// Intermediate values are coalesced (only the latest per stream_id is delivered). /// Intermediate values are coalesced (only the latest per stream_id is delivered).
public var onLevel: ((UInt32, Float) -> Void)? public var onLevel: ((UInt32, Float) -> Void)?
/// Lock-protected buffers, written from the core's event thread, drained on main.
private let bufferLock = NSLock() private let bufferLock = NSLock()
private var eventBuffer: [VoiceCatEvent] = [] private var eventBuffer: [VoiceCatEvent] = []
private var levelSamples: [UInt32: Float] = [:] private var levelSamples: [UInt32: Float] = [:]
@@ -98,19 +55,12 @@ public final class VoiceCatClient {
// MARK: - Init / deinit // MARK: - Init / deinit
/// Create a client. `config.clientName`/`clientVersion`/`tofuStorePath` are copied to
/// native CString storage held for the client's entire lifetime (the core reads them
/// later, e.g. when `connect()` runs on the io thread).
public init(config: VoiceCatConfig) { public init(config: VoiceCatConfig) {
// Allocate native C strings must persist until after vc_client_destroy in deinit.
// These don't need `self`, so they're safe to set first.
self.clientNamePtr = strdup(config.clientName) self.clientNamePtr = strdup(config.clientName)
self.clientVersionPtr = strdup(config.clientVersion) self.clientVersionPtr = strdup(config.clientVersion)
self.tofuStorePathPtr = config.tofuStorePath.flatMap { strdup($0) } self.tofuStorePathPtr = config.tofuStorePath.flatMap { strdup($0) }
self.handle = nil // placeholder set below after callbacks are wired self.handle = nil // placeholder set below after callbacks are wired
// All stored properties are now initialized `self` is fully available, so we can
// call `selfPointer` (the computed property) to build the callbacks struct.
var nativeConfig = vc_config() var nativeConfig = vc_config()
nativeConfig.client_name = UnsafePointer(clientNamePtr) nativeConfig.client_name = UnsafePointer(clientNamePtr)
nativeConfig.client_version = UnsafePointer(clientVersionPtr) nativeConfig.client_version = UnsafePointer(clientVersionPtr)
@@ -219,7 +169,7 @@ public final class VoiceCatClient {
VoiceCatResult(vc_authenticate_user(handle, username, password)) VoiceCatResult(vc_authenticate_user(handle, username, password))
} }
// MARK: - TOFU server-identity gate (M4) // MARK: - TOFU server-identity gate
/// Accept or reject the pending server-identity check. Call after a `.serverIdentity` /// Accept or reject the pending server-identity check. Call after a `.serverIdentity`
/// event. `accept=true` on firstConnect/mismatch updates the pin file and proceeds; /// event. `accept=true` on firstConnect/mismatch updates the pin file and proceeds;
@@ -256,6 +206,16 @@ public final class VoiceCatClient {
VoiceCatResult(vc_leave_channel(handle)) VoiceCatResult(vc_leave_channel(handle))
} }
@discardableResult
public func joinVoice() -> VoiceCatResult {
VoiceCatResult(vc_join_voice(handle))
}
@discardableResult
public func leaveVoice() -> VoiceCatResult {
VoiceCatResult(vc_leave_voice(handle))
}
/// Pull the current channel tree. Re-call after `.channelList`/`.userJoined`/`.userLeft`/ /// Pull the current channel tree. Re-call after `.channelList`/`.userJoined`/`.userLeft`/
/// `.userUpdated` events. The native list is freed inside this call callers never /// `.userUpdated` events. The native list is freed inside this call callers never
/// manage native lifetime. /// manage native lifetime.
@@ -403,12 +363,29 @@ public final class VoiceCatClient {
/// Global playback volume applied after mixing all remote streams. gain 0.0 = silent, /// Global playback volume applied after mixing all remote streams. gain 0.0 = silent,
/// 1.0 = unity (default), >1.0 amplifies. Always LOCAL no protocol traffic. Mirrors the /// 1.0 = unity (default), >1.0 amplifies. Always LOCAL no protocol traffic. Mirrors the
/// Windows client's `SetOutputVolume` and the C ABI `vc_set_output_volume` added in M5. /// Windows client's `SetOutputVolume` and the C ABI `vc_set_output_volume`.
@discardableResult @discardableResult
public func setOutputVolume(_ gain: Float) -> VoiceCatResult { public func setOutputVolume(_ gain: Float) -> VoiceCatResult {
VoiceCatResult(vc_set_output_volume(handle, gain < 0 ? 0 : gain)) VoiceCatResult(vc_set_output_volume(handle, gain < 0 ? 0 : gain))
} }
/// Send-side microphone input gain. Applied to captured MIC PCM before the VAD/PTT gate and
/// Opus encode (so boosting a quiet mic also helps it cross the VAD threshold). gain 0.0 =
/// silent, 1.0 = unity (default), >1.0 amplifies (clamped to int16). Always LOCAL.
@discardableResult
public func setInputGain(_ gain: Float) -> VoiceCatResult {
VoiceCatResult(vc_set_input_gain(handle, gain < 0 ? 0 : gain))
}
/// Send-side microphone noise suppression (RNNoise). Denoises captured MIC PCM before the
/// input gain and VAD/PTT gate, so everyone hears the cleaned signal (one pass for all
/// listeners). MIC stream only, mono only; always LOCAL no protocol traffic. Independent
/// of the per-listener receive-side NR in `setRemoteStream` (docs/voice.md §10).
@discardableResult
public func setInputNoiseReduction(_ enable: Bool) -> VoiceCatResult {
VoiceCatResult(vc_set_input_noise_reduction(handle, enable ? 1 : 0))
}
// MARK: - AVAudioSession interruption hooks (iOS) // MARK: - AVAudioSession interruption hooks (iOS)
/// Pause miniaudio device I/O. Call when AVAudioSession interruption begins. /// Pause miniaudio device I/O. Call when AVAudioSession interruption begins.
@@ -473,7 +450,7 @@ public final class VoiceCatClient {
return Marshaling.devices(&native) return Marshaling.devices(&native)
} }
// MARK: - M5: Moderation // MARK: - Moderation
@discardableResult @discardableResult
public func kickUser(_ userId: UInt32, reason: String? = nil) -> VoiceCatResult { public func kickUser(_ userId: UInt32, reason: String? = nil) -> VoiceCatResult {
@@ -508,7 +485,7 @@ public final class VoiceCatClient {
VoiceCatResult(vc_move_user(handle, userId, channelId)) VoiceCatResult(vc_move_user(handle, userId, channelId))
} }
// MARK: - M5: Channel admin // MARK: - Channel admin
@discardableResult @discardableResult
public func createChannel(_ info: ChannelEdit) -> VoiceCatResult { public func createChannel(_ info: ChannelEdit) -> VoiceCatResult {
@@ -531,7 +508,7 @@ public final class VoiceCatClient {
VoiceCatResult(vc_delete_channel(handle, channelId)) VoiceCatResult(vc_delete_channel(handle, channelId))
} }
// MARK: - M5: Account admin // MARK: - Account admin
@discardableResult @discardableResult
public func createAccount(_ username: String, password: String) -> VoiceCatResult { public func createAccount(_ username: String, password: String) -> VoiceCatResult {

View File

@@ -64,7 +64,7 @@ private final class ServerHarness {
} }
self.port = port self.port = port
// Provision a known admin account for moderation/admin tests (M5). // Provision a known admin account for moderation/admin tests.
let adminURL = URL(fileURLWithPath: repoRoot) let adminURL = URL(fileURLWithPath: repoRoot)
.appendingPathComponent("build/dev/bin/voicecat-admin") .appendingPathComponent("build/dev/bin/voicecat-admin")
guard FileManager.default.isExecutableFile(atPath: adminURL.path) else { guard FileManager.default.isExecutableFile(atPath: adminURL.path) else {
@@ -252,12 +252,12 @@ final class VoiceCatClientSmokeTests: XCTestCase {
XCTAssertTrue(channels.contains { $0.id == 1 && $0.name == "Lobby" }, XCTAssertTrue(channels.contains { $0.id == 1 && $0.name == "Lobby" },
"expected Lobby (channel 1) in \(channels.map { $0.name })") "expected Lobby (channel 1) in \(channels.map { $0.name })")
// M5: permissions getter round-trip. // Permissions getter round-trip.
let perms = client.getPermissions() let perms = client.getPermissions()
XCTAssertFalse(perms.isAdmin) XCTAssertFalse(perms.isAdmin)
XCTAssertFalse(perms.canKick) XCTAssertFalse(perms.canKick)
// M5: guest ListAccounts is rejected by the server with a GenericResult proves the // Guest ListAccounts is rejected by the server with a GenericResult proves the
// moderation wrapper path works end-to-end through the Swift interop layer. // moderation wrapper path works end-to-end through the Swift interop layer.
events.removeAll() events.removeAll()
XCTAssertEqual(client.requestAccountList(), .ok) XCTAssertEqual(client.requestAccountList(), .ok)

View File

@@ -4,7 +4,7 @@
<dict> <dict>
<key>com.apple.security.application-groups</key> <key>com.apple.security.application-groups</key>
<array> <array>
<string>group.cat.voice.VoiceCat</string> <string>group.me.iamtalon.voicecat</string>
</array> </array>
</dict> </dict>
</plist> </plist>

View File

@@ -7,6 +7,7 @@
objects = { objects = {
/* Begin PBXBuildFile section */ /* Begin PBXBuildFile section */
AAAA00000000000000000002 /* Assets.xcassets in Resources */ = {isa = PBXBuildFile; fileRef = AAAA00000000000000000001 /* Assets.xcassets */; };
BBBB00000000000000000030 /* VoiceCatiOSApp.swift in Sources */ = {isa = PBXBuildFile; fileRef = BBBB00000000000000000017 /* VoiceCatiOSApp.swift */; }; BBBB00000000000000000030 /* VoiceCatiOSApp.swift in Sources */ = {isa = PBXBuildFile; fileRef = BBBB00000000000000000017 /* VoiceCatiOSApp.swift */; };
BBBB00000000000000000031 /* AppState.swift in Sources */ = {isa = PBXBuildFile; fileRef = BBBB00000000000000000018 /* AppState.swift */; }; BBBB00000000000000000031 /* AppState.swift in Sources */ = {isa = PBXBuildFile; fileRef = BBBB00000000000000000018 /* AppState.swift */; };
BBBB00000000000000000032 /* SessionState.swift in Sources */ = {isa = PBXBuildFile; fileRef = BBBB00000000000000000019 /* SessionState.swift */; }; BBBB00000000000000000032 /* SessionState.swift in Sources */ = {isa = PBXBuildFile; fileRef = BBBB00000000000000000019 /* SessionState.swift */; };
@@ -61,6 +62,7 @@
/* End PBXTargetDependency section */ /* End PBXTargetDependency section */
/* Begin PBXFileReference section */ /* Begin PBXFileReference section */
AAAA00000000000000000001 /* Assets.xcassets */ = {isa = PBXFileReference; lastKnownFileType = folder.assetcatalog; path = Assets.xcassets; sourceTree = "<group>"; };
BBBB00000000000000000012 /* VoiceCatiOS.app */ = {isa = PBXFileReference; explicitFileType = wrapper.application; includeInIndex = 0; path = VoiceCatiOS.app; sourceTree = BUILT_PRODUCTS_DIR; }; BBBB00000000000000000012 /* VoiceCatiOS.app */ = {isa = PBXFileReference; explicitFileType = wrapper.application; includeInIndex = 0; path = VoiceCatiOS.app; sourceTree = BUILT_PRODUCTS_DIR; };
BBBB00000000000000000015 /* Info.plist */ = {isa = PBXFileReference; lastKnownFileType = text.plist.xml; path = Info.plist; sourceTree = "<group>"; }; BBBB00000000000000000015 /* Info.plist */ = {isa = PBXFileReference; lastKnownFileType = text.plist.xml; path = Info.plist; sourceTree = "<group>"; };
BBBB00000000000000000016 /* VoiceCatiOS.entitlements */ = {isa = PBXFileReference; lastKnownFileType = text.plist.entitlements; path = VoiceCatiOS.entitlements; sourceTree = "<group>"; }; BBBB00000000000000000016 /* VoiceCatiOS.entitlements */ = {isa = PBXFileReference; lastKnownFileType = text.plist.entitlements; path = VoiceCatiOS.entitlements; sourceTree = "<group>"; };
@@ -138,6 +140,7 @@
BBBB00000000000000000003 /* VoiceCatiOS */ = { BBBB00000000000000000003 /* VoiceCatiOS */ = {
isa = PBXGroup; isa = PBXGroup;
children = ( children = (
AAAA00000000000000000001 /* Assets.xcassets */,
BBBB00000000000000000015 /* Info.plist */, BBBB00000000000000000015 /* Info.plist */,
BBBB00000000000000000016 /* VoiceCatiOS.entitlements */, BBBB00000000000000000016 /* VoiceCatiOS.entitlements */,
BBBB00000000000000000017 /* VoiceCatiOSApp.swift */, BBBB00000000000000000017 /* VoiceCatiOSApp.swift */,
@@ -284,6 +287,7 @@
isa = PBXResourcesBuildPhase; isa = PBXResourcesBuildPhase;
buildActionMask = 2147483647; buildActionMask = 2147483647;
files = ( files = (
AAAA00000000000000000002 /* Assets.xcassets in Resources */,
); );
runOnlyForDeploymentPostprocessing = 0; runOnlyForDeploymentPostprocessing = 0;
}; };
@@ -473,7 +477,7 @@
"$(inherited)", "$(inherited)",
"-lc++", "-lc++",
); );
PRODUCT_BUNDLE_IDENTIFIER = cat.voice.VoiceCatiOS; PRODUCT_BUNDLE_IDENTIFIER = me.iamtalon.voicecat;
PRODUCT_NAME = "$(TARGET_NAME)"; PRODUCT_NAME = "$(TARGET_NAME)";
SWIFT_EMIT_LOC_STRINGS = YES; SWIFT_EMIT_LOC_STRINGS = YES;
SWIFT_VERSION = 5.9; SWIFT_VERSION = 5.9;
@@ -502,7 +506,7 @@
"$(inherited)", "$(inherited)",
"-lc++", "-lc++",
); );
PRODUCT_BUNDLE_IDENTIFIER = cat.voice.VoiceCatiOS; PRODUCT_BUNDLE_IDENTIFIER = me.iamtalon.voicecat;
PRODUCT_NAME = "$(TARGET_NAME)"; PRODUCT_NAME = "$(TARGET_NAME)";
SWIFT_EMIT_LOC_STRINGS = YES; SWIFT_EMIT_LOC_STRINGS = YES;
SWIFT_VERSION = 5.9; SWIFT_VERSION = 5.9;
@@ -526,7 +530,7 @@
"@executable_path/../../Frameworks", "@executable_path/../../Frameworks",
); );
MARKETING_VERSION = 0.0.1; MARKETING_VERSION = 0.0.1;
PRODUCT_BUNDLE_IDENTIFIER = cat.voice.VoiceCatiOS.broadcast; PRODUCT_BUNDLE_IDENTIFIER = me.iamtalon.voicecat.broadcast;
PRODUCT_NAME = "$(TARGET_NAME)"; PRODUCT_NAME = "$(TARGET_NAME)";
SKIP_INSTALL = YES; SKIP_INSTALL = YES;
SWIFT_OPTIMIZATION_LEVEL = "-Onone"; SWIFT_OPTIMIZATION_LEVEL = "-Onone";
@@ -551,7 +555,7 @@
"@executable_path/../../Frameworks", "@executable_path/../../Frameworks",
); );
MARKETING_VERSION = 0.0.1; MARKETING_VERSION = 0.0.1;
PRODUCT_BUNDLE_IDENTIFIER = cat.voice.VoiceCatiOS.broadcast; PRODUCT_BUNDLE_IDENTIFIER = me.iamtalon.voicecat.broadcast;
PRODUCT_NAME = "$(TARGET_NAME)"; PRODUCT_NAME = "$(TARGET_NAME)";
SKIP_INSTALL = YES; SKIP_INSTALL = YES;
SWIFT_COMPILATION_MODE = wholemodule; SWIFT_COMPILATION_MODE = wholemodule;

View File

@@ -1,4 +1,5 @@
import Foundation import Foundation
import Network
import VoiceCatCore import VoiceCatCore
struct PendingIdentity: Identifiable { struct PendingIdentity: Identifiable {
@@ -25,6 +26,33 @@ final class AppState {
private(set) var connectingServer: SavedServer? private(set) var connectingServer: SavedServer?
private var identityHandled = false private var identityHandled = false
/// Retained after authentication so an interrupted session can be restored.
private var connectedServer: SavedServer?
// MARK: - Reconnect state
/// Distinguishes an explicit disconnect from a transport failure.
private var userInitiatedDisconnect = false
private struct LastSession {
let server: SavedServer
let channelId: UInt32
let voiceSubscribed: Bool
let micMuted: Bool
let deafened: Bool
}
private var lastSession: LastSession?
private var reconnectAttempt = 0
private var reconnectTask: Task<Void, Never>?
/// Detects interface changes before TCP keepalive notices a dead path.
private var pathMonitor: NWPathMonitor?
private let pathQueue = DispatchQueue(label: "cat.voice.network.path")
private var lastPathSignature: String?
// MARK: - Server list management // MARK: - Server list management
func addServer(_ server: SavedServer, password: String?) { func addServer(_ server: SavedServer, password: String?) {
@@ -54,11 +82,19 @@ final class AppState {
// MARK: - Connect flow // MARK: - Connect flow
func connectTo(_ server: SavedServer) { func connectTo(_ server: SavedServer) {
connectTo(server, restoring: nil)
}
private func connectTo(_ server: SavedServer, restoring: LastSession?) {
guard !isConnecting else { return } guard !isConnecting else { return }
isConnecting = true isConnecting = true
connectStatus = "Connecting…" connectStatus = (restoring != nil) ? "Reconnecting…" : "Connecting…"
connectingServer = server connectingServer = server
identityHandled = false identityHandled = false
userInitiatedDisconnect = false
// Releasing the wrapper joins the core's I/O thread before freeing native strings.
connectingClient = nil
let config = VoiceCatConfig( let config = VoiceCatConfig(
clientName: "VoiceCat-iOS", clientName: "VoiceCat-iOS",
@@ -69,11 +105,14 @@ final class AppState {
connectingClient = client connectingClient = client
client.onEvent = { [weak self] ev in client.onEvent = { [weak self] ev in
Task { @MainActor [weak self] in self?.handleConnectEvent(ev, server: server) } Task { @MainActor [weak self] in
self?.handleConnectEvent(ev, server: server, restoring: restoring)
}
} }
// Authentication can start audio, so select the external path before connecting.
client.setExternalPlayback(true)
client.connect(host: server.host, port: server.port) client.connect(host: server.host, port: server.port)
// Auth is queued immediately the core serialises it behind TLS + TOFU.
switch server.authMode { switch server.authMode {
case .guest: case .guest:
let nick = (server.nickname?.isEmpty == false) ? server.nickname! : "iOS User" let nick = (server.nickname?.isEmpty == false) ? server.nickname! : "iOS User"
@@ -89,13 +128,19 @@ final class AppState {
} }
func disconnect() { func disconnect() {
session?.stopMicStream() // Set before disconnect so its event cannot arm reconnect.
userInitiatedDisconnect = true
cancelReconnect()
lastSession = nil
session?.leaveVoice()
session?.client.disconnect() session?.client.disconnect()
IOSAudioEngine.shared.stop()
AudioSessionManager.shared.deactivateSession() AudioSessionManager.shared.deactivateSession()
session = nil session = nil
connectingClient?.disconnect() connectingClient?.disconnect()
connectingClient = nil connectingClient = nil
connectingServer = nil connectingServer = nil
connectedServer = nil
isConnecting = false isConnecting = false
connectStatus = "" connectStatus = ""
showPasswordPrompt = false showPasswordPrompt = false
@@ -116,26 +161,143 @@ final class AppState {
} }
func cancelConnect() { func cancelConnect() {
// User explicitly cancelled no reconnect for the resulting .disconnected event.
userInitiatedDisconnect = true
cancelReconnect()
lastSession = nil
connectingClient?.disconnect() connectingClient?.disconnect()
connectingClient = nil connectingClient = nil
connectingServer = nil connectingServer = nil
connectedServer = nil
isConnecting = false isConnecting = false
connectStatus = "" connectStatus = ""
showPasswordPrompt = false showPasswordPrompt = false
pendingIdentity = nil pendingIdentity = nil
} }
// MARK: - Reconnect orchestration
private func cancelReconnect() {
reconnectTask?.cancel()
reconnectTask = nil
stopPathMonitor()
}
/// Schedules the next reconnect with exponential backoff capped at 30 seconds.
private func scheduleReconnect() {
guard !userInitiatedDisconnect, let last = lastSession else { return }
reconnectTask?.cancel()
reconnectAttempt = max(1, reconnectAttempt + 1)
let delaySec = min(pow(2.0, Double(reconnectAttempt - 1)), 30.0)
connectStatus = "Reconnecting (attempt \(reconnectAttempt))…"
startPathMonitor()
let task = Task { [weak self, last] in
guard let self else { return }
try? await Task.sleep(nanoseconds: UInt64(delaySec * 1_000_000_000))
if Task.isCancelled { return }
guard !self.userInitiatedDisconnect else { return }
guard self.lastSession != nil else { return }
guard self.session == nil else { return }
self.connectTo(last.server, restoring: last)
}
reconnectTask = task
}
private func startPathMonitor() {
guard pathMonitor == nil else { return }
let monitor = NWPathMonitor()
monitor.pathUpdateHandler = { [weak self] path in
Task { @MainActor [weak self] in
guard let self else { return }
guard !self.userInitiatedDisconnect else { return }
let sig = Self.pathSignature(path)
let prevSig = self.lastPathSignature
self.lastPathSignature = sig
if prevSig == nil { return }
if self.session != nil {
if path.status != .satisfied || sig != prevSig {
self.proactiveReconnect()
}
} else if self.lastSession != nil {
if path.status == .satisfied {
self.reconnectAttempt = 0
self.scheduleReconnect()
}
}
}
}
monitor.start(queue: pathQueue)
pathMonitor = monitor
}
private func stopPathMonitor() {
pathMonitor?.cancel()
pathMonitor = nil
lastPathSignature = nil
}
private static func pathSignature(_ path: NWPath) -> String {
guard path.status == .satisfied else { return "unsatisfied" }
var parts: [String] = []
if path.usesInterfaceType(.wifi) { parts.append("wifi") }
if path.usesInterfaceType(.cellular) { parts.append("cellular") }
if path.usesInterfaceType(.wiredEthernet) { parts.append("wired") }
if path.usesInterfaceType(.other) { parts.append("other") }
return parts.isEmpty ? "none" : parts.sorted().joined(separator: "+")
}
// MARK: - Live-session disconnect (called by SessionState)
/// Receives disconnects after `SessionState` takes ownership of authenticated events.
func onLiveSessionDisconnected() {
guard !userInitiatedDisconnect else { return }
teardownLiveSessionAndReconnect(sound: false)
}
private func proactiveReconnect() {
guard !userInitiatedDisconnect else { return }
guard session != nil else { return }
teardownLiveSessionAndReconnect(sound: true)
}
private func teardownLiveSessionAndReconnect(sound: Bool) {
if let s = session, let srv = connectedServer {
lastSession = LastSession(
server: srv,
channelId: s.currentChannelId,
voiceSubscribed: s.voiceState.voiceSubscribed,
micMuted: s.voiceState.selfMuted,
deafened: s.voiceState.selfDeafened)
}
IOSAudioEngine.shared.stop()
AudioSessionManager.shared.deactivateSession()
session = nil
isConnecting = false
connectingClient = nil
connectedServer = nil
if sound {
EventFeedback.shared.play(.connectionLost)
EventFeedback.shared.speak("Network changed — reconnecting")
}
reconnectAttempt = 0
scheduleReconnect()
}
// MARK: - Connect event handler // MARK: - Connect event handler
private func handleConnectEvent(_ ev: VoiceCatEvent, server: SavedServer) { private func handleConnectEvent(_ ev: VoiceCatEvent, server: SavedServer,
restoring: LastSession?) {
switch ev.type { switch ev.type {
case .connectionState: case .connectionState:
switch ev.connectionState { switch ev.connectionState {
case .connecting: connectStatus = "Connecting…" case .connecting: connectStatus = (restoring != nil) ? "Reconnecting…" : "Connecting…"
case .tlsHandshake: connectStatus = "TLS handshake…" case .tlsHandshake: connectStatus = "TLS handshake…"
case .authenticating: connectStatus = "Authenticating…" case .authenticating: connectStatus = "Authenticating…"
case .verifyingIdentity: connectStatus = "Verifying server identity…" case .verifyingIdentity: connectStatus = "Verifying server identity…"
case .connected: connectStatus = "Connected" case .connected: connectStatus = "Connected"
default: break default: break
} }
case .serverIdentity: case .serverIdentity:
@@ -153,35 +315,86 @@ final class AppState {
guard let client = connectingClient else { break } guard let client = connectingClient else { break }
let perms = client.getPermissions() let perms = client.getPermissions()
let newSession = SessionState(client: client, selfUserId: ev.userId, permissions: perms) let newSession = SessionState(client: client, selfUserId: ev.userId, permissions: perms)
newSession.appState = self
connectingClient = nil connectingClient = nil
isConnecting = false isConnecting = false
connectStatus = "" connectStatus = ""
showPasswordPrompt = false showPasswordPrompt = false
connectedServer = server
self.session = newSession self.session = newSession
// Activate the audio session now, while connected NOT lazily when the first EventFeedback.shared.play(.login)
// remote stream arrives. The core opens its miniaudio playback device the moment EventFeedback.shared.speak(restoring != nil ? "Reconnected" : "Connected")
// a remote stream starts and only THEN emits .streamStarted; if we waited for // External-playback mode was enabled before connect() so the core never opens a
// that event to activate, the playback device would open against an inactive // miniaudio device on iOS (the single ordering rule of the unified audio path).
// AVAudioSession and produce no sound (the "can't hear anyone" bug). Activating // Now activate the session and start the engine in listening mode so remote audio
// here guarantees the session is live before any device opens. // plays the moment someone talks, even before we join voice (no "can't hear anyone").
do { do {
try AudioSessionManager.shared.ensureSessionActive() try AudioSessionManager.shared.ensureSessionActive()
} catch { } catch {
print("Audio session activate on connect failed: \(error)") print("Audio session activate on connect failed: \(error)")
} }
IOSAudioEngine.shared.startListening(client: client)
// The path monitor runs the whole time we're connected so a network change fires
// proactiveReconnect immediately instead of waiting for the C core's TCP keepalive
// timeout (~30-60 s on a hard Wi-Fi drop). It stays armed across reconnects and is
// stopped only on user-initiated disconnect.
startPathMonitor()
// Reconnect restore: rejoin the prior channel and re-enable voice/mic if they
// were on. The session is fresh (server auto-places us in Lobby), so the restore
// is driven through SessionState.requestRestore, which issues a JoinChannel then
// (on the resulting .joinResult) re-arms voice + mute/deafen. A successful auth
// means the server is reachable, so the backoff counter resets and `lastSession`
// clears; the path monitor keeps watching for the next change.
if let restoring {
newSession.requestRestore(channelId: restoring.channelId,
voiceSubscribed: restoring.voiceSubscribed,
micMuted: restoring.micMuted,
deafened: restoring.deafened)
reconnectAttempt = 0
lastSession = nil
}
} else { } else {
connectStatus = "Auth failed: \(ev.result.description)" connectStatus = "Auth failed: \(ev.result.description)"
showPasswordPrompt = true showPasswordPrompt = true
} }
case .disconnected: case .disconnected:
if session == nil { cancelConnect() } // This handler runs ONLY during the connecting phase after auth success
else { // `SessionState.init` overwrites `client.onEvent`, so a live-session disconnect
AudioSessionManager.shared.deactivateSession() // reaches `SessionState.handleEvent` and comes back via
session = nil; isConnecting = false // `onLiveSessionDisconnected`, not here. Two outcomes for this branch:
// - A reconnect's connecting phase failed (`lastSession != nil`, set by a prior
// teardown) re-arm `scheduleReconnect` so the backoff loop continues.
// - A fresh connect failed before auth (`lastSession == nil`) show the error, do
// not auto-reconnect (the user should retry manually once the server is reachable).
connectingClient = nil
isConnecting = false
IOSAudioEngine.shared.stop()
AudioSessionManager.shared.deactivateSession()
if userInitiatedDisconnect {
connectStatus = ""
showPasswordPrompt = false
pendingIdentity = nil
lastSession = nil
connectedServer = nil
cancelReconnect()
} else if lastSession != nil {
// Mid-reconnect drop keep the backoff loop going.
EventFeedback.shared.play(.connectionLost)
EventFeedback.shared.speak("Connection lost — reconnecting")
scheduleReconnect()
} else {
// Fresh connect failed before auth. Surface the reason; no auto-reconnect.
connectStatus = ev.text ?? "Disconnected"
showPasswordPrompt = false
pendingIdentity = nil
connectedServer = nil
cancelReconnect()
} }
case .error: case .error:
connectStatus = ev.text ?? "Unknown error" connectStatus = ev.text ?? "Unknown error"
if session == nil { isConnecting = false } // Errors don't disconnect us; the .disconnected event handles teardown/reconnect.
default: default:
break break
} }

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.4 KiB

View File

@@ -0,0 +1,14 @@
{
"images" : [
{
"filename" : "AppIcon-1024.png",
"idiom" : "universal",
"platform" : "ios",
"size" : "1024x1024"
}
],
"info" : {
"author" : "xcode",
"version" : 1
}
}

View File

@@ -0,0 +1,6 @@
{
"info" : {
"author" : "xcode",
"version" : 1
}
}

View File

@@ -8,19 +8,6 @@ private let logger = Logger(subsystem: "cat.voice.VoiceCatiOS", category: "Audio
final class AudioSessionManager { final class AudioSessionManager {
static let shared = AudioSessionManager() static let shared = AudioSessionManager()
weak var client: VoiceCatClient?
/// The stream ID of the currently active local MIC stream, if any. Set by `SessionState`
/// when the user joins/leaves voice so `IOSAudioRouter` can reset the core's capture
/// channel count (e.g. when switching stereo mono) without going through `SessionState`.
var activeMicStreamId: UInt32?
/// Set by `SessionState`. Invoked by `IOSAudioRouter` after an audio-config change so the
/// voice path (native VPIO vs the core's miniaudio path) can be restarted to match the new
/// preset/route when the mic is active. No-op when not in voice. See
/// `SessionState.reconcileVoicePath()` and `IOSVoiceProcessingEngine`.
var reconcileVoicePath: (() -> Void)?
/// Tracks whether WE activated the session. The session must be active whenever the /// Tracks whether WE activated the session. The session must be active whenever the
/// AudioEngine is running (for capture OR playback), so it is activated when any audio /// AudioEngine is running (for capture OR playback), so it is activated when any audio
/// needs to play (a remote stream started OR the user joins voice) and only deactivated /// needs to play (a remote stream started OR the user joins voice) and only deactivated
@@ -28,6 +15,10 @@ final class AudioSessionManager {
/// want to hear remote audio. /// want to hear remote audio.
private var isSessionActive = false private var isSessionActive = false
/// Whether the AVAudioSession is currently active (we activated it). Read by `IOSAudioRouter`
/// to decide whether the post-activation A2DP speaker fallback can be applied.
var isActive: Bool { isSessionActive }
func configure() { func configure() {
// Load stored audio routing preferences and apply them before any audio session // Load stored audio routing preferences and apply them before any audio session
// activation. IOSAudioRouter drives all iOS audio route selection via AVAudioSession; // activation. IOSAudioRouter drives all iOS audio route selection via AVAudioSession;
@@ -44,6 +35,20 @@ final class AudioSessionManager {
name: AVAudioSession.routeChangeNotification, object: nil) name: AVAudioSession.routeChangeNotification, object: nil)
} }
/// Idempotently restores audio after an interruption or external route change.
func recoverAudio() {
guard IOSAudioEngine.shared.isConnected else { return }
do {
try ensureSessionActive()
} catch {
logger.error("recoverAudio — session activate failed: \(error.localizedDescription)")
}
IOSAudioRouter.shared.applyConfiguration()
if isSessionActive { IOSAudioRouter.shared.applyA2dpSpeakerFallback() }
IOSAudioEngine.shared.reconfigure()
logSessionState("after recoverAudio")
}
/// Activate the AVAudioSession if not already active. Call before any audio I/O: /// Activate the AVAudioSession if not already active. Call before any audio I/O:
/// when the user joins voice, or when a remote stream starts (so playback works even /// when the user joins voice, or when a remote stream starts (so playback works even
/// before the user has joined voice). Idempotent safe to call multiple times. /// before the user has joined voice). Idempotent safe to call multiple times.
@@ -114,22 +119,18 @@ final class AudioSessionManager {
switch type { switch type {
case .began: case .began:
// The system stops our AVAudioEngine and deactivates the session. Nothing to tear
// down `IOSAudioEngine` rebuilds on resume.
logger.info("interruption began — session suspended by system") logger.info("interruption began — session suspended by system")
isSessionActive = false // system deactivated us isSessionActive = false
client?.audioSuspend()
case .ended: case .ended:
let optionsValue = info[AVAudioSessionInterruptionOptionKey] as? UInt ?? 0 // Always attempt recovery when we have a live session. iOS sometimes ends an
let options = AVAudioSession.InterruptionOptions(rawValue: optionsValue) // interruption without the `.shouldResume` hint (e.g. Siri), and the previous
if options.contains(.shouldResume) { // behavior of only reactivating when `.shouldResume` was set left the session
do { // permanently dead audio never came back. `recoverAudio()` is intent-gated on
try AVAudioSession.sharedInstance().setActive(true) // `IOSAudioEngine.isConnected` and idempotent, so speculatively calling it is safe.
isSessionActive = true logger.info("interruption ended — recovery requested")
logger.info("interruption ended — session reactivated") recoverAudio()
client?.audioResume()
} catch {
logger.error("interruption ended — reactivation failed: \(error.localizedDescription)")
}
}
@unknown default: break @unknown default: break
} }
} }
@@ -139,34 +140,23 @@ final class AudioSessionManager {
let reasonValue = info[AVAudioSessionRouteChangeReasonKey] as? UInt, let reasonValue = info[AVAudioSessionRouteChangeReasonKey] as? UInt,
let reason = AVAudioSession.RouteChangeReason(rawValue: reasonValue) let reason = AVAudioSession.RouteChangeReason(rawValue: reasonValue)
else { else {
logger.warning("routeChange — unknown reason, refreshing only") logger.warning("routeChange — unknown reason, refreshing + recovery")
IOSAudioRouter.shared.refreshRoutes() IOSAudioRouter.shared.refreshRoutes()
NotificationCenter.default.post(name: .voiceCatDeviceListChanged, object: nil) NotificationCenter.default.post(name: .voiceCatDeviceListChanged, object: nil)
recoverAudio()
return return
} }
logger.info("routeChange reason=\(self.reasonLabel(reason))") logger.info("routeChange reason=\(self.reasonLabel(reason))")
// Re-apply preferences ONLY on external device plug/unplug. Do NOT re-apply on
// .categoryChange / .routeConfigurationChange those are triggered by our own
// applyConfiguration() calls (setCategory, setPreferredInput, etc.), and re-applying
// would create an infinite notification loop:
// handleRouteChange applyConfiguration setCategory routeChange ...
// That loop burns CPU and cycles the audio session on/off the "glitching" bug.
// IOSAudioRouter.applyConfiguration() also has a re-entrancy guard for synchronous
// notifications, but the reason check here is the primary defense.
if reason == .oldDeviceUnavailable || reason == .newDeviceAvailable {
logger.info("routeChange — external device change, re-applying config")
IOSAudioRouter.shared.applyConfiguration()
// Re-evaluate the A2DP-mode speaker fallback: a Bluetooth unplug should drop us onto
// the loud speaker (not the earpiece), and a replug should hand output back to A2DP.
if isSessionActive {
IOSAudioRouter.shared.applyA2dpSpeakerFallback()
}
}
IOSAudioRouter.shared.refreshRoutes() IOSAudioRouter.shared.refreshRoutes()
NotificationCenter.default.post(name: .voiceCatDeviceListChanged, object: nil) NotificationCenter.default.post(name: .voiceCatDeviceListChanged, object: nil)
// Ignore notifications caused by our own configuration calls; rebuilding for them
// recursively emits more route changes. Engine-configuration notifications remain
// the recovery path if a self-initiated change actually stops AVAudioEngine.
if reason != .categoryChange && reason != .routeConfigurationChange && reason != .override {
recoverAudio()
}
logSessionState("route change (\(reasonLabel(reason)))") logSessionState("route change (\(reasonLabel(reason)))")
} }
@@ -179,6 +169,7 @@ final class AudioSessionManager {
case .wakeFromSleep: return "wakeFromSleep" case .wakeFromSleep: return "wakeFromSleep"
case .noSuitableRouteForCategory: return "noSuitableRouteForCategory" case .noSuitableRouteForCategory: return "noSuitableRouteForCategory"
case .routeConfigurationChange: return "routeConfigurationChange" case .routeConfigurationChange: return "routeConfigurationChange"
case .unknown: return "unknown"
@unknown default: return "unknown" @unknown default: return "unknown"
} }
} }

View File

@@ -4,47 +4,8 @@ import VoiceCatCore
private let logger = Logger(subsystem: "cat.voice.VoiceCatiOS", category: "IOSAudioRouter") private let logger = Logger(subsystem: "cat.voice.VoiceCatiOS", category: "IOSAudioRouter")
/// iOS audio routing layer drives all iOS audio route selection via `AVAudioSession` /// Owns `AVAudioSession` routing for the iOS external-audio path.
/// *before* the core (miniaudio) opens its device. This class is the sole owner of the /// Route configuration and ordering constraints are documented in `docs/voice.md`.
/// session: miniaudio does NOT touch `AVAudioSession` on iOS, because the core opens its
/// devices through a `ma_context` configured with `sessionCategory = none` +
/// `noAudioSessionActivate/Deactivate` (see `AudioEngine::make_context_config` in
/// `core/src/audio/audio_engine.cpp`). Without that, miniaudio's default path resets the
/// category to `Record`/`Playback` with no options on every device open, wiping
/// `.allowBluetoothA2DP`/`.playAndRecord` and killing headphone/A2DP output so that
/// config must stay in place. All iOS audio routing (input port selection, mic
/// orientation/polar patterns, HFP vs A2DP, measurement/raw mode, stereo capture) must be
/// driven from here.
///
/// The three user-facing choices:
/// 1. **Input port** which physical input (built-in mic, Bluetooth HFP, headset,
/// USB, AirPlay). For the built-in mic, a sub-selection of **data source**
/// (orientation: front/back/top/bottom) and **polar pattern**
/// (omni/cardioid/subcardioid/bidirectional).
/// 2. **Bluetooth mode** how Bluetooth headsets are handled:
/// - "BT HFP voice" (`.allowBluetoothHFP` + `.allowBluetoothA2DP`): both profiles
/// allowed, iOS picks HFP for two-way mic or A2DP for output-only. Mono, AEC on.
/// - "Built-in Mic + BT A2DP stereo" (`.allowBluetoothA2DP` only): stereo output,
/// built-in mic, no HFP processing.
/// - "Built-in Mic + Speaker" (neither): no Bluetooth at all.
/// 3. **Mic processing mode** Standard (`.voiceChat`: AEC/AGC/HPF on) or
/// Raw/Studio (`.measurement`: all processing off). Raw mode is allowed always
/// but shows a warning when the output route is the speaker (echo risk, no AEC).
///
/// Additionally, **stereo capture** (2-channel built-in mic) is enabled by switching the
/// built-in mic's data source to the `.stereo` polar pattern. The recipe is:
/// `setPreferredDataSource(.stereo source)` + `setPreferredPolarPattern(.stereo)` +
/// `setPreferredInput(built-in mic)` + `setInputDataSource(stereo source)`. The channel
/// count itself must NOT be requested via `setPreferredInputNumberOfChannels(2)` that
/// session-level call collapses the A2DP output route. Instead the core is told to open the
/// device with 2 channels via `vc_set_capture_channels(streamId, 2)`, and the AVAudioSession
/// input anchor (`setPreferredInput` + `setInputDataSource`) keeps the route stable during
/// the HFPA2DP and monostereo reconfigurations.
///
/// Voice Isolation / Wide Spectrum (iOS 17+/18+) are user-toggleable in Control Center
/// for `.voiceChat` apps surfaced as a hint, not a programmatic toggle.
///
/// All choices are persisted in `UserDefaults` and re-applied on route changes.
@MainActor @MainActor
final class IOSAudioRouter: ObservableObject { final class IOSAudioRouter: ObservableObject {
@@ -64,77 +25,59 @@ final class IOSAudioRouter: ObservableObject {
@Published var selectedInputPortId: String? @Published var selectedInputPortId: String?
@Published var selectedDataSourceId: String? @Published var selectedDataSourceId: String?
@Published var selectedPolarPattern: String? @Published var selectedPolarPattern: String?
/// Master voice-processing switch (Apple VPIO: AEC + noise suppression bundled together).
/// iOS exposes no per-stage toggle, so this is the finest "echo cancellation / noise
/// reduction" control available. Only takes effect on a VPIO-capable config (mono + standard
/// + not A2DP); stereo / A2DP configs can't use VPIO regardless. Default on.
@Published var voiceProcessingEnabled: Bool = true
/// VPIO automatic gain control the one VPIO sub-stage iOS lets us toggle independently.
/// Only meaningful when voice processing is active. Default on.
@Published var agcEnabled: Bool = true
@Published var showsRawModeSpeakerWarning: Bool = false @Published var showsRawModeSpeakerWarning: Bool = false
@Published var showsA2dpNoAecWarning: Bool = false @Published var showsA2dpNoAecWarning: Bool = false
@Published var hasBluetoothDevice: Bool = false @Published var hasBluetoothDevice: Bool = false
@Published var hasWiredHeadset: Bool = false @Published var hasWiredHeadset: Bool = false
/// Audio presets sensible combinations of settings for common scenarios. /// Audio presets the four scenarios from the product spec. Pick a preset for a quick start,
/// The app is about choice: users can pick a preset for a quick start, then /// then fine-tune individual settings under "Advanced". HFP / wired headsets are not separate
/// fine-tune individual settings under "Advanced Audio". /// presets: Voice Chat lets the system route to them, and Advanced exposes manual selection.
enum AudioPreset: String, CaseIterable, Identifiable { enum AudioPreset: String, CaseIterable, Identifiable {
/// Standard iOS VoIP experience: AEC/AGC/HPF on, mono, system picks best route /// Voice chat: Apple VPIO does real AEC + noise suppression + AGC. Mono. The system picks
/// (BT HFP if connected, wired if connected, speaker if nothing). Always available. /// the best route (Bluetooth HFP / wired / speaker / earpiece). Always available.
case voiceChat = "Voice Chat" case voiceChat = "Voice Chat"
/// Stereo built-in mic capture (front+back capsules). A2DP output if BT is connected, /// Internal **stereo** built-in mic regardless of the output route. A2DP output when a
/// else built-in speaker / wired. Standard processing (no AEC stereo needs a non-VPIO /// Bluetooth headset is connected, else built-in speaker / wired. No VPIO (stereo can't
/// mode). Always available. /// use it). Always available.
case stereoMic = "Stereo Mic" case stereoMic = "Stereo Mic"
/// Maximum fidelity: stereo mic, no AEC/AGC/HPF (raw mode). A2DP output if BT connected, /// Internal **mono** built-in mic regardless of the output route. A2DP output when a
/// else speaker/wired. Always available. Echo risk on speaker. /// Bluetooth headset is connected, else built-in speaker / wired. No VPIO. Always available.
case studio = "Studio (No Processing)" case monoMic = "Mono Mic"
/// Bluetooth HFP: BT mic + BT output, AEC on, mono. Only when BT is connected. /// Everything manual input port, mic orientation / polar pattern, mono/stereo, Bluetooth
case bluetoothHeadset = "Bluetooth Headset (HFP)" /// mode, raw vs standard, and the VPIO / AGC toggles. Also the display state when the
/// A2DP stereo output + built-in mono mic, AEC off. Only when BT is connected. (For /// individual settings don't match a named preset.
/// A2DP output + stereo mic, use the Stereo Mic preset while BT is connected.) case advanced = "Advanced"
case btHeadphonesMonoMic = "BT Headphones + Mono Mic"
/// Wired headset/earpods: wired output + wired mic (or built-in), AEC on, mono.
/// Only when a wired audio device is connected.
case wiredHeadset = "Wired Headset"
/// Settings don't match any preset user has tweaked advanced controls.
case custom = "Custom"
var id: String { rawValue } var id: String { rawValue }
var requiresBluetooth: Bool {
switch self {
case .bluetoothHeadset, .btHeadphonesMonoMic: return true
default: return false
}
}
var requiresWired: Bool {
self == .wiredHeadset
}
var bluetoothMode: BluetoothMode { var bluetoothMode: BluetoothMode {
switch self { switch self {
case .voiceChat, .bluetoothHeadset: return .btHfpVoice case .voiceChat: return .btHfpVoice
// A2DP output when BT is connected; falls back to speaker/wired when it isn't. // Internal-mic presets: A2DP output when BT is connected; speaker/wired when not.
case .stereoMic, .studio, .btHeadphonesMonoMic: return .builtInMicBtA2dp case .stereoMic, .monoMic: return .builtInMicBtA2dp
case .wiredHeadset: return .builtInMicSpeaker case .advanced: return .builtInMicSpeaker // placeholder; Advanced sets it manually
case .custom: return .builtInMicSpeaker // placeholder
} }
} }
var captureChannels: CaptureChannels { var captureChannels: CaptureChannels {
switch self { self == .stereoMic ? .stereo : .mono
case .stereoMic, .studio: return .stereo
default: return .mono
}
} }
var micMode: MicMode { var micMode: MicMode { .standard }
switch self {
case .studio: return .raw
default: return .standard
}
}
/// Whether this preset explicitly selects the built-in mic port. /// Whether this preset explicitly pins the built-in mic port (the internal-mic presets).
var usesBuiltInMic: Bool { var usesBuiltInMic: Bool {
switch self { switch self {
case .stereoMic, .studio, .btHeadphonesMonoMic: return true case .stereoMic, .monoMic: return true
default: return false default: return false
} }
} }
@@ -170,14 +113,15 @@ final class IOSAudioRouter: ObservableObject {
private let kPolarPattern = "cat.voice.audio.polarPattern" private let kPolarPattern = "cat.voice.audio.polarPattern"
private let kPreset = "cat.voice.audio.preset" private let kPreset = "cat.voice.audio.preset"
private let kForceSpeaker = "cat.voice.audio.forceSpeaker" private let kForceSpeaker = "cat.voice.audio.forceSpeaker"
private let kVoiceProcessing = "cat.voice.audio.voiceProcessing"
private let kAgc = "cat.voice.audio.agc"
/// Re-entrancy guard: setCategory/setPreferredInput/etc. trigger route-change /// AVAudioSession setters can synchronously emit route-change notifications.
/// notifications synchronously on the same thread. Without this guard,
/// handleRouteChange applyConfiguration setCategory route-change notification
/// handleRouteChange applyConfiguration ... creates an infinite loop that
/// burns CPU and cycles the audio session on/off (the "glitching" bug).
private var isApplyingConfiguration = false private var isApplyingConfiguration = false
/// Prevents redundant overrides; `setCategory` invalidates the cached value.
private var lastAppliedOutputOverride: AVAudioSession.PortOverride?
private init() {} private init() {}
// MARK: - Load / refresh from AVAudioSession // MARK: - Load / refresh from AVAudioSession
@@ -271,55 +215,41 @@ final class IOSAudioRouter: ObservableObject {
} }
} }
/// The presets available given the current device connection state. /// The presets the user can pick. All four are always available the named presets simply
/// Always includes Voice Chat, Stereo Mic, Studio, and Custom. BT presets only when /// describe what to do "regardless of the output route", and Advanced is always offered.
/// a Bluetooth device is connected. Wired preset only when a wired device is connected. var availablePresets: [AudioPreset] { AudioPreset.allCases }
var availablePresets: [AudioPreset] {
AudioPreset.allCases.filter { preset in
if preset == .custom { return true }
if preset.requiresBluetooth && !hasBluetoothDevice { return false }
if preset.requiresWired && !hasWiredHeadset { return false }
return true
}
}
/// Which preset matches the current settings, or .custom if nothing matches. /// Which named preset matches the current settings, or `.advanced` if nothing matches.
/// Checks device-specific presets first (BT, wired) so that e.g. when BT is connected
/// and settings match "Bluetooth Headset", it returns that instead of the equivalent
/// "Voice Chat" (which has the same bluetoothMode/micMode/channels but is more general).
var activePreset: AudioPreset { var activePreset: AudioPreset {
// Check device-specific presets first (most specific least specific) for preset in [AudioPreset.voiceChat, .stereoMic, .monoMic] {
let order: [AudioPreset] = [
.bluetoothHeadset, .btHeadphonesMonoMic,
.wiredHeadset,
.voiceChat, .stereoMic, .studio,
]
for preset in order {
if bluetoothMode == preset.bluetoothMode if bluetoothMode == preset.bluetoothMode
&& captureChannels == preset.captureChannels && captureChannels == preset.captureChannels
&& micMode == preset.micMode { && micMode == preset.micMode {
// Don't match a BT preset if no BT is connected fall through to Voice Chat
if preset.requiresBluetooth && !hasBluetoothDevice { continue }
if preset.requiresWired && !hasWiredHeadset { continue }
return preset return preset
} }
} }
return .custom return .advanced
} }
/// Whether the current configuration should use the native iOS Voice-Processing path (VPIO: /// Whether the current configuration should engage Apple's Voice-Processing I/O unit (VPIO:
/// real AEC/NS/AGC via `IOSVoiceProcessingEngine`). True exactly when `applyConfiguration` /// real AEC + noise suppression + AGC, driven by `IOSAudioEngine`). VPIO forces mono and
/// selects the `.voiceChat` AVAudioSession mode mono + standard processing + not A2DP /// can't run on an A2DP route, so it is available only for a mono + standard + non-A2DP
/// (A2DP / stereo / raw modes can't use VPIO, so they keep the core's miniaudio path). /// config, and then only when the user hasn't disabled it via the Advanced master toggle.
var currentConfigUsesVoiceProcessing: Bool { var currentConfigUsesVoiceProcessing: Bool {
voiceProcessingEnabled && voiceProcessingAvailable
}
/// Whether the current config *could* use VPIO (mono + standard + non-A2DP), independent of
/// the user's master toggle. Drives whether the Advanced "Voice Processing" switch is shown.
var voiceProcessingAvailable: Bool {
captureChannels == .mono && micMode == .standard && bluetoothMode != .builtInMicBtA2dp captureChannels == .mono && micMode == .standard && bluetoothMode != .builtInMicBtA2dp
} }
// MARK: - Apply configuration // MARK: - Apply configuration
/// Apply the full audio configuration to AVAudioSession. Call this before the core /// Apply the full audio configuration to AVAudioSession. Call this before (re)building the
/// opens its capture device (i.e. before `startMicStream` `activateForStreaming`). /// `IOSAudioEngine` graph so the engine binds to the intended route (`applyAndReconfigure`
/// Re-entrant-safe: if a route-change notification fires synchronously during a /// does both). Re-entrant-safe: if a route-change notification fires synchronously during a
/// `setCategory`/`setPreferredInput` call, the guard prevents re-entry. /// `setCategory`/`setPreferredInput` call, the guard prevents re-entry.
func applyConfiguration() { func applyConfiguration() {
guard !isApplyingConfiguration else { guard !isApplyingConfiguration else {
@@ -327,6 +257,9 @@ final class IOSAudioRouter: ObservableObject {
return return
} }
isApplyingConfiguration = true isApplyingConfiguration = true
// setCategory below can reset the override out from under us, so drop our cached
// value applyA2dpSpeakerFallback will re-derive and re-apply it from scratch.
lastAppliedOutputOverride = nil
defer { isApplyingConfiguration = false } defer { isApplyingConfiguration = false }
let session = AVAudioSession.sharedInstance() let session = AVAudioSession.sharedInstance()
@@ -397,13 +330,8 @@ final class IOSAudioRouter: ObservableObject {
// 3. Input & mic-capsule configuration. // 3. Input & mic-capsule configuration.
if captureChannels == .stereo { if captureChannels == .stereo {
// Stereo: enable the built-in mic's .stereo polar pattern AND anchor the input // See configureStereoCapture's doc comment for the full stereo-capture recipe
// route explicitly via setPreferredInput + setInputDataSource. With HFP disabled // and why each step is necessary.
// the system routes input to the built-in mic, but without the explicit
// preferred-input anchor the route can collapse during the mode switch
// (.voiceChat .default) and the output dies. The channel count is requested by
// miniaudio at the audio-unit level (vc_set_capture_channels), NOT via
// setPreferredInputNumberOfChannels(2) that call collapses the A2DP output route.
configureStereoCapture(session: session) configureStereoCapture(session: session)
} else if let portId = selectedInputPortId, !portId.isEmpty, } else if let portId = selectedInputPortId, !portId.isEmpty,
let port = session.availableInputs?.first(where: { $0.uid == portId }) { let port = session.availableInputs?.first(where: { $0.uid == portId }) {
@@ -424,16 +352,8 @@ final class IOSAudioRouter: ObservableObject {
updateWarnings() updateWarnings()
} }
/// Enable 2-channel capture on the built-in mic. The recipe that achieves stereo mic + /// Anchors the built-in stereo data source without using
/// A2DP Bluetooth output simultaneously: /// `setPreferredInputNumberOfChannels`, which disrupts A2DP routing.
/// 1. `setPreferredDataSource(stereoSource)` on the built-in mic port
/// 2. `setPreferredPolarPattern(.stereo)` on that data source
/// 3. `setPreferredInput(builtIn)` anchor the input route explicitly. Without this
/// anchor the route can collapse during the mode switch (.voiceChat .default).
/// 4. `setInputDataSource(stereoSource)` commit the data source at the session level
/// The channel count itself is requested by miniaudio at the audio-unit level via
/// `vc_set_capture_channels(2)`. We must NOT call `setPreferredInputNumberOfChannels(2)`
/// that session-level call collapses the A2DP output route.
private func configureStereoCapture(session: AVAudioSession) { private func configureStereoCapture(session: AVAudioSession) {
guard let builtIn = session.availableInputs?.first(where: { $0.portType == .builtInMic }) guard let builtIn = session.availableInputs?.first(where: { $0.portType == .builtInMic })
else { else {
@@ -449,8 +369,6 @@ final class IOSAudioRouter: ObservableObject {
do { do {
try builtIn.setPreferredDataSource(stereoSource) try builtIn.setPreferredDataSource(stereoSource)
try stereoSource.setPreferredPolarPattern(.stereo) try stereoSource.setPreferredPolarPattern(.stereo)
// Anchor the input route explicitly; without it the route can collapse during
// the mode switch (.voiceChat .default) and the A2DP output dies.
try session.setPreferredInput(builtIn) try session.setPreferredInput(builtIn)
// Commit the data source at the session level. setPreferredDataSource alone only // Commit the data source at the session level. setPreferredDataSource alone only
// sets the port-level preference; setInputDataSource makes it the active source. // sets the port-level preference; setInputDataSource makes it the active source.
@@ -538,6 +456,9 @@ final class IOSAudioRouter: ObservableObject {
selectedDataSourceId = UserDefaults.standard.string(forKey: kDataSourceId) selectedDataSourceId = UserDefaults.standard.string(forKey: kDataSourceId)
selectedPolarPattern = UserDefaults.standard.string(forKey: kPolarPattern) selectedPolarPattern = UserDefaults.standard.string(forKey: kPolarPattern)
forceSpeaker = UserDefaults.standard.bool(forKey: kForceSpeaker) forceSpeaker = UserDefaults.standard.bool(forKey: kForceSpeaker)
// VPIO toggles default ON when never set (object(forKey:) is nil use true).
voiceProcessingEnabled = (UserDefaults.standard.object(forKey: kVoiceProcessing) as? Bool) ?? true
agcEnabled = (UserDefaults.standard.object(forKey: kAgc) as? Bool) ?? true
} }
/// Persist current selections to UserDefaults. /// Persist current selections to UserDefaults.
func savePreferences() { func savePreferences() {
@@ -548,86 +469,84 @@ final class IOSAudioRouter: ObservableObject {
UserDefaults.standard.set(selectedDataSourceId, forKey: kDataSourceId) UserDefaults.standard.set(selectedDataSourceId, forKey: kDataSourceId)
UserDefaults.standard.set(selectedPolarPattern, forKey: kPolarPattern) UserDefaults.standard.set(selectedPolarPattern, forKey: kPolarPattern)
UserDefaults.standard.set(forceSpeaker, forKey: kForceSpeaker) UserDefaults.standard.set(forceSpeaker, forKey: kForceSpeaker)
UserDefaults.standard.set(voiceProcessingEnabled, forKey: kVoiceProcessing)
UserDefaults.standard.set(agcEnabled, forKey: kAgc)
} }
// MARK: - Selection setters (called from SettingsView pickers) // MARK: - Selection setters (called from SettingsView pickers)
/// Persists the selection and rebuilds the engine against the resulting route.
private func applyAndReconfigure() {
savePreferences()
applyConfiguration()
if AudioSessionManager.shared.isActive { applyA2dpSpeakerFallback() }
refreshRoutes()
IOSAudioEngine.shared.reconfigure()
}
func selectInputPort(_ portId: String) { func selectInputPort(_ portId: String) {
selectedInputPortId = portId selectedInputPortId = portId
selectedDataSourceId = nil selectedDataSourceId = nil
selectedPolarPattern = nil selectedPolarPattern = nil
savePreferences() applyAndReconfigure()
applyConfiguration()
refreshRoutes()
} }
func selectDataSource(_ dataSourceId: String) { func selectDataSource(_ dataSourceId: String) {
selectedDataSourceId = dataSourceId selectedDataSourceId = dataSourceId
selectedPolarPattern = nil selectedPolarPattern = nil
savePreferences() applyAndReconfigure()
applyConfiguration()
refreshRoutes()
} }
func selectPolarPattern(_ pattern: String) { func selectPolarPattern(_ pattern: String) {
selectedPolarPattern = pattern selectedPolarPattern = pattern
savePreferences() applyAndReconfigure()
applyConfiguration()
refreshRoutes()
} }
func selectBluetoothMode(_ mode: BluetoothMode) { func selectBluetoothMode(_ mode: BluetoothMode) {
bluetoothMode = mode bluetoothMode = mode
savePreferences() applyAndReconfigure()
applyConfiguration()
refreshRoutes()
// VPIO class or route may have changed restart the voice path if mic is active.
AudioSessionManager.shared.reconcileVoicePath?()
} }
func setForceSpeaker(_ on: Bool) { func setForceSpeaker(_ on: Bool) {
forceSpeaker = on forceSpeaker = on
savePreferences() applyAndReconfigure()
applyConfiguration()
refreshRoutes()
// Route changed under a possibly-running VPIO engine reconcile if mic is active.
AudioSessionManager.shared.reconcileVoicePath?()
} }
func selectMicMode(_ mode: MicMode) { func selectMicMode(_ mode: MicMode) {
micMode = mode micMode = mode
applyAndReconfigure()
}
func setVoiceProcessingEnabled(_ on: Bool) {
voiceProcessingEnabled = on
applyAndReconfigure()
}
func setAgcEnabled(_ on: Bool) {
agcEnabled = on
// No session reconfigure needed just rebuild the engine so VPIO picks up the AGC flag.
savePreferences() savePreferences()
applyConfiguration() IOSAudioEngine.shared.reconfigure()
updateWarnings()
// StandardRaw flips the VPIO class reconcile if mic is active.
AudioSessionManager.shared.reconcileVoicePath?()
} }
func selectCaptureChannels(_ channels: CaptureChannels) { func selectCaptureChannels(_ channels: CaptureChannels) {
captureChannels = channels captureChannels = channels
savePreferences() savePreferences()
applyConfiguration() applyConfiguration()
// Update the core's stored capture channel count (does not restart the engine). if AudioSessionManager.shared.isActive { applyA2dpSpeakerFallback() }
if let streamId = AudioSessionManager.shared.activeMicStreamId { refreshRoutes()
_ = AudioSessionManager.shared.client?.setCaptureChannels( // Push the channel count into the core's MIC stream, then rebuild the engine graph so the
streamId: streamId, channels: channels.channelCount) // mic tap captures the right number of channels. The engine owns the route now, so there's
} // no stereo-vs-A2DP race to sequence around.
// Restart the engine AFTER AVAudioSession routing has settled and the channel IOSAudioEngine.shared.setCaptureChannels(channels.channelCount)
// count is stored. The engine reopens playback first (committing the A2DP/output
// route), then capture avoiding the race where stereo capture activation drops
// A2DP before the playback device has a chance to claim the route.
_ = AudioSessionManager.shared.client?.audioRestart()
// Monostereo flips the VPIO class (stereo can't use VPIO) reconcile if mic is active.
AudioSessionManager.shared.reconcileVoicePath?()
} }
// MARK: - Presets // MARK: - Presets
/// Apply a preset sets all individual audio settings to the preset's values, then /// Apply a named preset set all individual settings to the preset's values, then re-apply
/// applies the configuration. For presets that use the built-in mic (A2DP presets), /// the configuration and rebind the engine. The internal-mic presets pin the built-in mic.
/// finds the built-in mic port UID from availableInputs.
func applyPreset(_ preset: AudioPreset) { func applyPreset(_ preset: AudioPreset) {
guard preset != .custom else { return } // can't "apply" custom it's a display state guard preset != .advanced else { return } // Advanced is a display state, not "applied"
bluetoothMode = preset.bluetoothMode bluetoothMode = preset.bluetoothMode
micMode = preset.micMode micMode = preset.micMode
@@ -638,19 +557,17 @@ final class IOSAudioRouter: ObservableObject {
if preset == .voiceChat { forceSpeaker = true } if preset == .voiceChat { forceSpeaker = true }
if preset.usesBuiltInMic { if preset.usesBuiltInMic {
// Find the built-in mic port from available inputs and select it. // Pin the built-in mic. In stereo, iOS uses multiple capsules automatically; in mono
let session = AVAudioSession.sharedInstance() // the default orientation is fine so don't force a specific data source / pattern.
if let builtInMic = (session.availableInputs ?? []).first(where: { if let builtInMic = (AVAudioSession.sharedInstance().availableInputs ?? []).first(where: {
$0.portType == .builtInMic $0.portType == .builtInMic
}) { }) {
selectedInputPortId = builtInMic.uid selectedInputPortId = builtInMic.uid
} }
// Don't set a specific data source in stereo mode, iOS uses multiple mic
// capsules automatically. In mono, the default orientation is fine.
selectedDataSourceId = nil selectedDataSourceId = nil
selectedPolarPattern = nil selectedPolarPattern = nil
} else { } else {
// For Default and Bluetooth Headset presets, let the system pick the input. // Voice Chat: let the system pick the input (Bluetooth HFP / wired / built-in).
selectedInputPortId = nil selectedInputPortId = nil
selectedDataSourceId = nil selectedDataSourceId = nil
selectedPolarPattern = nil selectedPolarPattern = nil
@@ -659,18 +576,11 @@ final class IOSAudioRouter: ObservableObject {
UserDefaults.standard.set(preset.rawValue, forKey: kPreset) UserDefaults.standard.set(preset.rawValue, forKey: kPreset)
savePreferences() savePreferences()
applyConfiguration() applyConfiguration()
// Update the core's stored capture channel count (does not restart the engine). if AudioSessionManager.shared.isActive { applyA2dpSpeakerFallback() }
if let streamId = AudioSessionManager.shared.activeMicStreamId {
_ = AudioSessionManager.shared.client?.setCaptureChannels(
streamId: streamId, channels: preset.captureChannels.channelCount)
}
// Restart the engine AFTER AVAudioSession routing has settled and the channel
// count is stored. Playback opens first (commits A2DP route), then capture.
_ = AudioSessionManager.shared.client?.audioRestart()
refreshRoutes() refreshRoutes()
// The preset may have flipped the VPIO class (and/or the route) restart the voice path // Push the channel count to the core, then rebuild the engine graph (VPIO on/off + tap).
// if the mic is active so AEC/NS engage (or disengage) to match the new preset. IOSAudioEngine.shared.setCaptureChannels(preset.captureChannels.channelCount)
AudioSessionManager.shared.reconcileVoicePath?() IOSAudioEngine.shared.reconfigure()
logger.info("applyPreset — \(preset.rawValue)") logger.info("applyPreset — \(preset.rawValue)")
} }
@@ -689,15 +599,8 @@ final class IOSAudioRouter: ObservableObject {
showsA2dpNoAecWarning = (bluetoothMode == .builtInMicBtA2dp) showsA2dpNoAecWarning = (bluetoothMode == .builtInMicBtA2dp)
} }
/// Route fallback for the A2DP-output presets (Stereo Mic / Studio / BT Headphones + Mono /// Uses the speaker only when an A2DP-capable preset has no external output.
/// Mic, all `.builtInMicBtA2dp`). These presets deliberately omit `.defaultToSpeaker` (it /// The cached override avoids recursively generated route-change notifications.
/// breaks A2DP routing) and skip the `forceSpeaker` override, so when NO external output
/// (Bluetooth A2DP / wired / AirPlay) is connected `.playAndRecord` pins output to the quiet
/// built-in receiver (earpiece). This routes to the loud built-in speaker instead via a
/// post-activation `overrideOutputAudioPort(.speaker)` the documented "A2DP if connected,
/// else speaker" behavior. When an external output IS present we clear the override so A2DP /
/// headphones / AirPlay are honored. No-op outside `.builtInMicBtA2dp` mode (other modes pick
/// their route via category options). Must be called AFTER the session is active.
func applyA2dpSpeakerFallback() { func applyA2dpSpeakerFallback() {
guard bluetoothMode == .builtInMicBtA2dp else { return } guard bluetoothMode == .builtInMicBtA2dp else { return }
let session = AVAudioSession.sharedInstance() let session = AVAudioSession.sharedInstance()
@@ -706,19 +609,30 @@ final class IOSAudioRouter: ObservableObject {
let hasExternalOutput = session.currentRoute.outputs.contains { let hasExternalOutput = session.currentRoute.outputs.contains {
$0.portType != .builtInReceiver && $0.portType != .builtInSpeaker $0.portType != .builtInReceiver && $0.portType != .builtInSpeaker
} }
let desired: AVAudioSession.PortOverride = hasExternalOutput ? .none : .speaker
if desired == lastAppliedOutputOverride {
logger.debug("A2DP fallback — desired=\(self.overrideLabel(desired)) already applied, skipping")
return
}
do { do {
if hasExternalOutput { try session.overrideOutputAudioPort(desired)
try session.overrideOutputAudioPort(.none) lastAppliedOutputOverride = desired
logger.info("A2DP mode — external output present, clearing speaker override") logger.info("A2DP mode — override applied: \(self.overrideLabel(desired))")
} else {
try session.overrideOutputAudioPort(.speaker)
logger.info("A2DP mode — no external output, routing to built-in speaker")
}
} catch { } catch {
// Drop the cache so the next call re-derives from the live session state.
lastAppliedOutputOverride = nil
logger.error("A2DP speaker fallback failed: \(error.localizedDescription)") logger.error("A2DP speaker fallback failed: \(error.localizedDescription)")
} }
} }
private func overrideLabel(_ o: AVAudioSession.PortOverride) -> String {
switch o {
case .none: return "none"
case .speaker: return "speaker"
@unknown default: return "unknown"
}
}
/// The selected input port object, if any. /// The selected input port object, if any.
var selectedPort: IOSAudioInputPort? { var selectedPort: IOSAudioInputPort? {
inputPorts.first(where: { $0.id == selectedInputPortId }) inputPorts.first(where: { $0.id == selectedInputPortId })

View File

@@ -3,9 +3,9 @@ import Darwin
import os import os
import VoiceCatCore import VoiceCatCore
private let logger = Logger(subsystem: "cat.voice.VoiceCatiOS", category: "IOSVoiceProcessingEngine") private let logger = Logger(subsystem: "cat.voice.VoiceCatiOS", category: "IOSAudioEngine")
/// In-process single-producer/single-consumer int16 PCM ring for the VPIO playback path. /// In-process single-producer/single-consumer int16 PCM ring for the playback path.
/// ///
/// producer = the core's mixer-timer thread (the `vc_set_mixed_output_sink` callback) /// producer = the core's mixer-timer thread (the `vc_set_mixed_output_sink` callback)
/// consumer = the `AVAudioSourceNode` render thread /// consumer = the `AVAudioSourceNode` render thread
@@ -73,44 +73,88 @@ final class PCMRing {
return n return n
} }
/// Consumer-side snapshot of how many interleaved int16 samples are currently buffered. Lets a
/// paced consumer check for a full frame *before* calling `read`, so it never reads (and thus
/// discards) a partial frame. Single consumer only (same thread that calls `read`).
var availableSamples: Int {
let r = readIdx
OSMemoryBarrier()
let w = writeIdx
return Int(w &- r)
}
/// Discard everything buffered call before (re)starting so stale pre-roll isn't played. /// Discard everything buffered call before (re)starting so stale pre-roll isn't played.
func reset() { OSMemoryBarrier(); readIdx = writeIdx } func reset() { OSMemoryBarrier(); readIdx = writeIdx }
/// Diagnostics: monotonic total samples written / read since the ring was created. The /// Diagnostics: monotonic total samples written / read since the ring was created. The
/// indices are already cumulative, so these are free. Only read them when both threads are /// indices are already cumulative, so these are free. Only read them when both threads are
/// quiesced (e.g. at teardown after the engine + mixer sink are stopped) they are not /// quiesced (e.g. at teardown after the engine + mixer sink are stopped) they are not
/// synchronized for live cross-thread reads. Lets us tell "core never delivered PCM" (Bug 1 /// synchronized for live cross-thread reads. Lets us tell "core never delivered PCM" apart
/// core path) apart from "PCM arrived but produced no sound" (AVAudioEngine output graph). /// from "PCM arrived but produced no sound" (the AVAudioEngine output graph).
var debugTotalWritten: UInt64 { writeIdx } var debugTotalWritten: UInt64 { writeIdx }
var debugTotalRead: UInt64 { readIdx } var debugTotalRead: UInt64 { readIdx }
} }
/// Native iOS voice-processing audio path (docs/voice.md §8 "iOS voice processing"). /// The single iOS audio engine (docs/voice.md §8 "iOS audio engine").
/// ///
/// Real iOS echo cancellation / noise suppression / AGC come ONLY from Apple's Voice-Processing /// **One path, always external.** On iOS the core never opens a miniaudio device: a MIC stream is
/// I/O unit (VPIO), which `AVAudioEngine.setVoiceProcessingEnabled(true)` enables. For VPIO to /// always started with `external_feed=1`, `vc_set_external_playback(1)` is set once at connect, and
/// cancel echo it must own BOTH the mic capture and the remote-audio playback (it subtracts the /// this engine drives *both* directions through one `AVAudioEngine`:
/// played-back signal from the mic), so on the AEC presets this engine drives both directions and
/// the core runs in external mode (no hardware devices):
/// - **Mic core:** a tap on the VPIO input node 48 kHz int16 `client.feedPcm(micStreamId)`.
/// - **core speaker:** the core's mixed-output sink fills `ring`; an `AVAudioSourceNode` pulls /// - **core speaker:** the core's mixed-output sink fills `ring`; an `AVAudioSourceNode` pulls
/// from it and renders through the VPIO output, giving AEC its reference signal. /// from it and renders through the engine output. This runs the whole time we're connected,
/// so remote audio plays even before the user joins voice (no "can't hear anyone").
/// - **mic core:** when the mic is active a tap on the input node converts to 48 kHz int16 and
/// writes to a pacing ring; a 20 ms timer releases steady 960-sample frames to
/// `client.feedPcm(micStreamId)`. The core sends each captured frame synchronously, so this
/// steady cadence is what keeps packets from bursting and fluttering the receiver's playout.
/// ///
/// Lifecycle is driven by `SessionState` join/leave. The Stereo Mic / Studio / A2DP presets keep /// Echo cancellation / noise suppression / AGC come from Apple's Voice-Processing I/O unit (VPIO),
/// the core's miniaudio path instead (they want raw / stereo / no-AEC routing VPIO can't provide). /// which `inputNode.setVoiceProcessingEnabled(true)` enables. VPIO forces mono, so it is engaged
/// only when the active preset wants it (`IOSAudioRouter.currentConfigUsesVoiceProcessing`) the
/// Stereo Mic / A2DP configs run the same engine with VPIO off.
///
/// Every preset / route / interruption change funnels through `reconfigure()`: a single
/// deterministic stop AVAudioSession reconfigure rebuild graph start. There is no second
/// (miniaudio) audio path to hand off to, so a switch cannot leave one direction dropped.
@MainActor @MainActor
final class IOSVoiceProcessingEngine { final class IOSAudioEngine {
static let shared = IOSVoiceProcessingEngine() static let shared = IOSAudioEngine()
private(set) var isRunning = false /// True while connected (between `startListening` and `stop`) the playback graph should run.
private(set) var isConnected = false
/// True while a local mic stream is active the input tap should be installed.
private(set) var micActive = false
private let engine = AVAudioEngine() private let engine = AVAudioEngine()
private var sourceNode: AVAudioSourceNode? private var sourceNode: AVAudioSourceNode?
private weak var client: VoiceCatClient? private weak var client: VoiceCatClient?
private var micStreamId: UInt32 = 0 private var micStreamId: UInt32 = 0
private var captureChannels: UInt32 = 1
// 48 kHz stereo Float32 (deinterleaved) the format the source node renders and the engine // AVAudioEngine may deliver several codec frames per callback. Pace complete 20 ms frames
// processes in. The core delivers 48 kHz stereo int16 via the mixed-output sink. // through an SPSC ring; never consume partial frames, and recreate the timer when the channel
// count changes.
private let micRing = PCMRing(capacitySamples: 48000 * 2) // ~1 s stereo ample elastic slack
private var micTimer: DispatchSourceTimer?
private let micQueue = DispatchQueue(label: "cat.voice.mic.feedPump")
private let micDrainScratch: UnsafeMutablePointer<Int16>
private static let micFrameSamplesPerChannel = 960 // 20 ms @ 48 kHz core's frame size
/// Feed-pump state, touched only on `micQueue` (the pump's serial queue). A reference type so
/// the timer closure mutates it without capturing `self` (which is @MainActor). `targetFrames`
/// is the prebuffer depth: the pump fills this many frames before it starts releasing, so the
/// tap's bursty delivery (~2 frames at once) can't drain it to empty between bursts. It persists
/// across rebuilds and self-heals upward (capped) on an underrun, so it tunes to whatever IO
/// buffer size the active route/VPIO actually uses without a hard-coded guess.
private final class PumpState {
var primed = false
var targetFrames = 3 // ~60 ms initial cushion; grows on underrun up to maxTargetFrames
static let maxTargetFrames = 6 // ~120 ms cap bounds added latency
}
private let pumpState = PumpState()
// 48 kHz stereo Float32 (deinterleaved) the format the source node renders. The core
// delivers 48 kHz stereo int16 via the mixed-output sink; mainMixerNode adapts to the route.
private let outFormat = AVAudioFormat( private let outFormat = AVAudioFormat(
commonFormat: .pcmFormatFloat32, sampleRate: 48000, channels: 2, interleaved: false)! commonFormat: .pcmFormatFloat32, sampleRate: 48000, channels: 2, interleaved: false)!
@@ -121,33 +165,189 @@ final class IOSVoiceProcessingEngine {
private let renderScratchFrames = 8192 private let renderScratchFrames = 8192
private let renderScratch: UnsafeMutablePointer<Int16> private let renderScratch: UnsafeMutablePointer<Int16>
// Mic-feed converter (input-node format 48 kHz int16) and its target buffer. Owned here so
// the (background) tap block reuses them instead of allocating per callback.
private var micConverter: AVAudioConverter?
private var micTargetFormat: AVAudioFormat?
private init() { private init() {
renderScratch = UnsafeMutablePointer<Int16>.allocate(capacity: renderScratchFrames * 2) renderScratch = UnsafeMutablePointer<Int16>.allocate(capacity: renderScratchFrames * 2)
renderScratch.initialize(repeating: 0, count: renderScratchFrames * 2) renderScratch.initialize(repeating: 0, count: renderScratchFrames * 2)
micDrainScratch = UnsafeMutablePointer<Int16>.allocate(capacity: 960 * 2)
micDrainScratch.initialize(repeating: 0, count: 960 * 2)
// AVAudioEngine stops itself on a mid-session route/configuration change (it stops
// if its I/O graph no longer matches the active route). Our route-change handler in
// AudioSessionManager normally rebuilds us before the user notices, but if the engine
// stops itself AFTER our recovery (because the route-change notification raced ahead
// of the engine's own self-stop), nothing restarts it. Catch that case here.
NotificationCenter.default.addObserver(
self, selector: #selector(handleEngineConfigurationChange),
name: .AVAudioEngineConfigurationChange, object: engine)
} }
/// Start the VPIO engine for an active mic stream. The caller must have already enabled /// The engine stopped itself because its configuration no longer matches the active AVAudio
/// external playback on the core (`client.setExternalPlayback(true)` + `audioRestart()`) and /// route (this fires after a route change that the route-change handler can't always outrun).
/// started the MIC stream with `externalFeed: true`. /// Dispatch to main and call the unified `recoverAudio()` it's intent-gated on
func start(client: VoiceCatClient, micStreamId: UInt32, captureChannels: UInt32) { /// `isConnected`, idempotent, and no-ops if the engine is already running (the common case
guard !isRunning else { return } /// where our route-change handler got there first).
@objc private func handleEngineConfigurationChange(_ notification: Notification) {
Task { @MainActor [weak self] in
guard let self else { return }
guard self.isConnected, !self.engine.isRunning else { return }
logger.info("engine configuration-change — engine stopped itself, recovering")
AudioSessionManager.shared.recoverAudio()
}
}
// MARK: - Lifecycle
/// Begin playback-only (listening) operation. Called once at connect, after
/// `client.setExternalPlayback(true)` and `AudioSessionManager.ensureSessionActive()`. Attaches
/// the source node, wires the core's mixed-output sink into the ring, and starts the engine so
/// remote audio plays immediately.
func startListening(client: VoiceCatClient) {
self.client = client self.client = client
self.micStreamId = micStreamId guard !isConnected else { return }
isConnected = true
ring.reset() ring.reset()
// Enable the voice-processing I/O unit (AEC/NS/AGC) on the shared input+output unit. // Wire the core's mixed-output sink into the ring (C function pointer, no captures). Stays
// registered for the whole connection; the ring is drained by the source-node render block.
let ringPtr = Unmanaged.passUnretained(self.ring).toOpaque()
client.setMixedOutputSink({ user, pcm, spc, ch, _ in
guard let user, let pcm else { return }
let ring = Unmanaged<PCMRing>.fromOpaque(user).takeUnretainedValue()
ring.write(pcm, count: spc * Int(ch))
}, user: ringPtr)
rebuild()
}
/// Tear down the engine and unhook the core sink. Called on disconnect.
func stop() {
guard isConnected else { return }
micActive = false
stopMicTimer()
isConnected = false
client?.setMixedOutputSink(nil, user: nil)
engine.inputNode.removeTap(onBus: 0)
if engine.isRunning { engine.stop() }
logger.info("audio engine stopped — ring written=\(self.ring.debugTotalWritten) read=\(self.ring.debugTotalRead) samples")
try? engine.inputNode.setVoiceProcessingEnabled(false)
if let src = sourceNode {
engine.detach(src)
sourceNode = nil
}
ring.reset()
micRing.reset()
client = nil
}
// MARK: - Mic transitions
/// Engage the mic: install the input tap and (if the preset wants it) VPIO. Called when the
/// user joins voice, after the MIC stream (external_feed) is started.
func startMic(streamId: UInt32, channels: UInt32) {
micStreamId = streamId
captureChannels = channels
micActive = true
rebuild()
}
/// Disengage the mic: remove the tap and VPIO, keep playback running for remaining remote audio.
func stopMic() {
guard micActive else { return }
micActive = false
rebuild()
}
/// Update the capture channel count (monostereo) for the active mic and rebuild.
func setCaptureChannels(_ channels: UInt32) {
captureChannels = channels
if let client, micStreamId != 0 {
client.setCaptureChannels(streamId: micStreamId, channels: channels)
}
if micActive { rebuild() }
}
/// Re-apply the engine graph against the current AVAudioSession config (preset / route change).
/// Safe to call when only listening it just rebuilds the playback graph against the new route.
func reconfigure() {
guard isConnected else { return }
rebuild()
}
// MARK: - Graph (re)build
/// The single place that (re)builds and starts the engine graph. Deterministic: stop set
/// VPIO (re)install the mic tap start. The caller is responsible for having applied the
/// AVAudioSession config (category/mode/route) first (`IOSAudioRouter.applyConfiguration`).
private func rebuild() {
guard isConnected else { return }
// Stop the feed pump before touching the tap / ring so the timer (on micQueue) can't race
// the ring reset in installMicTap. It is restarted at the end with the current channel count.
stopMicTimer()
if engine.isRunning { engine.stop() }
engine.inputNode.removeTap(onBus: 0)
let useVPIO = micActive && IOSAudioRouter.shared.currentConfigUsesVoiceProcessing
do { do {
try engine.inputNode.setVoiceProcessingEnabled(true) try engine.inputNode.setVoiceProcessingEnabled(useVPIO)
} catch { } catch {
logger.error("setVoiceProcessingEnabled failed: \(error.localizedDescription) — AEC unavailable") logger.error("setVoiceProcessingEnabled(\(useVPIO)) failed: \(error.localizedDescription)")
}
if useVPIO {
// AGC is the one VPIO sub-stage iOS exposes; AEC+NS are bundled into the master switch.
engine.inputNode.isVoiceProcessingAGCEnabled = IOSAudioRouter.shared.agcEnabled
} }
// Playback: source node pulls mixed PCM from the ring through the VPIO output. // The source node must bind to the selected voice-processing output unit.
rebuildSourceNode()
if micActive { installMicTap() }
engine.prepare()
do {
try engine.start()
let inFmt = engine.inputNode.outputFormat(forBus: 0)
let outFmt = engine.outputNode.outputFormat(forBus: 0)
let route = AVAudioSession.sharedInstance().currentRoute.outputs
.map { "\($0.portName)[\($0.portType.rawValue)]" }.joined(separator: ", ")
logger.info("""
engine started — mic=\(self.micActive) vpio=\(useVPIO) captureCh=\(self.captureChannels) \
inFormat=\(inFmt) outputNode=\(outFmt) outputRoute=[\(route)]
""")
} catch {
// Route changes can leave AVAudioSession inactive; retry once after reactivation.
logger.error("engine start failed: \(error.localizedDescription) — attempting one-shot recovery")
do {
try AudioSessionManager.shared.ensureSessionActive()
} catch {
logger.error("recovery — session re-activate failed: \(error.localizedDescription)")
}
IOSAudioRouter.shared.applyConfiguration()
if AudioSessionManager.shared.isActive {
IOSAudioRouter.shared.applyA2dpSpeakerFallback()
}
do {
try engine.start()
logger.info("engine start succeeded after one-shot recovery")
} catch {
logger.error("engine start failed after recovery: \(error.localizedDescription)")
// Not fatal a subsequent route-change or AVAudioEngine configuration-change
// notification will trigger recoverAudio() and re-attempt the rebuild.
}
}
// Start the feed pump last, with the current channel count, so it never carries a stale
// (frozen) channel count across a monostereo switch.
if micActive { startMicTimer() }
}
/// Detach any previous source node and attach a fresh one pulling mixed PCM from the ring.
/// Rebuilt on every graph rebuild so it always connects against the current output unit (the
/// VPIO state can change the output between rebuilds). Its format is route-independent
/// `mainMixerNode` adapts 48 kHz stereo to whatever the output route is.
private func rebuildSourceNode() {
if let old = sourceNode {
engine.detach(old)
sourceNode = nil
}
let ring = self.ring let ring = self.ring
let scratch = self.renderScratch let scratch = self.renderScratch
let scratchFrames = self.renderScratchFrames let scratchFrames = self.renderScratchFrames
@@ -156,17 +356,11 @@ final class IOSVoiceProcessingEngine {
let abl = UnsafeMutableAudioBufferListPointer(ablPtr) let abl = UnsafeMutableAudioBufferListPointer(ablPtr)
let n = min(frames, scratchFrames) let n = min(frames, scratchFrames)
let got = ring.read(into: scratch, count: n * 2) / 2 // interleaved stereo frames let got = ring.read(into: scratch, count: n * 2) / 2 // interleaved stereo frames
// Deinterleave int16 Float32 per channel; silence-fill any underrun tail.
let scale: Float = 1.0 / 32768.0 let scale: Float = 1.0 / 32768.0
for ch in 0..<abl.count { for ch in 0..<abl.count {
guard let base = abl[ch].mData?.assumingMemoryBound(to: Float.self) else { continue } guard let base = abl[ch].mData?.assumingMemoryBound(to: Float.self) else { continue }
for i in 0..<frames { for i in 0..<frames {
if i < got { base[i] = i < got ? Float(scratch[i * 2 + min(ch, 1)]) * scale : 0
let s = scratch[i * 2 + min(ch, 1)]
base[i] = Float(s) * scale
} else {
base[i] = 0
}
} }
} }
return noErr return noErr
@@ -174,29 +368,39 @@ final class IOSVoiceProcessingEngine {
sourceNode = src sourceNode = src
engine.attach(src) engine.attach(src)
engine.connect(src, to: engine.mainMixerNode, format: outFormat) engine.connect(src, to: engine.mainMixerNode, format: outFormat)
}
// Mic: tap the VPIO input node, convert to 48 kHz int16, feed the core. /// Install the mic tap: convert the input node's native format to 48 kHz int16 (mono or
/// stereo per `captureChannels`) and write it to the pacing ring. The 20 ms feed pump
/// (`startMicTimer`) releases steady 960-sample frames to `feedPcm` see the mic-feed comment
/// above for why the tap must NOT call feedPcm directly (it bursts packets receiver flutter).
/// Rebuilds the converter each time because the input format depends on the VPIO state and the
/// active route.
private func installMicTap() {
guard client != nil else { return }
// Fresh ring on every (re)install a rebuild must not feed stale pre-roll into the new tap.
// Safe here: the feed pump was stopped at the top of rebuild(), so no consumer is running.
micRing.reset()
let ring = micRing // captured by the closure as a `let` no self capture (see mic-feed comment)
let inFormat = engine.inputNode.outputFormat(forBus: 0) let inFormat = engine.inputNode.outputFormat(forBus: 0)
guard inFormat.sampleRate > 0 else {
logger.error("input format unavailable (\(inFormat)) — mic will not transmit")
return
}
let targetCh = max(1, min(2, captureChannels)) let targetCh = max(1, min(2, captureChannels))
let target = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 48000, guard let target = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 48000,
channels: AVAudioChannelCount(targetCh), interleaved: true) channels: AVAudioChannelCount(targetCh), interleaved: true),
micTargetFormat = target let converter = AVAudioConverter(from: inFormat, to: target) else {
micConverter = (target != nil && inFormat.sampleRate > 0) logger.error("mic converter unavailable (in=\(inFormat), ch=\(targetCh)) — mic will not transmit")
? AVAudioConverter(from: inFormat, to: target!) : nil return
if micConverter == nil {
logger.error("mic converter unavailable (in=\(inFormat)) — mic will not transmit")
} }
let c = client let chInt = Int(targetCh)
let sid = micStreamId
let converter = micConverter
let tgt = micTargetFormat
engine.inputNode.installTap(onBus: 0, bufferSize: 960, format: inFormat) { buffer, _ in engine.inputNode.installTap(onBus: 0, bufferSize: 960, format: inFormat) { buffer, _ in
guard let converter, let tgt else { return }
// Convert this tap buffer to 48 kHz int16. Output capacity scaled for any upsample. // Convert this tap buffer to 48 kHz int16. Output capacity scaled for any upsample.
let ratio = tgt.sampleRate / buffer.format.sampleRate let ratio = target.sampleRate / buffer.format.sampleRate
let outCap = AVAudioFrameCount(Double(buffer.frameLength) * ratio + 16) let outCap = AVAudioFrameCount(Double(buffer.frameLength) * ratio + 16)
guard let outBuf = AVAudioPCMBuffer(pcmFormat: tgt, frameCapacity: outCap) else { return } guard let outBuf = AVAudioPCMBuffer(pcmFormat: target, frameCapacity: outCap) else { return }
var fed = false var fed = false
let status = converter.convert(to: outBuf, error: nil) { _, outStatus in let status = converter.convert(to: outBuf, error: nil) { _, outStatus in
if fed { outStatus.pointee = .noDataNow; return nil } if fed { outStatus.pointee = .noDataNow; return nil }
@@ -206,64 +410,66 @@ final class IOSVoiceProcessingEngine {
} }
guard status != .error, outBuf.frameLength > 0, guard status != .error, outBuf.frameLength > 0,
let chData = outBuf.int16ChannelData else { return } let chData = outBuf.int16ChannelData else { return }
let spc = Int(outBuf.frameLength) // int16 interleaved channelData[0] is the interleaved buffer. Write the converter's
// int16 interleaved channelData[0] is the interleaved buffer for interleaved formats. // variable-length output to the pacing ring; the 20 ms feed pump releases steady
c.feedPcm(streamId: sid, pcm: chData[0], samplesPerChannel: spc, channels: targetCh) // 960-sample frames to feedPcm so packets leave the core at a steady 20 ms cadence.
} ring.write(chData[0], count: Int(outBuf.frameLength) * chInt)
// Wire the core's mixed-output sink into the ring (C function pointer, no captures).
let ringPtr = Unmanaged.passUnretained(self.ring).toOpaque()
client.setMixedOutputSink({ user, pcm, spc, ch, _ in
guard let user, let pcm else { return }
let ring = Unmanaged<PCMRing>.fromOpaque(user).takeUnretainedValue()
ring.write(pcm, count: spc * Int(ch))
}, user: ringPtr)
engine.prepare()
do {
try engine.start()
isRunning = true
// Diagnostics: capture the negotiated graph formats and the live output route so a
// silent-playback report can be triaged (format/rate mismatch vs. routing vs. the
// core not delivering PCM see the ring stats logged in teardown()).
let outFmt = engine.outputNode.outputFormat(forBus: 0)
let mixFmt = engine.mainMixerNode.outputFormat(forBus: 0)
let route = AVAudioSession.sharedInstance().currentRoute.outputs
.map { "\($0.portName)[\($0.portType.rawValue)]" }.joined(separator: ", ")
logger.info("""
VPIO engine started — inFormat=\(inFormat), captureCh=\(targetCh), \
outputNode=\(outFmt), mainMixer=\(mixFmt), outputRoute=[\(route)]
""")
} catch {
logger.error("VPIO engine start failed: \(error.localizedDescription)")
teardown()
} }
} }
/// Stop the VPIO engine. The caller is responsible for restoring the core's hardware playback // MARK: - Mic feed pump (paces feedPcm at a steady 20 ms cadence)
/// afterwards (`client.setExternalPlayback(false)` + `audioRestart()`).
func stop() { /// Start the 20 ms feed pump. After priming a small cushion (`pumpState.targetFrames`), it
guard isRunning else { return } /// releases ONE 960-sample frame per tick from `micRing` to `feedPcm`, so the core (which sends
teardown() /// synchronously per captured frame) emits packets at a steady 20 ms the cadence its receivers
logger.info("VPIO engine stopped") /// expect. The cushion is essential: the receiver's playout deliberately keeps near-zero
/// buffering (low latency), so it tolerates a steady stream but not bursts; the iOS tap delivers
/// ~2 frames at once, and without the cushion the pump runs at ~0 depth and underruns on every
/// tap/timer phase beat (crackle). Recreated on every `rebuild()` so `ch` always reflects the
/// current `captureChannels` (monostereo switches). Captures only locals + the reference-type
/// ring/client/state (no `self`, which is @MainActor).
private func startMicTimer() {
stopMicTimer()
guard let client else { return }
let ring = micRing
let scratch = micDrainScratch
let state = pumpState
let sid = micStreamId
let ch = max(1, min(2, Int(captureChannels)))
let frameSamples = Self.micFrameSamplesPerChannel
let full = frameSamples * ch
let chU32 = UInt32(ch)
// The ring was just reset in installMicTap, so the cushion must be refilled before sending.
state.primed = false
let feed: () -> Void = {
_ = ring.read(into: scratch, count: full) // caller guarantees a full frame is present
_ = client.feedPcm(streamId: sid, pcm: scratch,
samplesPerChannel: frameSamples, channels: chU32)
}
let t = DispatchSource.makeTimerSource(queue: micQueue)
t.schedule(deadline: .now(), repeating: .milliseconds(20), leeway: .milliseconds(2))
t.setEventHandler {
let frames = ring.availableSamples / full // whole frames currently buffered
if !state.primed {
if frames < state.targetFrames { return } // still filling the cushion (into silence)
state.primed = true
} else if frames == 0 {
// Re-prime with a larger cushion; consuming a partial frame would lose samples.
if state.targetFrames < PumpState.maxTargetFrames { state.targetFrames += 1 }
state.primed = false
return
}
feed() // one steady frame per tick (frames >= 1 here)
// Catch-up: if the backlog grew past the cushion (pump descheduled, or producer ran
// ahead via a burst), release one extra frame to drain it and keep latency bounded.
if frames - 1 > state.targetFrames + 1 { feed() }
}
t.resume()
micTimer = t
} }
private func teardown() { private func stopMicTimer() {
client?.setMixedOutputSink(nil, user: nil) micTimer?.cancel()
engine.inputNode.removeTap(onBus: 0) micTimer = nil
if engine.isRunning { engine.stop() }
// Diagnostics (threads now quiesced): how much mixed PCM the core delivered into the ring
// vs. how much the render thread consumed. written==0 the core never delivered (Bug 1
// core/lifecycle path); written>0 with no audible output the AVAudioEngine output graph.
logger.info("VPIO ring stats — written=\(self.ring.debugTotalWritten) read=\(self.ring.debugTotalRead) samples")
try? engine.inputNode.setVoiceProcessingEnabled(false)
if let src = sourceNode {
engine.detach(src)
sourceNode = nil
}
micConverter = nil
micTargetFormat = nil
ring.reset()
isRunning = false
} }
} }

View File

@@ -26,8 +26,8 @@
<array> <array>
<string>audio</string> <string>audio</string>
</array> </array>
<key>UILaunchStoryboardName</key> <key>UILaunchScreen</key>
<string>LaunchScreen</string> <dict/>
<key>UIRequiresFullScreen</key> <key>UIRequiresFullScreen</key>
<false/> <false/>
<key>UISupportedInterfaceOrientations</key> <key>UISupportedInterfaceOrientations</key>

View File

@@ -20,12 +20,15 @@ struct ActivityEntry: Identifiable {
struct VoiceState { struct VoiceState {
var micActive = false var micActive = false
var voiceSubscribed = false
var selfMuted = false var selfMuted = false
var selfDeafened = false var selfDeafened = false
var serverMuted = false var serverMuted = false
var serverDeafened = false var serverDeafened = false
var inputMode: VoiceCatInputMode = .voiceActivation var inputMode: VoiceCatInputMode = .voiceActivation
var vadThreshold: Float = 0.025 var vadThreshold: Float = 0.025
var inputGain: Float = 1.0
var inputNoiseReduction: Bool = false
var level: Float = 0.0 var level: Float = 0.0
var currentDeviceId: String? var currentDeviceId: String?
var localStreamId: UInt32 = 0 var localStreamId: UInt32 = 0
@@ -51,6 +54,30 @@ final class SessionState {
var accounts: [Account] = [] var accounts: [Account] = []
var devices: [Device] = [] var devices: [Device] = []
/// Back-reference to the app state. Once `SessionState.init` overwrites `client.onEvent`,
/// `AppState.handleConnectEvent` no longer receives per-session events so the
/// `.disconnected` event for a LIVE session arrives here in `handleEvent`, not in AppState.
/// This weak ref lets us hand the disconnect back to AppState (which owns the reconnect
/// state machine) so the auto-reconnect fires. Set by AppState on auth success.
weak var appState: AppState?
// MARK: - Reconnect restore state
//
// When the iOS client auto-reconnects after a network drop, AppState captures the prior
// session's channel + voice/mic state and asks the new SessionState (created on auth success)
// to restore it. We rejoin the channel explicitly (the server auto-placed us in Lobby on
// auth), and on the resulting `.joinResult` we re-arm voice subscription + mute/deafen. The
// drive is here, not in AppState, because once SessionState is created it owns
// `client.onEvent` and AppState no longer sees per-session events.
private struct RestoreRequest {
let channelId: UInt32
let voiceSubscribed: Bool
let micMuted: Bool
let deafened: Bool
}
private var pendingRestore: RestoreRequest?
private var didIssueRestoreJoin = false
/// Host side of iOS screen-audio sharing drains the broadcast extension's App Group ring /// Host side of iOS screen-audio sharing drains the broadcast extension's App Group ring
/// and feeds the SCREEN_AUDIO stream this session owns. See BroadcastAudioPump. /// and feeds the SCREEN_AUDIO stream this session owns. See BroadcastAudioPump.
private let broadcastPump = BroadcastAudioPump() private let broadcastPump = BroadcastAudioPump()
@@ -59,7 +86,7 @@ final class SessionState {
self.client = client self.client = client
self.selfUserId = selfUserId self.selfUserId = selfUserId
self.permissions = permissions self.permissions = permissions
AudioSessionManager.shared.client = client loadAndApplyVoiceSettings()
refreshChannels() refreshChannels()
refreshUsers() refreshUsers()
syncSelfChannel() syncSelfChannel()
@@ -73,16 +100,10 @@ final class SessionState {
broadcastPump.onBroadcastStarted = { [weak self] in self?.startScreenShare() } broadcastPump.onBroadcastStarted = { [weak self] in self?.startScreenShare() }
broadcastPump.onBroadcastFinished = { [weak self] in self?.stopScreenShare() } broadcastPump.onBroadcastFinished = { [weak self] in self?.stopScreenShare() }
broadcastPump.start() broadcastPump.start()
// When IOSAudioRouter changes the audio config, restart the voice path if needed so the
// native VPIO engine (AEC/NS/AGC) engages or disengages to match the new preset/route.
AudioSessionManager.shared.reconcileVoicePath = { [weak self] in self?.reconcileVoicePath() }
} }
deinit { deinit {
broadcastPump.stop() broadcastPump.stop()
MainActor.assumeIsolated {
AudioSessionManager.shared.client = nil
}
} }
// MARK: - Event dispatch // MARK: - Event dispatch
@@ -92,7 +113,22 @@ final class SessionState {
case .channelList: case .channelList:
refreshChannels() refreshChannels()
syncSelfChannel() syncSelfChannel()
case .userJoined, .userLeft: case .userJoined:
// ev.text = nickname, ev.channelId = the channel they joined (per voicecat.h).
if ev.userId != selfUserId && ev.channelId == currentChannelId {
EventFeedback.shared.play(.channelJoin)
EventFeedback.shared.speak("\(ev.text ?? "Someone") joined")
}
refreshUsers()
syncSelfChannel()
case .userLeft:
// Capture the leaving user's prior nickname/channel before refreshUsers() drops them.
if ev.userId != selfUserId,
let gone = users.first(where: { $0.id == ev.userId }),
gone.channelId == currentChannelId {
EventFeedback.shared.play(.channelLeave)
EventFeedback.shared.speak("\(gone.nickname) left")
}
refreshUsers() refreshUsers()
syncSelfChannel() syncSelfChannel()
case .userUpdated: case .userUpdated:
@@ -103,13 +139,27 @@ final class SessionState {
} }
case .textMessage: case .textMessage:
let sender = users.first(where: { $0.id == ev.userId })?.nickname ?? "Unknown" let sender = users.first(where: { $0.id == ev.userId })?.nickname ?? "Unknown"
let body = ev.text ?? ""
let isSelf = ev.userId == selfUserId
let isPrivate = ev.textScope == .private
messages.append(ChatMessage( messages.append(ChatMessage(
timestamp: Date(timeIntervalSince1970: Double(ev.timestampUnixMs) / 1000), timestamp: Date(timeIntervalSince1970: Double(ev.timestampUnixMs) / 1000),
senderName: sender, senderName: sender,
text: ev.text ?? "", text: body,
scope: ev.textScope)) scope: ev.textScope))
EventFeedback.shared.play(isPrivate
? (isSelf ? .pmSent : .pmRecv)
: (isSelf ? .channelSent : .channelRecv))
if !isSelf {
EventFeedback.shared.speak(isPrivate
? "Private message from \(sender): \(body)"
: "\(sender): \(body)")
}
case .talkState: case .talkState:
let talking = ev.u32a != 0 let talking = ev.u32a != 0
if ev.userId == selfUserId {
EventFeedback.shared.play(talking ? .vaStart : .vaStop)
}
let who = users.first(where: { $0.id == ev.userId })?.nickname ?? "user \(ev.userId)" let who = users.first(where: { $0.id == ev.userId })?.nickname ?? "user \(ev.userId)"
addActivity(talking ? "\(who) started talking" : "\(who) stopped talking") addActivity(talking ? "\(who) started talking" : "\(who) stopped talking")
case .streamStarted: case .streamStarted:
@@ -126,8 +176,6 @@ final class SessionState {
addActivity("Sharing screen audio (\(channels == 2 ? "stereo" : "mono"))") addActivity("Sharing screen audio (\(channels == 2 ? "stereo" : "mono"))")
break break
} }
// A remote user started a stream ensure the audio session is active so we can
// hear them even if we haven't joined voice ourselves.
if ev.userId != selfUserId { if ev.userId != selfUserId {
do { do {
try AudioSessionManager.shared.ensureSessionActive() try AudioSessionManager.shared.ensureSessionActive()
@@ -138,14 +186,48 @@ final class SessionState {
AudioSessionManager.shared.logSessionState("stream started (user \(ev.userId))") AudioSessionManager.shared.logSessionState("stream started (user \(ev.userId))")
addActivity("Stream started (user \(ev.userId))") addActivity("Stream started (user \(ev.userId))")
case .streamStopped: case .streamStopped:
addActivity("Stream stopped (user \(ev.userId))") if ev.userId == selfUserId {
if ev.streamId == voiceState.localStreamId {
voiceState.localStreamId = 0
voiceState.micActive = false
voiceState.level = 0
}
} else {
addActivity("Stream stopped (user \(ev.userId))")
}
case .voiceState:
let subscribed = ev.u32a != 0
voiceState.voiceSubscribed = subscribed
if subscribed {
doStartMicStream()
} else {
voiceState.micActive = false
voiceState.level = 0
EventFeedback.shared.play(.voiceOff)
}
case .joinResult: case .joinResult:
if ev.result == .ok { if ev.result == .ok {
currentChannelId = ev.channelId currentChannelId = ev.channelId
addActivity("Joined channel") addActivity("Joined channel")
refreshUsers() refreshUsers()
// Reconnect restore: this was our restore-join. Now that the server has
// processed the channel move, re-arm voice subscription (if the user was
// transmitting before the drop) and re-apply the local mute/deafen state.
// The server returns ok even when joining the channel we're already in, so
// this fires reliably for the Lobby-too case.
if didIssueRestoreJoin, let r = pendingRestore, r.channelId == ev.channelId {
didIssueRestoreJoin = false
completeRestore()
}
} else { } else {
addActivity("Join failed: \(ev.result.description)") addActivity("Join failed: \(ev.result.description)")
// Restore-join failed (channel was deleted, became password-protected or
// full while we were away). Give up on the voice/mute restore cleanly so we
// don't leave dangling state or attempt voice without being in a channel.
if didIssueRestoreJoin {
didIssueRestoreJoin = false
pendingRestore = nil
}
} }
case .error: case .error:
addActivity("Error: \(ev.text ?? ev.result.description)") addActivity("Error: \(ev.text ?? ev.result.description)")
@@ -155,6 +237,16 @@ final class SessionState {
} }
case .accountList: case .accountList:
accounts = client.listAccounts() accounts = client.listAccounts()
case .disconnected:
// Audible cue, then hand the disconnect back to AppState so its reconnect state
// machine fires. This is the ONLY way AppState learns a live session dropped
// after auth success, `SessionState.init` overwrites `client.onEvent`, so
// `AppState.handleConnectEvent` never sees this event. (Without this callback, a
// network drop on a live session would just play the cue and leave the session as a
// zombie the user would have to tap Disconnect manually.)
EventFeedback.shared.play(ev.result == .ok ? .logout : .connectionLost)
EventFeedback.shared.speak(ev.result == .ok ? "Disconnected" : "Connection lost")
appState?.onLiveSessionDisconnected()
default: default:
break break
} }
@@ -168,17 +260,18 @@ final class SessionState {
// MARK: - Self-channel / server-mute sync // MARK: - Self-channel / server-mute sync
/// Sync currentChannelId from the self user's channelId in the user list. Mirrors macOS /// Sync currentChannelId from the self user's channelId in the user list. Mirrors macOS
/// MainWindowController.swift:461,491,522. The server auto-places every authed user into /// MainWindowController's bootstrap/event-handling sync. The server auto-places every
/// the Lobby (channel 1) on connect, but without this sync currentChannelId stays 0 and /// authed user into the Lobby (channel 1) on connect, but without this sync
/// the mic button (gated on currentChannelId == 0) stays permanently dimmed. /// currentChannelId stays 0 and the mic button (gated on currentChannelId == 0) stays
/// permanently dimmed.
private func syncSelfChannel() { private func syncSelfChannel() {
if let me = users.first(where: { $0.id == selfUserId }) { if let me = users.first(where: { $0.id == selfUserId }) {
currentChannelId = me.channelId currentChannelId = me.channelId
} }
} }
/// Apply server-side mute/deafen state mirrors macOS MainWindowController.swift:693-700. /// Apply server-side mute/deafen state mirrors macOS MainWindowController's handling
/// iOS was previously ignoring server mute/deafen entirely. /// of UserEvent.UPDATED for the self user.
private func applyServerMuteState(muted: Bool, deafened: Bool) { private func applyServerMuteState(muted: Bool, deafened: Bool) {
if muted && !voiceState.serverMuted { addActivity("You have been server-muted") } if muted && !voiceState.serverMuted { addActivity("You have been server-muted") }
if deafened && !voiceState.serverDeafened { addActivity("You have been server-deafened") } if deafened && !voiceState.serverDeafened { addActivity("You have been server-deafened") }
@@ -205,12 +298,15 @@ final class SessionState {
currentChannelId = 0 currentChannelId = 0
} }
func startMicStream() { func joinVoice() {
AVAudioApplication.requestRecordPermission { [weak self] granted in AVAudioApplication.requestRecordPermission { [weak self] granted in
DispatchQueue.main.async { DispatchQueue.main.async {
guard let self else { return } guard let self else { return }
if granted { if granted {
self.doStartMicStream() let result = self.client.joinVoice()
if result != .ok {
self.addActivity("Failed to join voice: \(result.description)")
}
} else { } else {
self.addActivity("Microphone permission denied — grant in Settings > Privacy > Microphone") self.addActivity("Microphone permission denied — grant in Settings > Privacy > Microphone")
} }
@@ -226,94 +322,72 @@ final class SessionState {
return return
} }
// VPIO path: on the AEC presets, the native AVAudioEngine does AEC/NS/AGC and the core
// runs in external mode (no hardware mic/playback). The mic stream is started with
// externalFeed so the core skips the hardware capture device; setExternalPlayback makes
// it skip the hardware playback device and deliver the mix to IOSVoiceProcessingEngine.
//
// ORDER MATTERS: set the external-playback flag now, but defer audioRestart() until
// AFTER startStream (below) so the MIC LocalStream which carries external_feed=true
// already exists when ensure_audio_running() derives external_capture. Restarting before
// the stream exists makes the core reopen a hardware capture device that is never dropped
// (the announce-result restart early-returns because the engine is already running); that
// lingering miniaudio capture unit then fights the AVAudioEngine VPIO unit on the same
// .voiceChat session and silences VPIO playback.
let useVPIO = IOSAudioRouter.shared.currentConfigUsesVoiceProcessing
client.setExternalPlayback(useVPIO)
let desc = StreamDescriptor(kind: .mic, deviceId: voiceState.currentDeviceId, label: "Mic", let desc = StreamDescriptor(kind: .mic, deviceId: voiceState.currentDeviceId, label: "Mic",
externalFeed: useVPIO) externalFeed: true)
let (result, streamId) = client.startStream(desc) let (result, streamId) = client.startStream(desc)
if result == .ok { guard result == .ok else {
voiceState.micActive = true
voiceState.localStreamId = streamId
// Publish the active mic stream ID so IOSAudioRouter can reset the core's capture
// channel count when the user switches monostereo (selectCaptureChannels /
// applyPreset). Without this, switching stereomono leaves the LocalStream's
// capture_channels field at 2 and the next engine start still opens stereo.
AudioSessionManager.shared.activeMicStreamId = streamId
// Store the user's capture channel selection before the server acknowledges
// the stream. The engine hasn't started yet at this point (it starts when
// handle_stream_announce_result fires), so vc_set_capture_channels just
// stores the value no restart. ensure_audio_running() picks it up when
// the stream is confirmed and opens the device with the right channel count.
let channels = IOSAudioRouter.shared.captureChannels.channelCount
if channels != 1 {
client.setCaptureChannels(streamId: streamId, channels: channels)
}
if useVPIO {
// The external-feed MIC stream now exists, so restart the core into full
// external mode (no hardware capture/playback, mixer-timer only) mic and
// speaker are owned entirely by the VPIO engine, which we start right after.
client.audioRestart()
IOSVoiceProcessingEngine.shared.start(
client: client, micStreamId: streamId, captureChannels: channels)
}
} else {
addActivity("Failed to start mic: \(result.description)") addActivity("Failed to start mic: \(result.description)")
if useVPIO { // revert external-playback mode so remote audio still plays return
client.setExternalPlayback(false) }
client.audioRestart() voiceState.micActive = true
} voiceState.localStreamId = streamId
EventFeedback.shared.play(.voiceOn)
let channels = IOSAudioRouter.shared.captureChannels.channelCount
if channels != 1 {
client.setCaptureChannels(streamId: streamId, channels: channels)
}
IOSAudioEngine.shared.startMic(streamId: streamId, channels: channels)
}
func leaveVoice() {
if voiceState.screenStreamId != 0 { stopScreenShare() }
client.setPushToTalk(false)
IOSAudioEngine.shared.stopMic()
client.leaveVoice()
}
// MARK: - Reconnect restore
/// Called by `AppState` after a reconnect's auth success to rejoin the prior channel and
/// re-enable the prior voice/mic state. Drives the restore through the `.joinResult` event
/// so we re-arm voice only AFTER the server processed the join joining voice before the
/// channel move would be rejected server-side. `micMuted`/`deafened` are the user's LOCAL
/// mute/deafen state at the moment of the drop; the server resets those on a fresh auth, so
/// we re-push them via `setMute` after the channel is restored.
func requestRestore(channelId: UInt32, voiceSubscribed: Bool,
micMuted: Bool, deafened: Bool) {
pendingRestore = RestoreRequest(channelId: channelId,
voiceSubscribed: voiceSubscribed,
micMuted: micMuted,
deafened: deafened)
didIssueRestoreJoin = false
if channelId != 0 {
// The server auto-placed us in the Lobby on auth; join our prior channel explicitly.
// `vc_join_channel` is idempotent server-side (joining the channel you're already in
// returns ok), so this is safe even if the prior channel was the Lobby.
client.joinChannel(channelId)
didIssueRestoreJoin = true
} else {
// No prior channel go straight to the voice/mute restore. (voiceSubscribed with
// channelId == 0 is contradictory; `completeRestore` further guards on
// currentChannelId != 0 before subscribing to voice.)
completeRestore()
} }
} }
func stopMicStream() { /// Finish the restore after the channel is in place (or there was no channel to restore):
// Tear down the VPIO engine first (removes the mic tap + unregisters the mixed sink), /// re-subscribe to voice if the user was transmitting, and re-apply the local mute/deafen
// then stop the mic stream, then restore the core's hardware playback for any remaining /// state. Safe to call once per `pendingRestore`; clears it.
// remote audio. Order matters: the mic stream must be gone before audioRestart so the private func completeRestore() {
// core opens a normal playback device (and no capture device there's no mic stream). guard let r = pendingRestore else { return }
let wasVPIO = IOSVoiceProcessingEngine.shared.isRunning if r.voiceSubscribed && currentChannelId != 0 {
if wasVPIO { joinVoice()
IOSVoiceProcessingEngine.shared.stop()
} }
if voiceState.localStreamId != 0 { setMute(r.micMuted, deafened: r.deafened)
client.stopStream(voiceState.localStreamId) addActivity("Restored to channel \(currentChannelId)"
voiceState.localStreamId = 0 + (r.voiceSubscribed ? " with voice" : ""))
AudioSessionManager.shared.activeMicStreamId = nil pendingRestore = nil
}
if wasVPIO {
client.setExternalPlayback(false)
client.audioRestart() // reopen hardware playback (no mic stream no hw capture)
}
voiceState.micActive = false
voiceState.level = 0
// Do NOT deactivate the AVAudioSession here the user may still want to hear
// remote audio (other people talking). The session is deactivated only when
// disconnecting from the server (see AppState.disconnect / .disconnected event).
}
/// Restart the voice path when the audio config changes mid-call (driven by IOSAudioRouter).
/// If VPIO is involved on either the current or desired side, restart the mic so the native
/// voice-processing engine engages/disengages and re-binds to the new route. Pure miniaudio
/// config tweaks need no restart the core's own audioRestart (already issued) handles them.
private func reconcileVoicePath() {
guard voiceState.micActive else { return }
let want = IOSAudioRouter.shared.currentConfigUsesVoiceProcessing
let have = IOSVoiceProcessingEngine.shared.isRunning
guard want || have else { return }
stopMicStream()
doStartMicStream()
} }
// MARK: - Screen audio share // MARK: - Screen audio share
@@ -360,15 +434,65 @@ final class SessionState {
func setInputMode(_ mode: VoiceCatInputMode) { func setInputMode(_ mode: VoiceCatInputMode) {
client.setInputMode(mode) client.setInputMode(mode)
voiceState.inputMode = mode voiceState.inputMode = mode
UserDefaults.standard.set(Int(mode.rawValue), forKey: DefaultsKey.inputMode)
} }
func setVadThreshold(_ threshold: Float) { func setVadThreshold(_ threshold: Float) {
client.setVadThreshold(threshold) client.setVadThreshold(threshold)
voiceState.vadThreshold = threshold voiceState.vadThreshold = threshold
UserDefaults.standard.set(threshold, forKey: DefaultsKey.vadThreshold)
} }
func setInputGain(_ gain: Float) {
client.setInputGain(gain)
voiceState.inputGain = gain
UserDefaults.standard.set(gain, forKey: DefaultsKey.inputGain)
}
func setInputNoiseReduction(_ on: Bool) {
client.setInputNoiseReduction(on)
voiceState.inputNoiseReduction = on
UserDefaults.standard.set(on, forKey: DefaultsKey.inputNoiseReduction)
}
// MARK: - Persisted input settings
private enum DefaultsKey {
static let inputMode = "voice.inputMode"
static let vadThreshold = "voice.vadThreshold"
static let inputGain = "voice.inputGain"
static let inputNoiseReduction = "voice.inputNoiseReduction"
}
/// Restore the saved input mode / VAD threshold / mic gain and push them into the core so a
/// relaunch keeps the user's transmission settings instead of resetting to VAD defaults.
private func loadAndApplyVoiceSettings() {
let d = UserDefaults.standard
if d.object(forKey: DefaultsKey.inputMode) != nil {
let raw = UInt32(d.integer(forKey: DefaultsKey.inputMode))
voiceState.inputMode = VoiceCatInputMode(rawValue: raw) ?? .voiceActivation
}
if d.object(forKey: DefaultsKey.vadThreshold) != nil {
voiceState.vadThreshold = d.float(forKey: DefaultsKey.vadThreshold)
}
if d.object(forKey: DefaultsKey.inputGain) != nil {
voiceState.inputGain = d.float(forKey: DefaultsKey.inputGain)
}
if d.object(forKey: DefaultsKey.inputNoiseReduction) != nil {
voiceState.inputNoiseReduction = d.bool(forKey: DefaultsKey.inputNoiseReduction)
}
client.setInputMode(voiceState.inputMode)
client.setVadThreshold(voiceState.vadThreshold)
client.setInputGain(voiceState.inputGain)
client.setInputNoiseReduction(voiceState.inputNoiseReduction)
}
private var pttEngaged = false
func setPushToTalk(_ active: Bool) { func setPushToTalk(_ active: Bool) {
client.setPushToTalk(active) client.setPushToTalk(active)
// Play the PTT cue only on the press transition (the gesture fires repeatedly while held).
if active && !pttEngaged { EventFeedback.shared.play(.ptt) }
pttEngaged = active
} }
// MARK: - Text // MARK: - Text

View File

@@ -15,8 +15,7 @@ struct ChannelEditView: View {
@State private var maxUsers = "0" @State private var maxUsers = "0"
@State private var sortOrder = "0" @State private var sortOrder = "0"
// Audio (Opus). Note: the channel list does not carry the current audio config, so when // Audio (Opus) populated from the channel's current config when editing.
// editing an existing channel these start from the codec defaults (same as macOS/Windows).
@State private var stereo = false @State private var stereo = false
@State private var bitrate = "64000" @State private var bitrate = "64000"
@State private var sampleRate = "48000" @State private var sampleRate = "48000"
@@ -122,6 +121,17 @@ struct ChannelEditView: View {
parentId = ch.parentId parentId = ch.parentId
passwordProtected = ch.passwordProtected passwordProtected = ch.passwordProtected
maxUsers = "\(ch.maxUsers)" maxUsers = "\(ch.maxUsers)"
sortOrder = "\(ch.sortOrder)"
stereo = ch.audio.stereo
bitrate = "\(ch.audio.bitrateBps)"
sampleRate = "\(ch.audio.sampleRate)"
frameMs = ch.audio.frameMs
application = ch.audio.application
packetLoss = "\(ch.audio.expectedPacketLoss)"
complexity = Int(ch.audio.complexity)
fec = ch.audio.fec
dtx = ch.audio.dtx
dred = ch.audio.dred
} }
private func save() { private func save() {

View File

@@ -91,7 +91,7 @@ struct ChatView: View {
private func sendMessage() { private func sendMessage() {
let text = composeText.trimmingCharacters(in: .whitespacesAndNewlines) let text = composeText.trimmingCharacters(in: .whitespacesAndNewlines)
guard !text.isEmpty else { return } guard !text.isEmpty else { return }
session.sendText(text, scope: .channel) session.sendText(text, scope: .channel, targetId: session.currentChannelId)
composeText = "" composeText = ""
} }
} }

View File

@@ -20,7 +20,7 @@ struct PerUserTuningView: View {
Section("Volume") { Section("Volume") {
HStack { HStack {
Text("Gain") Text("Gain")
Slider(value: $gain, in: 0...2, step: 0.05) { _ in Slider(value: $gain, in: 0...4, step: 0.05) { _ in
applyToAllStreams() applyToAllStreams()
} }
.accessibilityLabel("Volume gain for \(user.nickname)") .accessibilityLabel("Volume gain for \(user.nickname)")
@@ -48,7 +48,6 @@ struct PerUserTuningView: View {
} }
} }
.onAppear { .onAppear {
// Load from first stream if available
let streams = streamsForUser let streams = streamsForUser
if let first = streams.first { if let first = streams.first {
let (_, state) = session.client.getRemoteStream(userId: user.id, streamId: first.id) let (_, state) = session.client.getRemoteStream(userId: user.id, streamId: first.id)

View File

@@ -8,6 +8,14 @@ struct SettingsView: View {
@StateObject private var router = IOSAudioRouter.shared @StateObject private var router = IOSAudioRouter.shared
@State private var showAdvanced = false @State private var showAdvanced = false
// Notification feedback prefs keys shared with VoiceCatCore's FeedbackSettings, so the
// EventFeedback player reads the same values these toggles write.
@AppStorage("feedback.sounds") private var soundsEnabled = true
@AppStorage("feedback.speech") private var speechEnabled = false
@AppStorage("feedback.volume") private var soundsVolume = 1.0
@AppStorage("feedback.selfTalk") private var selfTalkEnabled = false
@AppStorage("feedback.ptt") private var pttSoundEnabled = false
var body: some View { var body: some View {
NavigationStack { NavigationStack {
Form { Form {
@@ -30,9 +38,9 @@ struct SettingsView: View {
.accessibilityLabel("Speaker output") .accessibilityLabel("Speaker output")
.accessibilityHint("Routes audio to the speaker instead of the earpiece when no headphones are connected.") .accessibilityHint("Routes audio to the speaker instead of the earpiece when no headphones are connected.")
// Surface the voice-processing state. On the AEC presets the native iOS // Surface the voice-processing state. On Voice Chat the native iOS
// Voice-Processing unit (VPIO) does echo cancellation, noise suppression and // Voice-Processing unit (VPIO) does echo cancellation, noise suppression and
// automatic gain control; the other presets (stereo/studio/A2DP) can't use it. // automatic gain control; the stereo / mono-mic / A2DP configs can't use it.
if router.currentConfigUsesVoiceProcessing { if router.currentConfigUsesVoiceProcessing {
Label("Echo cancellation & noise suppression on (iOS voice processing)", Label("Echo cancellation & noise suppression on (iOS voice processing)",
systemImage: "waveform.badge.mic") systemImage: "waveform.badge.mic")
@@ -40,18 +48,11 @@ struct SettingsView: View {
.foregroundStyle(.secondary) .foregroundStyle(.secondary)
.accessibilityLabel("Echo cancellation and noise suppression are on") .accessibilityLabel("Echo cancellation and noise suppression are on")
} else { } else {
Label("No echo cancellation in this preset (stereo / studio / A2DP)", Label("No echo cancellation in this configuration (stereo / A2DP / off)",
systemImage: "waveform.slash") systemImage: "waveform.slash")
.font(.caption) .font(.caption)
.foregroundStyle(.secondary) .foregroundStyle(.secondary)
.accessibilityLabel("Echo cancellation is off in this preset") .accessibilityLabel("Echo cancellation is off in this configuration")
}
if !router.hasBluetoothDevice && !router.hasWiredHeadset {
Text("Connect Bluetooth headphones or a wired headset for more presets.")
.font(.caption)
.foregroundStyle(.secondary)
.accessibilityLabel("No external audio device connected")
} }
} }
@@ -119,6 +120,27 @@ struct SettingsView: View {
} }
.accessibilityLabel("Microphone processing mode") .accessibilityLabel("Microphone processing mode")
// Voice-processing (VPIO) controls only meaningful on a VPIO-capable
// config (mono + standard + non-A2DP). iOS bundles echo cancellation and
// noise suppression into one master switch (no per-stage toggle); AGC is
// the one sub-stage it lets us control independently.
if router.voiceProcessingAvailable {
Toggle("Voice Processing (AEC + noise suppression)", isOn: Binding(
get: { router.voiceProcessingEnabled },
set: { router.setVoiceProcessingEnabled($0) }
))
.accessibilityLabel("Voice processing")
.accessibilityHint("Echo cancellation and noise suppression, bundled together by iOS.")
if router.voiceProcessingEnabled {
Toggle("Automatic Gain Control", isOn: Binding(
get: { router.agcEnabled },
set: { router.setAgcEnabled($0) }
))
.accessibilityLabel("Automatic gain control")
}
}
if router.showsRawModeSpeakerWarning { if router.showsRawModeSpeakerWarning {
Label( Label(
"Raw mode on speaker — echo risk (no AEC)", "Raw mode on speaker — echo risk (no AEC)",
@@ -215,6 +237,48 @@ struct SettingsView: View {
.accessibilityLabel("Voice activation threshold") .accessibilityLabel("Voice activation threshold")
} }
} }
VStack(alignment: .leading, spacing: 4) {
Text("Mic Volume: \(Int((session.voiceState.inputGain * 100).rounded()))%")
.font(.caption)
Slider(
value: Binding(
get: { Double(session.voiceState.inputGain) },
set: { session.setInputGain(Float($0)) }
),
in: 0...4, step: 0.05
)
.accessibilityLabel("Microphone volume")
.accessibilityValue("\(Int((session.voiceState.inputGain * 100).rounded())) percent")
}
Toggle("Noise Reduction (RNNoise)", isOn: Binding(
get: { session.voiceState.inputNoiseReduction },
set: { session.setInputNoiseReduction($0) }
))
.accessibilityLabel("Microphone noise reduction")
.accessibilityHint("Denoises your microphone signal for everyone listening.")
}
// MARK: - Notifications
Section("Notifications") {
Toggle("Event sounds", isOn: $soundsEnabled)
.accessibilityLabel("Play event sounds")
if soundsEnabled {
VStack(alignment: .leading, spacing: 4) {
Text("Sound volume")
.font(.caption)
Slider(value: $soundsVolume, in: 0...1)
.accessibilityLabel("Sound volume")
}
}
Toggle("Speak events (text-to-speech)", isOn: $speechEnabled)
.accessibilityLabel("Speak events")
.accessibilityHint("Announces joins and leaves and reads message text aloud.")
Toggle("Your own voice-activity sounds", isOn: $selfTalkEnabled)
.accessibilityLabel("Voice activity sounds")
Toggle("Push-to-talk cue", isOn: $pttSoundEnabled)
.accessibilityLabel("Push to talk cue")
} }
// MARK: - Admin // MARK: - Admin
@@ -230,7 +294,7 @@ struct SettingsView: View {
// MARK: - Server // MARK: - Server
Section("Server") { Section("Server") {
Button(role: .destructive) { Button(role: .destructive) {
session.stopMicStream() session.leaveVoice()
appState.disconnect() appState.disconnect()
} label: { } label: {
Label("Disconnect", systemImage: "phone.down") Label("Disconnect", systemImage: "phone.down")

View File

@@ -18,6 +18,10 @@ struct UserRow: View {
var body: some View { var body: some View {
UserRowView(user: user, isSelf: user.id == session.selfUserId) UserRowView(user: user, isSelf: user.id == session.selfUserId)
.contextMenu { contextMenu } .contextMenu { contextMenu }
// The context menu is long-press only, which VoiceOver doesn't surface mirror the
// same buttons as accessibility actions so VoiceOver users can reach per-user tuning
// (and the admin actions) via the actions rotor on the focused row.
.accessibilityActions { contextMenu }
.sheet(item: $activeSheet) { sheet in .sheet(item: $activeSheet) { sheet in
switch sheet { switch sheet {
case .tuning: PerUserTuningView(user: user, session: session) case .tuning: PerUserTuningView(user: user, session: session)

View File

@@ -14,9 +14,9 @@ struct VoiceControlsView: View {
} else { } else {
Button { Button {
if session.voiceState.micActive { if session.voiceState.micActive {
session.stopMicStream() session.leaveVoice()
} else { } else {
session.startMicStream() session.joinVoice()
} }
} label: { } label: {
Text(session.voiceState.micActive ? "Leave Voice" : "Join Voice") Text(session.voiceState.micActive ? "Leave Voice" : "Join Voice")
@@ -28,7 +28,7 @@ struct VoiceControlsView: View {
.foregroundStyle(session.voiceState.micActive ? .green : .accentColor) .foregroundStyle(session.voiceState.micActive ? .green : .accentColor)
} }
.disabled(session.currentChannelId == 0) .disabled(session.currentChannelId == 0)
.accessibilityLabel(session.voiceState.micActive ? "Leave Voice — stop sending microphone audio" : "Join Voice — start sending microphone audio") .accessibilityLabel(session.voiceState.micActive ? "Leave Voice" : "Join Voice")
} }
// Level meter // Level meter
@@ -81,7 +81,7 @@ struct VoiceControlsView: View {
// Disconnect // Disconnect
Button(role: .destructive) { Button(role: .destructive) {
session.stopMicStream() session.leaveVoice()
session.client.disconnect() session.client.disconnect()
} label: { } label: {
Image(systemName: "phone.down.fill") Image(systemName: "phone.down.fill")
@@ -115,13 +115,13 @@ private struct PTTButton: View {
.updating($isPressing) { _, state, _ in state = true } .updating($isPressing) { _, state, _ in state = true }
.onChanged { _ in .onChanged { _ in
if !isPressing { return } if !isPressing { return }
if !session.voiceState.micActive { session.startMicStream() } if !session.voiceState.micActive { session.joinVoice() }
session.setPushToTalk(true) session.setPushToTalk(true)
UIImpactFeedbackGenerator(style: .medium).impactOccurred() UIImpactFeedbackGenerator(style: .medium).impactOccurred()
} }
.onEnded { _ in .onEnded { _ in
session.setPushToTalk(false) session.setPushToTalk(false)
session.stopMicStream() session.leaveVoice()
} }
) )
.accessibilityLabel("Push to talk, hold to transmit") .accessibilityLabel("Push to talk, hold to transmit")

View File

@@ -4,7 +4,7 @@
<dict> <dict>
<key>com.apple.security.application-groups</key> <key>com.apple.security.application-groups</key>
<array> <array>
<string>group.cat.voice.VoiceCat</string> <string>group.me.iamtalon.voicecat</string>
</array> </array>
</dict> </dict>
</plist> </plist>

View File

@@ -31,6 +31,7 @@
AAAA00000000000000000048 /* UserPickerSheet.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA00000000000000000047 /* UserPickerSheet.swift */; }; AAAA00000000000000000048 /* UserPickerSheet.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA00000000000000000047 /* UserPickerSheet.swift */; };
AAAA0000000000000000004A /* SettingsWindowController.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA00000000000000000049 /* SettingsWindowController.swift */; }; AAAA0000000000000000004A /* SettingsWindowController.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA00000000000000000049 /* SettingsWindowController.swift */; };
AAAA0000000000000000004C /* ScreenAudioCapture.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA0000000000000000004B /* ScreenAudioCapture.swift */; }; AAAA0000000000000000004C /* ScreenAudioCapture.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA0000000000000000004B /* ScreenAudioCapture.swift */; };
AAAA00000000000000000051 /* InputDeviceCapture.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA00000000000000000050 /* InputDeviceCapture.swift */; };
AAAA0000000000000000004E /* ScreenSharePickerSheet.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA0000000000000000004F /* ScreenSharePickerSheet.swift */; }; AAAA0000000000000000004E /* ScreenSharePickerSheet.swift in Sources */ = {isa = PBXBuildFile; fileRef = AAAA0000000000000000004F /* ScreenSharePickerSheet.swift */; };
/* End PBXBuildFile section */ /* End PBXBuildFile section */
@@ -45,6 +46,7 @@
AAAA00000000000000000019 /* ConnectWindowController.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = ConnectWindowController.swift; sourceTree = "<group>"; }; AAAA00000000000000000019 /* ConnectWindowController.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = ConnectWindowController.swift; sourceTree = "<group>"; };
AAAA0000000000000000001A /* MainWindowController.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = MainWindowController.swift; sourceTree = "<group>"; }; AAAA0000000000000000001A /* MainWindowController.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = MainWindowController.swift; sourceTree = "<group>"; };
AAAA0000000000000000004B /* ScreenAudioCapture.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = ScreenAudioCapture.swift; sourceTree = "<group>"; }; AAAA0000000000000000004B /* ScreenAudioCapture.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = ScreenAudioCapture.swift; sourceTree = "<group>"; };
AAAA00000000000000000050 /* InputDeviceCapture.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = InputDeviceCapture.swift; sourceTree = "<group>"; };
AAAA0000000000000000001B /* AddServerSheet.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = AddServerSheet.swift; sourceTree = "<group>"; }; AAAA0000000000000000001B /* AddServerSheet.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = AddServerSheet.swift; sourceTree = "<group>"; };
AAAA0000000000000000001C /* ServerIdentitySheet.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = ServerIdentitySheet.swift; sourceTree = "<group>"; }; AAAA0000000000000000001C /* ServerIdentitySheet.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = ServerIdentitySheet.swift; sourceTree = "<group>"; };
AAAA0000000000000000001D /* PasswordPromptSheet.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = PasswordPromptSheet.swift; sourceTree = "<group>"; }; AAAA0000000000000000001D /* PasswordPromptSheet.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = PasswordPromptSheet.swift; sourceTree = "<group>"; };
@@ -104,6 +106,7 @@
isa = PBXGroup; isa = PBXGroup;
children = ( children = (
AAAA0000000000000000004B /* ScreenAudioCapture.swift */, AAAA0000000000000000004B /* ScreenAudioCapture.swift */,
AAAA00000000000000000050 /* InputDeviceCapture.swift */,
); );
path = Audio; path = Audio;
sourceTree = "<group>"; sourceTree = "<group>";
@@ -247,6 +250,7 @@
AAAA0000000000000000004E /* ScreenSharePickerSheet.swift in Sources */, AAAA0000000000000000004E /* ScreenSharePickerSheet.swift in Sources */,
AAAA0000000000000000004A /* SettingsWindowController.swift in Sources */, AAAA0000000000000000004A /* SettingsWindowController.swift in Sources */,
AAAA0000000000000000004C /* ScreenAudioCapture.swift in Sources */, AAAA0000000000000000004C /* ScreenAudioCapture.swift in Sources */,
AAAA00000000000000000051 /* InputDeviceCapture.swift in Sources */,
); );
runOnlyForDeploymentPostprocessing = 0; runOnlyForDeploymentPostprocessing = 0;
}; };

View File

@@ -0,0 +1,229 @@
import AVFoundation
import CoreAudio
// InputDeviceCapture captures a single hardware INPUT device (mic / line-in / aux) on macOS and
// emits 20 ms (960 samples/channel @ 48 kHz, interleaved int16) frames for the aux outgoing stream.
//
// The input-device analogue of ScreenAudioCapture (which captures system audio via
// ScreenCaptureKit). The core already owns ONE capture device (the mic) and can't open a second
// arbitrary input, so for the aux stream the client captures the chosen device here and feeds PCM
// into the core via `vc_stream_feed_pcm` the same external-feed pipeline screen audio uses.
//
// Device selection: an AVAudioEngine's input node wraps an AUHAL audio unit; setting
// kAudioOutputUnitProperty_CurrentDevice on it pins capture to a specific Core Audio device. We
// pass the device's *UID* (stable across reboots/replugs, unlike AudioDeviceID) and resolve it at
// start. The tap delivers Float32; AVAudioConverter resamples/quantises to 48 kHz int16.
final class InputDeviceCapture {
/// Receives a full 20 ms frame: (interleaved int16 PCM, samplesPerChannel = 960, channels).
typealias PcmHandler = (UnsafePointer<Int16>, Int, UInt32) -> Void
enum CaptureError: Error { case deviceNotFound, engineStartFailed(OSStatus) }
private static let frameSamplesPerChannel = 960 // 20 ms @ 48 kHz
private let deviceUID: String? // nil = system default input device
private let onPcm: PcmHandler
private let engine = AVAudioEngine()
private var converter: AVAudioConverter?
private var outFormat: AVAudioFormat?
private var channels = 1
/// Interleaved int16 carry-over between tap callbacks (the tap buffer doesn't align to 20 ms),
/// drained in whole `frameSamplesPerChannel * channels` chunks. Only touched on the tap queue.
private var pending: [Int16] = []
init(deviceUID: String?, onPcm: @escaping PcmHandler) {
self.deviceUID = deviceUID
self.onPcm = onPcm
}
/// Begin capture. Throws if the device can't be resolved or the engine fails to start.
func start() throws {
let input = engine.inputNode
// Pin the engine's AUHAL to the chosen device (skip for default the engine already uses
// the system default input). Must happen before reading inputFormat, which changes with
// the selected device.
if let deviceUID, let devId = Self.deviceID(forUID: deviceUID) {
var dev = devId
let status = AudioUnitSetProperty(input.audioUnit!,
kAudioOutputUnitProperty_CurrentDevice,
kAudioUnitScope_Global, 0,
&dev, UInt32(MemoryLayout<AudioDeviceID>.size))
if status != noErr { throw CaptureError.engineStartFailed(status) }
} else if deviceUID != nil {
throw CaptureError.deviceNotFound
}
let inFormat = input.inputFormat(forBus: 0)
channels = max(1, min(2, Int(inFormat.channelCount)))
guard let out = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 48000,
channels: AVAudioChannelCount(channels), interleaved: true)
else { throw CaptureError.deviceNotFound }
outFormat = out
converter = AVAudioConverter(from: inFormat, to: out)
input.installTap(onBus: 0, bufferSize: 960, format: inFormat) { [weak self] buf, _ in
self?.process(buf)
}
engine.prepare()
do { try engine.start() }
catch { throw CaptureError.engineStartFailed(-1) }
}
/// Stop capture and tear down the engine. Safe to call multiple times.
func stop() {
engine.inputNode.removeTap(onBus: 0)
if engine.isRunning { engine.stop() }
converter = nil
pending.removeAll(keepingCapacity: false)
}
// MARK: - Conversion (called on the tap's realtime thread)
private func process(_ inBuf: AVAudioPCMBuffer) {
guard let converter, let outFormat else { return }
// Output capacity must cover up-sampling (e.g. 44.1 48 kHz) plus slack.
let ratio = outFormat.sampleRate / inBuf.format.sampleRate
let cap = AVAudioFrameCount(Double(inBuf.frameLength) * ratio + 32)
guard cap > 0, let outBuf = AVAudioPCMBuffer(pcmFormat: outFormat, frameCapacity: cap)
else { return }
var fed = false
var err: NSError?
let status = converter.convert(to: outBuf, error: &err) { _, outStatus in
if fed { outStatus.pointee = .noDataNow; return nil }
fed = true
outStatus.pointee = .haveData
return inBuf
}
guard status != .error, outBuf.frameLength > 0,
let ch = outBuf.int16ChannelData else { return }
// Interleaved int16: all channels live in the first buffer (ch[0]).
let n = Int(outBuf.frameLength) * channels
pending.append(contentsOf: UnsafeBufferPointer(start: ch[0], count: n))
emit()
}
/// Fire `onPcm` for every whole 20 ms frame accumulated.
private func emit() {
let full = Self.frameSamplesPerChannel * channels
while pending.count >= full {
pending.withUnsafeBufferPointer { buf in
onPcm(buf.baseAddress!, Self.frameSamplesPerChannel, UInt32(channels))
}
pending.removeFirst(full)
}
}
// MARK: - UID AudioDeviceID resolution
private static func deviceID(forUID uid: String) -> AudioDeviceID? {
var addr = AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyTranslateUIDToDevice,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var deviceID = AudioDeviceID(0)
var cfUID = uid as CFString
var size = UInt32(MemoryLayout<AudioDeviceID>.size)
let status = withUnsafeMutablePointer(to: &cfUID) { uidPtr -> OSStatus in
AudioObjectGetPropertyData(AudioObjectID(kAudioObjectSystemObject), &addr,
UInt32(MemoryLayout<CFString>.size), uidPtr,
&size, &deviceID)
}
return (status == noErr && deviceID != 0) ? deviceID : nil
}
}
// InputDeviceInfo / InputDeviceEnumerator Core Audio input-device enumeration for the aux-stream
// picker. Separate from the core's vc_list_devices (whose ids are miniaudio-opaque and can't be
// passed to Core Audio); the aux device is client-captured, so the picker uses device UIDs.
struct InputDeviceInfo: Equatable {
let uid: String // stable across reboots/replugs what we persist
let name: String
let isDefault: Bool
}
enum InputDeviceEnumerator {
/// All Core Audio devices that expose at least one input channel.
static func list() -> [InputDeviceInfo] {
let defaultUID = defaultInputUID()
var result: [InputDeviceInfo] = []
for devID in allDeviceIDs() {
guard inputChannelCount(devID) > 0 else { continue }
guard let uid = stringProperty(devID, kAudioDevicePropertyDeviceUID) else { continue }
let name = stringProperty(devID, kAudioObjectPropertyName)
?? stringProperty(devID, kAudioDevicePropertyDeviceNameCFString)
?? "Unknown input device"
result.append(InputDeviceInfo(uid: uid, name: name, isDefault: uid == defaultUID))
}
return result.sorted { $0.name.localizedCaseInsensitiveCompare($1.name) == .orderedAscending }
}
// MARK: - Core Audio helpers
private static func allDeviceIDs() -> [AudioDeviceID] {
var addr = AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyDevices,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var size = UInt32(0)
guard AudioObjectGetPropertyDataSize(AudioObjectID(kAudioObjectSystemObject), &addr, 0, nil,
&size) == noErr, size > 0 else { return [] }
let count = Int(size) / MemoryLayout<AudioDeviceID>.size
var ids = [AudioDeviceID](repeating: 0, count: count)
let status = AudioObjectGetPropertyData(AudioObjectID(kAudioObjectSystemObject), &addr, 0,
nil, &size, &ids)
return status == noErr ? ids : []
}
private static func inputChannelCount(_ devID: AudioDeviceID) -> Int {
var addr = AudioObjectPropertyAddress(
mSelector: kAudioDevicePropertyStreamConfiguration,
mScope: kAudioObjectPropertyScopeInput,
mElement: kAudioObjectPropertyElementMain)
var size = UInt32(0)
guard AudioObjectGetPropertyDataSize(devID, &addr, 0, nil, &size) == noErr, size > 0
else { return 0 }
let bufList = UnsafeMutableRawPointer.allocate(byteCount: Int(size),
alignment: MemoryLayout<AudioBufferList>.alignment)
defer { bufList.deallocate() }
guard AudioObjectGetPropertyData(devID, &addr, 0, nil, &size, bufList) == noErr
else { return 0 }
let abl = UnsafeMutableAudioBufferListPointer(
bufList.assumingMemoryBound(to: AudioBufferList.self))
return abl.reduce(0) { $0 + Int($1.mNumberChannels) }
}
private static func defaultInputUID() -> String? {
var addr = AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyDefaultInputDevice,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var devID = AudioDeviceID(0)
var size = UInt32(MemoryLayout<AudioDeviceID>.size)
guard AudioObjectGetPropertyData(AudioObjectID(kAudioObjectSystemObject), &addr, 0, nil,
&size, &devID) == noErr, devID != 0 else { return nil }
return stringProperty(devID, kAudioDevicePropertyDeviceUID)
}
private static func stringProperty(_ devID: AudioDeviceID,
_ selector: AudioObjectPropertySelector) -> String? {
var addr = AudioObjectPropertyAddress(
mSelector: selector,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var cf: CFString? = nil
var size = UInt32(MemoryLayout<CFString?>.size)
let status = withUnsafeMutablePointer(to: &cf) { ptr in
AudioObjectGetPropertyData(devID, &addr, 0, nil, &size, ptr)
}
guard status == noErr else { return nil }
return cf as String?
}
}

View File

@@ -1,14 +1,11 @@
import AVFoundation import AVFoundation
import ScreenCaptureKit import ScreenCaptureKit
// Which apps' audio the SCREEN_AUDIO stream captures. ScreenCaptureKit filters audio at the // ScreenCaptureKit filters audio by application bundle identifier.
// *application* level (not per-window), so the selection is expressed as bundle IDs. The
// picker UI (ScreenSharePickerSheet) produces a `ScreenAudioSelection`; `start()` turns it
// into the matching `SCContentFilter`.
enum ScreenAudioScope: Equatable { enum ScreenAudioScope: Equatable {
case entireDesktop // whole display the original behaviour case entireDesktop
case onlyApps([String]) // capture only these bundle IDs case onlyApps([String])
case allExcept([String]) // capture everything except these bundle IDs case allExcept([String])
} }
struct ScreenAudioSelection: Equatable { struct ScreenAudioSelection: Equatable {
@@ -20,20 +17,8 @@ struct ScreenAudioSelection: Equatable {
static let `default` = ScreenAudioSelection() static let `default` = ScreenAudioSelection()
} }
// ScreenAudioCapture macOS system/desktop audio capture for the SCREEN_AUDIO stream. // SCStream requires a minimal video configuration even for audio-only capture. Only its audio
// // output is registered, and current-process audio is excluded to prevent feedback.
// The macOS analog of the Windows WASAPI loopback path (docs/voice.md §9). ScreenCaptureKit
// (macOS 13+) captures whatever the system is playing; we convert each audio CMSampleBuffer
// (Float32) int16 interleaved and push 20 ms frames (960 samples/channel @ 48 kHz) into the
// core via `vc_stream_feed_pcm` (exposed as `VoiceCatClient.feedPcm`). The core then runs the
// same Opus-encode media-AEAD UDP path as any other stream only the *source* is
// platform-specific (architecture.md §4).
//
// Audio-only: we request a 2×2 video plane at 1 fps purely because SCStream needs a video
// configuration, and we never add a `.screen` output only `.audio`. `excludesCurrentProcess
// Audio` prevents the self-echo loop of re-capturing our own incoming voice mix.
//
// `feedPcm` is thread-safe (any thread), so we forward straight from the sample-handler queue.
final class ScreenAudioCapture: NSObject, SCStreamOutput, SCStreamDelegate { final class ScreenAudioCapture: NSObject, SCStreamOutput, SCStreamDelegate {
/// Receives a full 20 ms frame: (interleaved int16 PCM, samplesPerChannel = 960, channels). /// Receives a full 20 ms frame: (interleaved int16 PCM, samplesPerChannel = 960, channels).

View File

@@ -49,7 +49,7 @@ final class PerUserTuningSheet: NSViewController {
let muted = state?.muted ?? false let muted = state?.muted ?? false
let nr = state?.noiseReduction ?? false let nr = state?.noiseReduction ?? false
let gainSlider = NSSlider(value: Double(gain * 100), minValue: 0, maxValue: 200, target: self, action: #selector(sliderChanged)) let gainSlider = NSSlider(value: Double(gain * 100), minValue: 0, maxValue: 400, target: self, action: #selector(sliderChanged))
gainSlider.tag = Int(stream.id) gainSlider.tag = Int(stream.id)
gainSlider.numberOfTickMarks = 0 gainSlider.numberOfTickMarks = 0
gainSlider.setAccessibilityLabel("Volume for \(stream.label): \(Int(gain * 100)) percent") gainSlider.setAccessibilityLabel("Volume for \(stream.label): \(Int(gain * 100)) percent")

View File

@@ -44,7 +44,6 @@ final class ConnectWindowController: NSWindowController, NSWindowDelegate {
private func buildUI() { private func buildUI() {
guard let contentView = window?.contentView else { return } guard let contentView = window?.contentView else { return }
// Server list
let col = NSTableColumn(identifier: NSUserInterfaceItemIdentifier("server")) let col = NSTableColumn(identifier: NSUserInterfaceItemIdentifier("server"))
col.title = "Saved Servers" col.title = "Saved Servers"
serverTableView.addTableColumn(col) serverTableView.addTableColumn(col)
@@ -61,7 +60,6 @@ final class ConnectWindowController: NSWindowController, NSWindowDelegate {
serverScrollView.translatesAutoresizingMaskIntoConstraints = false serverScrollView.translatesAutoresizingMaskIntoConstraints = false
contentView.addSubview(serverScrollView) contentView.addSubview(serverScrollView)
// Buttons row
configureButton(addButton, title: "Add…", action: #selector(addClicked)) configureButton(addButton, title: "Add…", action: #selector(addClicked))
configureButton(editButton, title: "Edit…", action: #selector(editClicked)) configureButton(editButton, title: "Edit…", action: #selector(editClicked))
configureButton(removeButton, title: "Remove", action: #selector(removeClicked)) configureButton(removeButton, title: "Remove", action: #selector(removeClicked))
@@ -72,13 +70,11 @@ final class ConnectWindowController: NSWindowController, NSWindowDelegate {
buttonStack.translatesAutoresizingMaskIntoConstraints = false buttonStack.translatesAutoresizingMaskIntoConstraints = false
contentView.addSubview(buttonStack) contentView.addSubview(buttonStack)
// Status
statusLabel.translatesAutoresizingMaskIntoConstraints = false statusLabel.translatesAutoresizingMaskIntoConstraints = false
statusLabel.textColor = .secondaryLabelColor statusLabel.textColor = .secondaryLabelColor
statusLabel.setAccessibilityLabel("Connection status") statusLabel.setAccessibilityLabel("Connection status")
contentView.addSubview(statusLabel) contentView.addSubview(statusLabel)
// Connect button
connectButton.title = "Connect" connectButton.title = "Connect"
connectButton.bezelStyle = .rounded connectButton.bezelStyle = .rounded
connectButton.keyEquivalent = "\r" connectButton.keyEquivalent = "\r"
@@ -376,7 +372,3 @@ extension ConnectWindowController: NSTableViewDataSource, NSTableViewDelegate {
return cell return cell
} }
} }
// MARK: - Helper

View File

@@ -31,10 +31,15 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
internal var micStreamId: UInt32 = 0 internal var micStreamId: UInt32 = 0
private var screenStreamId: UInt32 = 0 private var screenStreamId: UInt32 = 0
private var screenCapture: ScreenAudioCapture? private var screenCapture: ScreenAudioCapture?
private var auxStreamId: UInt32 = 0 // 0 = aux (second input device) stream not active
private var auxCapture: InputDeviceCapture? // client-side capture feeding the aux stream
// Last app/exclusion choice from the share picker; reused as the default next time. // Last app/exclusion choice from the share picker; reused as the default next time.
private var screenAudioSelection: ScreenAudioSelection = .default private var screenAudioSelection: ScreenAudioSelection = .default
internal var pttKeyCode: UInt16 = 0x60 // F8 internal var pttKeyCode: UInt16 = 0x60 { // F8
didSet { UserDefaults.standard.set(Int(pttKeyCode), forKey: AudioDefaults.pttKeyCode) }
}
private var pttMonitor: Any? private var pttMonitor: Any?
private var pttEngaged = false // guards the PTT cue against key-repeat
private var serverMuted = false private var serverMuted = false
private var serverDeafened = false private var serverDeafened = false
private var channelTree: [ChannelNode] = [] private var channelTree: [ChannelNode] = []
@@ -48,11 +53,56 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
private var settingsWindowController: SettingsWindowController? private var settingsWindowController: SettingsWindowController?
// MARK: - Audio settings state (source of truth read/written by SettingsWindowController) // MARK: - Audio settings state (source of truth read/written by SettingsWindowController)
// The input mode / VAD threshold / mic gain / PTT key persist via UserDefaults (didSet below)
// so they survive relaunch; loadPersistedAudioSettings() restores them at startup and they are
// pushed into the core when the mic stream starts (micToggleClicked).
internal var selectedInputMode: VoiceCatInputMode = .voiceActivation enum AudioDefaults {
internal var vadThresholdValue: Float = 0.05 static let inputMode = "voice.inputMode"
static let vadThreshold = "voice.vadThreshold"
static let inputGain = "voice.inputGain"
static let inputNoiseReduction = "voice.inputNoiseReduction"
static let stereoMic = "voice.stereoMic"
static let pttKeyCode = "voice.pttKeyCode"
static let auxEnabled = "voice.auxEnabled"
static let auxDeviceUID = "voice.auxDeviceUID"
static let auxGain = "voice.auxGain"
}
internal var selectedInputMode: VoiceCatInputMode = .voiceActivation {
didSet { UserDefaults.standard.set(Int(selectedInputMode.rawValue), forKey: AudioDefaults.inputMode) }
}
internal var vadThresholdValue: Float = 0.05 {
didSet { UserDefaults.standard.set(vadThresholdValue, forKey: AudioDefaults.vadThreshold) }
}
internal var inputGain: Float = 1.0 {
didSet { UserDefaults.standard.set(inputGain, forKey: AudioDefaults.inputGain) }
}
internal var inputNoiseReduction: Bool = false {
didSet { UserDefaults.standard.set(inputNoiseReduction, forKey: AudioDefaults.inputNoiseReduction) }
}
// Capture the mic in stereo (interleaved L/R) instead of mono. Real stereo only reaches the
// wire on a stereo channel; the core folds a stereo mic to mono on a mono channel. Applied to
// the core when the mic stream starts (micToggleClicked) and live via SettingsWindowController.
internal var stereoMic: Bool = false {
didSet { UserDefaults.standard.set(stereoMic, forKey: AudioDefaults.stereoMic) }
}
internal var selectedInputDeviceId: String? internal var selectedInputDeviceId: String?
// Aux outgoing stream: a second hardware input device the client captures itself and feeds to
// the core (kind = AUX_DEVICE, external_feed). Device + volume only aux is always-on (the
// core never gates AUX_DEVICE on VAD/PTT). auxDeviceUID is a Core Audio device UID (stable),
// NOT a core/miniaudio id. auxGain is read live by the capture feed, so the slider is instant.
internal var auxEnabled: Bool = false {
didSet { UserDefaults.standard.set(auxEnabled, forKey: AudioDefaults.auxEnabled) }
}
internal var auxDeviceUID: String? {
didSet { UserDefaults.standard.set(auxDeviceUID, forKey: AudioDefaults.auxDeviceUID) }
}
internal var auxGain: Float = 1.0 {
didSet { UserDefaults.standard.set(auxGain, forKey: AudioDefaults.auxGain) }
}
// MARK: - UI components // MARK: - UI components
private let channelOutlineView = NSOutlineView() private let channelOutlineView = NSOutlineView()
private let userTableView = NSTableView() private let userTableView = NSTableView()
@@ -96,17 +146,46 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
super.init(window: window) super.init(window: window)
window.delegate = self window.delegate = self
loadPersistedAudioSettings()
buildUI() buildUI()
buildToolbar() buildToolbar()
wireEvents() wireEvents()
bootstrap() bootstrap()
} }
/// Restore the saved input mode / VAD threshold / mic gain / PTT key from UserDefaults so a
/// relaunch keeps the user's transmission settings instead of resetting to VAD defaults.
private func loadPersistedAudioSettings() {
let d = UserDefaults.standard
if d.object(forKey: AudioDefaults.inputMode) != nil {
let raw = UInt32(d.integer(forKey: AudioDefaults.inputMode))
selectedInputMode = VoiceCatInputMode(rawValue: raw) ?? .voiceActivation
}
if d.object(forKey: AudioDefaults.vadThreshold) != nil {
vadThresholdValue = d.float(forKey: AudioDefaults.vadThreshold)
}
if d.object(forKey: AudioDefaults.inputGain) != nil {
inputGain = d.float(forKey: AudioDefaults.inputGain)
}
if d.object(forKey: AudioDefaults.inputNoiseReduction) != nil {
inputNoiseReduction = d.bool(forKey: AudioDefaults.inputNoiseReduction)
}
stereoMic = d.bool(forKey: AudioDefaults.stereoMic)
if d.object(forKey: AudioDefaults.pttKeyCode) != nil {
pttKeyCode = UInt16(d.integer(forKey: AudioDefaults.pttKeyCode))
}
auxEnabled = d.bool(forKey: AudioDefaults.auxEnabled)
if d.object(forKey: AudioDefaults.auxDeviceUID) != nil {
auxDeviceUID = d.string(forKey: AudioDefaults.auxDeviceUID)
}
if d.object(forKey: AudioDefaults.auxGain) != nil {
auxGain = d.float(forKey: AudioDefaults.auxGain)
}
}
required init?(coder: NSCoder) { fatalError() } required init?(coder: NSCoder) { fatalError() }
deinit { deinit {}
NSLog("[VoiceCatMac] MainWindowController deinit — client and event handlers are gone")
}
// MARK: - UI construction // MARK: - UI construction
@@ -343,7 +422,11 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
pttMonitor = NSEvent.addLocalMonitorForEvents(matching: [.keyDown, .keyUp]) { [weak self] event in pttMonitor = NSEvent.addLocalMonitorForEvents(matching: [.keyDown, .keyUp]) { [weak self] event in
guard let self, self.selectedInputMode == .pushToTalk, guard let self, self.selectedInputMode == .pushToTalk,
self.micStreamId != 0, event.keyCode == self.pttKeyCode else { return event } self.micStreamId != 0, event.keyCode == self.pttKeyCode else { return event }
self.client.setPushToTalk(event.type == .keyDown) let down = event.type == .keyDown
self.client.setPushToTalk(down)
// Cue only on the press transition key-down auto-repeats while held.
if down && !self.pttEngaged { EventFeedback.shared.play(.ptt) }
self.pttEngaged = down
return nil return nil
} }
@@ -370,10 +453,6 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
} }
ownPermissions = client.getPermissions() ownPermissions = client.getPermissions()
NSLog("[VoiceCatMac] bootstrap: channels=%d users=%d perms{admin=%d kick=%d} currentChannelId=%u",
channels.count, allUsers.count,
ownPermissions.isAdmin, ownPermissions.canKick, currentChannelId)
// Apply initial output volume (default 80% matches Windows client) // Apply initial output volume (default 80% matches Windows client)
client.setOutputVolume(0.8) client.setOutputVolume(0.8)
@@ -385,14 +464,13 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
buildMessagesMenu() buildMessagesMenu()
buildAdminMenu() buildAdminMenu()
addActivity("Connected to server as \(nickname)") addActivity("Connected to server as \(nickname)")
EventFeedback.shared.play(.login)
EventFeedback.shared.speak("Connected")
} }
// MARK: - Event handling // MARK: - Event handling
private func handleEvent(_ event: VoiceCatEvent) { private func handleEvent(_ event: VoiceCatEvent) {
NSLog("[VoiceCatMac] event type=%d result=%d userId=%u channelId=%u streamId=%u text=%@",
event.type.rawValue, event.result.rawValue, event.userId, event.channelId,
event.streamId, event.text ?? "(nil)")
switch event.type { switch event.type {
case .channelList: case .channelList:
channels = client.listChannels() channels = client.listChannels()
@@ -410,11 +488,14 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
isGuest: true, isGuest: true,
channelId: event.channelId, channelId: event.channelId,
selfMicMuted: false, selfDeafened: false, selfMicMuted: false, selfDeafened: false,
serverMuted: false, serverDeafened: false) serverMuted: false, serverDeafened: false,
voiceSubscribed: false)
users[event.userId] = u users[event.userId] = u
refreshChannelTree(); refreshUserList() refreshChannelTree(); refreshUserList()
if event.channelId == currentChannelId && event.userId != selfUserId { if event.channelId == currentChannelId && event.userId != selfUserId {
addActivity("\(u.nickname) joined the channel") addActivity("\(u.nickname) joined the channel")
EventFeedback.shared.play(.channelJoin)
EventFeedback.shared.speak("\(u.nickname) joined")
} }
case .userLeft: case .userLeft:
@@ -423,7 +504,11 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
users.removeValue(forKey: event.userId) users.removeValue(forKey: event.userId)
talkingUsers.remove(event.userId) talkingUsers.remove(event.userId)
refreshChannelTree(); refreshUserList() refreshChannelTree(); refreshUserList()
if wasHere { addActivity("\(nick) left the channel") } if wasHere {
addActivity("\(nick) left the channel")
EventFeedback.shared.play(.channelLeave)
EventFeedback.shared.speak("\(nick) left")
}
if let pmWin = pmWindows[event.userId] { if let pmWin = pmWindows[event.userId] {
pmWin.appendActivity("\(nick) disconnected from server") pmWin.appendActivity("\(nick) disconnected from server")
} }
@@ -447,7 +532,8 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
users[selfUserId] = User(id: self_.id, nickname: self_.nickname, isGuest: self_.isGuest, users[selfUserId] = User(id: self_.id, nickname: self_.nickname, isGuest: self_.isGuest,
channelId: event.channelId, channelId: event.channelId,
selfMicMuted: self_.selfMicMuted, selfDeafened: self_.selfDeafened, selfMicMuted: self_.selfMicMuted, selfDeafened: self_.selfDeafened,
serverMuted: self_.serverMuted, serverDeafened: self_.serverDeafened) serverMuted: self_.serverMuted, serverDeafened: self_.serverDeafened,
voiceSubscribed: self_.voiceSubscribed)
} }
refreshChannelTree(); refreshUserList(); updateStatusLabel() refreshChannelTree(); refreshUserList(); updateStatusLabel()
let name = channels.first(where: { $0.id == event.channelId })?.name ?? "Channel #\(event.channelId)" let name = channels.first(where: { $0.id == event.channelId })?.name ?? "Channel #\(event.channelId)"
@@ -473,6 +559,9 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
let talking = event.u32a == 1 let talking = event.u32a == 1
if talking { talkingUsers.insert(event.userId) } else { talkingUsers.remove(event.userId) } if talking { talkingUsers.insert(event.userId) } else { talkingUsers.remove(event.userId) }
refreshUserList() refreshUserList()
if event.userId == selfUserId {
EventFeedback.shared.play(talking ? .vaStart : .vaStop)
}
if talking && event.userId != selfUserId, if talking && event.userId != selfUserId,
let u = users[event.userId], u.channelId == currentChannelId { let u = users[event.userId], u.channelId == currentChannelId {
addActivity("\(u.nickname) started talking") addActivity("\(u.nickname) started talking")
@@ -493,10 +582,18 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
addActivity("\(u.nickname) started \(kindStr) stream") addActivity("\(u.nickname) started \(kindStr) stream")
case .streamStopped: case .streamStopped:
if let u = users[event.userId], u.channelId == currentChannelId { if event.userId == selfUserId {
if event.streamId == micStreamId {
micStreamId = 0
settingsWindowController?.resetLevel()
}
} else if let u = users[event.userId], u.channelId == currentChannelId {
addActivity("\(u.nickname) stopped a stream") addActivity("\(u.nickname) stopped a stream")
} }
case .voiceState:
handleVoiceState(event)
case .disconnected: case .disconnected:
handleDisconnected(event) handleDisconnected(event)
@@ -521,21 +618,28 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
time = DateFormatter.localizedString(from: Date(), dateStyle: .none, timeStyle: .short) time = DateFormatter.localizedString(from: Date(), dateStyle: .none, timeStyle: .short)
} }
let sender = nickname(for: event.userId) let sender = nickname(for: event.userId)
let isSelf = event.userId == selfUserId
let body = event.text ?? ""
if event.textScope == .private { if event.textScope == .private {
// Route PMs to per-conversation windows. For our own outgoing PM, ev.channelId // Route PMs to per-conversation windows. For our own outgoing PM, ev.channelId
// carries the recipient user ID; for incoming, ev.userId is the sender. // carries the recipient user ID; for incoming, ev.userId is the sender.
let otherId = event.userId == selfUserId ? event.channelId : event.userId let otherId = isSelf ? event.channelId : event.userId
let win = getOrOpenPmWindow(otherId) let win = getOrOpenPmWindow(otherId)
win.appendMessage(time: time, isSelf: event.userId == selfUserId, sender: sender, win.appendMessage(time: time, isSelf: isSelf, sender: sender, text: body)
text: event.text ?? "") EventFeedback.shared.play(isSelf ? .pmSent : .pmRecv)
if event.userId != selfUserId { if !isSelf {
addActivity("Private message from \(sender)") addActivity("Private message from \(sender)")
EventFeedback.shared.speak("Private message from \(sender): \(body)")
} }
} else { } else {
let line = "[\(time)] \(sender): \(event.text ?? "")\n" let line = "[\(time)] \(sender): \(body)\n"
logTextView.textStorage?.append(NSAttributedString(string: line)) logTextView.textStorage?.append(NSAttributedString(string: line))
logTextView.scrollToEndOfDocument(nil) logTextView.scrollToEndOfDocument(nil)
EventFeedback.shared.play(isSelf ? .channelSent : .channelRecv)
if !isSelf {
EventFeedback.shared.speak("\(sender): \(body)")
}
} }
} }
@@ -607,11 +711,14 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
let msg = event.text.map { "Disconnected: \($0)" } ?? "Disconnected from server." let msg = event.text.map { "Disconnected: \($0)" } ?? "Disconnected from server."
statusLabel.stringValue = msg statusLabel.stringValue = msg
addActivity(msg) addActivity(msg)
EventFeedback.shared.play(event.result == .ok ? .logout : .connectionLost)
EventFeedback.shared.speak(event.result == .ok ? "Disconnected" : "Connection lost")
channelTree = []; channelOutlineView.reloadData() channelTree = []; channelOutlineView.reloadData()
displayedUsers = []; userTableView.reloadData() displayedUsers = []; userTableView.reloadData()
users.removeAll(); talkingUsers.removeAll() users.removeAll(); talkingUsers.removeAll()
stopScreenCapture() stopScreenCapture()
currentChannelId = 0; micStreamId = 0; screenStreamId = 0 auxCapture?.stop(); auxCapture = nil // connection gone drop capture, no stopStream
currentChannelId = 0; micStreamId = 0; screenStreamId = 0; auxStreamId = 0
composeField.isEnabled = false; sendButton.isEnabled = false composeField.isEnabled = false; sendButton.isEnabled = false
joinVoiceButton?.isEnabled = false joinVoiceButton?.isEnabled = false
shareScreenButton?.isEnabled = false shareScreenButton?.isEnabled = false
@@ -677,30 +784,127 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
@objc private func micToggleClicked() { @objc private func micToggleClicked() {
if micStreamId == 0 { if micStreamId == 0 {
let result = client.joinVoice()
if result != .ok {
addActivity("Failed to join voice: \(result)")
}
} else {
stopAuxStream()
if screenStreamId != 0 { stopScreenAudio() }
client.setPushToTalk(false)
client.leaveVoice()
}
}
private func handleVoiceState(_ event: VoiceCatEvent) {
let subscribed = event.u32a != 0
if subscribed {
let (result, streamId) = client.startStream(StreamDescriptor(kind: .mic, deviceId: nil, label: "Microphone")) let (result, streamId) = client.startStream(StreamDescriptor(kind: .mic, deviceId: nil, label: "Microphone"))
if result == .ok { if result == .ok {
micStreamId = streamId micStreamId = streamId
if let devId = selectedInputDeviceId { if let devId = selectedInputDeviceId {
client.setInputDevice(streamId: streamId, deviceId: devId) client.setInputDevice(streamId: streamId, deviceId: devId)
} }
client.setCaptureChannels(streamId: streamId, channels: stereoMic ? 2 : 1)
client.setInputMode(selectedInputMode) client.setInputMode(selectedInputMode)
if selectedInputMode == .voiceActivation { if selectedInputMode == .voiceActivation {
client.setVadThreshold(vadThresholdValue) client.setVadThreshold(vadThresholdValue)
} }
client.setInputGain(inputGain)
client.setInputNoiseReduction(inputNoiseReduction)
setVoiceJoinedState(true) setVoiceJoinedState(true)
addActivity("Joined voice — microphone active") addActivity("Joined voice — microphone active")
EventFeedback.shared.play(.voiceOn)
NSAccessibility.post(element: logTextView, notification: .announcementRequested, NSAccessibility.post(element: logTextView, notification: .announcementRequested,
userInfo: [.announcement: "Joined voice", .priority: NSAccessibilityPriorityLevel.medium]) userInfo: [.announcement: "Joined voice", .priority: NSAccessibilityPriorityLevel.medium])
startAuxStream()
} else { } else {
addActivity("Failed to start microphone: \(result)") addActivity("Failed to start microphone: \(result)")
} }
} else { } else {
client.setPushToTalk(false)
client.stopStream(micStreamId)
micStreamId = 0 micStreamId = 0
settingsWindowController?.resetLevel() settingsWindowController?.resetLevel()
setVoiceJoinedState(false) setVoiceJoinedState(false)
addActivity("Left voice") addActivity("Left voice")
EventFeedback.shared.play(.voiceOff)
}
}
// MARK: - Aux input stream (second hardware input device)
/// Called by SettingsWindowController when the user toggles the aux checkbox.
func applyAuxEnabled(_ on: Bool) {
auxEnabled = on
guard micStreamId != 0 else { return } // not in voice applied on next Join Voice
if on { startAuxStream() } else { stopAuxStream() }
}
/// Called by SettingsWindowController when the user picks a different aux device.
func applyAuxDevice(_ uid: String?) {
auxDeviceUID = uid
if auxStreamId != 0 { restartAuxCapture() }
}
private func startAuxStream() {
guard auxStreamId == 0, auxEnabled else { return }
let (result, streamId) = client.startStream(
StreamDescriptor(kind: .auxDevice, deviceId: nil, label: "Aux device", externalFeed: true))
guard result == .ok else {
addActivity("Failed to start aux stream: \(result)")
return
}
auxStreamId = streamId
startAuxCapture()
addActivity("Aux input stream active")
}
private func startAuxCapture() {
let capture = InputDeviceCapture(deviceUID: auxDeviceUID) { [weak self] ptr, spc, ch in
self?.feedAux(ptr, samplesPerChannel: spc, channels: ch)
}
auxCapture = capture
do { try capture.start() }
catch {
addActivity("Failed to open aux input device")
stopAuxStream()
}
}
// Re-open the capture on a different device while the aux stream stays up (the core stream id
// is unchanged only the client-side capture source changes).
private func restartAuxCapture() {
guard auxStreamId != 0 else { return }
auxCapture?.stop()
auxCapture = nil
startAuxCapture()
}
private func stopAuxStream() {
auxCapture?.stop()
auxCapture = nil
if auxStreamId != 0 {
client.stopStream(auxStreamId)
auxStreamId = 0
}
}
// Fired on the capture's realtime thread. feedPcm is thread-safe. Gain is read live from
// auxGain each frame so the volume slider takes effect immediately.
private func feedAux(_ pcm: UnsafePointer<Int16>, samplesPerChannel: Int, channels: UInt32) {
guard auxStreamId != 0 else { return }
let gain = auxGain
if gain != 1.0 {
let n = samplesPerChannel * Int(channels)
var scaled = [Int16](repeating: 0, count: n)
for i in 0..<n {
let v = (Float(pcm[i]) * gain).rounded()
scaled[i] = Int16(max(-32768, min(32767, v)))
}
client.feedPcm(streamId: auxStreamId, pcm: scaled,
samplesPerChannel: samplesPerChannel, channels: channels)
} else {
client.feedPcm(streamId: auxStreamId, pcm: pcm,
samplesPerChannel: samplesPerChannel, channels: channels)
} }
} }
@@ -740,6 +944,15 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
} }
} }
private func stopScreenAudio() {
guard screenStreamId != 0 else { return }
stopScreenCapture()
client.stopStream(screenStreamId)
screenStreamId = 0
setShareScreenButton(active: false)
addActivity("Stopped sharing screen audio")
}
/// Announce the SCREEN_AUDIO stream with the chosen selection in hand. ScreenCaptureKit /// Announce the SCREEN_AUDIO stream with the chosen selection in hand. ScreenCaptureKit
/// capture starts once the server's StreamAnnounceResult lands (the .streamStarted event), /// capture starts once the server's StreamAnnounceResult lands (the .streamStarted event),
/// when the effective audio config and thus the channel count is known. See /// when the effective audio config and thus the channel count is known. See
@@ -845,7 +1058,7 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
composeField.stringValue = "" composeField.stringValue = ""
} }
// MARK: - M5: Moderation helpers // MARK: - Moderation helpers
private func moveUser(_ user: User) { private func moveUser(_ user: User) {
let sheet = MoveUserSheet(channels: channels, currentChannelId: user.channelId) let sheet = MoveUserSheet(channels: channels, currentChannelId: user.channelId)
@@ -1019,7 +1232,6 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
presentSheet(sheet) presentSheet(sheet)
} }
/// Open a PM window from the user context menu.
private func openPmWindow(_ user: User) { private func openPmWindow(_ user: User) {
getOrOpenPmWindow(user.id) getOrOpenPmWindow(user.id)
} }
@@ -1066,17 +1278,13 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
func windowWillClose(_ notification: Notification) { func windowWillClose(_ notification: Notification) {
if let mon = pttMonitor { NSEvent.removeMonitor(mon) } if let mon = pttMonitor { NSEvent.removeMonitor(mon) }
NotificationCenter.default.removeObserver(self) NotificationCenter.default.removeObserver(self)
// Close all PM windows
for (_, pmWin) in pmWindows { pmWin.close() } for (_, pmWin) in pmWindows { pmWin.close() }
pmWindows.removeAll() pmWindows.removeAll()
// Close settings window
settingsWindowController?.close() settingsWindowController?.close()
settingsWindowController = nil settingsWindowController = nil
// Remove app menus we added
if let item = voiceMenuItem { NSApp.mainMenu?.removeItem(item) } if let item = voiceMenuItem { NSApp.mainMenu?.removeItem(item) }
if let item = messagesMenuItem { NSApp.mainMenu?.removeItem(item) } if let item = messagesMenuItem { NSApp.mainMenu?.removeItem(item) }
if let item = adminMenuItem { NSApp.mainMenu?.removeItem(item) } if let item = adminMenuItem { NSApp.mainMenu?.removeItem(item) }
// Remove Settings menu item + separator from app menu
if let appMenu = NSApp.mainMenu?.item(at: 0)?.submenu { if let appMenu = NSApp.mainMenu?.item(at: 0)?.submenu {
if let item = settingsMenuItem { appMenu.removeItem(item) } if let item = settingsMenuItem { appMenu.removeItem(item) }
// Remove the separator we inserted before Quit // Remove the separator we inserted before Quit
@@ -1091,6 +1299,7 @@ final class MainWindowController: NSWindowController, NSWindowDelegate {
} }
client.setPushToTalk(false) client.setPushToTalk(false)
stopScreenCapture() stopScreenCapture()
if auxStreamId != 0 { stopAuxStream() }
if screenStreamId != 0 { client.stopStream(screenStreamId) } if screenStreamId != 0 { client.stopStream(screenStreamId) }
if micStreamId != 0 { client.stopStream(micStreamId) } if micStreamId != 0 { client.stopStream(micStreamId) }
client.onLevel = nil client.onLevel = nil
@@ -1255,7 +1464,7 @@ extension MainWindowController: NSMenuDelegate {
let info = ChannelEdit(id: ch.id, parentId: ch.parentId, name: ch.name, let info = ChannelEdit(id: ch.id, parentId: ch.parentId, name: ch.name,
topic: ch.topic, passwordProtected: ch.passwordProtected, topic: ch.topic, passwordProtected: ch.passwordProtected,
password: nil, maxUsers: ch.maxUsers, password: nil, maxUsers: ch.maxUsers,
sortOrder: 0, audio: AudioConfig()) sortOrder: ch.sortOrder, audio: ch.audio)
let sheet = ChannelEditSheet(channels: channels, editing: info) let sheet = ChannelEditSheet(channels: channels, editing: info)
sheet.onComplete = { [weak self] edited in sheet.onComplete = { [weak self] edited in
guard let edited else { return } guard let edited else { return }

View File

@@ -50,6 +50,56 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
// Cached VAD slider position so we can restore it when the window reopens. // Cached VAD slider position so we can restore it when the window reopens.
private var vadSliderValue: Double = 50 private var vadSliderValue: Double = 50
// Mic input gain: 0400 % (100 = unity). Persisted via MainWindowController.inputGain.
private let inputGainSlider: NSSlider = {
let s = NSSlider(value: 100, minValue: 0, maxValue: 400, target: nil, action: nil)
s.numberOfTickMarks = 0
return s
}()
private let inputGainValueLabel = NSTextField(labelWithString: "100%")
// Send-side mic noise reduction (RNNoise). MIC-only. Persisted via
// MainWindowController.inputNoiseReduction.
private let nrCheckbox = NSButton(checkboxWithTitle: "Noise reduction (RNNoise)",
target: nil, action: nil)
// Capture the mic in stereo (interleaved L/R) instead of mono. Real stereo only reaches the
// wire on a stereo channel; the core folds a stereo mic to mono on a mono channel. Persisted
// via MainWindowController.stereoMic.
private let stereoMicCheckbox = NSButton(checkboxWithTitle: "Stereo microphone",
target: nil, action: nil)
// Aux input stream: a second outgoing stream from another hardware input device (e.g. line-in
// / aux), captured client-side. Device + volume only aux is always-on. Persisted via
// MainWindowController.auxEnabled / auxDeviceUID / auxGain.
private let auxCheckbox = NSButton(checkboxWithTitle: "Aux input stream (second device)",
target: nil, action: nil)
private let auxDevicePicker = NSPopUpButton()
private let auxRefreshButton = NSButton()
private let auxGainSlider: NSSlider = {
let s = NSSlider(value: 100, minValue: 0, maxValue: 400, target: nil, action: nil)
s.numberOfTickMarks = 0
return s
}()
private let auxGainValueLabel = NSTextField(labelWithString: "100%")
private let auxDeviceLabel = NSTextField(labelWithString: "Aux device:")
private let auxGainLabel = NSTextField(labelWithString: "Aux volume:")
// Notification feedback controls. Read/write UserDefaults with the same keys VoiceCatCore's
// FeedbackSettings reads, so EventFeedback honours these immediately.
private let soundsCheckbox = NSButton(checkboxWithTitle: "Event sounds", target: nil, action: nil)
private let soundsVolumeSlider: NSSlider = {
let s = NSSlider(value: 1, minValue: 0, maxValue: 1, target: nil, action: nil)
s.numberOfTickMarks = 0
return s
}()
private let speechCheckbox = NSButton(checkboxWithTitle: "Speak events (text-to-speech)",
target: nil, action: nil)
private let selfTalkCheckbox = NSButton(checkboxWithTitle: "Your own voice-activity sounds",
target: nil, action: nil)
private let pttSoundCheckbox = NSButton(checkboxWithTitle: "Push-to-talk cue",
target: nil, action: nil)
// MARK: - Init // MARK: - Init
init(client: VoiceCatClient, mainController: MainWindowController) { init(client: VoiceCatClient, mainController: MainWindowController) {
@@ -57,13 +107,13 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
self.mainController = mainController self.mainController = mainController
let window = NSWindow( let window = NSWindow(
contentRect: NSRect(x: 0, y: 0, width: 380, height: 260), contentRect: NSRect(x: 0, y: 0, width: 380, height: 420),
styleMask: [.titled, .closable, .miniaturizable], styleMask: [.titled, .closable, .miniaturizable],
backing: .buffered, backing: .buffered,
defer: false defer: false
) )
window.title = "Audio Settings" window.title = "Settings"
window.minSize = NSSize(width: 340, height: 220) window.minSize = NSSize(width: 340, height: 380)
window.center() window.center()
super.init(window: window) super.init(window: window)
window.delegate = self window.delegate = self
@@ -71,6 +121,7 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
buildUI() buildUI()
syncFromMainController() syncFromMainController()
loadInputDevices() loadInputDevices()
loadAuxDevices()
} }
required init?(coder: NSCoder) { fatalError() } required init?(coder: NSCoder) { fatalError() }
@@ -120,6 +171,24 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
levelMeter.setAccessibilityLabel("Microphone input level") levelMeter.setAccessibilityLabel("Microphone input level")
levelMeter.setAccessibilityHelp("Shows current microphone volume level") levelMeter.setAccessibilityHelp("Shows current microphone volume level")
let inputGainLabel = NSTextField(labelWithString: "Mic volume:")
inputGainLabel.setAccessibilityLabel("Microphone volume")
inputGainSlider.target = self
inputGainSlider.action = #selector(inputGainChanged)
inputGainSlider.setAccessibilityLabel("Microphone volume")
inputGainSlider.setAccessibilityHelp("Boost a quiet microphone. 100% is unity.")
inputGainValueLabel.setAccessibilityLabel("Microphone volume value")
nrCheckbox.target = self
nrCheckbox.action = #selector(nrChanged)
nrCheckbox.setAccessibilityLabel("Microphone noise reduction")
nrCheckbox.setAccessibilityHelp("RNNoise denoising of your microphone. Cleans your signal for everyone.")
stereoMicCheckbox.target = self
stereoMicCheckbox.action = #selector(stereoMicChanged)
stereoMicCheckbox.setAccessibilityLabel("Stereo microphone")
stereoMicCheckbox.setAccessibilityHelp("Capture your microphone in stereo. Only transmitted in stereo on a stereo channel.")
let inputModeRow = NSStackView(views: [inputModeLabel, inputModeControl]) let inputModeRow = NSStackView(views: [inputModeLabel, inputModeControl])
inputModeRow.orientation = .horizontal inputModeRow.orientation = .horizontal
inputModeRow.spacing = 8 inputModeRow.spacing = 8
@@ -140,7 +209,71 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
levelRow.orientation = .horizontal levelRow.orientation = .horizontal
levelRow.spacing = 8 levelRow.spacing = 8
let stack = NSStackView(views: [inputModeRow, vadRow, pttRow, deviceRow, levelRow]) let inputGainRow = NSStackView(views: [inputGainLabel, inputGainSlider, inputGainValueLabel])
inputGainRow.orientation = .horizontal
inputGainRow.spacing = 8
// Aux input stream
let auxHeader = NSTextField(labelWithString: "Aux input stream")
auxHeader.font = .boldSystemFont(ofSize: NSFont.systemFontSize)
auxCheckbox.target = self
auxCheckbox.action = #selector(auxEnabledChanged)
auxCheckbox.setAccessibilityLabel("Enable aux input stream")
auxCheckbox.setAccessibilityHelp("Transmit a second hardware input device alongside your microphone.")
auxDeviceLabel.setAccessibilityLabel("Aux input device")
auxDevicePicker.setAccessibilityLabel("Aux input device")
auxDevicePicker.target = self
auxDevicePicker.action = #selector(auxDeviceChanged)
auxRefreshButton.title = ""
auxRefreshButton.bezelStyle = .rounded
auxRefreshButton.target = self
auxRefreshButton.action = #selector(refreshAuxDevicesClicked)
auxRefreshButton.setAccessibilityLabel("Refresh aux device list")
auxRefreshButton.toolTip = "Refresh"
auxGainLabel.setAccessibilityLabel("Aux volume")
auxGainSlider.target = self
auxGainSlider.action = #selector(auxGainChanged)
auxGainSlider.setAccessibilityLabel("Aux volume")
auxGainSlider.setAccessibilityHelp("Volume of the aux input stream. 100% is unity.")
auxGainValueLabel.setAccessibilityLabel("Aux volume value")
let auxDeviceRow = NSStackView(views: [auxDeviceLabel, auxDevicePicker, auxRefreshButton])
auxDeviceRow.orientation = .horizontal
auxDeviceRow.spacing = 8
let auxGainRow = NSStackView(views: [auxGainLabel, auxGainSlider, auxGainValueLabel])
auxGainRow.orientation = .horizontal
auxGainRow.spacing = 8
// Notifications
let notificationsHeader = NSTextField(labelWithString: "Notifications")
notificationsHeader.font = .boldSystemFont(ofSize: NSFont.systemFontSize)
for box in [soundsCheckbox, speechCheckbox, selfTalkCheckbox, pttSoundCheckbox] {
box.target = self
box.action = #selector(notificationSettingChanged)
}
soundsCheckbox.setAccessibilityLabel("Play event sounds")
speechCheckbox.setAccessibilityLabel("Speak events")
selfTalkCheckbox.setAccessibilityLabel("Your own voice-activity sounds")
pttSoundCheckbox.setAccessibilityLabel("Push-to-talk cue")
let volumeLabel = NSTextField(labelWithString: "Sound volume:")
volumeLabel.setAccessibilityLabel("Sound volume")
soundsVolumeSlider.target = self
soundsVolumeSlider.action = #selector(notificationSettingChanged)
soundsVolumeSlider.setAccessibilityLabel("Sound volume")
let volumeRow = NSStackView(views: [volumeLabel, soundsVolumeSlider])
volumeRow.orientation = .horizontal
volumeRow.spacing = 8
let stack = NSStackView(views: [inputModeRow, vadRow, inputGainRow, nrCheckbox, stereoMicCheckbox, pttRow, deviceRow,
levelRow, auxHeader, auxCheckbox, auxDeviceRow, auxGainRow,
notificationsHeader, soundsCheckbox, volumeRow,
speechCheckbox, selfTalkCheckbox, pttSoundCheckbox])
stack.orientation = .vertical stack.orientation = .vertical
stack.spacing = 12 stack.spacing = 12
stack.alignment = .leading stack.alignment = .leading
@@ -155,9 +288,35 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
stack.bottomAnchor.constraint(equalTo: contentView.bottomAnchor), stack.bottomAnchor.constraint(equalTo: contentView.bottomAnchor),
vadSlider.widthAnchor.constraint(greaterThanOrEqualToConstant: 200), vadSlider.widthAnchor.constraint(greaterThanOrEqualToConstant: 200),
inputGainSlider.widthAnchor.constraint(greaterThanOrEqualToConstant: 180),
levelMeter.widthAnchor.constraint(equalToConstant: 200), levelMeter.widthAnchor.constraint(equalToConstant: 200),
devicePicker.widthAnchor.constraint(greaterThanOrEqualToConstant: 180), devicePicker.widthAnchor.constraint(greaterThanOrEqualToConstant: 180),
soundsVolumeSlider.widthAnchor.constraint(greaterThanOrEqualToConstant: 200),
auxDevicePicker.widthAnchor.constraint(greaterThanOrEqualToConstant: 180),
auxGainSlider.widthAnchor.constraint(greaterThanOrEqualToConstant: 180),
]) ])
syncNotificationControls()
}
/// Load the notification checkbox/slider states from UserDefaults. Touches EventFeedback.shared
/// first so its default values are registered before we read them.
private func syncNotificationControls() {
let s = FeedbackSettings.current
soundsCheckbox.state = s.sounds ? .on : .off
speechCheckbox.state = s.speech ? .on : .off
selfTalkCheckbox.state = s.selfTalkSounds ? .on : .off
pttSoundCheckbox.state = s.pttSound ? .on : .off
soundsVolumeSlider.doubleValue = Double(s.volume)
}
@objc private func notificationSettingChanged() {
let d = UserDefaults.standard
d.set(soundsCheckbox.state == .on, forKey: "feedback.sounds")
d.set(speechCheckbox.state == .on, forKey: "feedback.speech")
d.set(selfTalkCheckbox.state == .on, forKey: "feedback.selfTalk")
d.set(pttSoundCheckbox.state == .on, forKey: "feedback.ptt")
d.set(soundsVolumeSlider.doubleValue, forKey: "feedback.volume")
} }
// MARK: - Sync from MainWindowController // MARK: - Sync from MainWindowController
@@ -173,9 +332,22 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
case .alwaysOn: inputModeControl.selectedSegment = 2 case .alwaysOn: inputModeControl.selectedSegment = 2
} }
// Restore the slider from the persisted threshold (invert vadThresholdFromSlider) so a
// relaunch shows the saved sensitivity, not the default mid-point.
vadSliderValue = vadSliderFromThreshold(mc.vadThresholdValue)
vadSlider.doubleValue = vadSliderValue vadSlider.doubleValue = vadSliderValue
pttKeyLabel.stringValue = "(\(keyCodeName(mc.pttKeyCode)))" pttKeyLabel.stringValue = "(\(keyCodeName(mc.pttKeyCode)))"
inputGainSlider.doubleValue = Double(mc.inputGain * 100)
updateInputGainLabel()
nrCheckbox.state = mc.inputNoiseReduction ? .on : .off
stereoMicCheckbox.state = mc.stereoMic ? .on : .off
auxCheckbox.state = mc.auxEnabled ? .on : .off
auxGainSlider.doubleValue = Double(mc.auxGain * 100)
updateAuxGainLabel()
updateAuxControlsEnabled()
updateConditionalControls() updateConditionalControls()
} }
@@ -213,6 +385,39 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
} }
} }
@objc private func inputGainChanged() {
let gain = Float(inputGainSlider.doubleValue) / 100.0
mainController?.inputGain = gain
updateInputGainLabel()
if let mc = mainController, mc.micStreamId != 0 {
client.setInputGain(gain)
}
}
@objc private func nrChanged() {
let on = nrCheckbox.state == .on
mainController?.inputNoiseReduction = on
if let mc = mainController, mc.micStreamId != 0 {
client.setInputNoiseReduction(on)
}
}
@objc private func stereoMicChanged() {
let on = stereoMicCheckbox.state == .on
mainController?.stereoMic = on
// Channel count only takes effect when the capture device (re)starts, so restart it live.
if let mc = mainController, mc.micStreamId != 0 {
client.setCaptureChannels(streamId: mc.micStreamId, channels: on ? 2 : 1)
client.audioRestart()
}
}
private func updateInputGainLabel() {
let pct = Int(inputGainSlider.doubleValue.rounded())
inputGainValueLabel.stringValue = "\(pct)%"
inputGainSlider.setAccessibilityValue("\(pct) percent")
}
@objc private func changePttClicked() { @objc private func changePttClicked() {
guard let mc = mainController else { return } guard let mc = mainController else { return }
let sheet = PttKeyCaptureSheet(currentKeyCode: mc.pttKeyCode) let sheet = PttKeyCaptureSheet(currentKeyCode: mc.pttKeyCode)
@@ -234,6 +439,65 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
} }
} }
// MARK: - Aux input stream actions
@objc private func auxEnabledChanged() {
let on = auxCheckbox.state == .on
updateAuxControlsEnabled()
mainController?.applyAuxEnabled(on)
}
@objc private func refreshAuxDevicesClicked() { loadAuxDevices() }
@objc private func auxDeviceChanged() {
let uid = auxDevicePicker.selectedItem?.representedObject as? String
mainController?.applyAuxDevice(uid)
}
@objc private func auxGainChanged() {
mainController?.auxGain = Float(auxGainSlider.doubleValue) / 100.0
updateAuxGainLabel()
}
private func updateAuxGainLabel() {
let pct = Int(auxGainSlider.doubleValue.rounded())
auxGainValueLabel.stringValue = "\(pct)%"
auxGainSlider.setAccessibilityValue("\(pct) percent")
}
private func updateAuxControlsEnabled() {
let on = auxCheckbox.state == .on
auxDeviceLabel.isEnabled = on
auxDevicePicker.isEnabled = on
auxRefreshButton.isEnabled = on
auxGainLabel.isEnabled = on
auxGainSlider.isEnabled = on
auxGainValueLabel.isEnabled = on
}
private func loadAuxDevices() {
let devices = InputDeviceEnumerator.list()
let prevSelected = (auxDevicePicker.selectedItem?.representedObject as? String)
?? mainController?.auxDeviceUID
auxDevicePicker.removeAllItems()
for d in devices {
let item = NSMenuItem(title: d.name, action: nil, keyEquivalent: "")
item.representedObject = d.uid
auxDevicePicker.menu?.addItem(item)
}
if let prev = prevSelected,
let item = auxDevicePicker.itemArray.first(where: { ($0.representedObject as? String) == prev }) {
auxDevicePicker.select(item)
} else if let def = devices.first(where: { $0.isDefault }),
let item = auxDevicePicker.itemArray.first(where: { ($0.representedObject as? String) == def.uid }) {
auxDevicePicker.select(item)
} else if auxDevicePicker.numberOfItems > 0 {
auxDevicePicker.selectItem(at: 0)
}
// Persist the resolved selection so a relaunch (or Join Voice) opens the same device.
mainController?.auxDeviceUID = auxDevicePicker.selectedItem?.representedObject as? String
}
// MARK: - Level meter (called by MainWindowController) // MARK: - Level meter (called by MainWindowController)
func updateLevel(rms: Float) { func updateLevel(rms: Float) {
@@ -284,6 +548,12 @@ final class SettingsWindowController: NSWindowController, NSWindowDelegate {
0.1 * (1.0 - Float(vadSlider.doubleValue - 1.0) / 99.0) 0.1 * (1.0 - Float(vadSlider.doubleValue - 1.0) / 99.0)
} }
/// Inverse of vadThresholdFromSlider: map a stored threshold back to a 1100 slider position.
private func vadSliderFromThreshold(_ threshold: Float) -> Double {
let clamped = min(max(threshold, 0.0), 0.1)
return Double(1.0 + (1.0 - clamped / 0.1) * 99.0)
}
private func presentSheet(_ vc: NSViewController) { private func presentSheet(_ vc: NSViewController) {
if let cvc = window?.contentViewController { if let cvc = window?.contentViewController {
cvc.presentAsSheet(vc) cvc.presentAsSheet(vc)

View File

@@ -54,7 +54,8 @@ done
# ── Resolve VCPKG_ROOT ─────────────────────────────────────────────────────────── # ── Resolve VCPKG_ROOT ───────────────────────────────────────────────────────────
# The apple-dev CMake cache records the vcpkg root it was configured with (Z_VCPKG_ROOT_DIR); # The apple-dev CMake cache records the vcpkg root it was configured with (Z_VCPKG_ROOT_DIR);
# reuse that so a developer who already configured `cmake --preset dev` doesn't need VCPKG_ROOT # reuse that so a developer who already configured `cmake --preset dev` doesn't need VCPKG_ROOT
# in their shell env to run this script. # in their shell env to run this script. Falls back to the bundled submodule at
# <repo-root>/vcpkg if neither the env var nor the cache resolve it.
if [[ -z "${VCPKG_ROOT:-}" ]]; then if [[ -z "${VCPKG_ROOT:-}" ]]; then
cache="$REPO_ROOT/build/apple-dev/CMakeCache.txt" cache="$REPO_ROOT/build/apple-dev/CMakeCache.txt"
if [[ -f "$cache" ]]; then if [[ -f "$cache" ]]; then
@@ -64,9 +65,16 @@ if [[ -z "${VCPKG_ROOT:-}" ]]; then
fi fi
fi fi
fi fi
if [[ -z "${VCPKG_ROOT:-}" ]]; then
bundled="$REPO_ROOT/vcpkg"
if [[ -f "$bundled/scripts/buildsystems/vcpkg.cmake" ]]; then
export VCPKG_ROOT="$bundled"
fi
fi
if [[ -z "${VCPKG_ROOT:-}" || ! -d "$VCPKG_ROOT" ]]; then if [[ -z "${VCPKG_ROOT:-}" || ! -d "$VCPKG_ROOT" ]]; then
echo "error: VCPKG_ROOT is not set or does not exist." >&2 echo "error: VCPKG_ROOT is not set or does not exist." >&2
echo " bootstrap vcpkg (https://vcpkg.io) then: export VCPKG_ROOT=/path/to/vcpkg" >&2 echo " either init the bundled submodule: git submodule update --init vcpkg" >&2
echo " or point at an external checkout: export VCPKG_ROOT=/path/to/vcpkg" >&2
exit 1 exit 1
fi fi
echo "[build-xcframework] VCPKG_ROOT=$VCPKG_ROOT" echo "[build-xcframework] VCPKG_ROOT=$VCPKG_ROOT"
@@ -104,14 +112,24 @@ build_slice() {
# to build protoc/etc at configure time. Merging those macOS .a files into the iOS # to build protoc/etc at configure time. Merging those macOS .a files into the iOS
# fat library triggers xcodebuild's "binaries with multiple platforms" rejection. # fat library triggers xcodebuild's "binaries with multiple platforms" rejection.
local vcpkg_target_lib_dir="$REPO_ROOT/build/$preset/vcpkg_installed/$vcpkg_triplet/lib" local vcpkg_target_lib_dir="$REPO_ROOT/build/$preset/vcpkg_installed/$vcpkg_triplet/lib"
local local_lib_dir="$REPO_ROOT/build/$preset/lib"
local fat_lib="$REPO_ROOT/build/$preset/lib/libvoicecat-fat.a" local fat_lib="$REPO_ROOT/build/$preset/lib/libvoicecat-fat.a"
echo "[build-xcframework] $slice_name: merging vcpkg deps into fat static lib (triplet: $vcpkg_triplet)" echo "[build-xcframework] $slice_name: merging vcpkg + vendored deps into fat static lib (triplet: $vcpkg_triplet)"
# Collect all .a files (libvoicecat.a + every vcpkg target .a). libtool -static concatenates # Collect all .a files (libvoicecat.a + every vcpkg target .a + locally-built vendored static
# object files from all input archives; duplicate-object warnings are benign (the linker # libs). libtool -static concatenates object files from all input archives; duplicate-object
# resolves duplicates at final link time). "no symbols" warnings are for empty AVX2/AVX512 # warnings are benign (the linker resolves duplicates at final link time). "no symbols"
# objects on arm64 — also benign. # warnings are for empty AVX2/AVX512 objects on arm64 — also benign.
#
# Two source dirs:
# - vcpkg_target_lib_dir: vcpkg-installed deps (protobuf, mbedtls, sodium, opus, …).
# - local_lib_dir: CMake static-lib targets built in this tree that are NOT vcpkg deps —
# e.g. the vendored RNNoise lib (third_party/rnnoise, VOICECAT_HAS_NS). Without it the
# fat lib references _rnnoise_* symbols that nothing defines → undefined-symbol link
# errors in Xcode. Exclude libvoicecat* (the main lib is $lib; libvoicecat-fat.a is the
# output we're building here).
local all_libs=( "$lib" ) local all_libs=( "$lib" )
while IFS= read -r f; do all_libs+=( "$f" ); done < <(find "$vcpkg_target_lib_dir" -name '*.a' -not -name 'libvoicecat*' | sort) while IFS= read -r f; do all_libs+=( "$f" ); done < <(find "$vcpkg_target_lib_dir" -name '*.a' -not -name 'libvoicecat*' | sort)
while IFS= read -r f; do all_libs+=( "$f" ); done < <(find "$local_lib_dir" -maxdepth 1 -name '*.a' -not -name 'libvoicecat*' | sort)
libtool -static -o "$fat_lib" "${all_libs[@]}" 2>&1 | grep -v 'has no symbols' || true libtool -static -o "$fat_lib" "${all_libs[@]}" 2>&1 | grep -v 'has no symbols' || true
echo "[build-xcframework] $slice_name -> $fat_lib ($(stat -f%z "$fat_lib") bytes, fat)" echo "[build-xcframework] $slice_name -> $fat_lib ($(stat -f%z "$fat_lib") bytes, fat)"
} }

View File

@@ -6,91 +6,161 @@ namespace VoiceCat.App.Audio;
// ── Scope types (mirror macOS ScreenAudioScope / ScreenAudioSelection) ──────── // ── Scope types (mirror macOS ScreenAudioScope / ScreenAudioSelection) ────────
public abstract record AppAudioScope; public abstract record AppAudioScope;
public sealed record EntireDesktop : AppAudioScope; // ExcludeSelf=true captures the whole system render mix minus VoiceCat's own process tree
// (kills the self-echo loop). Routed through the external-feed mixer as a single EXCLUDE
// capture; ExcludeSelf=false keeps the core's whole-device loopback path.
public sealed record EntireDesktop(bool ExcludeSelf = false) : AppAudioScope;
public sealed record OnlyApps(IReadOnlyList<int> Pids, IReadOnlyList<string> Names) : AppAudioScope; public sealed record OnlyApps(IReadOnlyList<int> Pids, IReadOnlyList<string> Names) : AppAudioScope;
public sealed record AllExceptApps(IReadOnlyList<int> Pids, IReadOnlyList<string> Names) : AppAudioScope; public sealed record AllExceptApps(IReadOnlyList<int> Pids, IReadOnlyList<string> Names) : AppAudioScope;
public sealed record AudioAppInfo(int Pid, string DisplayName); // IsPlaying = a process in this app's executable group currently has an *active* audio
// session on some render endpoint. Purely informational for the picker — capture works
// on any PID regardless (process-loopback yields silence until the app plays).
public sealed record AudioAppInfo(int Pid, string DisplayName, bool IsPlaying);
// ── Enumerator ───────────────────────────────────────────────────────────────── // ── Enumerator ─────────────────────────────────────────────────────────────────
public static class AudioSessionEnumerator public static class AudioSessionEnumerator
{ {
// Returns all user-facing apps: visible-window processes (primary, like macOS // Returns every app in the user's interactive session — windowed or not — so any of
// SCShareableContent.current) plus any background audio-session-only processes // them can be picked for capture even before it starts producing audio. Process
// (e.g. Spotify in mini-player). Excludes VoiceCat itself and system processes. // loopback (see ProcessLoopbackCapture) targets a PID and its child tree, so a silent
// selection simply starts working the moment that app plays.
//
// Apps are deduped by executable (multiple PIDs of the same program collapse to one
// row whose PID is the process-tree root). Session-0 services and VoiceCat itself are
// excluded. Apps currently producing audio are flagged IsPlaying and sorted first.
public static IReadOnlyList<AudioAppInfo> GetAudioApps() public static IReadOnlyList<AudioAppInfo> GetAudioApps()
{ {
var seen = new HashSet<int>(); int selfPid = Environment.ProcessId;
var result = new List<AudioAppInfo>(); int sessionId = SafeCurrentSessionId();
int selfPid = Environment.ProcessId; var playingPids = GetActiveAudioPids();
// ── 1. Visible-window processes (EnumWindows) ───────────────────────── // Group by executable name; keep the best representative PID per group.
// Same set macOS ScreenCaptureKit exposes: all apps with at least one var groups = new Dictionary<string, AppGroup>(StringComparer.OrdinalIgnoreCase);
// visible top-level window. Shows apps even when not currently producing audio.
EnumWindows((hWnd, _) => foreach (var proc in Process.GetProcesses())
{ {
if (!IsWindowVisible(hWnd)) return true;
GetWindowThreadProcessId(hWnd, out uint pid);
if (pid == 0 || pid == (uint)selfPid || !seen.Add((int)pid)) return true;
try try
{ {
var proc = Process.GetProcessById((int)pid); if (proc.Id == selfPid) continue;
string name = proc.MainWindowTitle.Length > 0 if (proc.SessionId != sessionId) continue; // drop session-0 services
? $"{proc.ProcessName} — {proc.MainWindowTitle}" string name = proc.ProcessName;
: proc.ProcessName; if (string.IsNullOrEmpty(name)) continue;
if (!string.IsNullOrEmpty(proc.ProcessName))
result.Add(new AudioAppInfo((int)pid, name)); bool hasWindow = proc.MainWindowHandle != IntPtr.Zero;
bool playing = playingPids.Contains(proc.Id);
string title = hasWindow ? SafeWindowTitle(proc) : "";
string display = title.Length > 0 ? $"{name} — {title}" : name;
if (groups.TryGetValue(name, out var g))
{
g.AnyPlaying |= playing;
// Prefer a windowed PID (the process-tree root) as the capture target;
// among non-windowed, prefer a currently-playing PID.
bool better = (hasWindow && !g.HasWindow)
|| (!g.HasWindow && !hasWindow && playing && !g.RepPlaying);
if (better)
{
g.Pid = proc.Id;
g.Display = display;
g.HasWindow = hasWindow;
g.RepPlaying = playing;
}
}
else
{
groups[name] = new AppGroup
{
Pid = proc.Id,
Display = display,
HasWindow = hasWindow,
RepPlaying = playing,
AnyPlaying = playing,
};
}
} }
catch { /* process exited between EnumWindows and GetProcessById */ } catch { /* protected/exited process — skip */ }
finally { proc.Dispose(); }
}
return true; // continue enumeration return groups.Values
}, IntPtr.Zero); .Select(g => new AudioAppInfo(g.Pid, g.Display, g.AnyPlaying))
.OrderByDescending(a => a.IsPlaying)
// ── 2. Background audio-session processes (WASAPI, supplement) ──────── .ThenBy(a => a.DisplayName, StringComparer.OrdinalIgnoreCase)
// Catches apps that produce audio but have no visible window (screen reader, .ToList();
// background music player, etc.). Silently skipped if WASAPI is unavailable.
AppendAudioSessionApps(seen, selfPid, result);
return result.OrderBy(a => a.DisplayName, StringComparer.OrdinalIgnoreCase).ToList();
} }
// ── EnumWindows P/Invoke ────────────────────────────────────────────────── private sealed class AppGroup
private delegate bool EnumWindowsProc(IntPtr hWnd, IntPtr lParam);
[DllImport("user32.dll")]
private static extern bool EnumWindows(EnumWindowsProc lpEnumFunc, IntPtr lParam);
[DllImport("user32.dll")]
private static extern bool IsWindowVisible(IntPtr hWnd);
[DllImport("user32.dll")]
private static extern uint GetWindowThreadProcessId(IntPtr hWnd, out uint lpdwProcessId);
// ── WASAPI audio session supplement ──────────────────────────────────────
private static void AppendAudioSessionApps(HashSet<int> seen, int selfPid,
List<AudioAppInfo> result)
{ {
public int Pid;
public string Display = "";
public bool HasWindow;
public bool RepPlaying; // representative PID is playing
public bool AnyPlaying; // any PID in the group is playing
}
private static int SafeCurrentSessionId()
{
try { using var me = Process.GetCurrentProcess(); return me.SessionId; }
catch { return 1; } // typical interactive session fallback
}
private static string SafeWindowTitle(Process proc)
{
try { return proc.MainWindowTitle; }
catch { return ""; }
}
// ── WASAPI: PIDs with an active render session (any endpoint) ─────────────
// Scans ALL active render endpoints, not just the default — an app routed to a
// secondary device still counts as "playing now".
private static HashSet<int> GetActiveAudioPids()
{
var pids = new HashSet<int>();
IMMDeviceEnumerator? enumerator = null; IMMDeviceEnumerator? enumerator = null;
IMMDevice? device = null; IMMDeviceCollection? devices = null;
IAudioSessionManager2? manager = null;
IAudioSessionEnumerator? sessions = null;
try try
{ {
enumerator = (IMMDeviceEnumerator)Activator.CreateInstance( enumerator = (IMMDeviceEnumerator)Activator.CreateInstance(
Type.GetTypeFromCLSID(new Guid("BCDE0395-E52F-467C-8E3D-C4579291692E"))!)!; Type.GetTypeFromCLSID(new Guid("BCDE0395-E52F-467C-8E3D-C4579291692E"))!)!;
enumerator.GetDefaultAudioEndpoint(0 /*eRender*/, 1 /*eMultimedia*/, out device); if (enumerator.EnumAudioEndpoints(0 /*eRender*/, 0x1 /*DEVICE_STATE_ACTIVE*/,
out devices) < 0 || devices == null)
return pids;
devices.GetCount(out int devCount);
for (int d = 0; d < devCount; d++)
CollectActivePids(devices, d, pids);
}
catch { /* no audio device or WASAPI unavailable — ignore */ }
finally
{
if (devices != null) Marshal.ReleaseComObject(devices);
if (enumerator != null) Marshal.ReleaseComObject(enumerator);
}
return pids;
}
private static void CollectActivePids(IMMDeviceCollection devices, int index, HashSet<int> pids)
{
IMMDevice? device = null;
IAudioSessionManager2? manager = null;
IAudioSessionEnumerator? sessions = null;
try
{
if (devices.Item(index, out device) < 0 || device == null) return;
var mgr2Iid = new Guid("77AA99A0-1BD6-484F-8BC7-2C654C9A9B6F"); var mgr2Iid = new Guid("77AA99A0-1BD6-484F-8BC7-2C654C9A9B6F");
device.Activate(ref mgr2Iid, 0x17 /*CLSCTX_ALL*/, IntPtr.Zero, out object mgr); if (device.Activate(ref mgr2Iid, 0x17 /*CLSCTX_ALL*/, IntPtr.Zero, out object mgr) < 0)
return;
manager = (IAudioSessionManager2)mgr; manager = (IAudioSessionManager2)mgr;
manager.GetSessionEnumerator(out sessions); if (manager.GetSessionEnumerator(out sessions) < 0 || sessions == null) return;
sessions.GetCount(out int count); sessions.GetCount(out int count);
for (int i = 0; i < count; i++) for (int i = 0; i < count; i++)
@@ -98,32 +168,24 @@ public static class AudioSessionEnumerator
IAudioSessionControl? ctrl = null; IAudioSessionControl? ctrl = null;
try try
{ {
sessions.GetSession(i, out ctrl); if (sessions.GetSession(i, out ctrl) < 0 || ctrl == null) continue;
ctrl.GetState(out int state);
if (state != 1 /*AudioSessionStateActive*/) continue;
var ctrl2 = (IAudioSessionControl2)ctrl; var ctrl2 = (IAudioSessionControl2)ctrl;
ctrl2.GetProcessId(out uint pid); ctrl2.GetProcessId(out uint pid);
if (pid != 0) pids.Add((int)pid);
int ipid = (int)pid;
if (pid == 0 || ipid == selfPid || !seen.Add(ipid)) continue;
try
{
var proc = Process.GetProcessById(ipid);
if (!string.IsNullOrEmpty(proc.ProcessName))
result.Add(new AudioAppInfo(ipid, proc.ProcessName));
}
catch { /* exited */ }
} }
catch { /* stale session */ } catch { /* stale session */ }
finally { if (ctrl != null) Marshal.ReleaseComObject(ctrl); } finally { if (ctrl != null) Marshal.ReleaseComObject(ctrl); }
} }
} }
catch { /* no audio device or WASAPI unavailable — ignore */ } catch { /* device went away */ }
finally finally
{ {
if (sessions != null) Marshal.ReleaseComObject(sessions); if (sessions != null) Marshal.ReleaseComObject(sessions);
if (manager != null) Marshal.ReleaseComObject(manager); if (manager != null) Marshal.ReleaseComObject(manager);
if (device != null) Marshal.ReleaseComObject(device); if (device != null) Marshal.ReleaseComObject(device);
if (enumerator != null) Marshal.ReleaseComObject(enumerator);
} }
} }
@@ -133,13 +195,21 @@ public static class AudioSessionEnumerator
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)] InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IMMDeviceEnumerator internal interface IMMDeviceEnumerator
{ {
[PreserveSig] int EnumAudioEndpoints(int dataFlow, int stateMask, out IntPtr devices); [PreserveSig] int EnumAudioEndpoints(int dataFlow, int stateMask, out IMMDeviceCollection devices);
[PreserveSig] int GetDefaultAudioEndpoint(int dataFlow, int role, out IMMDevice endpoint); [PreserveSig] int GetDefaultAudioEndpoint(int dataFlow, int role, out IMMDevice endpoint);
[PreserveSig] int GetDevice([MarshalAs(UnmanagedType.LPWStr)] string id, out IMMDevice device); [PreserveSig] int GetDevice([MarshalAs(UnmanagedType.LPWStr)] string id, out IMMDevice device);
[PreserveSig] int RegisterEndpointNotificationCallback(IntPtr client); [PreserveSig] int RegisterEndpointNotificationCallback(IntPtr client);
[PreserveSig] int UnregisterEndpointNotificationCallback(IntPtr client); [PreserveSig] int UnregisterEndpointNotificationCallback(IntPtr client);
} }
[ComImport, Guid("0BD7A1BE-7A1A-44DB-8397-CC5392387B5E"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IMMDeviceCollection
{
[PreserveSig] int GetCount(out int count);
[PreserveSig] int Item(int index, out IMMDevice device);
}
[ComImport, Guid("D666063F-1587-4E43-81F1-B948E807363F"), [ComImport, Guid("D666063F-1587-4E43-81F1-B948E807363F"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)] InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IMMDevice internal interface IMMDevice

View File

@@ -0,0 +1,443 @@
using System.Runtime.InteropServices;
// Captures a second hardware input for AUX_DEVICE and feeds it through vc_stream_feed_pcm.
// Endpoint identifiers are WASAPI-specific and cannot be exchanged with the core's miniaudio ids.
namespace VoiceCat.App.Audio;
/// <summary>An audio input (capture) endpoint for the aux-stream device picker. <see cref="Id"/>
/// is a WASAPI endpoint id (round-trip only; never construct by hand) — pass null to capture the
/// system default. <see cref="ToString"/> returns the friendly name for ComboBox display.</summary>
public sealed record InputDeviceInfo(string Id, string Name, bool IsDefault)
{
public override string ToString() => Name;
}
/// <summary>Enumerates WASAPI capture endpoints. Separate from the core's vc_list_devices because
/// the aux device is opened client-side and needs a WASAPI id, not a miniaudio one.</summary>
public static class InputDeviceEnumerator
{
public static IReadOnlyList<InputDeviceInfo> List()
{
var result = new List<InputDeviceInfo>();
InputDeviceCapture.IMMDeviceEnumerator? enumerator = null;
IntPtr collectionPtr = IntPtr.Zero;
string? defaultId = null;
try
{
enumerator = (InputDeviceCapture.IMMDeviceEnumerator)Activator.CreateInstance(
Type.GetTypeFromCLSID(new Guid("BCDE0395-E52F-467C-8E3D-C4579291692E"))!)!;
// Resolve the default capture endpoint id so the picker can flag it.
if (enumerator.GetDefaultAudioEndpoint(1 /*eCapture*/, 0 /*eConsole*/,
out var defDev) == 0 && defDev != null)
{
try { if (defDev.GetId(out string id) == 0) defaultId = id; }
finally { Marshal.ReleaseComObject(defDev); }
}
// DEVICE_STATE_ACTIVE = 0x1 — only currently-usable endpoints.
if (enumerator.EnumAudioEndpoints(1 /*eCapture*/, 0x1, out collectionPtr) != 0
|| collectionPtr == IntPtr.Zero)
return result;
var collection = (InputDeviceCapture.IMMDeviceCollection)
Marshal.GetObjectForIUnknown(collectionPtr);
collection.GetCount(out int count);
for (int i = 0; i < count; i++)
{
if (collection.Item(i, out var dev) != 0 || dev == null) continue;
try
{
if (dev.GetId(out string id) != 0) continue;
string name = ReadFriendlyName(dev) ?? "Unknown input device";
result.Add(new InputDeviceInfo(id, name, id == defaultId));
}
finally { Marshal.ReleaseComObject(dev); }
}
}
catch { /* no audio subsystem / WASAPI unavailable — return what we have */ }
finally
{
if (collectionPtr != IntPtr.Zero) Marshal.Release(collectionPtr);
if (enumerator != null) Marshal.ReleaseComObject(enumerator);
}
return result.OrderBy(d => d.Name, StringComparer.OrdinalIgnoreCase).ToList();
}
private static string? ReadFriendlyName(InputDeviceCapture.IMMDevice dev)
{
if (dev.OpenPropertyStore(0 /*STGM_READ*/, out IntPtr storePtr) != 0
|| storePtr == IntPtr.Zero)
return null;
try
{
var store = (InputDeviceCapture.IPropertyStore)Marshal.GetObjectForIUnknown(storePtr);
// PKEY_Device_FriendlyName = {a45c254e-df1c-4efd-8020-67d146a850e0}, pid 14.
var key = new InputDeviceCapture.PropertyKey
{
fmtid = new Guid("a45c254e-df1c-4efd-8020-67d146a850e0"),
pid = 14,
};
if (store.GetValue(ref key, out var pv) != 0) return null;
try
{
// VT_LPWSTR = 31.
return pv.vt == 31 ? Marshal.PtrToStringUni(pv.pointerValue) : null;
}
finally { InputDeviceCapture.PropVariantClear(ref pv); }
}
finally { Marshal.Release(storePtr); }
}
}
/// <summary>Captures a single hardware input device in WASAPI shared mode and raises a 20 ms
/// (960 samples/channel @ 48 kHz, interleaved s16) frame event. The caller feeds these to the
/// core's external-feed stream. Threading mirrors ProcessLoopbackCapture: all WASAPI work runs on
/// a dedicated background (MTA) thread; the event fires on that thread.</summary>
public sealed class InputDeviceCapture : IDisposable
{
/// <summary>Fired on the capture thread every 20 ms: (interleaved s16 PCM, samplesPerChannel
/// = 960, channels).</summary>
public event Action<short[], int /*samplesPerChannel*/, int /*channels*/>? PcmFrameReady;
private const int SampleRate = 48000;
private const int FrameSamples = 960; // 20 ms
private readonly string? _deviceId; // null = system default capture endpoint
private IAudioClient? _audioClient;
private IAudioCaptureClient? _captureClient;
private AutoResetEvent? _bufferEvent;
private Thread? _captureThread;
private volatile bool _running;
private int _channels;
// Accumulator: assembles driver-callback-sized fragments into FrameSamples chunks.
private short[] _accumBuf = [];
private int _accumCount;
// Init-done signal: Set() by the capture thread after activation completes.
private readonly ManualResetEventSlim _initDone = new(false);
private bool _initOk;
public InputDeviceCapture(string? deviceId) => _deviceId = deviceId;
/// <summary>Starts capture. Blocks until WASAPI activation completes (typically &lt;100 ms).
/// Returns false if the device cannot be opened.</summary>
public bool Start()
{
if (_running) return false;
_running = true;
_captureThread = new Thread(CaptureThreadProc)
{
IsBackground = true,
Name = "AuxInputCapture",
};
_captureThread.Start();
bool ok = _initDone.Wait(5000) && _initOk;
if (!ok) _running = false;
return ok;
}
public void Stop()
{
_running = false;
_bufferEvent?.Set();
_captureThread?.Join(500);
try { _audioClient?.Stop(); } catch { /* device already gone */ }
}
public void Dispose()
{
Stop();
if (_captureClient != null) { Marshal.ReleaseComObject(_captureClient); _captureClient = null; }
if (_audioClient != null) { Marshal.ReleaseComObject(_audioClient); _audioClient = null; }
_bufferEvent?.Dispose();
_initDone.Dispose();
}
// ── Capture thread (MTA) ──────────────────────────────────────────────────
private void CaptureThreadProc()
{
_initOk = ActivateAndStart();
_initDone.Set();
if (!_initOk) return;
CaptureLoop();
}
private bool ActivateAndStart()
{
if (!ActivateClient()) return false;
_bufferEvent = new AutoResetEvent(false);
if (_audioClient!.SetEventHandle(_bufferEvent.SafeWaitHandle.DangerousGetHandle()) < 0)
return false;
return _audioClient.Start() >= 0;
}
private bool ActivateClient()
{
IMMDeviceEnumerator? enumerator = null;
IMMDevice? device = null;
try
{
enumerator = (IMMDeviceEnumerator)Activator.CreateInstance(
Type.GetTypeFromCLSID(new Guid("BCDE0395-E52F-467C-8E3D-C4579291692E"))!)!;
int hr = _deviceId is null
? enumerator.GetDefaultAudioEndpoint(1 /*eCapture*/, 0 /*eConsole*/, out device)
: enumerator.GetDevice(_deviceId, out device);
if (hr != 0 || device == null) return false;
var iidAudioClient = new Guid("1CB9AD4C-DBFA-4c32-B178-C2F568A703B2");
if (device.Activate(ref iidAudioClient, 0x17 /*CLSCTX_ALL*/, IntPtr.Zero,
out object acObj) != 0 || acObj is not IAudioClient ac)
return false;
_audioClient = ac;
return InitializeStream();
}
catch { return false; }
finally
{
if (device != null) Marshal.ReleaseComObject(device);
if (enumerator != null) Marshal.ReleaseComObject(enumerator);
}
}
private bool InitializeStream()
{
// Try s16 stereo first; fall back to s16 mono. AUTOCONVERTPCM lets the audio engine
// resample/convert the device's native format to our requested 48 kHz s16; EVENTCALLBACK
// drives the buffer-ready event. Shared mode (0), no LOOPBACK (this is a capture device).
// AUDCLNT_STREAMFLAGS_EVENTCALLBACK = 0x00040000
// AUDCLNT_STREAMFLAGS_AUTOCONVERTPCM = 0x80000000
// AUDCLNT_STREAMFLAGS_SRC_DEFAULT_QUALITY = 0x08000000
const uint streamFlags = 0x00040000u | 0x80000000u | 0x08000000u;
foreach (int ch in new[] { 2, 1 })
{
var fmt = new WaveFormatEx
{
wFormatTag = 1, // WAVE_FORMAT_PCM
nChannels = (ushort)ch,
nSamplesPerSec = SampleRate,
wBitsPerSample = 16,
nBlockAlign = (ushort)(ch * 2),
nAvgBytesPerSec = (uint)(SampleRate * ch * 2),
cbSize = 0,
};
IntPtr pFmt = Marshal.AllocHGlobal(Marshal.SizeOf<WaveFormatEx>());
try
{
Marshal.StructureToPtr(fmt, pFmt, false);
int hr = _audioClient!.Initialize(0 /*AUDCLNT_SHAREMODE_SHARED*/, streamFlags,
2_000_000 /*200 ms hns*/, 0, pFmt, IntPtr.Zero);
if (hr >= 0)
{
_channels = ch;
_accumBuf = new short[FrameSamples * ch];
_accumCount = 0;
var iidCapture = new Guid("C8ADBD64-E71E-48a0-A4DE-185C395CD317");
if (_audioClient.GetService(ref iidCapture, out object ccObj) != 0
|| ccObj is not IAudioCaptureClient cc)
return false;
_captureClient = cc;
return true;
}
if (ch == 1) return false;
}
finally { Marshal.FreeHGlobal(pFmt); }
}
return false;
}
// ── Capture loop ─────────────────────────────────────────────────────────
private void CaptureLoop()
{
while (_running)
{
_bufferEvent!.WaitOne(100);
if (!_running) break;
while (_running)
{
if (_captureClient!.GetNextPacketSize(out uint packetSize) < 0 || packetSize == 0)
break;
if (_captureClient.GetBuffer(out IntPtr dataPtr, out uint framesAvailable,
out uint flags, out _, out _) < 0)
break;
bool silent = (flags & 2) != 0; // AUDCLNT_BUFFERFLAGS_SILENT
if (framesAvailable > 0)
{
if (silent) AccumulateSilence((int)framesAvailable);
else AccumulatePcm(dataPtr, (int)framesAvailable);
}
_captureClient.ReleaseBuffer(framesAvailable);
}
}
}
private unsafe void AccumulatePcm(IntPtr data, int frames)
{
var src = (short*)data.ToPointer();
int total = frames * _channels;
int idx = 0;
while (idx < total)
{
int space = _accumBuf.Length - _accumCount;
int copy = Math.Min(total - idx, space);
fixed (short* dst = _accumBuf)
Buffer.MemoryCopy(src + idx, dst + _accumCount, copy * 2L, copy * 2L);
_accumCount += copy;
idx += copy;
if (_accumCount == _accumBuf.Length)
FlushFrame();
}
}
private void AccumulateSilence(int frames)
{
int total = frames * _channels;
int idx = 0;
while (idx < total)
{
int space = _accumBuf.Length - _accumCount;
int fill = Math.Min(total - idx, space);
Array.Clear(_accumBuf, _accumCount, fill);
_accumCount += fill;
idx += fill;
if (_accumCount == _accumBuf.Length)
FlushFrame();
}
}
private void FlushFrame()
{
var copy = new short[_accumBuf.Length];
_accumBuf.AsSpan().CopyTo(copy);
PcmFrameReady?.Invoke(copy, FrameSamples, _channels);
_accumCount = 0;
}
// ── COM declarations ───────────────────────────────────────────────────────
//
// Declared internal here so InputDeviceEnumerator can share them. These are standard MMDevice
// / WASAPI interfaces; a normal capture endpoint honours QueryInterface, so RCW marshalling is
// safe (unlike ProcessLoopbackCapture's process-loopback objects, which need raw vtable calls).
[ComImport, Guid("A95664D2-9614-4F35-A746-DE8DB63617E6"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IMMDeviceEnumerator
{
[PreserveSig] int EnumAudioEndpoints(int dataFlow, int stateMask, out IntPtr devices);
[PreserveSig] int GetDefaultAudioEndpoint(int dataFlow, int role, out IMMDevice endpoint);
[PreserveSig] int GetDevice([MarshalAs(UnmanagedType.LPWStr)] string id, out IMMDevice device);
[PreserveSig] int RegisterEndpointNotificationCallback(IntPtr client);
[PreserveSig] int UnregisterEndpointNotificationCallback(IntPtr client);
}
[ComImport, Guid("0BD7A1BE-7A1A-44DB-8397-CC5392387B5E"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IMMDeviceCollection
{
[PreserveSig] int GetCount(out int count);
[PreserveSig] int Item(int index, out IMMDevice device);
}
[ComImport, Guid("D666063F-1587-4E43-81F1-B948E807363F"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IMMDevice
{
[PreserveSig] int Activate(ref Guid iid, int clsCtx, IntPtr activationParams,
[MarshalAs(UnmanagedType.IUnknown)] out object ppInterface);
[PreserveSig] int OpenPropertyStore(int stgmAccess, out IntPtr propStore);
[PreserveSig] int GetId([MarshalAs(UnmanagedType.LPWStr)] out string id);
[PreserveSig] int GetState(out int state);
}
[ComImport, Guid("886D8EEB-8CF2-4446-8D02-CDBA1DBDCF99"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IPropertyStore
{
[PreserveSig] int GetCount(out int count);
[PreserveSig] int GetAt(int index, out PropertyKey key);
[PreserveSig] int GetValue(ref PropertyKey key, out PropVariant value);
[PreserveSig] int SetValue(ref PropertyKey key, ref PropVariant value);
[PreserveSig] int Commit();
}
[ComImport, Guid("1CB9AD4C-DBFA-4C32-B178-C2F568A703B2"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IAudioClient
{
[PreserveSig] int Initialize(int shareMode, uint streamFlags, long hnsBufferDuration,
long hnsPeriodicity, IntPtr pFormat, IntPtr audioSessionGuid);
[PreserveSig] int GetBufferSize(out uint numBufferFrames);
[PreserveSig] int GetStreamLatency(out long latency);
[PreserveSig] int GetCurrentPadding(out uint numPaddingFrames);
[PreserveSig] int IsFormatSupported(int shareMode, IntPtr pFormat, out IntPtr closestMatch);
[PreserveSig] int GetMixFormat(out IntPtr deviceFormat);
[PreserveSig] int GetDevicePeriod(out long defaultDevicePeriod, out long minimumDevicePeriod);
[PreserveSig] int Start();
[PreserveSig] int Stop();
[PreserveSig] int Reset();
[PreserveSig] int SetEventHandle(IntPtr eventHandle);
[PreserveSig] int GetService(ref Guid riid,
[MarshalAs(UnmanagedType.IUnknown)] out object ppv);
}
[ComImport, Guid("C8ADBD64-E71E-48A0-A4DE-185C395CD317"),
InterfaceType(ComInterfaceType.InterfaceIsIUnknown)]
internal interface IAudioCaptureClient
{
[PreserveSig] int GetBuffer(out IntPtr data, out uint numFramesToRead, out uint flags,
out ulong devicePosition, out ulong qpcPosition);
[PreserveSig] int ReleaseBuffer(uint numFramesRead);
[PreserveSig] int GetNextPacketSize(out uint numFramesInNextPacket);
}
[StructLayout(LayoutKind.Sequential, Pack = 2)]
private struct WaveFormatEx
{
public ushort wFormatTag;
public ushort nChannels;
public uint nSamplesPerSec;
public uint nAvgBytesPerSec;
public ushort nBlockAlign;
public ushort wBitsPerSample;
public ushort cbSize;
}
[StructLayout(LayoutKind.Sequential)]
internal struct PropertyKey
{
public Guid fmtid;
public int pid;
}
// Minimal PROPVARIANT: we only ever read VT_LPWSTR (friendly name). x64 layout — the value
// union starts at offset 8 after vt(2)+reserved(6).
[StructLayout(LayoutKind.Explicit)]
internal struct PropVariant
{
[FieldOffset(0)] public ushort vt;
[FieldOffset(8)] public IntPtr pointerValue;
}
[DllImport("ole32.dll")]
internal static extern int PropVariantClear(ref PropVariant pvar);
}

View File

@@ -22,15 +22,14 @@ public sealed class ProcessAudioMixer : IDisposable
private VoiceCatClient? _client; private VoiceCatClient? _client;
private uint _streamId; private uint _streamId;
public void Start(AppAudioScope scope, VoiceCatClient client, uint streamId, public void Start(AppAudioScope scope, VoiceCatClient client, uint streamId)
IReadOnlyList<AudioAppInfo> allApps)
{ {
if (_running) return; if (_running) return;
_client = client; _client = client;
_streamId = streamId; _streamId = streamId;
var pids = ResolvePids(scope, allApps); var specs = ResolveCaptures(scope);
if (pids.Count == 0) if (specs.Count == 0)
{ {
// nothing to capture — scope resolved to empty set // nothing to capture — scope resolved to empty set
return; return;
@@ -38,14 +37,15 @@ public sealed class ProcessAudioMixer : IDisposable
lock (_frameLock) lock (_frameLock)
{ {
_latestFrames = new List<short[]>(new short[pids.Count][]); _latestFrames = new List<short[]>(new short[specs.Count][]);
_activeChannels = Channels; _activeChannels = Channels;
} }
for (int i = 0; i < pids.Count; i++) for (int i = 0; i < specs.Count; i++)
{ {
int captureIndex = i; int captureIndex = i;
var cap = new ProcessLoopbackCapture(pids[i], ProcessLoopbackCapture.Mode.Include); var (pid, mode) = specs[i];
var cap = new ProcessLoopbackCapture(pid, mode);
cap.PcmFrameReady += (pcm, spc, ch) => OnCaptureFrame(captureIndex, pcm, ch); cap.PcmFrameReady += (pcm, spc, ch) => OnCaptureFrame(captureIndex, pcm, ch);
_captures.Add(cap); _captures.Add(cap);
} }
@@ -105,7 +105,6 @@ public sealed class ProcessAudioMixer : IDisposable
for (int i = 0; i < frameLen; i++) for (int i = 0; i < frameLen; i++)
{ {
int sum = mix[i] + frame[i]; int sum = mix[i] + frame[i];
// Saturating clamp
mix[i] = (short)Math.Clamp(sum, short.MinValue, short.MaxValue); mix[i] = (short)Math.Clamp(sum, short.MinValue, short.MaxValue);
} }
} }
@@ -117,18 +116,21 @@ public sealed class ProcessAudioMixer : IDisposable
// ── Helpers ─────────────────────────────────────────────────────────────── // ── Helpers ───────────────────────────────────────────────────────────────
// For "AllExcept": enumerate all running audio apps and exclude the specified ones. // Resolve the scope to the WASAPI captures to open:
// For "OnlyApps": use their PIDs directly. // OnlyApps → one INCLUDE capture per selected process tree.
private static List<int> ResolvePids(AppAudioScope scope, IReadOnlyList<AudioAppInfo> allApps) // AllExceptApps → one EXCLUDE capture of the single selected process tree, which
// natively captures the whole system render mix minus that tree
// (dynamic — apps launched later are included automatically).
private static List<(int pid, ProcessLoopbackCapture.Mode mode)> ResolveCaptures(AppAudioScope scope)
{ {
return scope switch return scope switch
{ {
OnlyApps o => [.. o.Pids], OnlyApps o => o.Pids.Select(p => (p, ProcessLoopbackCapture.Mode.Include)).ToList(),
AllExceptApps a => AllExceptApps a when a.Pids.Count > 0 =>
allApps [(a.Pids[0], ProcessLoopbackCapture.Mode.Exclude)],
.Where(app => !a.Pids.Contains(app.Pid)) // Entire desktop minus VoiceCat itself: EXCLUDE our own process tree.
.Select(app => app.Pid) EntireDesktop { ExcludeSelf: true } =>
.ToList(), [(Environment.ProcessId, ProcessLoopbackCapture.Mode.Exclude)],
_ => [], _ => [],
}; };
} }

View File

@@ -1,17 +1,8 @@
using System.Runtime.InteropServices; using System.Runtime.InteropServices;
// Single-process WASAPI loopback capture via AUDIOCLIENT_ACTIVATION_PARAMS // Process loopback requires MTA activation; blocking activation from the WinForms STA
// (Windows 10 2004+ / Build 19041+). // deadlocks COM completion. The returned interfaces also reject RCW QueryInterface, so audio
// // calls use explicitly owned raw pointers and vtable dispatch.
// Threading: ALL WASAPI init runs on the capture thread (MTA). If called from the
// WinForms UI thread (STA), ActivateAudioInterfaceAsync fires ActivateCompleted on
// an MTA pool thread; COM marshals that back to the STA pump — but the STA thread is
// blocked on CompletionEvent.Wait → deadlock. MTA capture thread avoids this.
//
// COM QI policy: the COM objects returned by the process-loopback activation path
// reject QueryInterface for their own IIDs under .NET's RCW mechanism. Every call
// to IAudioClient and IAudioCaptureClient is therefore dispatched via raw vtable
// pointer arithmetic, bypassing .NET COM interop entirely.
namespace VoiceCat.App.Audio; namespace VoiceCat.App.Audio;
public sealed class ProcessLoopbackCapture : IDisposable public sealed class ProcessLoopbackCapture : IDisposable

View File

@@ -12,12 +12,18 @@ public sealed class AppAudioPickerDialog : Form
private readonly RadioButton _rdoAll; private readonly RadioButton _rdoAll;
private readonly RadioButton _rdoOnly; private readonly RadioButton _rdoOnly;
private readonly RadioButton _rdoExcept; private readonly RadioButton _rdoExcept;
private readonly CheckBox _chkExcludeSelf;
private readonly TextBox _txtFilter;
private readonly ListView _appList; private readonly ListView _appList;
private readonly Label _lblApps; private readonly Label _lblApps;
// Snapshot taken when the dialog opens (refresh on open, not on every check change). // Snapshot taken when the dialog opens (refresh on open, not on every check change).
private IReadOnlyList<AudioAppInfo> _apps = []; private IReadOnlyList<AudioAppInfo> _apps = [];
// Checked apps survive filtering (a filtered-out row keeps its check here).
// pid → clean display name (used for the shared scope's Names).
private readonly Dictionary<int, string> _checked = new();
public AppAudioScope? ChosenScope { get; private set; } public AppAudioScope? ChosenScope { get; private set; }
public AppAudioPickerDialog() public AppAudioPickerDialog()
@@ -50,51 +56,78 @@ public sealed class AppAudioPickerDialog : Form
_rdoOnly.CheckedChanged += OnModeChanged; _rdoOnly.CheckedChanged += OnModeChanged;
_rdoExcept.CheckedChanged += OnModeChanged; _rdoExcept.CheckedChanged += OnModeChanged;
// ── App list ──────────────────────────────────────────────────────── // ── Self-exclude (echo prevention) ───────────────────────────────────
_lblApps = new Label // Captures the whole desktop minus VoiceCat's own playback, so the channel
// doesn't echo back the voices you're already hearing. Only applies to the
// entire-desktop mode (the per-app modes already exclude this app's tree).
_chkExcludeSelf = new CheckBox
{ {
Text = "Apps with active audio sessions:", Text = "E&xclude VoiceCat's own audio (prevents echo)",
AutoSize = true, Checked = true,
Location = new Point(12, 90), Location = new Point(28, 84),
Visible = false, Size = new Size(344, 20),
TabIndex = 3, TabIndex = 3,
}; };
// ── App list ────────────────────────────────────────────────────────
_lblApps = new Label
{
Text = "Apps (▶ = currently playing):",
AutoSize = true,
Location = new Point(12, 112),
Visible = false,
TabIndex = 4,
};
_txtFilter = new TextBox
{
Location = new Point(12, 132),
Size = new Size(360, 23),
PlaceholderText = "Filter apps…",
Visible = false,
TabIndex = 5,
};
_txtFilter.TextChanged += (_, _) => PopulateList();
_appList = new ListView _appList = new ListView
{ {
Location = new Point(12, 112), Location = new Point(12, 160),
Size = new Size(360, 180), Size = new Size(360, 150),
CheckBoxes = true, CheckBoxes = true,
View = View.List, View = View.List,
Visible = false, Visible = false,
TabIndex = 4, TabIndex = 6,
FullRowSelect = true, FullRowSelect = true,
}; };
// Exclude mode supports only one target process tree (WASAPI EXCLUDE takes a single
// PID), so enforce a single check while "All apps except selected" is active.
_appList.ItemCheck += OnAppItemCheck;
// ── Buttons ───────────────────────────────────────────────────────── // ── Buttons ─────────────────────────────────────────────────────────
var btnShare = new Button var btnShare = new Button
{ {
Text = "&Share", Text = "&Share",
DialogResult = DialogResult.OK, DialogResult = DialogResult.OK,
Location = new Point(216, 308), Location = new Point(216, 320),
Size = new Size(75, 27), Size = new Size(75, 27),
TabIndex = 5, TabIndex = 7,
}; };
var btnCancel = new Button var btnCancel = new Button
{ {
Text = "&Cancel", Text = "&Cancel",
DialogResult = DialogResult.Cancel, DialogResult = DialogResult.Cancel,
Location = new Point(297, 308), Location = new Point(297, 320),
Size = new Size(75, 27), Size = new Size(75, 27),
TabIndex = 6, TabIndex = 8,
}; };
btnShare.Click += OnShareClick; btnShare.Click += OnShareClick;
AcceptButton = btnShare; AcceptButton = btnShare;
CancelButton = btnCancel; CancelButton = btnCancel;
AutoScaleMode = AutoScaleMode.Font; AutoScaleMode = AutoScaleMode.Font;
ClientSize = new Size(384, 348); ClientSize = new Size(384, 360);
Controls.AddRange([_rdoAll, _rdoOnly, _rdoExcept, _lblApps, _appList, btnShare, btnCancel]); Controls.AddRange([_rdoAll, _rdoOnly, _rdoExcept, _chkExcludeSelf, _lblApps, _txtFilter,
_appList, btnShare, btnCancel]);
FormBorderStyle = FormBorderStyle.FixedDialog; FormBorderStyle = FormBorderStyle.FixedDialog;
MaximizeBox = false; MaximizeBox = false;
MinimizeBox = false; MinimizeBox = false;
@@ -111,41 +144,105 @@ public sealed class AppAudioPickerDialog : Form
private void RefreshAppList() private void RefreshAppList()
{ {
_apps = AudioSessionEnumerator.GetAudioApps(); _apps = AudioSessionEnumerator.GetAudioApps();
PopulateList();
}
// Rebuild the visible rows from _apps + the current filter text, restoring check
// state from _checked so a selection persists while the user filters.
private void PopulateList()
{
string filter = _txtFilter.Text.Trim();
_suppressItemCheck = true;
_appList.BeginUpdate();
_appList.Items.Clear(); _appList.Items.Clear();
foreach (var app in _apps) foreach (var app in _apps)
_appList.Items.Add(new ListViewItem($"{app.DisplayName} (PID {app.Pid})") { Tag = app.Pid }); {
if (filter.Length > 0 &&
app.DisplayName.IndexOf(filter, StringComparison.OrdinalIgnoreCase) < 0)
continue;
string marker = app.IsPlaying ? "▶ " : "";
_appList.Items.Add(new ListViewItem($"{marker}{app.DisplayName} (PID {app.Pid})")
{
Tag = app,
Checked = _checked.ContainsKey(app.Pid),
});
}
_appList.EndUpdate();
_suppressItemCheck = false;
} }
private void OnModeChanged(object? sender, EventArgs e) private void OnModeChanged(object? sender, EventArgs e)
{ {
bool showList = _rdoOnly.Checked || _rdoExcept.Checked; bool showList = _rdoOnly.Checked || _rdoExcept.Checked;
_lblApps.Visible = showList; _lblApps.Visible = showList;
_appList.Visible = showList; _txtFilter.Visible = showList;
_appList.Visible = showList;
_lblApps.Text = _rdoExcept.Checked
? "App to exclude (everything else is shared):"
: "Apps (▶ = currently playing):";
// Self-exclude only applies to entire-desktop; the per-app modes already exclude
// this app's own tree (Include) or spend the single EXCLUDE slot on the chosen app.
_chkExcludeSelf.Enabled = _rdoAll.Checked;
// Switching into exclude mode: collapse any multi-selection down to a single item.
if (_rdoExcept.Checked)
TrimCheckedToOne();
if (showList && _appList.Items.Count == 0) if (showList && _appList.Items.Count == 0)
RefreshAppList(); RefreshAppList();
else
PopulateList(); // re-render markers/labels and restore check state
}
// In exclude mode the list behaves like radio buttons: checking one item clears the rest.
private bool _suppressItemCheck;
private void OnAppItemCheck(object? sender, ItemCheckEventArgs e)
{
if (_suppressItemCheck) return;
if (_appList.Items[e.Index].Tag is not AudioAppInfo app) return;
if (e.NewValue == CheckState.Checked)
{
if (_rdoExcept.Checked)
{
// Radio behavior: clear every other visible check and the persisted set.
_suppressItemCheck = true;
foreach (ListViewItem item in _appList.Items)
if (item.Index != e.Index && item.Checked) item.Checked = false;
_suppressItemCheck = false;
_checked.Clear();
}
_checked[app.Pid] = app.DisplayName;
}
else
{
_checked.Remove(app.Pid);
}
}
// Reduce the persisted selection to a single (first) entry — used when entering
// exclude mode, whose single EXCLUDE slot can only target one process tree.
private void TrimCheckedToOne()
{
if (_checked.Count <= 1) return;
var first = _checked.First();
_checked.Clear();
_checked[first.Key] = first.Value;
} }
private void OnShareClick(object? sender, EventArgs e) private void OnShareClick(object? sender, EventArgs e)
{ {
if (_rdoAll.Checked) if (_rdoAll.Checked)
{ {
ChosenScope = new EntireDesktop(); ChosenScope = new EntireDesktop(_chkExcludeSelf.Checked);
return; return;
} }
var checkedPids = new List<int>(); if (_checked.Count == 0)
var checkedNames = new List<string>();
foreach (ListViewItem item in _appList.CheckedItems)
{
if (item.Tag is int pid)
{
checkedPids.Add(pid);
checkedNames.Add(item.Text);
}
}
if (checkedPids.Count == 0)
{ {
MessageBox.Show( MessageBox.Show(
"Select at least one app, or choose 'Entire desktop'.", "Select at least one app, or choose 'Entire desktop'.",
@@ -156,11 +253,11 @@ public sealed class AppAudioPickerDialog : Form
return; return;
} }
ChosenScope = _rdoOnly.Checked var pids = _checked.Keys.ToList();
? new OnlyApps(checkedPids, checkedNames) var names = _checked.Values.ToList();
: new AllExceptApps(checkedPids, checkedNames);
}
/// <summary>All apps that were visible in the list when the user clicked Share.</summary> ChosenScope = _rdoOnly.Checked
public IReadOnlyList<AudioAppInfo> VisibleApps => _apps; ? new OnlyApps(pids, names)
: new AllExceptApps(pids, names);
}
} }

View File

@@ -0,0 +1,574 @@
using System.ComponentModel;
using VoiceCat.App.Audio;
using VoiceCat.App.Models;
using VoiceCat.Interop;
namespace VoiceCat.App.Forms;
/// <summary>
/// Edits audio input settings: device selection, transmission mode (VAD/PTT/always-on),
/// VAD sensitivity, mic gain, and PTT key. Changes are applied live to the client for
/// immediate feedback; Cancel reverts them.
/// </summary>
public sealed class AudioSettingsForm : Form
{
private readonly VoiceCatClient _client;
private readonly VoiceSettings _settings;
private readonly uint _micStreamId;
// Aux-stream live-apply callbacks, supplied by MainForm (which owns the aux capture). Null
// when not connected — the controls still edit settings, just without live effect.
private readonly Action<bool>? _applyAuxEnabled;
private readonly Action<string?>? _applyAuxDevice;
private readonly Action<float>? _applyAuxGain;
// Snapshot of original values so Cancel can restore them
private readonly string? _origDeviceId;
private readonly VcInputMode _origMode;
private readonly int _origVadSlider;
private readonly int _origMicGain;
private readonly bool _origMicNoiseReduction;
private readonly bool _origStereoMic;
private readonly Keys _origPttKey;
private readonly bool _origSystemWidePtt;
private readonly bool _origAuxEnabled;
private readonly string? _origAuxDeviceId;
private readonly int _origAuxGain;
private readonly ComboBox _cboDevice;
private readonly Button _btnRefresh;
private readonly RadioButton _radioVad;
private readonly RadioButton _radioPtt;
private readonly RadioButton _radioAlwaysOn;
private readonly Label _lblPttKey;
private readonly Button _btnChangePtt;
private readonly CheckBox _chkSystemWidePtt;
private readonly Label _lblSensitivity;
private readonly TrackBar _trkVad;
private readonly TrackBar _trkGain;
private readonly CheckBox _chkNoiseReduction;
private readonly CheckBox _chkStereoMic;
private readonly CheckBox _chkAux;
private readonly Label _lblAuxDevice;
private readonly ComboBox _cboAuxDevice;
private readonly Button _btnAuxRefresh;
private readonly Label _lblAuxGain;
private readonly TrackBar _trkAuxGain;
private Keys _pttKey;
public AudioSettingsForm(VoiceCatClient client, VoiceSettings settings, uint micStreamId,
Action<bool>? applyAuxEnabled = null, Action<string?>? applyAuxDevice = null,
Action<float>? applyAuxGain = null)
{
_client = client;
_settings = settings;
_micStreamId = micStreamId;
_applyAuxEnabled = applyAuxEnabled;
_applyAuxDevice = applyAuxDevice;
_applyAuxGain = applyAuxGain;
_pttKey = (Keys)settings.PttKey;
_origDeviceId = settings.InputDeviceId;
_origMode = (VcInputMode)settings.InputMode;
_origVadSlider = settings.VadThresholdSlider;
_origMicGain = settings.MicGain;
_origMicNoiseReduction = settings.MicNoiseReduction;
_origStereoMic = settings.StereoMic;
_origPttKey = _pttKey;
_origSystemWidePtt = settings.SystemWidePtt;
_origAuxEnabled = settings.AuxEnabled;
_origAuxDeviceId = settings.AuxDeviceId;
_origAuxGain = settings.AuxGain;
Text = "Audio settings";
FormBorderStyle = FormBorderStyle.FixedDialog;
MaximizeBox = false;
MinimizeBox = false;
StartPosition = FormStartPosition.CenterParent;
AutoScaleMode = AutoScaleMode.Font;
ClientSize = new Size(420, 639);
// ── Device row ────────────────────────────────────────────────────────
var lblDevice = new Label
{
Text = "Input &device:",
Location = new Point(12, 16),
AutoSize = true,
};
_cboDevice = new ComboBox
{
Location = new Point(12, 36),
Width = 300,
DropDownStyle = ComboBoxStyle.DropDownList,
DisplayMember = "Name",
ValueMember = "Id",
AccessibleName = "Input device",
AccessibleDescription = "Select which microphone or audio input device to use.",
};
_cboDevice.SelectedIndexChanged += CboDevice_SelectedIndexChanged;
_btnRefresh = new Button
{
Text = "&Refresh",
Location = new Point(320, 34),
Size = new Size(80, 26),
};
_btnRefresh.Click += (_, _) => LoadDevices();
// ── Transmission mode ─────────────────────────────────────────────────
var lblMode = new Label
{
Text = "Transmission mode:",
Location = new Point(12, 76),
AutoSize = true,
};
_radioVad = new RadioButton
{
Text = "&Voice activation",
Location = new Point(20, 96),
AutoSize = true,
};
_radioVad.CheckedChanged += RadioMode_CheckedChanged;
_lblSensitivity = new Label
{
Text = "Sensitivity:",
Location = new Point(36, 122),
AutoSize = true,
};
_trkVad = new TrackBar
{
Location = new Point(36, 140),
Size = new Size(200, 45),
Minimum = 1,
Maximum = 100,
TickFrequency = 10,
SmallChange = 1,
LargeChange = 10,
Value = Math.Clamp(settings.VadThresholdSlider, 1, 100),
};
_trkVad.AccessibleName = "VAD sensitivity";
_trkVad.AccessibleDescription =
"Voice detection sensitivity. Higher = more sensitive. Range 1100.";
_trkVad.Scroll += TrkVad_Scroll;
_radioPtt = new RadioButton
{
Text = "&Push to talk",
Location = new Point(20, 192),
AutoSize = true,
};
_radioPtt.CheckedChanged += RadioMode_CheckedChanged;
_lblPttKey = new Label
{
Location = new Point(140, 194),
AutoSize = true,
};
_btnChangePtt = new Button
{
Text = "Change key...",
Location = new Point(200, 189),
Size = new Size(105, 26),
};
_btnChangePtt.Click += BtnChangePtt_Click;
// Indented PTT sub-option: observe the key system-wide (Raw Input) so it works while
// another app is focused. Visible only in PTT mode (RadioMode_CheckedChanged).
_chkSystemWidePtt = new CheckBox
{
Text = "Wor&ks in the background (system-wide)",
Location = new Point(36, 218),
AutoSize = true,
Checked = settings.SystemWidePtt,
AccessibleName = "System-wide push-to-talk",
AccessibleDescription =
"Let push-to-talk work while another application is focused. Uses the Windows Raw " +
"Input API, not a keyboard hook.",
};
_chkSystemWidePtt.CheckedChanged += ChkSystemWidePtt_CheckedChanged;
_radioAlwaysOn = new RadioButton
{
Text = "A&lways on",
Location = new Point(20, 252),
AutoSize = true,
};
_radioAlwaysOn.CheckedChanged += RadioMode_CheckedChanged;
// ── Mic gain ──────────────────────────────────────────────────────────
var lblGain = new Label
{
Text = "Microphone &volume:",
Location = new Point(12, 296),
AutoSize = true,
};
_trkGain = new TrackBar
{
Location = new Point(12, 316),
Size = new Size(200, 45),
Minimum = 0,
Maximum = 400,
TickFrequency = 25,
SmallChange = 5,
LargeChange = 25,
Value = Math.Clamp(settings.MicGain, 0, 400),
};
_trkGain.AccessibleName = "Microphone volume";
_trkGain.AccessibleDescription =
"Boost a quiet microphone. 100 is unity gain; range 0400 percent.";
_trkGain.Scroll += TrkGain_Scroll;
// ── Mic noise reduction ───────────────────────────────────────────────
// Send-side RNNoise denoise of the mic stream. MIC-only (the core's NR runs before
// the input gain and VAD/PTT gate; aux/screen are excluded). One pass for all listeners.
_chkNoiseReduction = new CheckBox
{
Text = "Noise &reduction (RNNoise)",
Location = new Point(12, 368),
AutoSize = true,
Checked = settings.MicNoiseReduction,
AccessibleName = "Microphone noise reduction",
AccessibleDescription =
"RNNoise denoising of your microphone. Cleans your signal for everyone listening.",
};
_chkNoiseReduction.CheckedChanged += ChkNoiseReduction_CheckedChanged;
// ── Stereo microphone ─────────────────────────────────────────────────
// Opens the mic capture device in stereo (interleaved L/R) instead of mono. Real stereo
// only reaches the wire on a stereo channel; the core folds a stereo mic to mono on a mono
// channel. Toggling while connected restarts the capture device (vc_audio_restart) so the
// new channel count takes effect immediately.
_chkStereoMic = new CheckBox
{
Text = "&Stereo microphone",
Location = new Point(12, 392),
AutoSize = true,
Checked = settings.StereoMic,
AccessibleName = "Stereo microphone",
AccessibleDescription =
"Capture your microphone in stereo. Only transmitted in stereo on a stereo channel.",
};
_chkStereoMic.CheckedChanged += ChkStereoMic_CheckedChanged;
// ── Aux input stream ──────────────────────────────────────────────────
// A second outgoing stream from another hardware input device (e.g. line-in / aux),
// captured client-side and fed to the core. Device + volume only — aux is always-on
// (the core never gates AUX_DEVICE on VAD/PTT).
_chkAux = new CheckBox
{
Text = "&Aux stream (second input device)",
Location = new Point(12, 432),
AutoSize = true,
Checked = settings.AuxEnabled,
AccessibleName = "Enable aux input stream",
AccessibleDescription =
"Transmit a second hardware input device alongside your microphone.",
};
_chkAux.CheckedChanged += ChkAux_CheckedChanged;
_lblAuxDevice = new Label
{
Text = "Aux d&evice:",
Location = new Point(12, 462),
AutoSize = true,
};
_cboAuxDevice = new ComboBox
{
Location = new Point(12, 482),
Width = 300,
DropDownStyle = ComboBoxStyle.DropDownList,
DisplayMember = "Name",
ValueMember = "Id",
AccessibleName = "Aux input device",
AccessibleDescription = "Select the second audio input device to transmit.",
};
_cboAuxDevice.SelectedIndexChanged += CboAuxDevice_SelectedIndexChanged;
_btnAuxRefresh = new Button
{
Text = "Re&fresh",
Location = new Point(320, 480),
Size = new Size(80, 26),
};
_btnAuxRefresh.Click += (_, _) => LoadAuxDevices();
_lblAuxGain = new Label
{
Text = "Aux vo&lume:",
Location = new Point(12, 518),
AutoSize = true,
};
_trkAuxGain = new TrackBar
{
Location = new Point(12, 538),
Size = new Size(200, 45),
Minimum = 0,
Maximum = 400,
TickFrequency = 25,
SmallChange = 5,
LargeChange = 25,
Value = Math.Clamp(settings.AuxGain, 0, 400),
};
_trkAuxGain.AccessibleName = "Aux volume";
_trkAuxGain.AccessibleDescription =
"Volume of the aux input stream. 100 is unity gain; range 0400 percent.";
_trkAuxGain.Scroll += TrkAuxGain_Scroll;
// ── OK / Cancel ───────────────────────────────────────────────────────
var btnOk = new Button
{
Text = "&OK",
DialogResult = DialogResult.OK,
Location = new Point(228, 600),
Size = new Size(80, 27),
};
var btnCancel = new Button
{
Text = "&Cancel",
DialogResult = DialogResult.Cancel,
Location = new Point(316, 600),
Size = new Size(80, 27),
};
AcceptButton = btnOk;
CancelButton = btnCancel;
btnOk.Click += (_, _) =>
{
_settings.Save();
};
FormClosing += AudioSettingsForm_FormClosing;
Controls.AddRange([
lblDevice, _cboDevice, _btnRefresh,
lblMode, _radioVad, _lblSensitivity, _trkVad,
_radioPtt, _lblPttKey, _btnChangePtt, _chkSystemWidePtt, _radioAlwaysOn,
lblGain, _trkGain, _chkNoiseReduction, _chkStereoMic,
_chkAux, _lblAuxDevice, _cboAuxDevice, _btnAuxRefresh, _lblAuxGain, _trkAuxGain,
btnOk, btnCancel,
]);
// Apply saved mode (fires RadioMode_CheckedChanged which sets visibility)
switch ((VcInputMode)settings.InputMode)
{
case VcInputMode.PushToTalk: _radioPtt.Checked = true; break;
case VcInputMode.AlwaysOn: _radioAlwaysOn.Checked = true; break;
default: _radioVad.Checked = true; break;
}
LoadDevices();
LoadAuxDevices();
UpdateAuxControlsEnabled();
}
private void LoadDevices()
{
string? currentId = (_cboDevice.SelectedItem as DeviceInfo)?.Id ?? _settings.InputDeviceId;
var devices = _client.ListDevices(VcDeviceKind.Input);
_cboDevice.SelectedIndexChanged -= CboDevice_SelectedIndexChanged;
_cboDevice.DataSource = new BindingList<DeviceInfo>(devices);
_cboDevice.SelectedIndexChanged += CboDevice_SelectedIndexChanged;
// Try to restore previous selection
int idx = -1;
if (currentId is not null)
idx = devices.FindIndex(d => d.Id == currentId);
if (idx < 0)
idx = devices.FindIndex(d => d.IsDefault);
_cboDevice.SelectedIndex = idx >= 0 ? idx : (devices.Count > 0 ? 0 : -1);
}
private void CboDevice_SelectedIndexChanged(object? sender, EventArgs e)
{
if (_cboDevice.SelectedItem is not DeviceInfo dev) return;
_settings.InputDeviceId = dev.IsDefault ? null : dev.Id;
if (_micStreamId != 0)
_client.SetInputDevice(_micStreamId, dev.IsDefault ? null : dev.Id);
}
// ── Aux input stream ──────────────────────────────────────────────────────
private void LoadAuxDevices()
{
string? currentId = (_cboAuxDevice.SelectedItem as InputDeviceInfo)?.Id
?? _settings.AuxDeviceId;
var devices = InputDeviceEnumerator.List().ToList();
// Set the data source AND restore the selection while unsubscribed so neither the initial
// populate nor a Refresh fires a spurious device-change (which would restart the capture).
_cboAuxDevice.SelectedIndexChanged -= CboAuxDevice_SelectedIndexChanged;
_cboAuxDevice.DataSource = new BindingList<InputDeviceInfo>(devices);
int idx = -1;
if (currentId is not null)
idx = devices.FindIndex(d => d.Id == currentId);
if (idx < 0)
idx = devices.FindIndex(d => d.IsDefault);
_cboAuxDevice.SelectedIndex = idx >= 0 ? idx : (devices.Count > 0 ? 0 : -1);
_cboAuxDevice.SelectedIndexChanged += CboAuxDevice_SelectedIndexChanged;
// Persist the resolved selection (null = default) so Join Voice opens the same device.
if (_cboAuxDevice.SelectedItem is InputDeviceInfo dev)
_settings.AuxDeviceId = dev.IsDefault ? null : dev.Id;
}
private void ChkAux_CheckedChanged(object? sender, EventArgs e)
{
_settings.AuxEnabled = _chkAux.Checked;
UpdateAuxControlsEnabled();
_applyAuxEnabled?.Invoke(_chkAux.Checked);
}
private void CboAuxDevice_SelectedIndexChanged(object? sender, EventArgs e)
{
if (_cboAuxDevice.SelectedItem is not InputDeviceInfo dev) return;
_settings.AuxDeviceId = dev.IsDefault ? null : dev.Id;
if (_chkAux.Checked)
_applyAuxDevice?.Invoke(dev.IsDefault ? null : dev.Id);
}
private void TrkAuxGain_Scroll(object? sender, EventArgs e)
{
_settings.AuxGain = _trkAuxGain.Value;
if (_chkAux.Checked)
_applyAuxGain?.Invoke(_trkAuxGain.Value / 100f);
}
private void UpdateAuxControlsEnabled()
{
bool on = _chkAux.Checked;
_lblAuxDevice.Enabled = on;
_cboAuxDevice.Enabled = on;
_btnAuxRefresh.Enabled = on;
_lblAuxGain.Enabled = on;
_trkAuxGain.Enabled = on;
}
private void RadioMode_CheckedChanged(object? sender, EventArgs e)
{
var mode = CurrentMode();
bool isVad = mode == VcInputMode.VoiceActivation;
bool isPtt = mode == VcInputMode.PushToTalk;
_lblSensitivity.Visible = isVad;
_trkVad.Visible = isVad;
_lblPttKey.Visible = isPtt;
_btnChangePtt.Visible = isPtt;
_chkSystemWidePtt.Visible = isPtt;
UpdatePttKeyLabel();
_settings.InputMode = (int)mode;
if (_micStreamId != 0)
{
_client.SetInputMode(mode);
if (isVad) _client.SetVadThreshold(VadThresholdFromSlider());
if (isPtt) _client.SetPushToTalk(false);
}
}
private void TrkVad_Scroll(object? sender, EventArgs e)
{
_settings.VadThresholdSlider = _trkVad.Value;
if (_micStreamId != 0 && CurrentMode() == VcInputMode.VoiceActivation)
_client.SetVadThreshold(VadThresholdFromSlider());
}
private void TrkGain_Scroll(object? sender, EventArgs e)
{
_settings.MicGain = _trkGain.Value;
if (_micStreamId != 0)
_client.SetInputGain(_trkGain.Value / 100f);
}
private void ChkNoiseReduction_CheckedChanged(object? sender, EventArgs e)
{
_settings.MicNoiseReduction = _chkNoiseReduction.Checked;
if (_micStreamId != 0)
_client.SetInputNoiseReduction(_chkNoiseReduction.Checked);
}
private void ChkStereoMic_CheckedChanged(object? sender, EventArgs e)
{
_settings.StereoMic = _chkStereoMic.Checked;
// Channel count only takes effect when the capture device (re)starts, so restart it live.
if (_micStreamId != 0)
{
_client.SetCaptureChannels(_micStreamId, _chkStereoMic.Checked ? 2u : 1u);
_client.AudioRestart();
}
}
private void ChkSystemWidePtt_CheckedChanged(object? sender, EventArgs e) =>
// No live effect here — MainForm reads SystemWidePtt after the dialog closes and
// registers/unregisters the Raw Input keyboard sink accordingly.
_settings.SystemWidePtt = _chkSystemWidePtt.Checked;
private void BtnChangePtt_Click(object? sender, EventArgs e)
{
using var dlg = new PttKeyCaptureDialog(_pttKey);
if (dlg.ShowDialog(this) == DialogResult.OK)
{
_pttKey = dlg.CapturedKey;
_settings.PttKey = (int)_pttKey;
UpdatePttKeyLabel();
}
}
private void AudioSettingsForm_FormClosing(object? sender, FormClosingEventArgs e)
{
if (DialogResult == DialogResult.OK) return;
// Cancel: restore originals to settings and live client
_settings.InputDeviceId = _origDeviceId;
_settings.InputMode = (int)_origMode;
_settings.VadThresholdSlider = _origVadSlider;
_settings.MicGain = _origMicGain;
_settings.MicNoiseReduction = _origMicNoiseReduction;
_settings.StereoMic = _origStereoMic;
_settings.PttKey = (int)_origPttKey;
_settings.SystemWidePtt = _origSystemWidePtt;
if (_micStreamId != 0)
{
_client.SetInputDevice(_micStreamId, _origDeviceId);
_client.SetInputMode(_origMode);
if (_origMode == VcInputMode.VoiceActivation)
_client.SetVadThreshold(0.1f * (1f - (_origVadSlider - 1f) / 99f));
_client.SetInputGain(_origMicGain / 100f);
_client.SetInputNoiseReduction(_origMicNoiseReduction);
// Restore capture channel count; restart the device only if it actually changed.
if (_origStereoMic != _chkStereoMic.Checked)
{
_client.SetCaptureChannels(_micStreamId, _origStereoMic ? 2u : 1u);
_client.AudioRestart();
}
}
// Aux: restore originals to settings and re-apply live (order: device + gain first so a
// re-enable starts with the right configuration).
_settings.AuxEnabled = _origAuxEnabled;
_settings.AuxDeviceId = _origAuxDeviceId;
_settings.AuxGain = _origAuxGain;
_applyAuxDevice?.Invoke(_origAuxDeviceId);
_applyAuxGain?.Invoke(_origAuxGain / 100f);
_applyAuxEnabled?.Invoke(_origAuxEnabled);
}
private void UpdatePttKeyLabel() =>
_lblPttKey.Text = $"({_pttKey})";
private VcInputMode CurrentMode() =>
_radioPtt.Checked ? VcInputMode.PushToTalk :
_radioAlwaysOn.Checked ? VcInputMode.AlwaysOn :
VcInputMode.VoiceActivation;
private float VadThresholdFromSlider() =>
0.1f * (1f - (_trkVad.Value - 1f) / 99f);
}

View File

@@ -19,6 +19,7 @@ public partial class ConnectDialog : Form
public VoiceCatClient? ConnectedClient { get; private set; } public VoiceCatClient? ConnectedClient { get; private set; }
public uint SelfUserId { get; private set; } public uint SelfUserId { get; private set; }
public string Nickname { get; private set; } = ""; public string Nickname { get; private set; } = "";
public string ServerName { get; private set; } = "";
public ConnectDialog() public ConnectDialog()
{ {
@@ -91,13 +92,8 @@ public partial class ConnectDialog : Form
private void BtnConnect_Click(object? sender, EventArgs e) private void BtnConnect_Click(object? sender, EventArgs e)
{ {
Console.WriteLine("[ConnectDialog] BtnConnect_Click fired");
if (lstServers.SelectedItem is not SavedServer server) if (lstServers.SelectedItem is not SavedServer server)
{
Console.WriteLine("[ConnectDialog] no SavedServer selected — ignoring click");
return; return;
}
Console.WriteLine($"[ConnectDialog] selected server: Host={server.Host} Port={server.Port} AuthMode={server.AuthMode}");
try try
{ {
StartConnect(server); StartConnect(server);
@@ -115,22 +111,20 @@ public partial class ConnectDialog : Form
SetBusy(true); SetBusy(true);
lblStatus.Text = "Connecting..."; lblStatus.Text = "Connecting...";
ServerName = string.IsNullOrWhiteSpace(server.DisplayName)
? $"{server.Host}:{server.Port}"
: server.DisplayName;
string tofuDir = Path.GetDirectoryName(ServerListStore.TofuStorePath)!; string tofuDir = Path.GetDirectoryName(ServerListStore.TofuStorePath)!;
Console.WriteLine($"[ConnectDialog] tofu store dir: {tofuDir}");
Directory.CreateDirectory(tofuDir); Directory.CreateDirectory(tofuDir);
Console.WriteLine("[ConnectDialog] creating VoiceCatClient...");
_client = new VoiceCatClient("VoiceCat-Windows", VoiceCatClient.VersionString, _client = new VoiceCatClient("VoiceCat-Windows", VoiceCatClient.VersionString,
VcLogLevel.Info, ServerListStore.TofuStorePath); VcLogLevel.Info, ServerListStore.TofuStorePath);
Console.WriteLine("[ConnectDialog] VoiceCatClient created OK");
_client.EventReceived += OnEvent; _client.EventReceived += OnEvent;
_identityDialogShown = false; _identityDialogShown = false;
_pumpTimer.Start(); _pumpTimer.Start();
Console.WriteLine($"[ConnectDialog] pump timer started, Enabled={_pumpTimer.Enabled}, Interval={_pumpTimer.Interval}");
Console.WriteLine($"[ConnectDialog] calling Connect({server.Host}, {server.Port})...");
var connectResult = _client.Connect(server.Host, server.Port); var connectResult = _client.Connect(server.Host, server.Port);
Console.WriteLine($"[ConnectDialog] Connect() returned {connectResult}");
if (connectResult != VcResult.Ok) if (connectResult != VcResult.Ok)
{ {
lblStatus.Text = $"Connect failed: {connectResult}"; lblStatus.Text = $"Connect failed: {connectResult}";
@@ -141,9 +135,7 @@ public partial class ConnectDialog : Form
if (server.AuthMode == AuthMode.Guest) if (server.AuthMode == AuthMode.Guest)
{ {
Nickname = string.IsNullOrWhiteSpace(server.LastNickname) ? Environment.UserName : server.LastNickname; Nickname = string.IsNullOrWhiteSpace(server.LastNickname) ? Environment.UserName : server.LastNickname;
Console.WriteLine($"[ConnectDialog] calling AuthenticateGuest({Nickname})..."); _client.AuthenticateGuest(Nickname);
var authResult = _client.AuthenticateGuest(Nickname);
Console.WriteLine($"[ConnectDialog] AuthenticateGuest() returned {authResult}");
} }
else else
{ {
@@ -164,15 +156,12 @@ public partial class ConnectDialog : Form
password = pwDlg.Password; password = pwDlg.Password;
} }
Nickname = server.SavedUsername ?? ""; Nickname = server.SavedUsername ?? "";
Console.WriteLine($"[ConnectDialog] calling AuthenticateUser({Nickname})..."); _client.AuthenticateUser(server.SavedUsername ?? "", password);
var authResult = _client.AuthenticateUser(server.SavedUsername ?? "", password);
Console.WriteLine($"[ConnectDialog] AuthenticateUser() returned {authResult}");
} }
} }
private void OnEvent(VoiceCatEvent ev) private void OnEvent(VoiceCatEvent ev)
{ {
Console.WriteLine($"[ConnectDialog] event: {ev}");
switch (ev.Type) switch (ev.Type)
{ {
case VcEventType.ConnectionState: case VcEventType.ConnectionState:
@@ -220,24 +209,18 @@ public partial class ConnectDialog : Form
private void HandleServerIdentity(VcTofuStatus status, string certFingerprintHex) private void HandleServerIdentity(VcTofuStatus status, string certFingerprintHex)
{ {
Console.WriteLine($"[ConnectDialog] HandleServerIdentity status={status} fp={certFingerprintHex} alreadyShown={_identityDialogShown}");
if (_identityDialogShown) return; // one decision per connect attempt if (_identityDialogShown) return; // one decision per connect attempt
if (status == VcTofuStatus.Matched) if (status == VcTofuStatus.Matched)
{ {
// Silent success path — no dialog. See ServerIdentityDialog's doc comment. // Silent success path — no dialog. See ServerIdentityDialog's doc comment.
Console.WriteLine("[ConnectDialog] status=Matched -> auto-confirming, no dialog");
_client!.ConfirmServerIdentity(true); _client!.ConfirmServerIdentity(true);
return; return;
} }
_identityDialogShown = true; _identityDialogShown = true;
Console.WriteLine("[ConnectDialog] showing ServerIdentityDialog...");
using var dlg = new ServerIdentityDialog(status, certFingerprintHex, _client!.GetServerIdentityDisplay()); using var dlg = new ServerIdentityDialog(status, certFingerprintHex, _client!.GetServerIdentityDisplay());
var dlgResult = dlg.ShowDialog(this); bool accept = dlg.ShowDialog(this) == DialogResult.OK;
Console.WriteLine($"[ConnectDialog] ServerIdentityDialog closed with {dlgResult}"); _client.ConfirmServerIdentity(accept);
bool accept = dlgResult == DialogResult.OK;
var confirmResult = _client.ConfirmServerIdentity(accept);
Console.WriteLine($"[ConnectDialog] ConfirmServerIdentity({accept}) returned {confirmResult}");
if (!accept) lblStatus.Text = "Server identity rejected."; if (!accept) lblStatus.Text = "Server identity rejected.";
} }

View File

@@ -36,21 +36,10 @@ partial class MainForm
// Voice control panel (docked Bottom) // Voice control panel (docked Bottom)
private Panel pnlVoice = null!; private Panel pnlVoice = null!;
private FlowLayoutPanel flpVoiceTop = null!; private FlowLayoutPanel flpVoiceTop = null!;
private FlowLayoutPanel flpVoiceBottom = null!;
private CheckBox chkMute = null!; private CheckBox chkMute = null!;
private CheckBox chkDeafen = null!; private CheckBox chkDeafen = null!;
private RadioButton radioVad = null!;
private RadioButton radioPtt = null!;
private RadioButton radioAlwaysOn = null!;
private Label lblPttKey = null!;
private Button btnChangePtt = null!;
private Label lblInputDevice = null!;
private ComboBox cboInputDevice = null!;
private Button btnRefreshDevices = null!;
private Label lblLevel = null!; private Label lblLevel = null!;
private ProgressBar pbLevel = null!; private ProgressBar pbLevel = null!;
private Label lblVadThreshold = null!;
private TrackBar trkVadThreshold = null!;
protected override void Dispose(bool disposing) protected override void Dispose(bool disposing)
{ {
@@ -79,21 +68,10 @@ partial class MainForm
trkOutputVolume = new TrackBar(); trkOutputVolume = new TrackBar();
pnlVoice = new Panel(); pnlVoice = new Panel();
flpVoiceTop = new FlowLayoutPanel(); flpVoiceTop = new FlowLayoutPanel();
flpVoiceBottom = new FlowLayoutPanel();
chkMute = new CheckBox(); chkMute = new CheckBox();
chkDeafen = new CheckBox(); chkDeafen = new CheckBox();
radioVad = new RadioButton();
radioPtt = new RadioButton();
radioAlwaysOn = new RadioButton();
lblPttKey = new Label();
btnChangePtt = new Button();
lblInputDevice = new Label();
cboInputDevice = new ComboBox();
btnRefreshDevices = new Button();
lblLevel = new Label(); lblLevel = new Label();
pbLevel = new ProgressBar(); pbLevel = new ProgressBar();
lblVadThreshold = new Label();
trkVadThreshold = new TrackBar();
menuStrip = new MenuStrip(); menuStrip = new MenuStrip();
toolStrip = new ToolStrip(); toolStrip = new ToolStrip();
tsbJoinVoice = new ToolStripButton(); tsbJoinVoice = new ToolStripButton();
@@ -140,6 +118,13 @@ partial class MainForm
splitLeft.Panel1MinSize = 100; splitLeft.Panel1MinSize = 100;
splitLeft.Panel2MinSize = 80; splitLeft.Panel2MinSize = 80;
splitLeft.TabIndex = 0; splitLeft.TabIndex = 0;
// Keep the resize splitter out of the Tab cycle so focus moves control-to-control.
splitLeft.TabStop = false;
// Name the otherwise-anonymous "pane" containers so screen readers announce
// orientation instead of a stack of unnamed panes.
splitLeft.AccessibleName = "Channels and users";
splitLeft.Panel1.AccessibleName = "Channels";
splitLeft.Panel2.AccessibleName = "Users";
splitLeft.Panel1.Controls.Add(tvChannels); splitLeft.Panel1.Controls.Add(tvChannels);
splitLeft.Panel1.Controls.Add(lblChannels); splitLeft.Panel1.Controls.Add(lblChannels);
splitLeft.Panel2.Controls.Add(lstUsers); splitLeft.Panel2.Controls.Add(lstUsers);
@@ -214,6 +199,10 @@ partial class MainForm
splitMain.Dock = DockStyle.Fill; splitMain.Dock = DockStyle.Fill;
splitMain.Panel1MinSize = 150; splitMain.Panel1MinSize = 150;
splitMain.TabIndex = 1; splitMain.TabIndex = 1;
splitMain.TabStop = false;
splitMain.AccessibleName = "Main";
splitMain.Panel1.AccessibleName = "Channels and users";
splitMain.Panel2.AccessibleName = "Chat and activity";
splitMain.Panel1.Controls.Add(splitLeft); splitMain.Panel1.Controls.Add(splitLeft);
splitMain.Panel2.Controls.Add(tblRight); splitMain.Panel2.Controls.Add(tblRight);
@@ -230,71 +219,15 @@ partial class MainForm
chkDeafen.Margin = new Padding(0, 4, 12, 0); chkDeafen.Margin = new Padding(0, 4, 12, 0);
chkDeafen.TabIndex = 2; chkDeafen.TabIndex = 2;
var lblMode = new Label { Text = "Mode:", AutoSize = true, Margin = new Padding(0, 5, 4, 0) }; flpVoiceTop.Dock = DockStyle.Fill;
radioVad.Text = "&Voice activation";
radioVad.AutoSize = true;
radioVad.Checked = true;
radioVad.Enabled = false;
radioVad.Margin = new Padding(0, 4, 6, 0);
radioVad.TabIndex = 3;
radioPtt.Text = "&Push to talk";
radioPtt.AutoSize = true;
radioPtt.Enabled = false;
radioPtt.Margin = new Padding(0, 4, 4, 0);
radioPtt.TabIndex = 4;
radioAlwaysOn.Text = "A&lways on";
radioAlwaysOn.AutoSize = true;
radioAlwaysOn.Enabled = false;
radioAlwaysOn.Margin = new Padding(0, 4, 12, 0);
radioAlwaysOn.TabIndex = 5;
lblPttKey.Text = "(F8)";
lblPttKey.AutoSize = true;
lblPttKey.Margin = new Padding(2, 5, 4, 0);
lblPttKey.Visible = false;
btnChangePtt.Text = "Change key...";
btnChangePtt.AutoSize = true;
btnChangePtt.Margin = new Padding(0, 2, 0, 0);
btnChangePtt.Visible = false;
btnChangePtt.TabIndex = 6;
flpVoiceTop.Dock = DockStyle.Top;
flpVoiceTop.Height = 34;
flpVoiceTop.AutoSize = false; flpVoiceTop.AutoSize = false;
flpVoiceTop.Padding = new Padding(4, 2, 4, 0); flpVoiceTop.Padding = new Padding(4, 2, 4, 0);
flpVoiceTop.Controls.Add(chkMute); flpVoiceTop.Controls.Add(chkMute);
flpVoiceTop.Controls.Add(chkDeafen); flpVoiceTop.Controls.Add(chkDeafen);
flpVoiceTop.Controls.Add(lblMode);
flpVoiceTop.Controls.Add(radioVad);
flpVoiceTop.Controls.Add(radioPtt);
flpVoiceTop.Controls.Add(radioAlwaysOn);
flpVoiceTop.Controls.Add(lblPttKey);
flpVoiceTop.Controls.Add(btnChangePtt);
// Bottom row: device picker + level meter
lblInputDevice.Text = "Input:";
lblInputDevice.AutoSize = true;
lblInputDevice.Margin = new Padding(0, 5, 4, 0);
cboInputDevice.AccessibleName = "Input device";
cboInputDevice.AccessibleDescription = "Select which microphone or audio device to use.";
cboInputDevice.DropDownStyle = ComboBoxStyle.DropDownList;
cboInputDevice.Width = 200;
cboInputDevice.Margin = new Padding(0, 2, 4, 0);
cboInputDevice.TabIndex = 7;
btnRefreshDevices.Text = "Re&fresh";
btnRefreshDevices.AutoSize = true;
btnRefreshDevices.Margin = new Padding(0, 2, 12, 0);
btnRefreshDevices.TabIndex = 8;
lblLevel.Text = "Level:"; lblLevel.Text = "Level:";
lblLevel.AutoSize = true; lblLevel.AutoSize = true;
lblLevel.Margin = new Padding(0, 5, 4, 0); lblLevel.Margin = new Padding(12, 5, 4, 0);
pbLevel.AccessibleName = "Microphone level"; pbLevel.AccessibleName = "Microphone level";
pbLevel.AccessibleDescription = "Current input level from the microphone."; pbLevel.AccessibleDescription = "Current input level from the microphone.";
@@ -305,42 +238,14 @@ partial class MainForm
pbLevel.Style = ProgressBarStyle.Continuous; pbLevel.Style = ProgressBarStyle.Continuous;
pbLevel.TabStop = false; pbLevel.TabStop = false;
lblVadThreshold.Text = "Sensitivity:"; flpVoiceTop.Controls.Add(lblLevel);
lblVadThreshold.AutoSize = true; flpVoiceTop.Controls.Add(pbLevel);
lblVadThreshold.Margin = new Padding(12, 5, 4, 0);
lblVadThreshold.Visible = true;
trkVadThreshold.AccessibleName = "VAD sensitivity";
trkVadThreshold.AccessibleDescription =
"Voice detection sensitivity. Higher = more sensitive (triggers on quieter sounds). " +
"Range 1100; default 25.";
trkVadThreshold.Minimum = 1;
trkVadThreshold.Maximum = 100;
trkVadThreshold.Value = 76;
trkVadThreshold.TickFrequency = 10;
trkVadThreshold.SmallChange = 1;
trkVadThreshold.LargeChange = 10;
trkVadThreshold.Width = 120;
trkVadThreshold.Margin = new Padding(0, 2, 0, 0);
trkVadThreshold.TabIndex = 9;
trkVadThreshold.Visible = true;
flpVoiceBottom.Dock = DockStyle.Fill;
flpVoiceBottom.Padding = new Padding(4, 0, 4, 2);
flpVoiceBottom.Controls.Add(lblInputDevice);
flpVoiceBottom.Controls.Add(cboInputDevice);
flpVoiceBottom.Controls.Add(btnRefreshDevices);
flpVoiceBottom.Controls.Add(lblLevel);
flpVoiceBottom.Controls.Add(pbLevel);
flpVoiceBottom.Controls.Add(lblVadThreshold);
flpVoiceBottom.Controls.Add(trkVadThreshold);
pnlVoice.Dock = DockStyle.Bottom; pnlVoice.Dock = DockStyle.Bottom;
pnlVoice.Height = 68; pnlVoice.Height = 38;
pnlVoice.BorderStyle = BorderStyle.FixedSingle; pnlVoice.BorderStyle = BorderStyle.FixedSingle;
pnlVoice.Padding = new Padding(0); pnlVoice.Padding = new Padding(0);
pnlVoice.Controls.Add(flpVoiceBottom); // Fill — added first pnlVoice.Controls.Add(flpVoiceTop);
pnlVoice.Controls.Add(flpVoiceTop); // Top — added last
// ── Toolbar ─────────────────────────────────────────────────────────── // ── Toolbar ───────────────────────────────────────────────────────────
tsbJoinVoice.Text = "Join Voice"; tsbJoinVoice.Text = "Join Voice";

View File

@@ -1,4 +1,7 @@
using VoiceCat.App.Audio; using VoiceCat.App.Audio;
using VoiceCat.App.Models;
using VoiceCat.App.Native;
using VoiceCat.App.Notifications;
using VoiceCat.Interop; using VoiceCat.Interop;
namespace VoiceCat.App.Forms; namespace VoiceCat.App.Forms;
@@ -12,6 +15,8 @@ public partial class MainForm : Form
private readonly uint _selfUserId; private readonly uint _selfUserId;
private readonly string _nickname; private readonly string _nickname;
private readonly System.Windows.Forms.Timer _pumpTimer = new() { Interval = 30 }; private readonly System.Windows.Forms.Timer _pumpTimer = new() { Interval = 30 };
private readonly EventFeedback _feedback = new(FeedbackSettings.Load());
private readonly VoiceSettings _voiceSettings = VoiceSettings.Load();
// Channel / user state // Channel / user state
private uint _currentChannelId; private uint _currentChannelId;
@@ -24,7 +29,11 @@ public partial class MainForm : Form
private uint _micStreamId; // 0 = not started private uint _micStreamId; // 0 = not started
private uint _screenStreamId; // 0 = not sharing screen audio private uint _screenStreamId; // 0 = not sharing screen audio
private ProcessAudioMixer? _screenMixer; // non-null only in per-app capture mode private ProcessAudioMixer? _screenMixer; // non-null only in per-app capture mode
private uint _auxStreamId; // 0 = aux (second input device) stream not active
private InputDeviceCapture? _auxCapture; // client-side capture feeding the aux stream
private Keys _pttKey = Keys.F8; private Keys _pttKey = Keys.F8;
private bool _pttEngaged; // guards the PTT cue against key-repeat
private bool _rawInputRegistered; // true while the system-wide PTT keyboard sink is active
private bool _serverMuted; private bool _serverMuted;
private bool _serverDeafened; private bool _serverDeafened;
@@ -35,18 +44,21 @@ public partial class MainForm : Form
private ToolStripMenuItem _miJoinVoice = null!; private ToolStripMenuItem _miJoinVoice = null!;
private ToolStripMenuItem _miScreenShare = null!; private ToolStripMenuItem _miScreenShare = null!;
public MainForm(VoiceCatClient client, uint selfUserId, string nickname) public MainForm(VoiceCatClient client, uint selfUserId, string nickname, string serverName)
{ {
InitializeComponent(); InitializeComponent();
_client = client; _client = client;
_selfUserId = selfUserId; _selfUserId = selfUserId;
_nickname = nickname; _nickname = nickname;
Text = $"VoiceCat — {nickname}"; Text = string.IsNullOrWhiteSpace(serverName)
? $"VoiceCat — {nickname}"
: $"VoiceCat — {nickname} @ {serverName}";
_client.EventReceived += OnEvent; _client.EventReceived += OnEvent;
_client.LevelChanged += OnLevelChanged; _client.LevelChanged += OnLevelChanged;
_pumpTimer.Tick += (_, _) => _client.PumpEvents(); _pumpTimer.Tick += (_, _) => _client.PumpEvents();
_pumpTimer.Tick += (_, _) => PttWatchdog();
_ownPermissions = SafeGetPermissions(); _ownPermissions = SafeGetPermissions();
BuildMenus(); BuildMenus();
@@ -62,7 +74,7 @@ public partial class MainForm : Form
tvChannels.DoubleClick += TvChannels_DoubleClick; tvChannels.DoubleClick += TvChannels_DoubleClick;
tvChannels.KeyDown += TvChannels_KeyDown; tvChannels.KeyDown += TvChannels_KeyDown;
lstUsers.DoubleClick += (_, _) => OpenUserTuning(); lstUsers.DoubleClick += (_, _) => OpenUserTuning();
lstUsers.KeyDown += (_, e) => { if (e.KeyCode == Keys.Enter) OpenUserTuning(); }; lstUsers.KeyDown += (_, e) => { if (e.KeyCode == Keys.Enter) { OpenUserTuning(); e.Handled = e.SuppressKeyPress = true; } };
// Compose // Compose
txtCompose.KeyDown += TxtCompose_KeyDown; txtCompose.KeyDown += TxtCompose_KeyDown;
@@ -78,26 +90,28 @@ public partial class MainForm : Form
// Voice controls // Voice controls
chkMute.CheckedChanged += (_, _) => ApplySelfMute(); chkMute.CheckedChanged += (_, _) => ApplySelfMute();
chkDeafen.CheckedChanged += (_, _) => ApplySelfMute(); chkDeafen.CheckedChanged += (_, _) => ApplySelfMute();
radioVad.CheckedChanged += RadioVad_CheckedChanged;
radioPtt.CheckedChanged += RadioPtt_CheckedChanged;
radioAlwaysOn.CheckedChanged += RadioAlwaysOn_CheckedChanged;
trkVadThreshold.Scroll += TrkVadThreshold_Scroll;
btnChangePtt.Click += BtnChangePtt_Click;
btnRefreshDevices.Click += (_, _) => LoadInputDevices();
cboInputDevice.SelectedIndexChanged += CboInputDevice_SelectedIndexChanged;
// PTT and global hotkeys (focus-scoped — work only while this form has focus) // PTT and global hotkeys. The mute/deafen/screen-share toggles (MainForm_HotkeyDown) are
// always focus-scoped. PTT is focus-scoped via KeyDown/KeyUp when system-wide PTT is OFF;
// when ON, the WM_INPUT path (WndProc + Raw Input) owns PTT for both focused and unfocused
// cases and the KeyDown/KeyUp handlers early-return.
KeyDown += MainForm_KeyDown; KeyDown += MainForm_KeyDown;
KeyDown += MainForm_HotkeyDown; KeyDown += MainForm_HotkeyDown;
KeyUp += MainForm_KeyUp; KeyUp += MainForm_KeyUp;
Deactivate += (_, _) => Deactivate += (_, _) =>
{ {
if (_micStreamId != 0) _client.SetPushToTalk(false); // With system-wide PTT we WANT transmission to continue while unfocused, so don't
// release on deactivate — the Raw Input key-up (and PttWatchdog) handle release.
if (!_voiceSettings.SystemWidePtt && _micStreamId != 0) _client.SetPushToTalk(false);
}; };
ApplyPersistedVoiceSettings();
BootstrapFromServer(); BootstrapFromServer();
} }
private void ApplyPersistedVoiceSettings() =>
_pttKey = (Keys)_voiceSettings.PttKey;
// ── Startup ────────────────────────────────────────────────────────────── // ── Startup ──────────────────────────────────────────────────────────────
private void BootstrapFromServer() private void BootstrapFromServer()
@@ -118,6 +132,8 @@ public partial class MainForm : Form
RefreshUserList(); RefreshUserList();
UpdateStatusLabel(); UpdateStatusLabel();
AddActivity($"Connected to server as {_nickname}"); AddActivity($"Connected to server as {_nickname}");
_feedback.PlaySound(SoundEvent.Login);
_feedback.Speak("Connected");
} }
protected override void OnLoad(EventArgs e) protected override void OnLoad(EventArgs e)
@@ -125,7 +141,6 @@ public partial class MainForm : Form
base.OnLoad(e); base.OnLoad(e);
splitMain.SplitterDistance = Math.Min(220, splitMain.Width - 304); splitMain.SplitterDistance = Math.Min(220, splitMain.Width - 304);
splitLeft.SplitterDistance = Math.Min(260, splitLeft.Height - 84); splitLeft.SplitterDistance = Math.Min(260, splitLeft.Height - 84);
LoadInputDevices();
} }
// ── Menu / context menu builders ───────────────────────────────────────── // ── Menu / context menu builders ─────────────────────────────────────────
@@ -161,6 +176,35 @@ public partial class MainForm : Form
messagesMenu.DropDownItems.Add(miNewPm); messagesMenu.DropDownItems.Add(miNewPm);
menuStrip.Items.Add(messagesMenu); menuStrip.Items.Add(messagesMenu);
// Settings menu — always visible
var settingsMenu = new ToolStripMenuItem("&Settings");
var miAudio = new ToolStripMenuItem("&Audio...");
miAudio.Click += (_, _) =>
{
using var dlg = new AudioSettingsForm(_client, _voiceSettings, _micStreamId,
applyAuxEnabled: on =>
{
if (_micStreamId == 0) return; // not in voice — applied on next Join Voice
if (on) StartAuxStream(); else StopAuxStream();
},
applyAuxDevice: _ =>
{
if (_auxStreamId != 0) RestartAuxCapture(); // settings.AuxDeviceId already updated
});
dlg.ShowDialog(this);
_pttKey = (Keys)_voiceSettings.PttKey;
ApplySystemWidePtt(); // the system-wide toggle may have changed
};
settingsMenu.DropDownItems.Add(miAudio);
var miNotifications = new ToolStripMenuItem("&Notifications...");
miNotifications.Click += (_, _) =>
{
using var dlg = new NotificationSettingsForm(_feedback);
dlg.ShowDialog(this);
};
settingsMenu.DropDownItems.Add(miNotifications);
menuStrip.Items.Add(settingsMenu);
// Admin menu — only if permitted // Admin menu — only if permitted
if (_ownPermissions.CanAdminAccounts) if (_ownPermissions.CanAdminAccounts)
{ {
@@ -201,6 +245,7 @@ public partial class MainForm : Form
ctx.Items.Add("&Delete channel...", null, (_, _) => DeleteSelectedChannel()); ctx.Items.Add("&Delete channel...", null, (_, _) => DeleteSelectedChannel());
} }
}; };
ctx.Opened += (_, _) => SelectFirstMenuItem(ctx);
tvChannels.ContextMenuStrip = ctx; tvChannels.ContextMenuStrip = ctx;
} }
@@ -247,9 +292,26 @@ public partial class MainForm : Form
} }
} }
}; };
ctx.Opened += (_, _) => SelectFirstMenuItem(ctx);
lstUsers.ContextMenuStrip = ctx; lstUsers.ContextMenuStrip = ctx;
} }
// Work around the WinForms ContextMenuStrip accessibility bug: when opened by
// keyboard (Shift+F10 / Apps key) the focused item is not set, so screen readers
// stay silent until the first arrow key. Selecting the first item ourselves on
// Opened raises the UIA focus event immediately.
private static void SelectFirstMenuItem(ContextMenuStrip ctx)
{
foreach (ToolStripItem item in ctx.Items)
{
if (item is ToolStripMenuItem && item.Enabled)
{
item.Select();
break;
}
}
}
// ── Event dispatch ──────────────────────────────────────────────────────── // ── Event dispatch ────────────────────────────────────────────────────────
private void OnEvent(VoiceCatEvent ev) private void OnEvent(VoiceCatEvent ev)
@@ -287,10 +349,21 @@ public partial class MainForm : Form
HandleStreamStarted(ev); HandleStreamStarted(ev);
break; break;
case VcEventType.StreamStopped: case VcEventType.StreamStopped:
if (_users.TryGetValue(ev.UserId, out var stUser) && if (ev.UserId == _selfUserId)
{
if (ev.StreamId == _micStreamId)
{
_micStreamId = 0;
pbLevel.Value = 0;
}
}
else if (_users.TryGetValue(ev.UserId, out var stUser) &&
stUser.ChannelId == _currentChannelId) stUser.ChannelId == _currentChannelId)
AddActivity($"{stUser.Nickname} stopped a stream"); AddActivity($"{stUser.Nickname} stopped a stream");
break; break;
case VcEventType.VoiceState:
HandleVoiceState(ev);
break;
case VcEventType.Disconnected: case VcEventType.Disconnected:
HandleDisconnected(ev); HandleDisconnected(ev);
break; break;
@@ -316,12 +389,16 @@ public partial class MainForm : Form
private void HandleUserJoined(VoiceCatEvent ev) private void HandleUserJoined(VoiceCatEvent ev)
{ {
var user = new UserInfo(ev.UserId, ev.Text ?? $"User#{ev.UserId}", false, ev.ChannelId, var user = new UserInfo(ev.UserId, ev.Text ?? $"User#{ev.UserId}", false, ev.ChannelId,
false, false, false, false); false, false, false, false, false);
_users[ev.UserId] = user; _users[ev.UserId] = user;
RefreshChannelTree(); RefreshChannelTree();
RefreshUserList(); RefreshUserList();
if (ev.ChannelId == _currentChannelId && ev.UserId != _selfUserId) if (ev.ChannelId == _currentChannelId && ev.UserId != _selfUserId)
{
AddActivity($"{user.Nickname} joined the channel"); AddActivity($"{user.Nickname} joined the channel");
_feedback.PlaySound(SoundEvent.ChannelJoin);
_feedback.Speak($"{user.Nickname} joined");
}
} }
private void HandleUserLeft(VoiceCatEvent ev) private void HandleUserLeft(VoiceCatEvent ev)
@@ -332,7 +409,12 @@ public partial class MainForm : Form
_talkingUsers.Remove(ev.UserId); _talkingUsers.Remove(ev.UserId);
RefreshChannelTree(); RefreshChannelTree();
RefreshUserList(); RefreshUserList();
if (wasHere) AddActivity($"{user.Nickname} left the channel"); if (wasHere)
{
AddActivity($"{user.Nickname} left the channel");
_feedback.PlaySound(SoundEvent.ChannelLeave);
_feedback.Speak($"{user.Nickname} left");
}
if (_pmWindows.TryGetValue(ev.UserId, out var pmWin)) if (_pmWindows.TryGetValue(ev.UserId, out var pmWin))
pmWin.AppendActivity($"{user.Nickname} disconnected from server"); pmWin.AppendActivity($"{user.Nickname} disconnected from server");
} }
@@ -388,18 +470,23 @@ public partial class MainForm : Form
.LocalDateTime.ToString("HH:mm") .LocalDateTime.ToString("HH:mm")
: DateTime.Now.ToString("HH:mm"); : DateTime.Now.ToString("HH:mm");
string sender = GetNickname(ev.UserId); string sender = GetNickname(ev.UserId);
bool isSelf = ev.UserId == _selfUserId;
string body = ev.Text ?? "";
if (ev.TextScope == VcTextScope.Private) if (ev.TextScope == VcTextScope.Private)
{ {
// For our own outgoing PM, ev.ChannelId carries the recipient user ID. // For our own outgoing PM, ev.ChannelId carries the recipient user ID.
uint otherUserId = ev.UserId == _selfUserId ? ev.ChannelId : ev.UserId; uint otherUserId = isSelf ? ev.ChannelId : ev.UserId;
var win = GetOrOpenPmWindow(otherUserId); var win = GetOrOpenPmWindow(otherUserId);
bool isSelf = ev.UserId == _selfUserId; win.AppendMessage(time, isSelf, sender, body);
win.AppendMessage(time, isSelf, sender, ev.Text ?? ""); _feedback.PlaySound(isSelf ? SoundEvent.PmSent : SoundEvent.PmRecv);
if (!isSelf) _feedback.Speak($"Private message from {sender}: {body}");
} }
else else
{ {
AppendChat(time, sender, ev.Text ?? ""); AppendChat(time, sender, body);
_feedback.PlaySound(isSelf ? SoundEvent.ChannelSent : SoundEvent.ChannelRecv);
if (!isSelf) _feedback.Speak($"{sender}: {body}");
} }
} }
@@ -409,6 +496,8 @@ public partial class MainForm : Form
if (talking) _talkingUsers.Add(ev.UserId); if (talking) _talkingUsers.Add(ev.UserId);
else _talkingUsers.Remove(ev.UserId); else _talkingUsers.Remove(ev.UserId);
RefreshUserList(); RefreshUserList();
if (ev.UserId == _selfUserId)
_feedback.PlaySound(talking ? SoundEvent.VaStart : SoundEvent.VaStop);
if (talking && ev.UserId != _selfUserId && if (talking && ev.UserId != _selfUserId &&
_users.TryGetValue(ev.UserId, out var tUser) && _users.TryGetValue(ev.UserId, out var tUser) &&
tUser.ChannelId == _currentChannelId) tUser.ChannelId == _currentChannelId)
@@ -437,6 +526,16 @@ public partial class MainForm : Form
: $"Disconnected: {ev.Text}"; : $"Disconnected: {ev.Text}";
lblStatus.Text = msg; lblStatus.Text = msg;
AddActivity(msg); AddActivity(msg);
if (ev.Result == VcResult.Ok)
{
_feedback.PlaySound(SoundEvent.Logout);
_feedback.Speak("Disconnected");
}
else
{
_feedback.PlaySound(SoundEvent.ConnectionLost);
_feedback.Speak("Connection lost");
}
tvChannels.Nodes.Clear(); tvChannels.Nodes.Clear();
lstUsers.Items.Clear(); lstUsers.Items.Clear();
_users.Clear(); _users.Clear();
@@ -445,6 +544,7 @@ public partial class MainForm : Form
_micStreamId = 0; _micStreamId = 0;
_screenStreamId = 0; _screenStreamId = 0;
_screenMixer?.Stop(); _screenMixer?.Dispose(); _screenMixer = null; _screenMixer?.Stop(); _screenMixer?.Dispose(); _screenMixer = null;
DisposeAuxCapture(); _auxStreamId = 0; // connection gone — drop capture, no StopStream
txtCompose.Enabled = false; txtCompose.Enabled = false;
btnSend.Enabled = false; btnSend.Enabled = false;
tsbJoinVoice.Enabled = false; tsbJoinVoice.Enabled = false;
@@ -468,6 +568,7 @@ public partial class MainForm : Form
private void RefreshChannelTree() private void RefreshChannelTree()
{ {
uint toSelect = tvChannels.SelectedNode?.Tag is uint s ? s : _currentChannelId; uint toSelect = tvChannels.SelectedNode?.Tag is uint s ? s : _currentChannelId;
bool hadFocus = tvChannels.Focused;
tvChannels.BeginUpdate(); tvChannels.BeginUpdate();
tvChannels.Nodes.Clear(); tvChannels.Nodes.Clear();
@@ -493,6 +594,18 @@ public partial class MainForm : Form
tvChannels.ExpandAll(); tvChannels.ExpandAll();
SeekAndSelect(tvChannels.Nodes, toSelect); SeekAndSelect(tvChannels.Nodes, toSelect);
tvChannels.EndUpdate(); tvChannels.EndUpdate();
// A Nodes.Clear()/rebuild can drop keyboard focus and leave the screen reader
// without a current node. If the tree was focused before the refresh, restore
// focus and re-announce the now-current node (null-then-reselect forces UIA to
// fire a fresh focus event).
if (hadFocus && tvChannels.SelectedNode != null)
{
var node = tvChannels.SelectedNode;
tvChannels.Focus();
tvChannels.SelectedNode = null;
tvChannels.SelectedNode = node;
}
} }
private bool SeekAndSelect(TreeNodeCollection nodes, uint channelId) private bool SeekAndSelect(TreeNodeCollection nodes, uint channelId)
@@ -508,6 +621,11 @@ public partial class MainForm : Form
private void RefreshUserList() private void RefreshUserList()
{ {
// Preserve the keyboard selection across the rebuild: clearing the list resets
// SelectedIndex to -1, which would throw focus around every time a talking/mute
// indicator toggles. Capture the selected user id and reselect it afterwards.
uint? prevSel = (lstUsers.SelectedItem as UserListItem)?.UserId;
lstUsers.BeginUpdate(); lstUsers.BeginUpdate();
lstUsers.Items.Clear(); lstUsers.Items.Clear();
foreach (var user in _users.Values foreach (var user in _users.Values
@@ -521,6 +639,15 @@ public partial class MainForm : Form
if (user.SelfDeafened || user.ServerDeafened) label += " (deafened)"; if (user.SelfDeafened || user.ServerDeafened) label += " (deafened)";
lstUsers.Items.Add(new UserListItem(user.Id, label)); lstUsers.Items.Add(new UserListItem(user.Id, label));
} }
if (prevSel is uint sel)
{
for (int i = 0; i < lstUsers.Items.Count; i++)
if (lstUsers.Items[i] is UserListItem item && item.UserId == sel)
{
lstUsers.SelectedIndex = i;
break;
}
}
lstUsers.EndUpdate(); lstUsers.EndUpdate();
UpdateStatusLabel(); UpdateStatusLabel();
} }
@@ -542,68 +669,22 @@ public partial class MainForm : Form
lblStatus.Text = $"Connected as {_nickname}{suffix} — {chanName} ({count} user{(count == 1 ? "" : "s")})"; lblStatus.Text = $"Connected as {_nickname}{suffix} — {chanName} ({count} user{(count == 1 ? "" : "s")})";
} }
// ── Device management ─────────────────────────────────────────────────────
private void LoadInputDevices()
{
var devices = _client.ListDevices(VcDeviceKind.Input);
DeviceInfo? prevDevice = cboInputDevice.SelectedItem as DeviceInfo;
cboInputDevice.Items.Clear();
foreach (var d in devices) cboInputDevice.Items.Add(d);
if (prevDevice is not null)
{
for (int i = 0; i < cboInputDevice.Items.Count; i++)
{
if (cboInputDevice.Items[i] is DeviceInfo d && d.Id == prevDevice.Id)
{
cboInputDevice.SelectedIndex = i;
return;
}
}
}
for (int i = 0; i < cboInputDevice.Items.Count; i++)
{
if (cboInputDevice.Items[i] is DeviceInfo d && d.IsDefault)
{
cboInputDevice.SelectedIndex = i;
return;
}
}
if (cboInputDevice.Items.Count > 0) cboInputDevice.SelectedIndex = 0;
}
// ── Voice controls ──────────────────────────────────────────────────────── // ── Voice controls ────────────────────────────────────────────────────────
private void BtnMicToggle_Click(object? sender, EventArgs e) private void BtnMicToggle_Click(object? sender, EventArgs e)
{ {
if (_micStreamId == 0) if (_micStreamId == 0)
{ {
var (result, streamId) = _client.StartStream(VcStreamKind.Mic, "Microphone"); var result = _client.JoinVoice();
if (result == VcResult.Ok) if (result != VcResult.Ok)
{ AddActivity($"Failed to join voice: {result}");
_micStreamId = streamId;
if (cboInputDevice.SelectedItem is DeviceInfo { IsDefault: false } dev)
_client.SetInputDevice(streamId, dev.Id);
_client.SetInputMode(CurrentInputMode());
if (radioVad.Checked) _client.SetVadThreshold(VadThresholdFromSlider());
SetVoiceJoinedState(true);
AddActivity("Joined voice — microphone active");
}
else
{
AddActivity($"Failed to start microphone: {result}");
}
} }
else else
{ {
StopAuxStream();
if (_screenStreamId != 0) StopScreenAudio();
_client.SetPushToTalk(false); _client.SetPushToTalk(false);
_client.StopStream(_micStreamId); _client.LeaveVoice();
_micStreamId = 0;
pbLevel.Value = 0;
SetVoiceJoinedState(false);
AddActivity("Left voice");
} }
} }
@@ -613,9 +694,42 @@ public partial class MainForm : Form
_miJoinVoice.Text = joined ? "Leave &Voice" : "&Join Voice"; _miJoinVoice.Text = joined ? "Leave &Voice" : "&Join Voice";
chkMute.Enabled = joined; chkMute.Enabled = joined;
chkDeafen.Enabled = joined; chkDeafen.Enabled = joined;
radioVad.Enabled = joined; }
radioPtt.Enabled = joined;
radioAlwaysOn.Enabled = joined; private void HandleVoiceState(VoiceCatEvent ev)
{
bool subscribed = ev.U32a != 0;
if (subscribed)
{
var (result, streamId) = _client.StartStream(VcStreamKind.Mic, "Microphone");
if (result == VcResult.Ok)
{
_micStreamId = streamId;
var mode = (VcInputMode)_voiceSettings.InputMode;
if (_voiceSettings.InputDeviceId is string devId)
_client.SetInputDevice(streamId, devId);
_client.SetCaptureChannels(streamId, _voiceSettings.StereoMic ? 2u : 1u);
_client.SetInputMode(mode);
if (mode == VcInputMode.VoiceActivation)
_client.SetVadThreshold(VadThresholdFromSettings());
_client.SetInputGain(_voiceSettings.MicGain / 100f);
_client.SetInputNoiseReduction(_voiceSettings.MicNoiseReduction);
SetVoiceJoinedState(true);
AddActivity("Joined voice — microphone active");
_feedback.PlaySound(SoundEvent.VoiceOn);
StartAuxStream();
}
else
{
AddActivity($"Failed to start microphone: {result}");
}
}
else
{
SetVoiceJoinedState(false);
AddActivity("Left voice");
_feedback.PlaySound(SoundEvent.VoiceOff);
}
} }
private void ApplySelfMute() => private void ApplySelfMute() =>
@@ -640,7 +754,7 @@ public partial class MainForm : Form
var scope = picker.ChosenScope; var scope = picker.ChosenScope;
if (scope is EntireDesktop) if (scope is EntireDesktop { ExcludeSelf: false })
{ {
// Existing whole-device WASAPI loopback path — core handles it. // Existing whole-device WASAPI loopback path — core handles it.
var (result, streamId) = _client.StartStream(VcStreamKind.ScreenAudio, "Desktop audio"); var (result, streamId) = _client.StartStream(VcStreamKind.ScreenAudio, "Desktop audio");
@@ -654,7 +768,8 @@ public partial class MainForm : Form
} }
else else
{ {
// Per-app path: suppress core loopback, C# mixer feeds PCM. // External-feed path: suppress core loopback, C# mixer feeds PCM. Covers the
// per-app modes and "entire desktop except VoiceCat" (single EXCLUDE of self).
var (result, streamId) = _client.StartStreamExternalFeed(VcStreamKind.ScreenAudio, "App audio"); var (result, streamId) = _client.StartStreamExternalFeed(VcStreamKind.ScreenAudio, "App audio");
if (result != VcResult.Ok) if (result != VcResult.Ok)
{ {
@@ -664,12 +779,13 @@ public partial class MainForm : Form
_screenStreamId = streamId; _screenStreamId = streamId;
_screenMixer = new ProcessAudioMixer(); _screenMixer = new ProcessAudioMixer();
_screenMixer.Start(scope, _client, streamId, picker.VisibleApps); _screenMixer.Start(scope, _client, streamId);
string desc = scope switch string desc = scope switch
{ {
EntireDesktop => "entire desktop (excluding VoiceCat)",
OnlyApps o => $"only {o.Pids.Count} app(s)", OnlyApps o => $"only {o.Pids.Count} app(s)",
AllExceptApps a => $"all except {a.Pids.Count} app(s)", AllExceptApps a => $"all except {a.Names[0]}",
_ => "apps", _ => "apps",
}; };
AddActivity($"Sharing screen audio: {desc}"); AddActivity($"Sharing screen audio: {desc}");
@@ -692,88 +808,103 @@ public partial class MainForm : Form
AddActivity("Stopped sharing screen audio"); AddActivity("Stopped sharing screen audio");
} }
// ── Aux input stream (second hardware input device) ─────────────────────────
// A second outgoing stream (kind = AUX_DEVICE, external_feed). The core can't open a second
// capture device, so we capture the chosen device here and feed PCM in — the same external-
// feed pipeline as per-app screen audio. Tied to the voice session: started on Join Voice
// (when enabled) and stopped on Leave Voice. The aux is always-on (the core never gates
// AUX_DEVICE on VAD/PTT); volume is applied client-side before feeding.
private void StartAuxStream()
{
if (_auxStreamId != 0 || !_voiceSettings.AuxEnabled) return;
var (result, streamId) = _client.StartStreamExternalFeed(VcStreamKind.AuxDevice, "Aux device");
if (result != VcResult.Ok)
{
AddActivity($"Failed to start aux stream: {result}");
return;
}
_auxStreamId = streamId;
_auxCapture = new InputDeviceCapture(_voiceSettings.AuxDeviceId);
_auxCapture.PcmFrameReady += OnAuxPcmFrame;
if (!_auxCapture.Start())
{
AddActivity("Failed to open aux input device");
StopAuxStream();
return;
}
AddActivity("Aux input stream active");
}
private void StopAuxStream()
{
DisposeAuxCapture();
if (_auxStreamId != 0)
{
_client.StopStream(_auxStreamId);
_auxStreamId = 0;
}
}
// Re-open the capture on a different device while the aux stream stays up (the core stream id
// is unchanged — only the client-side capture source changes).
private void RestartAuxCapture()
{
if (_auxStreamId == 0) return;
DisposeAuxCapture();
_auxCapture = new InputDeviceCapture(_voiceSettings.AuxDeviceId);
_auxCapture.PcmFrameReady += OnAuxPcmFrame;
if (!_auxCapture.Start())
AddActivity("Failed to open aux input device");
}
private void DisposeAuxCapture()
{
if (_auxCapture == null) return;
_auxCapture.PcmFrameReady -= OnAuxPcmFrame;
_auxCapture.Stop();
_auxCapture.Dispose();
_auxCapture = null;
}
// Fired on the capture thread. vc_stream_feed_pcm is thread-safe, so feed directly. Gain is
// read live from settings each frame (so the volume slider takes effect immediately).
private void OnAuxPcmFrame(short[] pcm, int samplesPerChannel, int channels)
{
if (_auxStreamId == 0) return;
float gain = _voiceSettings.AuxGain / 100f;
if (gain != 1f)
{
for (int i = 0; i < pcm.Length; i++)
pcm[i] = (short)Math.Clamp((int)MathF.Round(pcm[i] * gain),
short.MinValue, short.MaxValue);
}
_client.StreamFeedPcm(_auxStreamId, pcm, samplesPerChannel, (uint)channels);
}
private void TrkOutputVolume_Scroll(object? sender, EventArgs e) => private void TrkOutputVolume_Scroll(object? sender, EventArgs e) =>
_client.SetOutputVolume(trkOutputVolume.Value / 100f); _client.SetOutputVolume(trkOutputVolume.Value / 100f);
private void RadioVad_CheckedChanged(object? sender, EventArgs e) private float VadThresholdFromSettings() =>
{ 0.1f * (1f - (_voiceSettings.VadThresholdSlider - 1f) / 99f);
if (!radioVad.Checked) return;
lblPttKey.Visible = false;
btnChangePtt.Visible = false;
lblVadThreshold.Visible = true;
trkVadThreshold.Visible = true;
if (_micStreamId != 0)
{
_client.SetInputMode(VcInputMode.VoiceActivation);
_client.SetVadThreshold(VadThresholdFromSlider());
}
}
private void RadioPtt_CheckedChanged(object? sender, EventArgs e)
{
if (!radioPtt.Checked) return;
lblPttKey.Text = $"({_pttKey})";
lblPttKey.Visible = true;
btnChangePtt.Visible = true;
lblVadThreshold.Visible = false;
trkVadThreshold.Visible = false;
if (_micStreamId != 0)
{
_client.SetInputMode(VcInputMode.PushToTalk);
_client.SetPushToTalk(false);
}
}
private void RadioAlwaysOn_CheckedChanged(object? sender, EventArgs e)
{
if (!radioAlwaysOn.Checked) return;
lblPttKey.Visible = false;
btnChangePtt.Visible = false;
lblVadThreshold.Visible = false;
trkVadThreshold.Visible = false;
if (_micStreamId != 0) _client.SetInputMode(VcInputMode.AlwaysOn);
}
private void TrkVadThreshold_Scroll(object? sender, EventArgs e)
{
if (_micStreamId != 0 && radioVad.Checked)
_client.SetVadThreshold(VadThresholdFromSlider());
}
private float VadThresholdFromSlider() =>
0.1f * (1f - (trkVadThreshold.Value - 1f) / 99f);
private VcInputMode CurrentInputMode() =>
radioPtt.Checked ? VcInputMode.PushToTalk :
radioAlwaysOn.Checked ? VcInputMode.AlwaysOn :
VcInputMode.VoiceActivation;
private void BtnChangePtt_Click(object? sender, EventArgs e)
{
using var dlg = new PttKeyCaptureDialog(_pttKey);
if (dlg.ShowDialog(this) == DialogResult.OK)
{
_pttKey = dlg.CapturedKey;
lblPttKey.Text = $"({_pttKey})";
}
}
private void CboInputDevice_SelectedIndexChanged(object? sender, EventArgs e)
{
if (_micStreamId == 0) return;
string? deviceId = (cboInputDevice.SelectedItem as DeviceInfo)?.Id;
_client.SetInputDevice(_micStreamId, deviceId);
}
// ── PTT key handling (focus-scoped) ─────────────────────────────────────── // ── PTT key handling (focus-scoped) ───────────────────────────────────────
private void MainForm_KeyDown(object? sender, KeyEventArgs e) private void MainForm_KeyDown(object? sender, KeyEventArgs e)
{ {
if (!radioPtt.Checked || e.KeyCode != _pttKey || _micStreamId == 0) return; if (_voiceSettings.SystemWidePtt) return; // handled by the WM_INPUT / Raw Input path
if ((VcInputMode)_voiceSettings.InputMode != VcInputMode.PushToTalk) return;
if (e.KeyCode != _pttKey || _micStreamId == 0) return;
if (ActiveControl is TextBox or RichTextBox) return; if (ActiveControl is TextBox or RichTextBox) return;
_client.SetPushToTalk(true); _client.SetPushToTalk(true);
lblPttKey.Text = $"({_pttKey} ▶)"; if (!_pttEngaged) // first key-down only, not auto-repeat
e.Handled = true; {
_pttEngaged = true;
_feedback.PlaySound(SoundEvent.Ptt);
}
e.Handled = e.SuppressKeyPress = true;
} }
private void MainForm_HotkeyDown(object? sender, KeyEventArgs e) private void MainForm_HotkeyDown(object? sender, KeyEventArgs e)
@@ -784,31 +915,120 @@ public partial class MainForm : Form
{ {
case Keys.V: case Keys.V:
BtnMicToggle_Click(null, EventArgs.Empty); BtnMicToggle_Click(null, EventArgs.Empty);
e.Handled = true;
break; break;
case Keys.S: case Keys.S:
BtnScreenShareToggle_Click(null, EventArgs.Empty); BtnScreenShareToggle_Click(null, EventArgs.Empty);
e.Handled = true;
break; break;
case Keys.M: case Keys.M:
chkMute.Checked = !chkMute.Checked; chkMute.Checked = !chkMute.Checked;
ApplySelfMute(); ApplySelfMute();
e.Handled = true;
break; break;
case Keys.D: case Keys.D:
chkDeafen.Checked = !chkDeafen.Checked; chkDeafen.Checked = !chkDeafen.Checked;
ApplySelfMute(); ApplySelfMute();
e.Handled = true;
break; break;
default:
return;
} }
// A handled hotkey: suppress the follow-on WM_CHAR so the focused control
// (channel tree / user list) doesn't emit the system "ding".
e.Handled = e.SuppressKeyPress = true;
} }
private void MainForm_KeyUp(object? sender, KeyEventArgs e) private void MainForm_KeyUp(object? sender, KeyEventArgs e)
{ {
if (!radioPtt.Checked || e.KeyCode != _pttKey || _micStreamId == 0) return; if (_voiceSettings.SystemWidePtt) return; // handled by the WM_INPUT / Raw Input path
if ((VcInputMode)_voiceSettings.InputMode != VcInputMode.PushToTalk) return;
if (e.KeyCode != _pttKey || _micStreamId == 0) return;
_client.SetPushToTalk(false); _client.SetPushToTalk(false);
lblPttKey.Text = $"({_pttKey})"; _pttEngaged = false;
e.Handled = true; e.Handled = e.SuppressKeyPress = true;
}
// ── System-wide PTT (Raw Input / WM_INPUT) ────────────────────────────────
// When the user enables "system-wide" PTT we observe the PTT key via the Raw Input API so it
// works while another app is focused. See VoiceCat.App.Native.RawInput for why this is used in
// preference to a low-level keyboard hook (antivirus keylogger heuristics).
protected override void OnHandleCreated(EventArgs e)
{
base.OnHandleCreated(e);
ApplySystemWidePtt();
}
protected override void OnHandleDestroyed(EventArgs e)
{
if (_rawInputRegistered)
{
RawInput.UnregisterKeyboardSink();
_rawInputRegistered = false;
}
base.OnHandleDestroyed(e);
}
protected override void WndProc(ref Message m)
{
if (m.Msg == RawInput.WM_INPUT && _voiceSettings.SystemWidePtt)
HandleRawInput(m.LParam);
base.WndProc(ref m);
}
/// <summary>Register or tear down the background keyboard sink to match the current
/// <see cref="VoiceSettings.SystemWidePtt"/> setting. Safe to call repeatedly.</summary>
private void ApplySystemWidePtt()
{
if (!IsHandleCreated) return; // OnHandleCreated will (re)apply once the handle exists
bool want = _voiceSettings.SystemWidePtt;
if (want && !_rawInputRegistered)
{
_rawInputRegistered = RawInput.RegisterKeyboardSink(Handle);
}
else if (!want && _rawInputRegistered)
{
RawInput.UnregisterKeyboardSink();
_rawInputRegistered = false;
// Release any PTT that was held via Raw Input so it can't stick after switching modes.
if (_micStreamId != 0) _client.SetPushToTalk(false);
_pttEngaged = false;
}
}
private void HandleRawInput(IntPtr lParam)
{
if ((VcInputMode)_voiceSettings.InputMode != VcInputMode.PushToTalk || _micStreamId == 0)
return;
if (!RawInput.TryParseKey(lParam, out ushort vkey, out bool keyUp)) return;
if (vkey != (ushort)_pttKey) return;
if (keyUp)
{
_client.SetPushToTalk(false);
_pttEngaged = false;
return;
}
// Key-down. Don't transmit while typing into our OWN text fields — matches the
// focus-scoped guard. When VoiceCat is in the background ContainsFocus is false, so the
// key still transmits (the whole point of system-wide PTT).
if (ContainsFocus && ActiveControl is TextBox or RichTextBox) return;
_client.SetPushToTalk(true);
if (!_pttEngaged) // first key-down only, not auto-repeat
{
_pttEngaged = true;
_feedback.PlaySound(SoundEvent.Ptt);
}
}
/// <summary>Watchdog (driven by the pump timer) that releases system-wide PTT if the key-up
/// was never observed — e.g. across an RDP or lock-screen focus switch — so PTT can't stick.</summary>
private void PttWatchdog()
{
if (!_pttEngaged || !_voiceSettings.SystemWidePtt || _micStreamId == 0) return;
if (!RawInput.IsKeyDown((int)_pttKey))
{
_client.SetPushToTalk(false);
_pttEngaged = false;
}
} }
// ── Channel navigation ──────────────────────────────────────────────────── // ── Channel navigation ────────────────────────────────────────────────────
@@ -858,13 +1078,17 @@ public partial class MainForm : Form
win = new PrivateMessageForm(_client, userId, nick, _selfUserId); win = new PrivateMessageForm(_client, userId, nick, _selfUserId);
win.FormClosed += (_, _) => _pmWindows.Remove(userId); win.FormClosed += (_, _) => _pmWindows.Remove(userId);
_pmWindows[userId] = win; _pmWindows[userId] = win;
win.Show(this); // Show without an owner: an owned form is forced to stay above MainForm and pulls
// focus back to itself, so the main window can't be worked in while a PM is open.
// OnFormClosed already closes any open PM windows, so this doesn't leak.
win.Show();
} }
else else
{ {
if (win.WindowState == FormWindowState.Minimized) if (win.WindowState == FormWindowState.Minimized)
win.WindowState = FormWindowState.Normal; win.WindowState = FormWindowState.Normal;
win.BringToFront(); // Raise it for this explicit user-initiated open without the owner-style focus trap.
win.Activate();
} }
return win; return win;
} }
@@ -890,7 +1114,7 @@ public partial class MainForm : Form
OpenPmWindow(dlg.SelectedUserId); OpenPmWindow(dlg.SelectedUserId);
} }
// ── M5: Moderation helpers ──────────────────────────────────────────────── // ── Moderation helpers ─────────────────────────────────────────────────────
private void UpdateSelfServerMuteState(bool muted, bool deafened) private void UpdateSelfServerMuteState(bool muted, bool deafened)
{ {
@@ -922,8 +1146,8 @@ public partial class MainForm : Form
var editInfo = new ChannelEditInfo( var editInfo = new ChannelEditInfo(
channel.Id, channel.ParentId, channel.Name, channel.Topic, channel.Id, channel.ParentId, channel.Name, channel.Topic,
channel.PasswordProtected, null, channel.MaxUsers, 0, channel.PasswordProtected, null, channel.MaxUsers, channel.SortOrder,
new AudioConfigInfo(0, false, 48000, 0, 20, 0, true, 0, false, 10, false)); channel.Audio);
using var dlg = new ChannelEditDialog(_channels, editInfo); using var dlg = new ChannelEditDialog(_channels, editInfo);
if (dlg.ShowDialog(this) != DialogResult.OK || dlg.Result is null) return; if (dlg.ShowDialog(this) != DialogResult.OK || dlg.Result is null) return;
@@ -1067,9 +1291,11 @@ public partial class MainForm : Form
foreach (var win in _pmWindows.Values.ToList()) win.Close(); foreach (var win in _pmWindows.Values.ToList()) win.Close();
_pmWindows.Clear(); _pmWindows.Clear();
if (_screenStreamId != 0) { _screenMixer?.Stop(); _screenMixer?.Dispose(); _screenMixer = null; _client.StopStream(_screenStreamId); } if (_screenStreamId != 0) { _screenMixer?.Stop(); _screenMixer?.Dispose(); _screenMixer = null; _client.StopStream(_screenStreamId); }
if (_auxStreamId != 0) StopAuxStream();
if (_micStreamId != 0) _client.StopStream(_micStreamId); if (_micStreamId != 0) _client.StopStream(_micStreamId);
_client.Disconnect(); _client.Disconnect();
_client.Dispose(); _client.Dispose();
_feedback.Dispose();
base.OnFormClosed(e); base.OnFormClosed(e);
} }

View File

@@ -0,0 +1,128 @@
using VoiceCat.App.Notifications;
namespace VoiceCat.App.Forms;
/// <summary>
/// Edits and persists <see cref="FeedbackSettings"/> (event sounds + spoken announcements).
/// On OK the supplied <see cref="EventFeedback"/> is updated live and the settings saved.
/// </summary>
public sealed class NotificationSettingsForm : Form
{
private readonly EventFeedback _feedback;
private readonly CheckBox _chkSounds;
private readonly CheckBox _chkSpeech;
private readonly TrackBar _trkVolume;
private readonly CheckBox _chkSelfTalk;
private readonly CheckBox _chkPtt;
public NotificationSettingsForm(EventFeedback feedback)
{
_feedback = feedback;
var s = feedback.Settings;
Text = "Notification settings";
FormBorderStyle = FormBorderStyle.FixedDialog;
MaximizeBox = false;
MinimizeBox = false;
StartPosition = FormStartPosition.CenterParent;
AutoScaleMode = AutoScaleMode.Font;
ClientSize = new Size(360, 300);
_chkSounds = new CheckBox
{
Text = "Play event &sounds",
Location = new Point(12, 12),
AutoSize = true,
Checked = s.Sounds,
};
var lblVolume = new Label
{
Text = "Sound &volume:",
Location = new Point(12, 42),
AutoSize = true,
};
_trkVolume = new TrackBar
{
Location = new Point(12, 64),
Size = new Size(330, 45),
Minimum = 0,
Maximum = 100,
TickFrequency = 10,
Value = (int)Math.Round(Math.Clamp(s.Volume, 0f, 1f) * 100),
};
_trkVolume.AccessibleName = "Sound volume";
_chkSpeech = new CheckBox
{
Text = "&Speak events (text-to-speech)",
Location = new Point(12, 118),
AutoSize = true,
Checked = s.Speech,
};
var lblSpeechHint = new Label
{
Text = feedback.SpeechAvailable
? "Announces joins/leaves and reads message text aloud."
: "No speech engine available on this machine.",
Location = new Point(30, 142),
AutoSize = true,
ForeColor = SystemColors.GrayText,
};
var lblOptional = new Label
{
Text = "Optional sounds:",
Location = new Point(12, 174),
AutoSize = true,
};
_chkSelfTalk = new CheckBox
{
Text = "Your own voice-&activity start/stop",
Location = new Point(12, 198),
AutoSize = true,
Checked = s.SelfTalkSounds,
};
_chkPtt = new CheckBox
{
Text = "&Push-to-talk cue",
Location = new Point(12, 224),
AutoSize = true,
Checked = s.PttSound,
};
var btnOk = new Button
{
Text = "&OK",
DialogResult = DialogResult.OK,
Location = new Point(192, 262),
Size = new Size(75, 27),
};
var btnCancel = new Button
{
Text = "&Cancel",
DialogResult = DialogResult.Cancel,
Location = new Point(273, 262),
Size = new Size(75, 27),
};
AcceptButton = btnOk;
CancelButton = btnCancel;
Controls.AddRange([_chkSounds, lblVolume, _trkVolume, _chkSpeech, lblSpeechHint,
lblOptional, _chkSelfTalk, _chkPtt, btnOk, btnCancel]);
btnOk.Click += (_, _) => Apply();
}
private void Apply()
{
var s = _feedback.Settings;
s.Sounds = _chkSounds.Checked;
s.Speech = _chkSpeech.Checked;
s.Volume = _trkVolume.Value / 100f;
s.SelfTalkSounds = _chkSelfTalk.Checked;
s.PttSound = _chkPtt.Checked;
s.Save();
}
}

View File

@@ -1,7 +1,7 @@
namespace VoiceCat.App.Forms; namespace VoiceCat.App.Forms;
/// <summary>Small modal for "type a password right now" — used when a saved server's /// <summary>Small modal for "type a password right now" — used when a saved server's
/// password wasn't remembered, and (Phase E) for password-protected channel joins.</summary> /// password wasn't remembered, and for password-protected channel joins.</summary>
public partial class PasswordPromptDialog : Form public partial class PasswordPromptDialog : Form
{ {
public string Password => txtPassword.Text; public string Password => txtPassword.Text;

View File

@@ -117,11 +117,11 @@ public sealed class PerUserTuningDialog : Form
var trkGain = new TrackBar var trkGain = new TrackBar
{ {
AccessibleName = $"Gain for {s.Label}", AccessibleName = $"Gain for {s.Label}",
AccessibleDescription = "Volume level for this stream. 100 is normal (1.0×), 200 is double.", AccessibleDescription = "Volume level for this stream. 100 is normal (1.0×); range 0400 percent.",
Location = new Point(12, y + 20), Location = new Point(12, y + 20),
Size = new Size(260, 45), Size = new Size(260, 45),
Minimum = 0, Minimum = 0,
Maximum = 200, Maximum = 400,
Value = ClampToTrack(gain0), Value = ClampToTrack(gain0),
TickFrequency = 25, TickFrequency = 25,
SmallChange = 5, SmallChange = 5,
@@ -150,7 +150,7 @@ public sealed class PerUserTuningDialog : Form
var chkNr = new CheckBox var chkNr = new CheckBox
{ {
Text = "&Noise reduction (planned — currently passthrough)", Text = "&Noise reduction",
AutoSize = true, AutoSize = true,
Location = new Point(96, y + 48), Location = new Point(96, y + 48),
Checked = nr0, Checked = nr0,
@@ -186,7 +186,7 @@ public sealed class PerUserTuningDialog : Form
private static int ClampToTrack(float gain) private static int ClampToTrack(float gain)
{ {
int v = (int)Math.Round(gain * 100f); int v = (int)Math.Round(gain * 100f);
return Math.Max(0, Math.Min(200, v)); return Math.Max(0, Math.Min(400, v));
} }
private sealed class StreamRow private sealed class StreamRow

View File

@@ -26,6 +26,11 @@ public sealed class PrivateMessageForm : Form
ClientSize = new Size(480, 360); ClientSize = new Size(480, 360);
MinimumSize = new Size(320, 240); MinimumSize = new Size(320, 240);
StartPosition = FormStartPosition.Manual; StartPosition = FormStartPosition.Manual;
KeyPreview = true;
KeyDown += (_, e) =>
{
if (e.KeyCode == Keys.Escape) { e.Handled = e.SuppressKeyPress = true; Close(); }
};
_rtbHistory = new RichTextBox _rtbHistory = new RichTextBox
{ {

View File

@@ -3,7 +3,7 @@ using VoiceCat.Interop;
namespace VoiceCat.App.Forms; namespace VoiceCat.App.Forms;
/// <summary> /// <summary>
/// TOFU server-identity confirmation (M4). Shown only for VcTofuStatus.FirstConnect/Mismatch /// TOFU server-identity confirmation. Shown only for VcTofuStatus.FirstConnect/Mismatch
/// — never Matched (that's the silent-success "subsequent connects verify the pin" path /// — never Matched (that's the silent-success "subsequent connects verify the pin" path
/// docs/security.md describes; showing a dialog on every routine reconnect would be exactly /// docs/security.md describes; showing a dialog on every routine reconnect would be exactly
/// the "overly chatty" experience this project avoids elsewhere too). /// the "overly chatty" experience this project avoids elsewhere too).

View File

@@ -0,0 +1,88 @@
using System.Text.Json;
namespace VoiceCat.App.Models;
/// <summary>
/// User's send-side voice input preferences — transmission mode, VAD sensitivity, mic input
/// gain, and the push-to-talk key. Persisted to %AppData%\VoiceCat\voice.json, same pattern as
/// <see cref="ServerListStore"/> and FeedbackSettings: a missing or corrupt file yields defaults
/// rather than throwing. Slider-position values are stored as-is so MainForm can restore the
/// TrackBars directly.
/// </summary>
public sealed class VoiceSettings
{
/// <summary>Transmission mode: 0 = voice activation, 1 = push-to-talk, 2 = always on
/// (matches Interop's VcInputMode).</summary>
public int InputMode { get; set; } = 0;
/// <summary>VAD sensitivity slider position, 1100 (default mirrors the designer's 76).</summary>
public int VadThresholdSlider { get; set; } = 76;
/// <summary>Microphone input gain slider position, 0400 percent (100 = unity).</summary>
public int MicGain { get; set; } = 100;
/// <summary>Send-side mic noise reduction (RNNoise) toggle. MIC stream only, mono only;
/// denoises captured mic PCM before input gain and VAD/PTT gate so everyone hears the
/// cleaned signal. Independent of the per-listener receive-side NR.</summary>
public bool MicNoiseReduction { get; set; } = false;
/// <summary>Capture the mic in stereo (interleaved L/R) instead of mono. Off by default. Real
/// stereo only reaches the wire on a stereo channel; on a mono channel the core folds the mic
/// to mono. Applied when the capture device next starts (Join Voice or an audio restart).</summary>
public bool StereoMic { get; set; } = false;
/// <summary>Push-to-talk key, stored as the integer value of System.Windows.Forms.Keys.</summary>
public int PttKey { get; set; } = (int)Keys.F8;
/// <summary>Make push-to-talk work system-wide — i.e. while another app is focused. When on,
/// MainForm registers a background Raw Input keyboard device (WM_INPUT + RIDEV_INPUTSINK) so the
/// PTT key is observed even when VoiceCat is not in the foreground; when off, PTT is focus-scoped
/// (only fires while the VoiceCat window has focus). Raw Input is used instead of a low-level
/// keyboard hook to avoid antivirus keylogger heuristics.</summary>
public bool SystemWidePtt { get; set; } = true;
/// <summary>Saved device ID from the last session; null means use the system default.</summary>
public string? InputDeviceId { get; set; } = null;
/// <summary>Whether the secondary "aux" outgoing stream is enabled. The aux stream is a
/// second hardware input device the client captures itself and feeds to the core via
/// vc_stream_feed_pcm (kind = AUX_DEVICE, external_feed). Lets a user transmit e.g. mic +
/// a line-in at once.</summary>
public bool AuxEnabled { get; set; } = false;
/// <summary>WASAPI endpoint id of the aux capture device; null = system default capture
/// device. NOTE: this is a WASAPI device id (from <see cref="Audio.InputDeviceEnumerator"/>),
/// NOT a core/miniaudio id — the aux device is opened client-side, so the two id spaces differ.</summary>
public string? AuxDeviceId { get; set; } = null;
/// <summary>Aux input volume slider position, 0400 percent (100 = unity). Applied client-side
/// to the captured PCM before feeding (the core's input gain is mic-only and global).</summary>
public int AuxGain { get; set; } = 100;
private static readonly JsonSerializerOptions JsonOptions = new() { WriteIndented = true };
private static string AppDataDir => Path.Combine(
Environment.GetFolderPath(Environment.SpecialFolder.ApplicationData), "VoiceCat");
private static string FilePath => Path.Combine(AppDataDir, "voice.json");
public static VoiceSettings Load()
{
try
{
if (!File.Exists(FilePath)) return new VoiceSettings();
string json = File.ReadAllText(FilePath);
return JsonSerializer.Deserialize<VoiceSettings>(json) ?? new VoiceSettings();
}
catch
{
return new VoiceSettings();
}
}
public void Save()
{
Directory.CreateDirectory(AppDataDir);
File.WriteAllText(FilePath, JsonSerializer.Serialize(this, JsonOptions));
}
}

View File

@@ -0,0 +1,147 @@
using System.Runtime.InteropServices;
namespace VoiceCat.App.Native;
/// <summary>
/// Thin wrapper over the Win32 Raw Input API (user32) used to observe the push-to-talk key
/// system-wide — i.e. while another application is focused.
///
/// We deliberately use Raw Input (<c>RegisterRawInputDevices</c> + <c>WM_INPUT</c> with
/// <c>RIDEV_INPUTSINK</c>) rather than a low-level keyboard hook (<c>SetWindowsHookEx</c> /
/// <c>WH_KEYBOARD_LL</c>). A low-level hook is the textbook keylogger pattern and is exactly
/// what antivirus heuristics flag — especially for an unsigned, statically linked MinGW binary
/// like ours. Raw Input involves no DLL injection and no global hook: the OS simply posts
/// <c>WM_INPUT</c> messages to our own window's message queue (it is the same input path games
/// use), so it does not trip those heuristics. It also passes keystrokes through to the focused
/// app rather than swallowing them.
///
/// Only the keyboard device is registered; <see cref="TryParseKey"/> filters to the single PTT
/// virtual-key code one level up (MainForm).
/// </summary>
internal static partial class RawInput
{
// ── Constants ──────────────────────────────────────────────────────────────────────────
public const int WM_INPUT = 0x00FF;
private const uint RID_INPUT = 0x10000003; // GetRawInputData: get the raw data
private const uint RIM_TYPEKEYBOARD = 1; // RAWINPUTHEADER.dwType for a keyboard
private const uint RIDEV_INPUTSINK = 0x00000100; // receive input even when not in foreground
private const uint RIDEV_REMOVE = 0x00000001; // stop receiving input from the device
private const ushort RI_KEY_BREAK = 0x01; // RAWKEYBOARD.Flags bit set on key-up
private const ushort HID_USAGE_PAGE_GENERIC = 0x01;
private const ushort HID_USAGE_GENERIC_KEYBOARD = 0x06;
// ── Structs (must match the Win32 layout exactly) ──────────────────────────────────────
[StructLayout(LayoutKind.Sequential)]
private struct RAWINPUTDEVICE
{
public ushort usUsagePage;
public ushort usUsage;
public uint dwFlags;
public IntPtr hwndTarget;
}
[StructLayout(LayoutKind.Sequential)]
private struct RAWINPUTHEADER
{
public uint dwType;
public uint dwSize;
public IntPtr hDevice;
public IntPtr wParam;
}
[StructLayout(LayoutKind.Sequential)]
private struct RAWKEYBOARD
{
public ushort MakeCode;
public ushort Flags;
public ushort Reserved;
public ushort VKey;
public uint Message;
public uint ExtraInformation;
}
// ── P/Invoke (source-generated via LibraryImport, matching VoiceCat.Interop) ────────────
[LibraryImport("user32.dll", SetLastError = true)]
[return: MarshalAs(UnmanagedType.Bool)]
private static partial bool RegisterRawInputDevices(
ref RAWINPUTDEVICE pRawInputDevices, uint uiNumDevices, uint cbSize);
[LibraryImport("user32.dll")]
private static partial uint GetRawInputData(
IntPtr hRawInput, uint uiCommand, IntPtr pData, ref uint pcbSize, uint cbSizeHeader);
[LibraryImport("user32.dll")]
private static partial short GetAsyncKeyState(int vKey);
// ── Public helpers ─────────────────────────────────────────────────────────────────────
/// <summary>Register a background keyboard sink targeting <paramref name="hwnd"/>. With
/// RIDEV_INPUTSINK the window receives WM_INPUT even when it is not in the foreground.</summary>
public static bool RegisterKeyboardSink(IntPtr hwnd)
{
var rid = new RAWINPUTDEVICE
{
usUsagePage = HID_USAGE_PAGE_GENERIC,
usUsage = HID_USAGE_GENERIC_KEYBOARD,
dwFlags = RIDEV_INPUTSINK,
hwndTarget = hwnd,
};
return RegisterRawInputDevices(ref rid, 1, (uint)Marshal.SizeOf<RAWINPUTDEVICE>());
}
/// <summary>Tear down the keyboard sink. For RIDEV_REMOVE the target handle must be NULL.</summary>
public static bool UnregisterKeyboardSink()
{
var rid = new RAWINPUTDEVICE
{
usUsagePage = HID_USAGE_PAGE_GENERIC,
usUsage = HID_USAGE_GENERIC_KEYBOARD,
dwFlags = RIDEV_REMOVE,
hwndTarget = IntPtr.Zero,
};
return RegisterRawInputDevices(ref rid, 1, (uint)Marshal.SizeOf<RAWINPUTDEVICE>());
}
/// <summary>Decode a WM_INPUT message's <paramref name="hRawInput"/> (the message's LParam).
/// Returns false for anything that isn't a keyboard event. On success, <paramref name="vkey"/>
/// is the virtual-key code and <paramref name="keyUp"/> distinguishes a release from a press
/// (auto-repeat arrives as repeated presses).</summary>
public static bool TryParseKey(IntPtr hRawInput, out ushort vkey, out bool keyUp)
{
vkey = 0;
keyUp = false;
uint headerSize = (uint)Marshal.SizeOf<RAWINPUTHEADER>();
uint size = 0;
// First call (pData == NULL) returns 0 on success and fills `size` with the buffer length.
if (GetRawInputData(hRawInput, RID_INPUT, IntPtr.Zero, ref size, headerSize) != 0 || size == 0)
return false;
IntPtr buf = Marshal.AllocHGlobal((int)size);
try
{
if (GetRawInputData(hRawInput, RID_INPUT, buf, ref size, headerSize) != size)
return false;
var header = Marshal.PtrToStructure<RAWINPUTHEADER>(buf);
if (header.dwType != RIM_TYPEKEYBOARD)
return false;
// The keyboard payload immediately follows the (8-byte-aligned) header.
var kb = Marshal.PtrToStructure<RAWKEYBOARD>(buf + (int)headerSize);
vkey = kb.VKey;
keyUp = (kb.Flags & RI_KEY_BREAK) != 0;
return true;
}
finally
{
Marshal.FreeHGlobal(buf);
}
}
/// <summary>True if the given virtual-key code is currently held down. Used as a watchdog to
/// catch a missed key-up (e.g. across an RDP / lock-screen focus switch) so PTT can't stick.</summary>
public static bool IsKeyDown(int vKey) => (GetAsyncKeyState(vKey) & 0x8000) != 0;
}

View File

@@ -0,0 +1,42 @@
namespace VoiceCat.App.Notifications;
/// <summary>
/// Combines the sound pool and the speech announcer behind the user's <see cref="FeedbackSettings"/>.
/// MainForm calls <see cref="PlaySound"/> / <see cref="Speak"/> from inside its existing event
/// handlers (which already resolve nicknames and channel membership), so this type only owns the
/// "should I, and how" policy — not the event routing.
/// </summary>
public sealed class EventFeedback : IDisposable
{
private readonly SoundPlayerPool _sounds = new();
private readonly SpeechAnnouncer _speech = new();
public FeedbackSettings Settings { get; set; }
public EventFeedback(FeedbackSettings settings) => Settings = settings;
/// <summary>True when speech is both enabled and actually available on this machine.</summary>
public bool SpeechAvailable => _speech.Available;
public void PlaySound(SoundEvent ev)
{
if (!Settings.Sounds || Settings.Volume <= 0f) return;
// The two opt-in categories are gated by their own flags.
if ((ev is SoundEvent.VaStart or SoundEvent.VaStop) && !Settings.SelfTalkSounds) return;
if (ev is SoundEvent.Ptt && !Settings.PttSound) return;
_sounds.Play(ev);
}
/// <summary>Announce <paramref name="text"/> when spoken feedback is enabled.</summary>
public void Speak(string text)
{
if (!Settings.Speech) return;
_speech.Speak(text);
}
public void Dispose()
{
_sounds.Dispose();
_speech.Dispose();
}
}

View File

@@ -0,0 +1,55 @@
using System.Text.Json;
namespace VoiceCat.App.Notifications;
/// <summary>
/// User preferences for event sounds and spoken (text-to-speech) feedback. Persisted to
/// %AppData%\VoiceCat\feedback.json — same pattern as <see cref="Models.ServerListStore"/>:
/// a missing or corrupt file yields defaults rather than throwing.
/// </summary>
public sealed class FeedbackSettings
{
/// <summary>Master switch for event sound effects.</summary>
public bool Sounds { get; set; } = true;
/// <summary>Master switch for spoken announcements (Prismatoid TTS). Off by default.</summary>
public bool Speech { get; set; } = false;
/// <summary>Sound-effects volume, 0..1. 0 mutes effects (SoundPlayer has no gain control,
/// so this is honoured as a mute gate — see <see cref="SoundPlayerPool"/>).</summary>
public float Volume { get; set; } = 1.0f;
/// <summary>Play va_start/va_stop for your own voice-activity transitions. Off by default
/// (fires on every utterance — noisy).</summary>
public bool SelfTalkSounds { get; set; } = false;
/// <summary>Play the ptt cue when push-to-talk is engaged. Off by default.</summary>
public bool PttSound { get; set; } = false;
private static readonly JsonSerializerOptions JsonOptions = new() { WriteIndented = true };
private static string AppDataDir => Path.Combine(
Environment.GetFolderPath(Environment.SpecialFolder.ApplicationData), "VoiceCat");
private static string FilePath => Path.Combine(AppDataDir, "feedback.json");
public static FeedbackSettings Load()
{
try
{
if (!File.Exists(FilePath)) return new FeedbackSettings();
string json = File.ReadAllText(FilePath);
return JsonSerializer.Deserialize<FeedbackSettings>(json) ?? new FeedbackSettings();
}
catch
{
return new FeedbackSettings();
}
}
public void Save()
{
Directory.CreateDirectory(AppDataDir);
File.WriteAllText(FilePath, JsonSerializer.Serialize(this, JsonOptions));
}
}

View File

@@ -0,0 +1,47 @@
namespace VoiceCat.App.Notifications;
/// <summary>
/// The set of audible event cues. Each maps to a WAV in the app's <c>sounds\</c> folder
/// (copied from <c>assets/sounds/</c> at build time). The same logical set is mirrored in the
/// macOS/iOS clients so feedback stays consistent across platforms.
/// </summary>
public enum SoundEvent
{
ChannelJoin,
ChannelLeave,
ChannelRecv, // channel text message from someone else
ChannelSent, // channel text message I sent
PmRecv,
PmSent,
Login,
Logout,
ConnectionLost,
VoiceOn,
VoiceOff,
VaStart, // my voice-activity began (off by default)
VaStop, // my voice-activity ended (off by default)
Ptt, // push-to-talk engaged (off by default)
}
internal static class SoundEventExtensions
{
/// <summary>WAV file name (without directory) for each event, matching assets/sounds/.</summary>
public static string FileName(this SoundEvent ev) => ev switch
{
SoundEvent.ChannelJoin => "channel_join.wav",
SoundEvent.ChannelLeave => "channel_leave.wav",
SoundEvent.ChannelRecv => "channel_recv.wav",
SoundEvent.ChannelSent => "channel_sent.wav",
SoundEvent.PmRecv => "pm_recv.wav",
SoundEvent.PmSent => "pm_sent.wav",
SoundEvent.Login => "login.wav",
SoundEvent.Logout => "logout.wav",
SoundEvent.ConnectionLost => "connection_lost.wav",
SoundEvent.VoiceOn => "voice_on.wav",
SoundEvent.VoiceOff => "voice_off.wav",
SoundEvent.VaStart => "va_start.wav",
SoundEvent.VaStop => "va_stop.wav",
SoundEvent.Ptt => "ptt.wav",
_ => throw new ArgumentOutOfRangeException(nameof(ev), ev, null),
};
}

View File

@@ -0,0 +1,58 @@
using System.Media;
namespace VoiceCat.App.Notifications;
/// <summary>
/// Plays short event cues from the app's <c>sounds\</c> folder. Each WAV is loaded once into a
/// cached <see cref="SoundPlayer"/> (built into the Windows Desktop framework — no extra
/// dependency). <see cref="Play"/> is non-blocking (<see cref="SoundPlayer.Play"/> renders on a
/// background thread).
///
/// Note: <see cref="SoundPlayer"/> exposes no gain control, so volume is honoured only as a mute
/// gate (volume == 0 → silent). If finer control or reliable overlapping playback is needed
/// later, swap this for NAudio.
/// </summary>
public sealed class SoundPlayerPool : IDisposable
{
private static readonly string SoundsDir =
Path.Combine(AppContext.BaseDirectory, "sounds");
private readonly Dictionary<SoundEvent, SoundPlayer?> _players = [];
/// <summary>Resolve, load and cache the player for an event. Returns null if the file is
/// missing or fails to load — playback then silently no-ops.</summary>
private SoundPlayer? Get(SoundEvent ev)
{
if (_players.TryGetValue(ev, out var cached)) return cached;
SoundPlayer? player = null;
try
{
string path = Path.Combine(SoundsDir, ev.FileName());
if (File.Exists(path))
{
player = new SoundPlayer(path);
player.Load();
}
}
catch
{
player = null; // unreadable/invalid WAV — degrade to silence
}
_players[ev] = player;
return player;
}
public void Play(SoundEvent ev)
{
try { Get(ev)?.Play(); }
catch { /* never let a notification sound surface as an error */ }
}
public void Dispose()
{
foreach (var p in _players.Values) p?.Dispose();
_players.Clear();
}
}

View File

@@ -0,0 +1,56 @@
using Prismatoid;
namespace VoiceCat.App.Notifications;
/// <summary>
/// Spoken event announcements via Prismatoid (bindings for the Prism speech library —
/// integrates with the active screen reader / system speech, no extra runtime deps).
///
/// All construction and calls are wrapped so a machine with no available speech backend simply
/// degrades to silence rather than crashing the app.
/// </summary>
public sealed class SpeechAnnouncer : IDisposable
{
private readonly PrismContext? _context;
private readonly object? _backend; // SpeechBackend; held as object to keep this resilient
public SpeechAnnouncer()
{
try
{
_context = new PrismContext();
_backend = _context.AcquireBestBackend();
}
catch
{
_context?.Dispose();
_context = null;
_backend = null;
}
}
/// <summary>True when a speech backend is available on this machine.</summary>
public bool Available => _backend is not null;
/// <summary>Speak <paramref name="text"/>. Queued (does not interrupt prior speech) so a
/// burst of events is read in order.</summary>
public void Speak(string text)
{
if (string.IsNullOrWhiteSpace(text)) return;
try
{
if (_backend is { } b)
((dynamic)b).Speak(text, interrupt: false);
}
catch
{
/* never surface a speech failure */
}
}
public void Dispose()
{
try { (_backend as IDisposable)?.Dispose(); } catch { /* ignore */ }
_context?.Dispose();
}
}

View File

@@ -7,30 +7,21 @@ internal static class Program
[STAThread] [STAThread]
private static void Main() private static void Main()
{ {
// Diagnostic-logging-only for now (manual debugging session) — every exception that // Surface exceptions that WinForms' default message-loop handling would otherwise
// would otherwise be silently caught by WinForms' default message-loop handling (or // swallow silently (or crash with no visible cause).
// crash with no visible cause) gets printed to stdout/stderr first.
Application.ThreadException += (_, e) => Application.ThreadException += (_, e) =>
Console.Error.WriteLine($"[UNHANDLED ThreadException] {e.Exception}"); Console.Error.WriteLine($"[UNHANDLED ThreadException] {e.Exception}");
AppDomain.CurrentDomain.UnhandledException += (_, e) => AppDomain.CurrentDomain.UnhandledException += (_, e) =>
Console.Error.WriteLine($"[UNHANDLED AppDomain exception] {e.ExceptionObject}"); Console.Error.WriteLine($"[UNHANDLED AppDomain exception] {e.ExceptionObject}");
Console.WriteLine("VoiceCat.App starting...");
ApplicationConfiguration.Initialize(); ApplicationConfiguration.Initialize();
using var connectDialog = new ConnectDialog(); using var connectDialog = new ConnectDialog();
Console.WriteLine("Showing ConnectDialog...");
var result = connectDialog.ShowDialog(); var result = connectDialog.ShowDialog();
Console.WriteLine($"ConnectDialog closed with DialogResult={result}, ConnectedClient={(connectDialog.ConnectedClient is null ? "null" : "set")}");
if (result != DialogResult.OK || connectDialog.ConnectedClient is null) if (result != DialogResult.OK || connectDialog.ConnectedClient is null)
{
Console.WriteLine("Exiting (cancelled or no connected client).");
return; return;
}
Console.WriteLine("Launching MainForm...");
Application.Run(new MainForm(connectDialog.ConnectedClient, connectDialog.SelfUserId, Application.Run(new MainForm(connectDialog.ConnectedClient, connectDialog.SelfUserId,
connectDialog.Nickname)); connectDialog.Nickname, connectDialog.ServerName));
Console.WriteLine("MainForm closed. Exiting.");
} }
} }

View File

@@ -4,6 +4,21 @@
<ProjectReference Include="..\VoiceCat.Interop\VoiceCat.Interop.csproj" /> <ProjectReference Include="..\VoiceCat.Interop\VoiceCat.Interop.csproj" />
</ItemGroup> </ItemGroup>
<!-- Prismatoid — .NET bindings for the Prism speech library, used for spoken event
announcements (text-to-speech). net10.0, no transitive dependencies. -->
<ItemGroup>
<PackageReference Include="Prismatoid" Version="0.3.0" />
</ItemGroup>
<!-- Event-cue WAVs (shared with the macOS/iOS clients) copied into sounds\ next to the exe;
SoundPlayerPool loads them from AppContext.BaseDirectory\sounds. -->
<ItemGroup>
<Content Include="..\..\..\assets\sounds\*.wav">
<Link>sounds\%(Filename)%(Extension)</Link>
<CopyToOutputDirectory>PreserveNewest</CopyToOutputDirectory>
</Content>
</ItemGroup>
<PropertyGroup> <PropertyGroup>
<OutputType>WinExe</OutputType> <OutputType>WinExe</OutputType>
<TargetFramework>net10.0-windows</TargetFramework> <TargetFramework>net10.0-windows</TargetFramework>

View File

@@ -23,10 +23,10 @@ public sealed class VoiceCatClientSmokeTests : IDisposable
_tempDir = Path.Combine(Path.GetTempPath(), "vc_csharp_smoke_" + Guid.NewGuid().ToString("N")); _tempDir = Path.Combine(Path.GetTempPath(), "vc_csharp_smoke_" + Guid.NewGuid().ToString("N"));
Directory.CreateDirectory(_tempDir); Directory.CreateDirectory(_tempDir);
string serverExe = Path.Combine(FindRepoRoot(), "build", "m1-dev", "bin", "voicecat-server.exe"); string serverExe = Path.Combine(FindRepoRoot(), "build", "dev", "bin", "voicecat-server.exe");
Assert.True(File.Exists(serverExe), Assert.True(File.Exists(serverExe),
$"voicecat-server.exe not found at '{serverExe}' — build the m1-dev preset first " + $"voicecat-server.exe not found at '{serverExe}' — build the dev preset first " +
"(cmake --preset m1-dev && cmake --build --preset m1-dev)."); "(cmake --preset dev && cmake --build --preset dev).");
var psi = new ProcessStartInfo(serverExe) var psi = new ProcessStartInfo(serverExe)
{ {
@@ -56,9 +56,9 @@ public sealed class VoiceCatClientSmokeTests : IDisposable
Assert.True(port is not null, "voicecat-server.exe did not report a bound TCP port within 10s."); Assert.True(port is not null, "voicecat-server.exe did not report a bound TCP port within 10s.");
_port = port!.Value; _port = port!.Value;
// M5: provision a known admin account so we can exercise moderation wrappers end-to-end. // Provision a known admin account so we can exercise moderation wrappers end-to-end.
string adminExe = Path.Combine(FindRepoRoot(), "build", "m1-dev", "bin", "voicecat-admin.exe"); string adminExe = Path.Combine(FindRepoRoot(), "build", "dev", "bin", "voicecat-admin.exe");
Assert.True(File.Exists(adminExe), "voicecat-admin.exe not found — build the m1-dev preset."); Assert.True(File.Exists(adminExe), "voicecat-admin.exe not found — build the dev preset.");
var adminPsi = new ProcessStartInfo(adminExe) var adminPsi = new ProcessStartInfo(adminExe)
{ {
Arguments = $"--data-dir \"{_tempDir}\" account add admin2 --admin --password testpassword123", Arguments = $"--data-dir \"{_tempDir}\" account add admin2 --admin --password testpassword123",
@@ -140,14 +140,14 @@ public sealed class VoiceCatClientSmokeTests : IDisposable
var channels = client.ListChannels(); var channels = client.ListChannels();
Assert.Contains(channels, c => c.Id == 1 && c.Name == "Lobby"); Assert.Contains(channels, c => c.Id == 1 && c.Name == "Lobby");
// M5: permissions getter round-trip. // Permissions getter round-trip.
var perms = client.GetPermissions(); var perms = client.GetPermissions();
Assert.False(perms.IsAdmin); Assert.False(perms.IsAdmin);
Assert.False(perms.CanKick); Assert.False(perms.CanKick);
// M5: moderation request wrappers queue without error. As a guest, account listing // Moderation request wrappers queue without error. As a guest, account listing
// is rejected by the server with a GenericResult, which proves the wrapper path works // is rejected by the server with a GenericResult, which proves the wrapper path works
// end-to-end and that the new event type is delivered through P/Invoke. // end-to-end and that the event type is delivered through P/Invoke.
Assert.Equal(VcResult.Ok, client.RequestAccountList()); Assert.Equal(VcResult.Ok, client.RequestAccountList());
Assert.True(PumpUntil(client, Assert.True(PumpUntil(client,
() => events.Any(e => e.Type == VcEventType.GenericResult), 3000), () => events.Any(e => e.Type == VcEventType.GenericResult), 3000),

View File

@@ -37,7 +37,7 @@ public enum VcConnectionState
TlsHandshake = 2, TlsHandshake = 2,
Authenticating = 3, Authenticating = 3,
Connected = 4, Connected = 4,
/// <summary>M4: handshake succeeded, waiting on vc_confirm_server_identity().</summary> /// <summary>Handshake succeeded, waiting on vc_confirm_server_identity().</summary>
VerifyingIdentity = 5, VerifyingIdentity = 5,
} }
@@ -84,14 +84,16 @@ public enum VcEventType
TalkState = 9, TalkState = 9,
Error = 10, Error = 10,
Disconnected = 11, Disconnected = 11,
/// <summary>M4: reply to VoiceCatClient.JoinChannelAsync's underlying vc_join_channel.</summary> /// <summary>Reply to VoiceCatClient.JoinChannelAsync's underlying vc_join_channel.</summary>
JoinResult = 12, JoinResult = 12,
/// <summary>M4: the TOFU server-identity gate — see VcTofuStatus.</summary> /// <summary>The TOFU server-identity gate — see VcTofuStatus.</summary>
ServerIdentity = 13, ServerIdentity = 13,
/// <summary>M5: async result for moderation/admin/channel operations.</summary> /// <summary>Async result for moderation/admin/channel operations.</summary>
GenericResult = 14, GenericResult = 14,
/// <summary>M5: reply to VoiceCatClient.RequestAccountList — call ListAccounts() to read.</summary> /// <summary>Reply to VoiceCatClient.RequestAccountList — call ListAccounts() to read.</summary>
AccountList = 15, AccountList = 15,
/// <summary>Voice-plane subscription state. u32a = 1 (subscribed) or 0 (unsubscribed).</summary>
VoiceState = 16,
} }
/// <summary> /// <summary>

View File

@@ -38,7 +38,9 @@ internal static class Marshaling
Marshal.PtrToStringUTF8(raw.Name) ?? string.Empty, Marshal.PtrToStringUTF8(raw.Name) ?? string.Empty,
Marshal.PtrToStringUTF8(raw.Topic) ?? string.Empty, Marshal.PtrToStringUTF8(raw.Topic) ?? string.Empty,
raw.PasswordProtected != 0, raw.PasswordProtected != 0,
raw.MaxUsers)); raw.MaxUsers,
raw.SortOrder,
ToManaged(in raw.Audio)));
} }
NativeMethods.vc_free_channel_list(ref native); NativeMethods.vc_free_channel_list(ref native);
return result; return result;
@@ -59,7 +61,8 @@ internal static class Marshaling
raw.SelfMicMuted != 0, raw.SelfMicMuted != 0,
raw.SelfDeafened != 0, raw.SelfDeafened != 0,
raw.ServerMuted != 0, raw.ServerMuted != 0,
raw.ServerDeafened != 0)); raw.ServerDeafened != 0,
raw.VoiceSubscribed != 0));
} }
NativeMethods.vc_free_user_list(ref native); NativeMethods.vc_free_user_list(ref native);
return result; return result;

View File

@@ -9,7 +9,9 @@ public sealed record ChannelInfo(
string Name, string Name,
string Topic, string Topic,
bool PasswordProtected, bool PasswordProtected,
uint MaxUsers); uint MaxUsers,
uint SortOrder,
AudioConfigInfo Audio);
public sealed record ChannelEditInfo( public sealed record ChannelEditInfo(
uint Id, uint Id,
@@ -30,7 +32,8 @@ public sealed record UserInfo(
bool SelfMicMuted, bool SelfMicMuted,
bool SelfDeafened, bool SelfDeafened,
bool ServerMuted, bool ServerMuted,
bool ServerDeafened); bool ServerDeafened,
bool VoiceSubscribed);
public sealed record PermissionsInfo( public sealed record PermissionsInfo(
bool CanCreateTempChannel, bool CanCreateTempChannel,

View File

@@ -54,6 +54,12 @@ internal static partial class NativeMethods
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_leave_channel(nint c); internal static partial VcResult vc_leave_channel(nint c);
[LibraryImport(LibName)]
internal static partial VcResult vc_join_voice(nint c);
[LibraryImport(LibName)]
internal static partial VcResult vc_leave_voice(nint c);
// ── Local media streams ───────────────────────────────────────────────────────────────── // ── Local media streams ─────────────────────────────────────────────────────────────────
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_stream_start(nint c, in VcStreamDescNative desc, internal static partial VcResult vc_stream_start(nint c, in VcStreamDescNative desc,
@@ -65,6 +71,12 @@ internal static partial class NativeMethods
[LibraryImport(LibName, StringMarshalling = StringMarshalling.Utf8)] [LibraryImport(LibName, StringMarshalling = StringMarshalling.Utf8)]
internal static partial VcResult vc_set_input_device(nint c, uint streamId, string? deviceId); internal static partial VcResult vc_set_input_device(nint c, uint streamId, string? deviceId);
[LibraryImport(LibName)]
internal static partial VcResult vc_set_capture_channels(nint c, uint streamId, uint channels);
[LibraryImport(LibName)]
internal static partial VcResult vc_audio_restart(nint c);
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_set_input_mode(nint c, VcInputMode mode); internal static partial VcResult vc_set_input_mode(nint c, VcInputMode mode);
@@ -80,6 +92,12 @@ internal static partial class NativeMethods
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_set_output_volume(nint c, float gain); internal static partial VcResult vc_set_output_volume(nint c, float gain);
[LibraryImport(LibName)]
internal static partial VcResult vc_set_input_gain(nint c, float gain);
[LibraryImport(LibName)]
internal static partial VcResult vc_set_input_noise_reduction(nint c, int enable);
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_set_remote_stream(nint c, uint userId, uint streamId, internal static partial VcResult vc_set_remote_stream(nint c, uint userId, uint streamId,
float gain, int muted, int noiseReduction); float gain, int muted, int noiseReduction);
@@ -130,7 +148,7 @@ internal static partial class NativeMethods
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial void vc_free_device_list(ref VcDeviceListNative list); internal static partial void vc_free_device_list(ref VcDeviceListNative list);
// ── M4: channel / user / stream snapshot getters ──────────────────────────────────────── // ── Channel / user / stream snapshot getters ────────────────────────────────────────────
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_list_channels(nint c, out VcChannelListNative outList); internal static partial VcResult vc_list_channels(nint c, out VcChannelListNative outList);
@@ -150,7 +168,7 @@ internal static partial class NativeMethods
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial void vc_free_stream_summary_list(ref VcStreamSummaryListNative list); internal static partial void vc_free_stream_summary_list(ref VcStreamSummaryListNative list);
// ── M4: TOFU server-identity gate ─────────────────────────────────────────────────────── // ── TOFU server-identity gate ───────────────────────────────────────────────────────────
[LibraryImport(LibName)] [LibraryImport(LibName)]
internal static partial VcResult vc_confirm_server_identity(nint c, int accept); internal static partial VcResult vc_confirm_server_identity(nint c, int accept);
@@ -158,7 +176,7 @@ internal static partial class NativeMethods
internal static partial VcResult vc_get_server_identity_display(nint c, nint outBuf, internal static partial VcResult vc_get_server_identity_display(nint c, nint outBuf,
nuint bufCap, out nuint outLen); nuint bufCap, out nuint outLen);
// ── M5: Moderation & admin ───────────────────────────────────────────────────────────── // ── Moderation & admin ───────────────────────────────────────────────────────────────────
[LibraryImport(LibName, StringMarshalling = StringMarshalling.Utf8)] [LibraryImport(LibName, StringMarshalling = StringMarshalling.Utf8)]
internal static partial VcResult vc_kick_user(nint c, uint userId, string? reason); internal static partial VcResult vc_kick_user(nint c, uint userId, string? reason);

Some files were not shown because too many files have changed in this diff Show More