Joining voice on the voice-chat preset was unreliable: audio arrived after
several seconds of the route flipping back and forth, sometimes not at all, and
VoiceOver went quiet while it happened. Device logs show why. A graph with
voice processing enabled reports a successful start and is then torn down
within a second, roughly three times in four; every configuration without voice
processing — both microphone presets, and voice chat with processing off — comes
up first time and runs indefinitely.
With voice processing the input and output are one IO unit, and it only stays up
while the input is part of the render chain. The input node carried a tap and no
connection, which leaves it out of that chain. Route the input through a silent
mixer so it is genuinely rendered.
The rest of this is the amplifier rather than the cause, and each part of it
turned one failed start into a storm:
The stall watchdog rebuilt on every missed tick, without bound. That converted a
graph that could not start into endless session reconfiguration, which is what
the user heard and what hid the reason from the log. It now backs off after each
failed attempt and stops after four, logging VC_WATCHDOG exhausted, so a
transient freeze still recovers and a graph that will not start fails visibly.
Nothing waited for a graph to start before judging it dead. Enabling voice
processing rebuilds both halves of the IO, which posts a configuration change
and reads as not running for several hundred milliseconds, so the
configuration-change handler and the watchdog both tore down graphs that were
about to run. A settling window holds them off for two seconds.
A route change forced a full rebuild, and every rebuild moves the route, so one
notification produced the next. Route changes now take the non-forcing path,
which rebuilds a stopped graph and leaves a healthy one alone; the hardware test
it uses reads the input node's format, not AVAudioSession, whose reported rate
and channel count do not settle until after the graph has started.
The input side is built once per session instead of being added when voice is
joined, so joining and leaving voice set a stream id rather than replacing the
graph, and a mono voice-chat apply no longer clears a stereo capsule
configuration it never applied.
Every rebuild now logs its cause, and VC_START/VC_START_CHECK record whether the
graph survived its start. The first-attempt failure is not fixed and is recorded
in PROGRESS.md as a release gate: capture still comes up on a watchdog rebuild
rather than immediately.
The changed logic sits on AVAudioSession and AVAudioEngine, which the net10.0
test project cannot reference, so the behaviour is covered by the existing
source assertions; verification is on device.
Toggling speaker output flipped the route back and forth indefinitely. Two
loops, both of which made a rebuild produce the condition for the next one.
Reconfiguring the session moves the route, and moving the route is reported
back through RouteChangeNotification. Forcing the speaker takes a headset out
of the route, which arrives as OldDeviceUnavailable, and releasing it brings
the headset back as NewDeviceAvailable; neither is among the reasons the
handler filters, so each rebuild answered its own echo with another rebuild.
Nothing compared the reported route against the route the live graph was
actually built on.
Record that route at the end of Apply, once the session is configured, and
rebuild only when a reported change differs from it; notifications that arrive
while Apply is still running describe the change Apply is itself making and are
ignored outright. The decision is AudioRouteWatcher in VoiceCat.Core, which is
platform-agnostic and tested, following ControlPathWatcher; the route identity
it compares is supplied by the caller, on iOS the UIDs of the current route's
ports. A graph whose route is unchanged but broken is still the stall
watchdog's to catch.
The port override was also re-asserted on every Apply, so where the system
wanted to hand output back to a connected headset each rebuild forced it to the
speaker again and the resulting route change drove the next rebuild. It is now
the one-shot request it should always have been, issued by the toggle alone;
the DefaultToSpeaker category option is the part that persists across rebuilds.
Also updates the route test from 724f7e9, which asserted the voice-chat preset
clearing the speaker flag and the absence of the port override. Both were
deliberately removed when speaker output became orthogonal to the preset, and
the test should have been updated with them.
Speaker output decides where audio goes, not how it is captured, which is why
it sits beside the preset row rather than inside Advanced audio. It was
nonetheless moving the preset to Advanced, and the voice-chat preset was
clearing it back, so the two controls overwrote each other and a user who
wanted the speaker lost the preset that describes their capture setup.
SetForceSpeaker no longer touches the preset, and Load and SelectPreset no
longer clear the flag. A speakerIsExplicit default records that the user
actually chose, so the automatic value older installs inherited from the
voice-chat preset is dropped once rather than pinning a headset user to the
speaker, and the toggle is honoured from then on.
Apply issues the speaker port override again when the flag is set. Without it
the toggle barely did anything with a headset attached, because
DefaultToSpeaker only decides the route when nothing else is connected. With
the toggle off it passes None, so a route left alone still follows an HFP
headset.
The changed logic sits on AVAudioSession, which the net10.0 test project
cannot reference, so this carries no tests; the toggle needs device
verification under the voice-chat preset with a Bluetooth headset connected.
A Wi-Fi to cellular switch left the session visibly dropping: the media transport rebound
itself within a few seconds, but nothing noticed the blackholed control connection until an
unanswered keepalive proved it, and the teardown that followed announced a lost connection and
waited another second before dialling again.
Watch the system path on iOS and fail the control connection the moment the carrying interface
changes, which is the only path change TCP cannot survive. Roaming between access points and a
link that is merely unusable for a while keep the same interface and the same source address,
so ControlPathWatcher reports neither; an unsatisfied path holds the last signature rather than
reporting, so a reconnect is never started into a route that cannot carry it. Tighten the
keepalive window on the phone as the backstop for what the monitor cannot see, run the first
reconnect attempt immediately, and defer the lost-connection announcement until an attempt has
actually failed, so a sub-second handover is silent and only a real outage is announced.
A control reconnect still re-authenticates and rejoins: the media keys come from the TLS
exporter of the connection that was lost, so seamless handover needs control-plane session
resumption rather than a faster reconnect.
Two iOS bugs with the same shape: unconditional rebuilds where a
conditional check belongs.
The audio graph was torn down on every unintentional disconnect and every
foreground transition. A lost connection ran the same teardown as an
explicit disconnect, deactivating the AVAudioSession and so dropping the
Bluetooth HFP link for a transport blip, and foregrounding always called
Reconfigure even though the `audio` background mode keeps the graph live.
Both cost seconds of dead audio on a headset.
Split "session ended" from "transport blipped". Detach unbinds the client
but keeps the session, graph, and route, so a reconnect rebinds to a live
HFP link; the route is parked on stream id 0 so capture cannot feed the
next connection a stream it never announced. StartListening reuses a
running graph, StartMicrophone reuses a running tap of the same width, and
Reconfigure gained a non-forcing mode that no-ops when tap presence,
channel width, and voice processing all still match. Foregrounding now
ensures the graph is running and only reconfigures if it actually stopped.
Route changes, media-services resets, and the stall watchdog still force a
full rebuild.
Every list also reloaded on a model event raised 20 times a second by the
microphone level timer. ReloadData recreates the accessibility element
tree, so VoiceOver explore mode re-announced the row under a dragging
finger and a double tap landed on an element that no longer existed. No
controller ever unsubscribed, so popped controllers kept reloading too.
Move the level to its own LevelChanged event, and reload lists through
ListRefresher, which subscribes only while on screen and only reloads when
the rendered content signature changed. The voice bar publishes its
accessibility value on 5% steps, MoveUserController reloads just its two
checkmark rows, and the chat transcripts skip reassigning identical text.
The changed logic sits on UIKit and AVFoundation types the net10.0 test
project cannot reference, so this carries no tests; the Bluetooth
reconnect and foreground paths need device verification.
The screen-audio pump opened and disposed a memory mapping plus its container
lookup and path strings on every 5 ms tick - about 200 mapping pairs per
second of steady allocation that churned the GC under long calls. The producer
opens the ring with O_CREAT and never replaces it, so the pump now keeps one
mapping across ticks and rebuilds it only when the ring file disappears or a
drain fails.
Sleep-paced threads stall for hundreds of milliseconds when iOS coalesces a
backgrounded app's wakeups, and the deadline resets that discarded the deficit
left the capture backlog queued until its ring overflowed: regular dropouts
that worsen the longer the app stays backgrounded. Pace both 20 ms hands-offs
from the AVAudioSourceNode render callback instead (mix via Audio.RunCycle
with deviceClockedAudio, capture handoff from the ring), wake the UDP sender
on a queue signal instead of a 1 ms poll, and let stalled consumers drop their
backlog to the buffer target instead of ratcheting it.
Declare XSAppIconAssets in the app manifest so actool receives --app-icon
and the bundle carries CFBundleIconName, and always pass both Xcode version
placeholders to the ReplayKit extension so an empty build number cannot
drop CFBundleVersion and fail installd.
Adds vcpkg as a submodule at vcpkg/, pinned to the exact commit vcpkg.json
already declares as builtin-baseline, so the bundled checkout and the
manifest's resolved port versions can never drift apart.
cmake/voicecat-toolchain.cmake, scripts/common.sh, and
clients/apple/scripts/build-xcframework.sh now resolve vcpkg as:
VCPKG_ROOT env var (external checkout) > bundled submodule. Docs updated
to describe the new one-time setup.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Removes leftover debug scaffolding (stray Console.WriteLine/NSLog traces,
dead nick_buf_ptr, a no-op --print-config flag now implemented for real),
fixes stale/misleading comments (channel passwords are no longer a "future
M5+" feature, a wrong cross-reference, a stale TlsContext::close() mention,
an incomplete BanRecord::subject_type doc, and a smoke test pointing at a
build/m1-dev preset that no longer exists), strips internal M1-M5 milestone
jargon from comments now that the roadmap is done, trims comments that just
restated the following line, and consolidates a few "why" explanations that
were duplicated 2-3 times in the same file.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds Assets.xcassets with a placeholder 1024x1024 AppIcon so Xcode
compiles CFBundleIconName and the required icon sizes into the bundle.
Replaces UILaunchStoryboardName (missing storyboard) with UILaunchScreen
dict, valid for iOS 14+ (min target is iOS 18). Fixes all four App Store
Connect validation errors blocking TestFlight upload.
Changes main app bundle ID from cat.voice.VoiceCatiOS, broadcast extension
from cat.voice.VoiceCatiOS.broadcast, and App Group from group.cat.voice.VoiceCat
to match the registered App Store identifier.
Recovery from the 99937c9 audio-device-change commit broadened the
route-change recovery set to "everything except categoryChange /
routeConfigurationChange", which added .override. But .override is
fired by our own applyA2dpSpeakerFallback() -> overrideOutputAudioPort,
which recoverAudio() calls on every recovery. On an A2DP preset
(Stereo Mic / Mono Mic), disconnecting AirPods ping-ponged:
oldDeviceUnavailable -> recoverAudio -> applyA2dpSpeakerFallback
-> overrideOutputAudioPort(.speaker) -> .override routeChange
-> recoverAudio -> applyConfiguration (setCategory resets override)
-> applyA2dpSpeakerFallback -> overrideOutputAudioPort -> .override -> ...
Each iteration also rebuilt the AVAudioEngine via reconfigure() ->
rebuild() -- the audible reinitialize loop + CPU spin. Voice Chat and
Built-in Mic + Speaker were unaffected (applyA2dpSpeakerFallback
early-returns for non-A2DP modes, so no overrideOutputAudioPort call).
Two-part fix (pure Swift iOS-app target; no C ABI / proto / docs changes):
1. AudioSessionManager.handleRouteChange: added .override to the skip
list alongside .categoryChange / .routeConfigurationChange. .override
is only ever fired by our own overrideOutputAudioPort call, so
treating it as a recovery reason is the loop by definition. The
AVAudioEngineConfigurationChange observer in IOSVoiceProcessingEngine
remains as the backstop if an override ever actually stops the engine.
2. IOSAudioRouter.applyA2dpSpeakerFallback: made idempotent via a
lastAppliedOutputOverride tracker that skips the redundant
overrideOutputAudioPort call when the desired state (.none for
external output present, .speaker otherwise) already matches. Reset
to nil at the top of applyConfiguration() (setCategory can reset the
override) and on a failed call. Defense-in-depth on top of fix 1.
Build: xcodebuild -project clients/apple/iOS/VoiceCatiOS.xcodeproj
-scheme VoiceCatiOS -destination 'generic/platform=iOS' build green
(Xcode 26.5 / iOS 18.0).
Network drops (e.g. Wi-Fi -> cellular) and audio-device plug/unplug (wired
headphones, AirPods) used to leave the iOS client in a dead/zombie state:
the engine went silent, no reconnect was attempted, and a live-session
disconnect waited 30-60 s for the C core's TCP keepalive/reaper timeout.
Reconnect (AppState.swift, SessionState.swift):
- Two-layer reconcile. Once SessionState overwrites client.onEvent at auth
success, AppState.handleConnectEvent no longer sees live-session events.
Added a weak SessionState.appState; SessionState.handleEvent .disconnected
calls appState.onLiveSessionDisconnected after the cue -- the single path
AppState learns a live session dropped. Shared teardownLiveSessionAndReconnect
snapshots lastSession, stops audio, releases session/VoiceCatClient (io-
thread join via vc_client_destroy), resets the backoff, and arms
scheduleReconnect (exponential 1s -> 30s cap, indefinite, restored on auth
success via existing TOFU_MATCHED auto-confirm + idempotent join_channel).
- NWPathMonitor now runs WHILE CONNECTED (not only mid-reconnect). On a Wi-Fi
<-> cellular interface change or path .unsatisfied it calls
proactiveReconnect: tearing the session down BEFORE the C core notices the
dead socket collapses the 30-60 s reaper wait into ~1 s + first backoff
tick. Same-interface refreshes (BSSID roams) are ignored via pathSignature.
While mid-reconnect a .satisfied path resets the backoff for a fast retry.
- User-initiated disconnect()/cancelConnect() set userInitiatedDisconnect
and cancel all reconnect state (task + monitor + lastSession + connectedServer).
Audio recovery (AudioSessionManager.swift, IOSVoiceProcessingEngine.swift):
- Intent-gated recoverAudio() replaces the narrow .oldDeviceUnavailable/
.newDeviceAvailable route-change guard; fires on every externally-initiated
route change reason except the ones we cause ourselves (.categoryChange/
.routeConfigurationChange) to avoid a notification loop. Interruption-end
now always recovers instead of only when .shouldResume is set.
- Added AVAudioEngineConfigurationChange observer on the engine so a system
self-stop after our route-change handler wins the race is caught.
- IOSAudioEngine.rebuild() does a one-shot reactivation-retry on
engine.start() failure (iOS sometimes refuses until the session is
re-reactivated -- the silent-death case).
No C ABI / voicecat.h / proto / core changes. Swift-only. iOS sim build green
via scripts/build-ios-client.sh --no-configure (Xcode 26.5 / iOS 18.0 sim).
Three bugs fixed across the full stack (proto/server/core/ABI/Win/macOS/iOS):
1. Join/Leave Voice now truly subscribes/unsubscribes from the voice plane.
Previously the button only toggled the local mic — receiving was always on
(gated by channel membership alone). Added a protocol-level voice subscription
concept: new SubscribeVoiceRequest/UnsubscribeVoiceRequest/VoiceSubscriptionResult
proto messages, User.voice_subscribed field, vc_join_voice/vc_leave_voice C ABI
functions, VC_EVENT_VOICE_STATE event, server-side voice_subscribed flag checked
by the SFU relay recipient filter, and core-client gating of remote-stream
decoder setup. All three clients rewired to subscribe+mic on Join / unsubscribe
on Leave. Text chat works regardless of voice subscription.
2. Channel edit dialog now shows the channel's actual current settings. The read
struct vc_channel was missing sort_order and audio fields — only the write
struct vc_channel_info had them. Extended vc_channel with both (additive, no
ABI break), updated the session model and list_channels marshaling to populate
them, and updated all three clients' edit callers to use actual channel info
instead of hardcoded defaults.
3. Channel parameter updates now automatically restart everyone's streams.
Previously editing a channel's audio config persisted and broadcast a
ChannelEvent::UPDATED, but no layer restarted streams — encoders/decoders are
frozen at announce time. handle_channel_event now detects audio-config changes
on the user's current channel and stop->starts each active local stream. The
server reads the updated config on re-announce; peers wire up fresh decoders
at the new ssrc.
All 29 CTest tests pass; Windows DLL + C# client build clean. Apple clients not
yet compile-verified (Windows environment).
PTT was focus-scoped (WinForms KeyDown/KeyUp) so it died the moment the
window lost focus. Add an optional system-wide path using the Raw Input
API (RegisterRawInputDevices + WM_INPUT with RIDEV_INPUTSINK) instead of
a WH_KEYBOARD_LL low-level hook -- the latter is the keylogger pattern AV
heuristics flag, which is worse for our unsigned MinGW binary. Raw Input
involves no DLL injection or global hook and passes keys through.
- New VoiceCat.App/Native/RawInput.cs: P/Invoke + structs; register the
keyboard sink, decode WM_INPUT to vkey/up-down, GetAsyncKeyState helper.
- MainForm overrides OnHandleCreated/OnHandleDestroyed/WndProc to manage
the sink and route WM_INPUT to PTT; gates the focus-scoped KeyDown/KeyUp
off when system-wide is on; makes the Deactivate force-release
conditional; adds a GetAsyncKeyState watchdog on the pump timer so a
missed key-up (RDP/lock-screen) cannot leave PTT stuck.
- VoiceSettings.SystemWidePtt (default ON) + system-wide checkbox in the
Audio settings PTT section.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Suppress the system ding on global hotkeys/PTT and the user-list Enter
key by setting SuppressKeyPress (Handled alone leaves WM_CHAR to beep).
- Preserve the user-list keyboard selection across talking/mute refreshes
instead of resetting it on every Items.Clear().
- Include the connected server name in the main window titlebar.
- Close the private-message window on Escape.
- Show the PM window without an owner so focus is no longer trapped to it
and the main window can be worked in while a PM is open.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Mic input gain (was 3x) and per-user receive gain (was 2x) had
asymmetric boost ceilings. Raise both, plus the desktop aux input
gain (was 3x), to a uniform 4x (400%) on macOS, iOS, and Windows.
The master Output volume slider is unchanged (still 1x). No core
changes needed: the C ABI only clamps negatives, so the ceilings
live entirely in the client UI sliders.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>