fix(ios): keep the voice-processing graph up instead of rebuilding it
Build and test / test (macos-latest) (push) Canceled after 0s
Build and test / test (ubuntu-24.04) (push) Canceled after 0s
Build and test / test (windows-latest) (push) Canceled after 0s
Build and test / apple-client (push) Canceled after 0s

Joining voice on the voice-chat preset was unreliable: audio arrived after
several seconds of the route flipping back and forth, sometimes not at all, and
VoiceOver went quiet while it happened. Device logs show why. A graph with
voice processing enabled reports a successful start and is then torn down
within a second, roughly three times in four; every configuration without voice
processing — both microphone presets, and voice chat with processing off — comes
up first time and runs indefinitely.

With voice processing the input and output are one IO unit, and it only stays up
while the input is part of the render chain. The input node carried a tap and no
connection, which leaves it out of that chain. Route the input through a silent
mixer so it is genuinely rendered.

The rest of this is the amplifier rather than the cause, and each part of it
turned one failed start into a storm:

The stall watchdog rebuilt on every missed tick, without bound. That converted a
graph that could not start into endless session reconfiguration, which is what
the user heard and what hid the reason from the log. It now backs off after each
failed attempt and stops after four, logging VC_WATCHDOG exhausted, so a
transient freeze still recovers and a graph that will not start fails visibly.

Nothing waited for a graph to start before judging it dead. Enabling voice
processing rebuilds both halves of the IO, which posts a configuration change
and reads as not running for several hundred milliseconds, so the
configuration-change handler and the watchdog both tore down graphs that were
about to run. A settling window holds them off for two seconds.

A route change forced a full rebuild, and every rebuild moves the route, so one
notification produced the next. Route changes now take the non-forcing path,
which rebuilds a stopped graph and leaves a healthy one alone; the hardware test
it uses reads the input node's format, not AVAudioSession, whose reported rate
and channel count do not settle until after the graph has started.

The input side is built once per session instead of being added when voice is
joined, so joining and leaving voice set a stream id rather than replacing the
graph, and a mono voice-chat apply no longer clears a stereo capsule
configuration it never applied.

Every rebuild now logs its cause, and VC_START/VC_START_CHECK record whether the
graph survived its start. The first-attempt failure is not fixed and is recorded
in PROGRESS.md as a release gate: capture still comes up on a watchdog rebuild
rather than immediately.

The changed logic sits on AVAudioSession and AVAudioEngine, which the net10.0
test project cannot reference, so the behaviour is covered by the existing
source assertions; verification is on device.
This commit is contained in:
2026-09-25 20:56:31 +02:00
parent 1a0ff957ae
commit 01bae734b8
6 changed files with 229 additions and 135 deletions
+15 -3
View File
@@ -37,9 +37,19 @@ the element a double tap was aimed at.
The iOS audio graph is rebuilt only when the audio configuration changed. A lost connection
unbinds the client but keeps the session, graph, and route alive, so a reconnect rebinds to a
live Bluetooth HFP link instead of renegotiating one; restoring voice reuses a running capture
tap of the same width; and foregrounding ensures the graph is running rather than rebuilding it.
Route changes, media-services resets, and the stall watchdog still force a full rebuild.
live Bluetooth HFP link instead of renegotiating one, and foregrounding ensures the graph is
running rather than rebuilding it. The input side — input node, voice processing, and capture tap
— is built once per session rather than added when voice is joined, so joining and leaving voice
name a stream instead of replacing the graph. The input is rendered through a silent mixer,
because the voice-processing IO unit only stays up while the input is in the render chain. Route
changes no longer force a rebuild; media-services resets still do, and the stall watchdog now
backs off and stops after four failed attempts instead of rebuilding without end.
Voice-chat capture is not yet reliable on the first attempt: a graph with voice processing still
sometimes fails to start and is only recovered by a watchdog rebuild, which costs seconds before
audio appears. Speaker output is an explicit choice that no longer changes the preset. Every
rebuild logs `VC_REBUILD cause=`, and `VC_START`/`VC_START_CHECK` record whether the graph
survived its start; that instrumentation is what the remaining investigation needs.
The media path now survives changing networks. A client whose source address changes proves
possession of its media key from the new address with an authenticated `Rebind` frame and the
@@ -63,6 +73,8 @@ cap the applied Opus hint at 30%, and feed it back over TLS.
## Release gates
- Find why an iOS graph with voice processing intermittently starts and then stops within a
second, so voice-chat capture comes up on the first attempt rather than after a watchdog rebuild.
- Run real multi-person calls on Windows, macOS, and physical iOS hardware, including adaptive
20/40/60 ms buffering, duration-aware DRED/FEC, automatic packet-loss feedback, and mismatched
input/output endpoints.