Files
voice-cat/PROGRESS.md
T
Talon 01bae734b8
Build and test / test (macos-latest) (push) Canceled after 0s
Build and test / test (ubuntu-24.04) (push) Canceled after 0s
Build and test / test (windows-latest) (push) Canceled after 0s
Build and test / apple-client (push) Canceled after 0s
fix(ios): keep the voice-processing graph up instead of rebuilding it
Joining voice on the voice-chat preset was unreliable: audio arrived after
several seconds of the route flipping back and forth, sometimes not at all, and
VoiceOver went quiet while it happened. Device logs show why. A graph with
voice processing enabled reports a successful start and is then torn down
within a second, roughly three times in four; every configuration without voice
processing — both microphone presets, and voice chat with processing off — comes
up first time and runs indefinitely.

With voice processing the input and output are one IO unit, and it only stays up
while the input is part of the render chain. The input node carried a tap and no
connection, which leaves it out of that chain. Route the input through a silent
mixer so it is genuinely rendered.

The rest of this is the amplifier rather than the cause, and each part of it
turned one failed start into a storm:

The stall watchdog rebuilt on every missed tick, without bound. That converted a
graph that could not start into endless session reconfiguration, which is what
the user heard and what hid the reason from the log. It now backs off after each
failed attempt and stops after four, logging VC_WATCHDOG exhausted, so a
transient freeze still recovers and a graph that will not start fails visibly.

Nothing waited for a graph to start before judging it dead. Enabling voice
processing rebuilds both halves of the IO, which posts a configuration change
and reads as not running for several hundred milliseconds, so the
configuration-change handler and the watchdog both tore down graphs that were
about to run. A settling window holds them off for two seconds.

A route change forced a full rebuild, and every rebuild moves the route, so one
notification produced the next. Route changes now take the non-forcing path,
which rebuilds a stopped graph and leaves a healthy one alone; the hardware test
it uses reads the input node's format, not AVAudioSession, whose reported rate
and channel count do not settle until after the graph has started.

The input side is built once per session instead of being added when voice is
joined, so joining and leaving voice set a stream id rather than replacing the
graph, and a mono voice-chat apply no longer clears a stereo capsule
configuration it never applied.

Every rebuild now logs its cause, and VC_START/VC_START_CHECK record whether the
graph survived its start. The first-attempt failure is not fixed and is recorded
in PROGRESS.md as a release gate: capture still comes up on a watchdog rebuild
rather than immediately.

The changed logic sits on AVAudioSession and AVAudioEngine, which the net10.0
test project cannot reference, so the behaviour is covered by the existing
source assertions; verification is on device.
2026-09-25 20:56:31 +02:00

101 lines
6.9 KiB
Markdown

# VoiceCat status
Updated: 2026-09-25
## Current state
VoiceCat's supported implementation is .NET 10. The managed protocol, crypto, TLS, server,
CLI, client state, audio engine, Windows client, macOS client, and iOS client are implemented.
Retired implementations and compatibility projects have been removed; this tree contains only
the supported product and its required native media boundaries.
The source-of-truth layout is:
- `proto/voicecat.proto` — wire schema.
- `src/` — protocol, crypto, codec/DSP bindings, server, client core, audio, and CLI.
- `tests/VoiceCat.Tests/` — managed behavior and integration tests.
- `clients/windows/` — WinForms application over the managed core.
- `clients/apple/` — AppKit and UIKit applications over the managed core.
- `native/media/` and `native/rnnoise/` — required Opus/RNNoise shim and vendored RNNoise.
- `native/apple/broadcast/` — required ReplayKit upload extension and shared ring producer.
The physical-device iOS voice path now supports ReplayKit fallback, stable Apple VPIO voice-chat
capture using a paced 20 ms handoff, and true built-in stereo microphone capture. Stereo was
verified on an iPhone 16 Pro Max with a two-channel AVAudioEngine input and distinct left/right
samples; the managed Apple binding requires native use of its otherwise-unmapped stereo polar
pattern constant. The voice-chat preset now leaves speaker routing off so a connected Bluetooth
headset can supply both input and output; speaker routing remains an explicit Advanced setting.
Device-selected inputs are not saved as preferences during route refresh. Bluetooth switching
still needs device validation.
The iOS user list now opens a remote-user detail view with independent tuning for each active
audio stream. Private messages are grouped into per-user conversations with direct access to the
same user and audio controls. Lists reload only when their rendered content actually changed and
only while on screen, and the 20 Hz microphone level is a separate signal from the general model
change, so VoiceOver explore mode no longer re-announces the row under a dragging finger or loses
the element a double tap was aimed at.
The iOS audio graph is rebuilt only when the audio configuration changed. A lost connection
unbinds the client but keeps the session, graph, and route alive, so a reconnect rebinds to a
live Bluetooth HFP link instead of renegotiating one, and foregrounding ensures the graph is
running rather than rebuilding it. The input side — input node, voice processing, and capture tap
— is built once per session rather than added when voice is joined, so joining and leaving voice
name a stream instead of replacing the graph. The input is rendered through a silent mixer,
because the voice-processing IO unit only stays up while the input is in the render chain. Route
changes no longer force a rebuild; media-services resets still do, and the stall watchdog now
backs off and stops after four failed attempts instead of rebuilding without end.
Voice-chat capture is not yet reliable on the first attempt: a graph with voice processing still
sometimes fails to start and is only recovered by a watchdog rebuild, which costs seconds before
audio appears. Speaker output is an explicit choice that no longer changes the preset. Every
rebuild logs `VC_REBUILD cause=`, and `VC_START`/`VC_START_CHECK` record whether the graph
survived its start; that instrumentation is what the remaining investigation needs.
The media path now survives changing networks. A client whose source address changes proves
possession of its media key from the new address with an authenticated `Rebind` frame and the
relay moves its endpoint, instead of the session dying silently in both directions; the client
rebuilds its UDP socket rather than retrying on one pinned to a vanished interface. The control
connection is judged live by server traffic rather than assumed live, so a blackholed TCP path
is detected in 30 s instead of waiting minutes for the OS, and 12 s on iOS. iOS also watches the
system path and fails the control connection the moment the carrying interface changes, so a
handover reconnects in about a second instead of waiting out the silence timeout; access-point
roaming and an unusable-but-unchanged link are deliberately not handovers and are ridden out.
The first reconnect attempt is immediate, and a loss is only announced once an attempt has
actually failed, so a handover reads as a hiccup rather than a dropped call. A control reconnect
still re-authenticates and rejoins: seamless handover needs control-plane session resumption,
because the media keys come from the TLS exporter of the connection that was lost. The receive jitter buffer keeps a
one-frame depth floor, measures late and reordered arrivals, and can deepen mid-call, and a
stalled consumer now costs bounded audio rather than the live talkspurt.
SQLite schema v4 persists DRED and the channel packet-loss mode. Manual loss remains the default;
automatic Fast/Balanced/Stable modes measure each sender's authenticated UDP uplink at the server,
cap the applied Opus hint at 30%, and feed it back over TLS.
## Release gates
- Find why an iOS graph with voice processing intermittently starts and then stops within a
second, so voice-chat capture comes up on the first attempt rather than after a watchdog rebuild.
- Run real multi-person calls on Windows, macOS, and physical iOS hardware, including adaptive
20/40/60 ms buffering, duration-aware DRED/FEC, automatic packet-loss feedback, and mismatched
input/output endpoints.
- Complete NVDA and VoiceOver navigation/announcement passes.
- Verify iOS remote-user tuning and private-conversation navigation with VoiceOver, including
multiple streams, users without active streams, and users who disconnect while a view is open.
- Exercise iOS background/lock, interruption, Bluetooth, route-change, ReplayKit, and iOS 27
ScreenCaptureKit paths on devices. The background/lock gate keeps a call active for 15+ minutes
backgrounded and screen-locked with no periodic glitches and flat `VC_AUDIO` `feedDrops`/`starved`
counters (the render callback now paces the mix and the 20 ms capture handoff, and a watchdog
rebuilds a graph that stops calling back). Take a Siri or phone-call interruption while
backgrounded and confirm audio resumes without foregrounding. Complete a
30-minute iOS call and Wi-Fi/cellular switching with voice restoration (the switch is covered
by simulation in `NetworkImpairmentTests`; hardware confirms the real route change), plus extended
mono/stereo/voice-chat switching while joined. Verify Windows desktop/per-app stereo sharing.
- Complete Developer ID signing/notarization. The iOS host and ReplayKit extension have been
distribution-signed and packaged locally; upload the IPA for Apple's server-side validation.
- Run the published Linux container and a 30-minute-or-longer server soak.
## Working rule
Keep this file short. It records only current state and open release gates. Git history is the
implementation diary.