Files
voice-cat/core/src/audio/apm_processor.h
Talon 63b241cc2e feat(M4): Windows WinForms client, TOFU identity pinning, VAD threshold + always-on mode
Core ABI extensions (voicecat.h):
- vc_list_channels / vc_list_users / vc_list_user_streams — pull-based snapshot getters
  for the channel-tree and user-list UI; session_model_mu_ guards cross-thread reads
- VC_EVENT_JOIN_RESULT / vc_join_channel — channel join with optional password
- VC_EVENT_SERVER_IDENTITY + vc_confirm_server_identity — TOFU gate that blocks io_thread_
  until the UI approves or rejects; pins TLS leaf-cert SHA-256 (not declared Ed25519)
- vc_get_server_identity_display — Ed25519 fingerprint for human-readable display only
- VC_INPUT_ALWAYS_ON = 2 in vc_input_mode — transmit unconditionally, no VAD gate
- vc_set_vad_threshold — live RMS threshold update (0.0–1.0); EnergyVadProcessor stores
  it atomically so the audio RT path reads without a lock

C++ implementation:
- SessionModel::apply_snapshot / apply_channel_event fixed to populate parent_id,
  password_protected, and max_users (were permanently zeroed)
- TlsContext::peer_cert_fingerprint — SHA-256 of peer leaf cert DER via mbedTLS
- TofuStore split into peek (read-only) + pin (write) so first-connect only persists
  after user approval; tofu_store_path in vc_config for per-user pin file location
- TcpAcceptor uses dual-stack IPv6+IPv4 fallback (fixes localhost → ::1 on Windows)
- windows-client CMake preset: Release shared DLL, static MinGW runtime, no tools/tests
- New C++ tests: test_channel_user_list_abi, test_tofu_flow (14/14 green)

Windows client (clients/windows/ — .NET 10 WinForms):
- VoiceCat.Interop: LibraryImport P/Invoke surface, UnmanagedCallersOnly callbacks,
  Channel<VoiceCatEvent> event delivery drained by 30ms WinForms Timer
- VoiceCat.App: ConnectDialog (saved servers, DPAPI password storage), ServerIdentity-
  Dialog (TOFU first-connect / mismatch warning), MainForm (channel TreeView, user
  ListBox, RichTextBox chat, voice controls, device pickers, VAD/PTT/always-on mode,
  per-user gain/mute/NR tuning, VAD sensitivity TrackBar, level meter ProgressBar)
- PttKeyCaptureDialog — focus-scoped PTT key capture (documented limitation)
- PerUserTuningDialog — real-time gain/mute/NR applied to all of a user's streams
- Accessibility: explicit AccessibleName/Description on every control, & mnemonics,
  Activity log ListBox as durable screen-reader record, AutomationNotification for
  curated live announcements

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 00:35:16 +02:00

54 lines
2.7 KiB
C++

/*
* audio/apm_processor.h — Send-side audio processing module (AEC/NS/AGC/VAD).
*
* Design: docs/voice.md §11. Uses webrtc-audio-processing when VOICECAT_HAS_APM is defined;
* falls back to a no-op passthrough (VAD always open, PCM unmodified) otherwise.
*/
#ifndef VOICECAT_AUDIO_APM_PROCESSOR_H
#define VOICECAT_AUDIO_APM_PROCESSOR_H
#include <cstdint>
#include <memory>
namespace voicecat::audio {
class ApmProcessor {
public:
virtual ~ApmProcessor() = default;
// Feed the most recent playback reference (for AEC). Call before process_capture().
virtual void process_render(const int16_t* pcm, int samples, int sample_rate) = 0;
// Process one capture frame in-place (AEC, NS, AGC).
// Returns true if VAD detects speech (or always true in passthrough mode).
// Returns false → caller should skip encode/send (silence gate).
virtual bool process_capture(int16_t* pcm, int samples, int sample_rate) = 0;
// Update the VAD RMS threshold in-place (used by EnergyVadProcessor; no-op in passthrough).
// Safe to call from any thread — EnergyVadProcessor stores it atomically.
virtual void set_threshold(float) {}
// Factory: returns a real APM if VOICECAT_HAS_APM is defined, else a passthrough. Used for
// recv-side per-stream noise reduction (docs/voice.md §10) — gating doesn't apply there, so
// this stays a passthrough until a real APM/NS backend exists (still inert; see
// PROGRESS.md). Do not use this for the send-side VAD gate — see create_vad() below.
static std::unique_ptr<ApmProcessor> create();
// Factory for the send-side input gate (docs/voice.md §11): a lightweight, dependency-free
// energy/RMS VAD with configurable threshold + hang-time. webrtc-audio-processing (the
// originally-planned APM) has no working Windows/MSVC build upstream (GCC-only Meson build,
// unfinished MinGW support, hard abseil-cpp dependency — see PROGRESS.md), so this is the
// real v1 implementation behind the same ApmProcessor interface, not a passthrough. No AEC
// — process_render() is a no-op here; that's a real limitation versus the originally-planned
// APM, not just a deferred VAD.
// rms_threshold: normalized 0.0-1.0 RMS-of-int16-range; default ~0.025.
// hang_time_ms: how long the gate stays open after the last loud frame; default 300 ms
// (matches AudioEngine's kTalkHangoverMs so "talking" and "gate open" agree).
static std::unique_ptr<ApmProcessor> create_vad(float rms_threshold = 0.025f,
int64_t hang_time_ms = 300);
};
} // namespace voicecat::audio
#endif // VOICECAT_AUDIO_APM_PROCESSOR_H