Core ABI extensions (voicecat.h): - vc_list_channels / vc_list_users / vc_list_user_streams — pull-based snapshot getters for the channel-tree and user-list UI; session_model_mu_ guards cross-thread reads - VC_EVENT_JOIN_RESULT / vc_join_channel — channel join with optional password - VC_EVENT_SERVER_IDENTITY + vc_confirm_server_identity — TOFU gate that blocks io_thread_ until the UI approves or rejects; pins TLS leaf-cert SHA-256 (not declared Ed25519) - vc_get_server_identity_display — Ed25519 fingerprint for human-readable display only - VC_INPUT_ALWAYS_ON = 2 in vc_input_mode — transmit unconditionally, no VAD gate - vc_set_vad_threshold — live RMS threshold update (0.0–1.0); EnergyVadProcessor stores it atomically so the audio RT path reads without a lock C++ implementation: - SessionModel::apply_snapshot / apply_channel_event fixed to populate parent_id, password_protected, and max_users (were permanently zeroed) - TlsContext::peer_cert_fingerprint — SHA-256 of peer leaf cert DER via mbedTLS - TofuStore split into peek (read-only) + pin (write) so first-connect only persists after user approval; tofu_store_path in vc_config for per-user pin file location - TcpAcceptor uses dual-stack IPv6+IPv4 fallback (fixes localhost → ::1 on Windows) - windows-client CMake preset: Release shared DLL, static MinGW runtime, no tools/tests - New C++ tests: test_channel_user_list_abi, test_tofu_flow (14/14 green) Windows client (clients/windows/ — .NET 10 WinForms): - VoiceCat.Interop: LibraryImport P/Invoke surface, UnmanagedCallersOnly callbacks, Channel<VoiceCatEvent> event delivery drained by 30ms WinForms Timer - VoiceCat.App: ConnectDialog (saved servers, DPAPI password storage), ServerIdentity- Dialog (TOFU first-connect / mismatch warning), MainForm (channel TreeView, user ListBox, RichTextBox chat, voice controls, device pickers, VAD/PTT/always-on mode, per-user gain/mute/NR tuning, VAD sensitivity TrackBar, level meter ProgressBar) - PttKeyCaptureDialog — focus-scoped PTT key capture (documented limitation) - PerUserTuningDialog — real-time gain/mute/NR applied to all of a user's streams - Accessibility: explicit AccessibleName/Description on every control, & mnemonics, Activity log ListBox as durable screen-reader record, AutomationNotification for curated live announcements Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
54 lines
2.7 KiB
C++
54 lines
2.7 KiB
C++
/*
|
|
* audio/apm_processor.h — Send-side audio processing module (AEC/NS/AGC/VAD).
|
|
*
|
|
* Design: docs/voice.md §11. Uses webrtc-audio-processing when VOICECAT_HAS_APM is defined;
|
|
* falls back to a no-op passthrough (VAD always open, PCM unmodified) otherwise.
|
|
*/
|
|
#ifndef VOICECAT_AUDIO_APM_PROCESSOR_H
|
|
#define VOICECAT_AUDIO_APM_PROCESSOR_H
|
|
|
|
#include <cstdint>
|
|
#include <memory>
|
|
|
|
namespace voicecat::audio {
|
|
|
|
class ApmProcessor {
|
|
public:
|
|
virtual ~ApmProcessor() = default;
|
|
|
|
// Feed the most recent playback reference (for AEC). Call before process_capture().
|
|
virtual void process_render(const int16_t* pcm, int samples, int sample_rate) = 0;
|
|
|
|
// Process one capture frame in-place (AEC, NS, AGC).
|
|
// Returns true if VAD detects speech (or always true in passthrough mode).
|
|
// Returns false → caller should skip encode/send (silence gate).
|
|
virtual bool process_capture(int16_t* pcm, int samples, int sample_rate) = 0;
|
|
|
|
// Update the VAD RMS threshold in-place (used by EnergyVadProcessor; no-op in passthrough).
|
|
// Safe to call from any thread — EnergyVadProcessor stores it atomically.
|
|
virtual void set_threshold(float) {}
|
|
|
|
// Factory: returns a real APM if VOICECAT_HAS_APM is defined, else a passthrough. Used for
|
|
// recv-side per-stream noise reduction (docs/voice.md §10) — gating doesn't apply there, so
|
|
// this stays a passthrough until a real APM/NS backend exists (still inert; see
|
|
// PROGRESS.md). Do not use this for the send-side VAD gate — see create_vad() below.
|
|
static std::unique_ptr<ApmProcessor> create();
|
|
|
|
// Factory for the send-side input gate (docs/voice.md §11): a lightweight, dependency-free
|
|
// energy/RMS VAD with configurable threshold + hang-time. webrtc-audio-processing (the
|
|
// originally-planned APM) has no working Windows/MSVC build upstream (GCC-only Meson build,
|
|
// unfinished MinGW support, hard abseil-cpp dependency — see PROGRESS.md), so this is the
|
|
// real v1 implementation behind the same ApmProcessor interface, not a passthrough. No AEC
|
|
// — process_render() is a no-op here; that's a real limitation versus the originally-planned
|
|
// APM, not just a deferred VAD.
|
|
// rms_threshold: normalized 0.0-1.0 RMS-of-int16-range; default ~0.025.
|
|
// hang_time_ms: how long the gate stays open after the last loud frame; default 300 ms
|
|
// (matches AudioEngine's kTalkHangoverMs so "talking" and "gate open" agree).
|
|
static std::unique_ptr<ApmProcessor> create_vad(float rms_threshold = 0.025f,
|
|
int64_t hang_time_ms = 300);
|
|
};
|
|
|
|
} // namespace voicecat::audio
|
|
|
|
#endif // VOICECAT_AUDIO_APM_PROCESSOR_H
|