Files
voice-cat/CLAUDE.md
Talon 5f6c223526 feat: device enumeration, VAD/PTT input gate, stereo playback, WASAPI loopback
Closes the three items PROGRESS.md's M3 section explicitly carried forward as
out of scope:

- Device enumeration (vc_list_devices) + input device selection
  (vc_set_input_device), backed by AudioEngine::enumerate_devices() via
  miniaudio's ma_context_get_devices. Device ids are opaque hex-encoded
  ma_device_id strings.
- VAD/PTT send-side input gate (vc_set_input_mode, vc_set_push_to_talk).
  webrtc-audio-processing (the originally-planned APM) has no working
  Windows/MSVC build upstream (GCC-only Meson, unfinished MinGW support, hard
  abseil-cpp dependency), so VAD is a new lightweight, dependency-free
  energy/RMS processor (EnergyVadProcessor) behind the existing ApmProcessor
  interface. Gating is MIC-only; SCREEN_AUDIO/AUX_DEVICE always bypass it.
- True stereo playback: AudioEngine's mixer and output device now carry
  stereo end-to-end (mono streams upmix L=R) instead of downmixing decoded
  stereo streams to mono before mixing.
- Real WASAPI loopback capture for SCREEN_AUDIO (Windows-only, via
  miniaudio's loopback device type), replacing test-only injection as the
  production capture path.

Also: vccli gains --list-devices, --input-device, --input-mode, and
--share-screen-audio flags, plus a stdin command loop (ptt on/off, mode
vad/ptt) for manual verification. New test_vad_ptt_devices.cpp covers all
four items (ABI-level + a white-box AudioEngine stereo-mix check).

Docs updated to match: voice.md, roadmap.md (decision-log entry superseding
the original webrtc-audio-processing choice), tech-stack.md, README.md,
architecture.md, CLAUDE.md, PROGRESS.md.

Still explicitly out of scope, documented not silently dropped: real
webrtc-audio-processing/AEC (no AEC/NS/AGC exists at all yet), macOS/iOS
SCREEN_AUDIO capture, process-specific loopback, and a pre-existing
RT-thread rule violation in the capture path that predates this work.

Verified: ctest 12/12 green across 3 consecutive full-suite runs (both dev
and m1-dev presets build clean); test_vad_ptt_devices passed 5 consecutive
standalone runs; manually verified live (vccli --list-devices against real
hardware, vccli --voice --input-mode vad streaming without incident).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 16:11:52 +02:00

7.0 KiB
Raw Blame History

CLAUDE.md — agent hub for VoiceCat

Auto-loaded each session. This is the map: build commands, architecture at a glance, and where everything is. For the working method read AGENTS.md; for what's done and what's next read PROGRESS.md; for design read docs/.

One-line status: M3 (multi-stream, per-channel tuning) is complete, plus a follow-up pass closing the device enumeration / VAD-PTT input gate / stereo playback / WASAPI loopback gaps it left open (ctest --test-dir build/m1-dev green — 12/12 tests, including test_vad_ptt_devices, and vccli --voice/--list-devices manually verified live). Real webrtc-audio-processing (AEC/NS/AGC) is still unbuilt — no working Windows/MSVC port upstream — so v1 ships a lightweight energy/RMS VAD instead. Next up is M4 (native clients). See PROGRESS.md.

VoiceCat = self-hosted native voice & text chat (TeamSpeak/Mumble-style). Plain TCP (control)

  • UDP (media), no WebRTC, encrypted by default. A shared C++ core (libvoicecat) drives native clients (Swift on macOS/iOS, C# on Windows) and the server.

Build & test commands

The M0 skeleton builds with no third-party dependencies — just CMake + Ninja + a C++20 compiler. Deps (vcpkg) are off until a subsystem needs them.

# Configure + build the skeleton (default; no vcpkg needed)
cmake --preset dev
cmake --build --preset dev

# Run the tests (behavior smoke test today; grows per milestone)
ctest --preset dev                      # or: ctest --test-dir build/dev --output-on-failure

# Run the binaries (Windows adds .exe; Linux/macOS no extension)
./build/dev/bin/vccli                    # headless test client
./build/dev/bin/voicecat-server --help
./build/dev/bin/voicecat-server --name "My Server"

# Build a single target / be verbose
cmake --build --preset dev --target vccli
cmake --build --preset dev --verbose

# Clean
rm -rf build/dev                         # nuke; or:
cmake --build --preset dev --target clean

When you start a subsystem that needs real libraries (mbedTLS, libsodium, opus, protobuf, …), turn vcpkg deps on:

# one-time: git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh  (.bat on Windows)
export VCPKG_ROOT=/path/to/vcpkg         # works on Linux / macOS / Windows
cmake --preset server-release            # auto-installs deps pinned in vcpkg.json
cmake --build --preset server-release

Other useful toggles (pass with -D at configure time):

cmake --preset dev -DVOICECAT_BUILD_SHARED=ON     # build libvoicecat as a .dll/.so/.dylib (for the C# client)
cmake --preset dev -DVOICECAT_BUILD_SERVER=OFF     # core + tools only
cmake --preset dev -DVOICECAT_BUILD_TESTS=OFF

Formatting: clang-format config is .clang-format (Google base, 100 cols, 4-space).

git ls-files '*.cpp' '*.h' | xargs clang-format -i

Architecture at a glance

Full detail: docs/architecture.md. The short version:

 Swift (macOS/iOS)  ─┐                                    ┌─ C# (Windows)
                     ├──▶  libvoicecat  (C ABI: voicecat.h) ◀──┤
 voicecat-server  ───┘   net · crypto · codec · protocol ·     └─ all UIs are thin
   (links core)          session · audio                          the core owns audio
  • One core, many faces. Protocol, Opus, crypto, networking, jitter buffer, and mixing live once in C++. Clients call the C ABI (core/include/voicecat.h); the server links the same core, so framing/crypto never drift between ends.
  • Two transports. TCP + TLS 1.3 (control, protobuf Envelope) and UDP + exported-key ChaCha20-Poly1305 AEAD (media, fixed binary voice frame). Encryption is mandatory.
  • Threading. Real-time audio threads never allocate/lock/block; they exchange data with the net thread via lock-free ring buffers; a worker pool absorbs blocking work.

Subsystem map (code ↔ design doc)

Path Subsystem Design
core/include/voicecat.h The C ABI (client/server contract) architecture.md §4
core/proto/voicecat.proto Control-plane wire format (source of truth) protocol.md
core/src/net/ Asio TCP/UDP transport, [u32 len][payload] framing protocol.md §1, voice.md §2
core/src/crypto/ TLS 1.3 (mbedTLS), media AEAD (libsodium), anti-replay security.md
core/src/codec/ Opus encode/decode, FEC/DTX voice.md §34
core/src/protocol/ Envelope (de)serialize, request/response, dispatch protocol.md
core/src/session/ Channels, users, streams, permissions, ephemeral text protocol.md §5
core/src/audio/ miniaudio I/O, APM DSP, jitter buffer, mixer voice.md §811
core/src/core/ vc_client — the handle behind the C ABI architecture.md §4
server/ Connection mgr, session registry, SFU relay, SQLite architecture.md §5
tools/vccli/ Headless client that drives/verifies the protocol
clients/apple/, clients/windows/ Native GUIs (M4) architecture.md §4

Documentation index (source of truth)

Read docs/ before changing behavior. Order:

  1. docs/README.md — overview, locked decisions, glossary
  2. docs/architecture.md — core, C ABI, threading, server
  3. docs/protocol.md — control plane, Envelope, message catalog
  4. docs/voice.md — UDP media, Opus, multi-stream, two-sided NR, VAD/PTT
  5. docs/security.md — mandatory encryption, TLS+AEAD, accounts, threat model
  6. docs/tech-stack.md — libraries, permissive-license rule, tooling
  7. docs/deployment.md — zero-config self-host (Docker / binary / source)
  8. docs/roadmap.md — milestones + resolved decisions

Keeping track of progress

PROGRESS.md is the living status file. When you finish a task, check it off there and note the next step, so the next agent can pick up instantly. Treat it as part of the work, not an afterthought — update it in the same commit as the code.


House rules (hard constraints)

  • A clean compile is the floor, not the goal. "Done" = the milestone's observable exit criterion in docs/roadmap.md passes (e.g. M1 = two vccli actually chat over TLS). Encode it as a test. See AGENTS.md.
  • No GPL/LGPL dependencies, ever (closed-source redistribution is a goal). docs/tech-stack.md §5.
  • Encryption is mandatory — never add a plaintext transport path. docs/security.md.
  • Real-time audio threads never allocate, lock, or block. docs/architecture.md §3.
  • Keep docs + code in sync. Changing a wire format (voicecat.proto) or the C ABI (voicecat.h) is a deliberate, versioned act — update the doc in the same commit (protocol.md §8).
  • Every commit must build (cmake --build --preset dev) and pass ctest --preset dev.