Files
voice-cat/CLAUDE.md
Talon bad9c7533a
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
feat(audio): real noise suppression via vendored RNNoise (send + receive)
The two-sided NR plumbing (RemoteStream::recv_ns + the per-listener
vc_set_remote_stream noise_reduction toggle) was wired but inert:
ApmProcessor::create() returned a no-op passthrough, because the
originally-planned webrtc-audio-processing has no working Windows/macOS
build. Drop in RNNoise as the real backend behind the same ApmProcessor
interface, lighting up both NR paths.

- Vendor RNNoise (BSD-3 + CC0) at third_party/rnnoise/ — the vcpkg port
  is !windows !arm, so it can't cover our primary targets. Shrunk int8
  model (78MB -> 11.7MB via upstream scripts/shrink_model.sh), built as a
  standalone C static lib with no RTCD (portable scalar path on x86,
  auto-NEON on arm64) under -DDISABLE_DEBUG_FLOAT. Model is baked in
  (rnnoise_create(NULL)); no runtime file.
- New RnnoiseProcessor (core/src/audio/apm_processor.cpp) selected by
  ApmProcessor::create() when VOICECAT_HAS_NS. Mono/48kHz/480-sample;
  our clock is fixed 48kHz and Opus frame sizes are multiples of 480, so
  no resampling. RT-safe: allocates at construction, lock-free in the
  capture/playback callbacks.
- Receive-side: lit up via the factory; gated to mono streams (a stereo
  stream is a screen-audio share, not voice).
- Send-side (new): vc_set_input_noise_reduction(client, enable) ABI +
  vc_client::mic_ns_, run before input gain/VAD in on_capture_frame. A
  stereo mic is downmixed to mono ONLY when NR is on — with NR off a
  stereo mic keeps full stereo (never collapse mic quality unasked).
- Enable C as a project language for the vendored lib.
- New noise_suppression test: white noise through ApmProcessor::create()
  drops ~99.9% RMS. ctest --preset dev green, 28/28. windows-client DLL
  builds clean with vc_set_input_noise_reduction exported, system-only deps.
- Docs synced: voice.md §10, tech-stack.md §1/§5, third_party/README.md,
  vcpkg.json note, PROGRESS.md, CLAUDE.md.

Client on/off UI toggles (Windows/macOS/iOS) are the remaining follow-up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 13:30:54 +02:00

8.3 KiB
Raw Blame History

CLAUDE.md — agent hub for VoiceCat

Auto-loaded each session. This is the map: build commands, architecture at a glance, and where everything is. For the working method read AGENTS.md; for what's done and what's next read PROGRESS.md; for design read docs/.

One-line status: M5 (moderation & admin UI) is complete — permissions, kick/ban/move, server-mute, channel CRUD, in-app account management, disconnect/keepalive/reaper. Windows WinForms C# client shipped (M4). macOS AppKit client shippedVoiceCatMac.xcodeproj at clients/apple/macOS/. iOS SwiftUI client shippedVoiceCatiOS.xcodeproj at clients/apple/iOS/. ctest --preset dev green — 28/28 tests. External PCM feed/tap API (vc_stream_feed_pcm + vc_set_pcm_sink) shipped. Screen-audio sharing shipped on macOS (ScreenCaptureKit) and iOS (ReplayKit Broadcast Upload Extension → host App Group ring → vc_stream_feed_pcm). Noise suppression shipped (RNNoise, vendored at third_party/rnnoise/) — both send-side mic NR (vc_set_input_noise_reduction) and per-listener receive NR; client on/off toggles TBD. See PROGRESS.md.

VoiceCat = self-hosted native voice & text chat (TeamSpeak/Mumble-style). Plain TCP (control)

  • UDP (media), no WebRTC, encrypted by default. A shared C++ core (libvoicecat) drives native clients (Swift on macOS/iOS, C# on Windows) and the server.

Build & test commands

The default development preset is dev — it builds everything (server + tools + tests) with real vcpkg deps. The skeleton preset (no deps, stubs only) is a fast smoke check; see docs/building.md for the full preset matrix.

# Configure + build (default development preset; needs VCPKG_ROOT)
cmake --preset dev
cmake --build --preset dev

# Run the tests (21 behavior tests — grows per milestone)
# NOTE on Windows: run ctest via PowerShell, NOT Git Bash — MinGW binaries fail in Git Bash
# with exit 0xc0000139 (STATUS_ENTRYPOINT_NOT_FOUND). PowerShell runs them correctly.
ctest --preset dev                      # or: ctest --test-dir build/dev --output-on-failure

# Run the binaries — same Windows rule: use PowerShell, not Git Bash
./build/dev/bin/vccli                    # headless test client
./build/dev/bin/voicecat-server --help
./build/dev/bin/voicecat-server --name "My Server"

# Build a single target / be verbose
cmake --build --preset dev --target vccli
cmake --build --preset dev --verbose

# Clean
rm -rf build/dev                         # nuke; or:
cmake --build --preset dev --target clean

Other presets (see docs/building.md for full detail):

cmake --preset skeleton                  # no-deps stub smoke (no VCPKG_ROOT needed) — 2 tests
cmake --preset release                   # optimized + tests on, symbols kept (profile/debug-friendly)
cmake --preset server-release            # optimized + stripped, no tests (deployment-shaped)
cmake --preset windows-client            # voicecat.dll for the C# WinForms client (Windows only)
cmake --preset apple-dev                 # libvoicecat.a for macOS Swift Package (scaffolding, macOS only)

Vcpkg triplet is auto-resolved from the host platform by cmake/voicecat-toolchain.cmakex64-mingw-static on Windows, x64-linux on Linux, arm64-osx on Apple Silicon. See docs/building.md §1 "Platform matrix" for details.

One-time vcpkg setup:

# one-time: git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh  (.bat on Windows)
export VCPKG_ROOT=/path/to/vcpkg         # works on Linux / macOS / Windows

Other useful toggles (pass with -D at configure time):

cmake --preset dev -DVOICECAT_BUILD_SHARED=ON     # build libvoicecat as a .dll/.so/.dylib (for the C# client)
cmake --preset dev -DVOICECAT_BUILD_SERVER=OFF     # core + tools only
cmake --preset dev -DVOICECAT_BUILD_TESTS=OFF

Formatting: clang-format config is .clang-format (Google base, 100 cols, 4-space).

git ls-files '*.cpp' '*.h' | xargs clang-format -i

Architecture at a glance

Full detail: docs/architecture.md. The short version:

 Swift (macOS/iOS)  ─┐                                    ┌─ C# (Windows)
                     ├──▶  libvoicecat  (C ABI: voicecat.h) ◀──┤
 voicecat-server  ───┘   net · crypto · codec · protocol ·     └─ all UIs are thin
   (links core)          session · audio                          the core owns audio
  • One core, many faces. Protocol, Opus, crypto, networking, jitter buffer, and mixing live once in C++. Clients call the C ABI (core/include/voicecat.h); the server links the same core, so framing/crypto never drift between ends.
  • Two transports. TCP + TLS 1.3 (control, protobuf Envelope) and UDP + exported-key ChaCha20-Poly1305 AEAD (media, fixed binary voice frame). Encryption is mandatory.
  • Threading. Real-time audio threads never allocate/lock/block; they exchange data with the net thread via lock-free ring buffers; a worker pool absorbs blocking work.

Subsystem map (code ↔ design doc)

Path Subsystem Design
core/include/voicecat.h The C ABI (client/server contract) architecture.md §4
core/proto/voicecat.proto Control-plane wire format (source of truth) protocol.md
core/src/net/ Asio TCP/UDP transport, [u32 len][payload] framing protocol.md §1, voice.md §2
core/src/crypto/ TLS 1.3 (mbedTLS), media AEAD (libsodium), anti-replay security.md
core/src/codec/ Opus encode/decode, FEC/DTX voice.md §34
core/src/protocol/ Envelope (de)serialize, request/response, dispatch protocol.md
core/src/session/ Channels, users, streams, permissions, ephemeral text protocol.md §5
core/src/audio/ miniaudio I/O, APM DSP, jitter buffer, mixer voice.md §811
core/src/core/ vc_client — the handle behind the C ABI architecture.md §4
server/ Connection mgr, session registry, SFU relay, SQLite architecture.md §5
tools/vccli/ Headless client that drives/verifies the protocol
clients/apple/, clients/windows/ Native GUIs (M4) architecture.md §4

Documentation index (source of truth)

Read docs/ before changing behavior. Order:

  1. docs/README.md — overview, locked decisions, glossary
  2. docs/architecture.md — core, C ABI, threading, server
  3. docs/protocol.md — control plane, Envelope, message catalog
  4. docs/voice.md — UDP media, Opus, multi-stream, two-sided NR, VAD/PTT
  5. docs/security.md — mandatory encryption, TLS+AEAD, accounts, threat model
  6. docs/tech-stack.md — libraries, permissive-license rule, tooling
  7. docs/deployment.md — zero-config self-host (Docker / binary / source)
  8. docs/roadmap.md — milestones + resolved decisions
  9. docs/building.md — what each CMake preset is for + manual server/vccli testing

Keeping track of progress

PROGRESS.md is the living status file. When you finish a task, check it off there and note the next step, so the next agent can pick up instantly. Treat it as part of the work, not an afterthought — update it in the same commit as the code.


House rules (hard constraints)

  • A clean compile is the floor, not the goal. "Done" = the milestone's observable exit criterion in docs/roadmap.md passes (e.g. M1 = two vccli actually chat over TLS). Encode it as a test. See AGENTS.md.
  • No GPL/LGPL dependencies, ever (closed-source redistribution is a goal). docs/tech-stack.md §5.
  • Encryption is mandatory — never add a plaintext transport path. docs/security.md.
  • Real-time audio threads never allocate, lock, or block. docs/architecture.md §3.
  • Keep docs + code in sync. Changing a wire format (voicecat.proto) or the C ABI (voicecat.h) is a deliberate, versioned act — update the doc in the same commit (protocol.md §8).
  • Every commit must build (cmake --build --preset dev) and pass ctest --preset dev.