Talon 5be869c61a fix(audio): decouple Opus decode cadence from playback callback period
on_playback() was passing miniaudio's hardware playback-callback frame
count to opus_decode()'s max_samples, instead of the decoder's fixed
frame size (960 samples @ 20ms/48kHz). Since real packets decode to
more samples than the (often smaller, e.g. ~480 on default low-latency
WASAPI) hardware period, opus_decode returned OPUS_BUFFER_TOO_SMALL on
nearly every callback -- packets were received/decrypted/jitter-buffered
correctly but never decoded into audible PCM. Result: control-plane
events and VAD worked, but zero audio in headphones.

mix_for_test()'s white-box test masked this since it always called
on_playback with frames == frame_samples, the one case where the bug
is invisible.

Fix: RemoteStream gained a small ring buffer (init_ring/push_ring/
pop_ring) that decouples decode cadence from playback-callback cadence.
on_playback now tops the ring up by decoding whole Opus frames (always
decoder.frame_samples(), never the hardware frame count) and drains
exactly what the callback asks for, silence-padding (PLC) on underrun.

Side effect: also fixes playout_ts, which was advancing by the wrong
unit (hardware frames instead of decoded samples) -- it now tracks
correctly against jitter-buffer timestamps.

ctest --test-dir build/m1-dev: 12/12 green.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-16 17:39:49 +02:00

VoiceCat

Self-hosted, native voice & text chat in the spirit of classic TeamSpeak / Mumble — channel-based voice, channel + private text, one server you run yourself. Plain TCP (control) and UDP (media), no WebRTC. Encrypted by default. A shared C++ core (libvoicecat) drives native clients (Swift on macOS/iOS, C# on Windows) and the server.

Status: pre-implementation. The design is complete in docs/. The code is an M0 skeleton — it compiles and links, but every subsystem is a stub. See AGENTS.md to start building, and docs/roadmap.md for the milestones.

Read the design first

The docs/ folder is the source of truth. Start at docs/README.md, then architectureprotocolvoicesecuritytech-stackdeploymentroadmap.

Build the skeleton (no dependencies needed yet)

The M0 skeleton builds with just a C++20 compiler + CMake + Ninja — no vcpkg, no third-party libraries, because every subsystem is currently a stub.

cmake --preset dev
cmake --build --preset dev
ctest --preset dev            # runs the smoke test (links the core, calls the C ABI)

Artifacts land in build/dev/bin/ (voicecat-server, vccli).

When you start implementing a subsystem that needs real libraries, build with vcpkg deps:

# one-time: git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh
export VCPKG_ROOT=/path/to/vcpkg            # set VCPKG_ROOT (works on Linux/macOS/Windows)
cmake --preset server-release               # auto-installs deps from vcpkg.json
cmake --build --preset server-release

Layout

docs/         design spec (read this)
core/         libvoicecat — the shared C++ core
  include/    voicecat.h  (the C ABI all clients call)
  proto/      voicecat.proto  (control-plane wire format, source of truth)
  src/        net/ crypto/ codec/ protocol/ session/ audio/  (stubs today)
server/       voicecat-server (headless; links the core)
tools/vccli/  headless test client — drives the protocol from M1 on
clients/      apple/ (Swift, M4)   windows/ (C#, M4)   — placeholders for now
tests/        CTest targets

License

Permissive-only dependencies (no GPL/LGPL) so the project can be redistributed freely, including closed-source. Project license: TBD (see docs/tech-stack.md §5).

Description
Native voice chat server and client
Readme 12 MiB
Languages
C++ 43.9%
Swift 30.6%
C# 17.5%
Shell 3.6%
CMake 2.1%
Other 2.3%