Latency between speakers grew to multiple seconds and only reset on rejoining voice. Root cause was the receiver playout logic, not the codec settings: the playout clock free-ran in real time while the sender omitted silence from its timestamps (and set no header flags), and the only correction snapped the clock to the *oldest* buffered frame — which could only ever add standing latency. target_depth_ms_ was computed but never enforced, so latency could only grow or reset. Fix: bound playout against the stream's leading edge (newest frame). (Re)seed to the leading edge on start/marker/starve (no prebuffer, so latency stays low), and frame-skip catch-up trims any backlog beyond target+hysteresis — the missing downward force. Hardening: sender now stamps kFlagMarker (talkspurt start) and kFlagDtx, consumed on recv for clean resync; adaptive late-drop window; EWMA outlier rejection so silence gaps/stragglers don't poison the estimate; duplicate counting and ring-underrun diagnostics. New test_jitter_depth asserts depth stays bounded (<200ms) while arrivals outrun playback. ctest --preset dev green (27/27). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
VoiceCat
Self-hosted, native voice & text chat in the spirit of classic TeamSpeak / Mumble —
channel-based voice, channel + private text, one server you run yourself. Plain TCP
(control) and UDP (media), no WebRTC. Encrypted by default. A shared C++ core
(libvoicecat) drives native clients (Swift on macOS/iOS, C# on Windows) and the server.
Status: Design complete in
docs/. M1–M5 are implemented — real TLS control plane, encrypted UDP voice (Opus), multi-stream, TOFU identity pinning, channel tree, permissions, moderation, disconnect/keepalive/reaper. Windows WinForms C# client shipped (M4). macOS/iOS Swift client is next. SeePROGRESS.mdanddocs/roadmap.md.
Read the design first
The docs/ folder is the source of truth. Start at docs/README.md,
then architecture → protocol → voice → security → tech-stack → deployment →
roadmap.
Build
The default development preset is dev — it builds everything (server + tools + tests)
with real vcpkg deps. It works on Windows, Linux, and macOS (vcpkg triplet auto-resolved).
# one-time vcpkg setup:
git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh # .bat on Windows
export VCPKG_ROOT=/path/to/vcpkg # Linux/macOS; or $env:VCPKG_ROOT on PowerShell
# configure + build + test:
cmake --preset dev
cmake --build --preset dev
ctest --preset dev # 21 behavior tests
Artifacts land in build/dev/bin/ (voicecat-server, vccli, voicecat-admin).
The skeleton preset (no vcpkg deps, stubs only) is a fast smoke check that needs no
third-party libraries:
cmake --preset skeleton && cmake --build --preset skeleton && ctest --preset skeleton
See docs/building.md for the full preset matrix (including release,
server-release, windows-client, and Apple platform scaffolding).
Layout
docs/ design spec (read this)
core/ libvoicecat — the shared C++ core
include/ voicecat.h (the C ABI all clients call)
proto/ voicecat.proto (control-plane wire format, source of truth)
src/ net/ crypto/ codec/ protocol/ session/ audio/ (stubs today)
server/ voicecat-server (headless; links the core)
tools/vccli/ headless test client — drives the protocol from M1 on
clients/ apple/ (Swift, M4) windows/ (C#, M4) — placeholders for now
tests/ CTest targets
License
Permissive-only dependencies (no GPL/LGPL) so the project can be redistributed freely,
including closed-source. Project license: TBD (see docs/tech-stack.md §5).