Talon ce2035f271
Some checks failed
Build Linux Binaries / linux/amd64 (push) Has been cancelled
Build Linux Binaries / linux/arm64 (push) Has been cancelled
fix(audio): bound playout depth to stop voice latency ratcheting up
Latency between speakers grew to multiple seconds and only reset on
rejoining voice. Root cause was the receiver playout logic, not the
codec settings: the playout clock free-ran in real time while the
sender omitted silence from its timestamps (and set no header flags),
and the only correction snapped the clock to the *oldest* buffered
frame — which could only ever add standing latency. target_depth_ms_
was computed but never enforced, so latency could only grow or reset.

Fix: bound playout against the stream's leading edge (newest frame).
(Re)seed to the leading edge on start/marker/starve (no prebuffer, so
latency stays low), and frame-skip catch-up trims any backlog beyond
target+hysteresis — the missing downward force.

Hardening: sender now stamps kFlagMarker (talkspurt start) and kFlagDtx,
consumed on recv for clean resync; adaptive late-drop window; EWMA
outlier rejection so silence gaps/stragglers don't poison the estimate;
duplicate counting and ring-underrun diagnostics.

New test_jitter_depth asserts depth stays bounded (<200ms) while
arrivals outrun playback. ctest --preset dev green (27/27).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 20:02:20 +02:00

VoiceCat

Self-hosted, native voice & text chat in the spirit of classic TeamSpeak / Mumble — channel-based voice, channel + private text, one server you run yourself. Plain TCP (control) and UDP (media), no WebRTC. Encrypted by default. A shared C++ core (libvoicecat) drives native clients (Swift on macOS/iOS, C# on Windows) and the server.

Status: Design complete in docs/. M1M5 are implemented — real TLS control plane, encrypted UDP voice (Opus), multi-stream, TOFU identity pinning, channel tree, permissions, moderation, disconnect/keepalive/reaper. Windows WinForms C# client shipped (M4). macOS/iOS Swift client is next. See PROGRESS.md and docs/roadmap.md.

Read the design first

The docs/ folder is the source of truth. Start at docs/README.md, then architectureprotocolvoicesecuritytech-stackdeploymentroadmap.

Build

The default development preset is dev — it builds everything (server + tools + tests) with real vcpkg deps. It works on Windows, Linux, and macOS (vcpkg triplet auto-resolved).

# one-time vcpkg setup:
git clone https://github.com/microsoft/vcpkg && ./vcpkg/bootstrap-vcpkg.sh  # .bat on Windows
export VCPKG_ROOT=/path/to/vcpkg            # Linux/macOS; or $env:VCPKG_ROOT on PowerShell

# configure + build + test:
cmake --preset dev
cmake --build --preset dev
ctest --preset dev                          # 21 behavior tests

Artifacts land in build/dev/bin/ (voicecat-server, vccli, voicecat-admin).

The skeleton preset (no vcpkg deps, stubs only) is a fast smoke check that needs no third-party libraries:

cmake --preset skeleton && cmake --build --preset skeleton && ctest --preset skeleton

See docs/building.md for the full preset matrix (including release, server-release, windows-client, and Apple platform scaffolding).

Layout

docs/         design spec (read this)
core/         libvoicecat — the shared C++ core
  include/    voicecat.h  (the C ABI all clients call)
  proto/      voicecat.proto  (control-plane wire format, source of truth)
  src/        net/ crypto/ codec/ protocol/ session/ audio/  (stubs today)
server/       voicecat-server (headless; links the core)
tools/vccli/  headless test client — drives the protocol from M1 on
clients/      apple/ (Swift, M4)   windows/ (C#, M4)   — placeholders for now
tests/        CTest targets

License

Permissive-only dependencies (no GPL/LGPL) so the project can be redistributed freely, including closed-source. Project license: TBD (see docs/tech-stack.md §5).

Description
Native voice chat server and client
Readme 12 MiB
Languages
C++ 43.9%
Swift 30.6%
C# 17.5%
Shell 3.6%
CMake 2.1%
Other 2.3%