feat(M3): multi-stream & per-channel tuning
Implements docs/roadmap.md M3: multiple concurrent streams per user (MIC + SCREEN_AUDIO + AUX_DEVICE), independent per-stream receiver gain/mute/noise- reduction, talk indicators, and enforced per-channel Opus configurability (mono/stereo, bitrate, frame size, FEC/DTX, application). Bugs fixed along the way (found while implementing, not pre-existing scope): - Server hard-coded stream_id=1 for every announce, so a second stream from the same user silently overwrote the first in SessionRegistry::set_user_stream. Now a per-session counter (ConnSession::next_stream_id_); handle_stream_stop validates against announced_stream_ids_ before clearing. - Client dropped mode/dtx/complexity/application from effective_audio even for the single M2 stream -- only sample_rate/bitrate_bps/frame_ms/fec were ever applied to OpusParams. Fixed on both the send (handle_stream_announce_result) and receive (sync_remote_streams) paths via a shared opus_params_from_audio_config() helper. - OpusEncoder always used OPUS_APPLICATION_VOIP; added OpusParams::application and wired it through. - on_playback's per-stream decode passed the wrong frame_size to opus_decode (total samples instead of samples-per-channel), which would have overflowed the decode buffer for any stereo stream. - teardown_voice() raced when called concurrently from run_io()'s own cleanup and from disconnect() on a different thread -- both could see udp_thread_/talk_timer_thread_ as joinable() at once and race to join() the same std::thread (intermittent std::system_error under ctest). Fixed with a teardown_mu_ guard instead of carrying the flake forward. New: - Per-channel AudioConfig: SessionRegistry now seeds Lobby (mono/24kbps/VOIP/ FEC+DTX) and a new "Music Room" channel (stereo/128kbps/AUDIO/no DTX); handle_stream_announce enforces the channel's config, clamping (not overriding) bitrate_bps to its ceiling. - core/src/core/client.h/.cpp: local-stream state is now a std::unordered_map<int, LocalStream> keyed by vc_stream_kind, with request_id-correlated announce/result handling (request_id already round-tripped on the wire; just wasn't read before). on_capture_frame is kind-aware and upmixes mono capture to stereo when a stream's config calls for it. set_self_mute's mic_muted now only gates the MIC kind. NS is wired through set_remote_stream. New run_talk_timer() thread emits VC_EVENT_TALK_STATE from both remote and local edge detection. - core/src/audio/audio_engine.h/.cpp: kind-keyed injection taps (inject_capture), stereo-to-mono downmix at the decode/mix boundary, RemoteStream gains recv_ns (lazy ApmProcessor) + noise_reduction_enabled and last_voice_ms/talking; new set_stream_noise_reduction() and poll_talk_transitions(). - core/src/session/session.h/.cpp: Stream now carries the full AudioConfig, not just sample_rate/frame_ms. - New additive C ABI (core/include/voicecat.h): vc_audio_config + vc_get_stream_audio_config (effective Opus config for any stream you own or a peer's); vc_test_inject_capture (test-only synthetic PCM injection, clearly marked, mirrors AudioEngine::inject_capture). - tests/test_m3_multistream.cpp: the M3 exit criterion through the real ABI (mirrors test_voice_client_abi.cpp's approach, not raw sockets) -- two concurrent local streams, independent gain/mute/NS control, per-channel config divergence via vc_get_stream_audio_config, talk indicators. Explicitly out of scope for this pass (tracked in PROGRESS.md, not silently dropped): VAD/PTT input gate + device enumeration; real WASAPI loopback capture for SCREEN_AUDIO (synthetic injection only); true stereo playback output (AudioEngine's mixer/output device stays mono -- Opus itself is fully stereo-correct on the wire). ctest --test-dir build/m1-dev: 11/11 green, verified across 3 consecutive full-suite runs plus 8 standalone runs of the new test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -6,10 +6,19 @@
|
||||
#include "audio/audio_engine.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
#include <cstring>
|
||||
|
||||
namespace voicecat::audio {
|
||||
|
||||
namespace {
|
||||
int64_t now_ms() {
|
||||
return std::chrono::duration_cast<std::chrono::milliseconds>(
|
||||
std::chrono::steady_clock::now().time_since_epoch())
|
||||
.count();
|
||||
}
|
||||
} // namespace
|
||||
|
||||
// ── JitterBuffer ─────────────────────────────────────────────────────────────
|
||||
|
||||
void JitterBuffer::push(Frame f) {
|
||||
@@ -66,9 +75,7 @@ void JitterBuffer::reset() {
|
||||
|
||||
// ── AudioEngine ──────────────────────────────────────────────────────────────
|
||||
|
||||
AudioEngine::AudioEngine() {
|
||||
inject_ring_.resize(kInjectCapSamples, 0);
|
||||
}
|
||||
AudioEngine::AudioEngine() = default;
|
||||
|
||||
AudioEngine::~AudioEngine() { stop(); }
|
||||
|
||||
@@ -135,31 +142,43 @@ void AudioEngine::stop() {
|
||||
#endif
|
||||
}
|
||||
|
||||
void AudioEngine::inject_capture(const int16_t* pcm, size_t n) {
|
||||
std::lock_guard lk(inject_mu_);
|
||||
size_t w = inject_write_.load(std::memory_order_relaxed);
|
||||
void AudioEngine::inject_capture(int kind, const int16_t* pcm, size_t n) {
|
||||
InjectTap* tap;
|
||||
{
|
||||
std::lock_guard lk(inject_mu_);
|
||||
auto& slot = inject_taps_[kind];
|
||||
if (!slot) {
|
||||
slot = std::make_unique<InjectTap>();
|
||||
slot->ring.resize(kInjectCapSamples, 0);
|
||||
}
|
||||
tap = slot.get();
|
||||
}
|
||||
|
||||
size_t w = tap->write.load(std::memory_order_relaxed);
|
||||
for (size_t i = 0; i < n; ++i)
|
||||
inject_ring_[(w + i) % kInjectCapSamples] = pcm[i];
|
||||
inject_write_.store(w + n, std::memory_order_release);
|
||||
tap->ring[(w + i) % kInjectCapSamples] = pcm[i];
|
||||
tap->write.store(w + n, std::memory_order_release);
|
||||
|
||||
// Fire capture_cb_ for each complete frame now available.
|
||||
while (true) {
|
||||
size_t r = inject_read_.load(std::memory_order_relaxed);
|
||||
size_t avail = inject_write_.load(std::memory_order_acquire) - r;
|
||||
size_t r = tap->read.load(std::memory_order_relaxed);
|
||||
size_t avail = tap->write.load(std::memory_order_acquire) - r;
|
||||
if (avail < static_cast<size_t>(frame_samples_)) break;
|
||||
|
||||
std::vector<int16_t> frame(frame_samples_);
|
||||
for (int i = 0; i < frame_samples_; ++i)
|
||||
frame[i] = inject_ring_[(r + i) % kInjectCapSamples];
|
||||
inject_read_.store(r + frame_samples_, std::memory_order_release);
|
||||
frame[i] = tap->ring[(r + i) % kInjectCapSamples];
|
||||
tap->read.store(r + frame_samples_, std::memory_order_release);
|
||||
|
||||
if (capture_cb_) capture_cb_(frame.data(), frame_samples_);
|
||||
if (capture_cb_) capture_cb_(kind, frame.data(), frame_samples_);
|
||||
}
|
||||
}
|
||||
|
||||
void AudioEngine::push_recv_frame(uint32_t ssrc, JitterBuffer::Frame f) {
|
||||
std::lock_guard lk(streams_mu_);
|
||||
streams_[ssrc].jitter.push(std::move(f));
|
||||
auto& s = streams_[ssrc];
|
||||
s.last_voice_ms.store(now_ms(), std::memory_order_relaxed);
|
||||
s.jitter.push(std::move(f));
|
||||
}
|
||||
|
||||
void AudioEngine::set_stream_gain(uint32_t ssrc, float gain) {
|
||||
@@ -172,11 +191,37 @@ void AudioEngine::set_stream_mute(uint32_t ssrc, bool mute) {
|
||||
streams_[ssrc].mute = mute;
|
||||
}
|
||||
|
||||
void AudioEngine::set_stream_noise_reduction(uint32_t ssrc, bool enable) {
|
||||
std::lock_guard lk(streams_mu_);
|
||||
auto& s = streams_[ssrc];
|
||||
s.noise_reduction_enabled = enable;
|
||||
if (enable) {
|
||||
if (!s.recv_ns) s.recv_ns = ApmProcessor::create();
|
||||
} else {
|
||||
s.recv_ns.reset();
|
||||
}
|
||||
}
|
||||
|
||||
void AudioEngine::remove_stream(uint32_t ssrc) {
|
||||
std::lock_guard lk(streams_mu_);
|
||||
streams_.erase(ssrc);
|
||||
}
|
||||
|
||||
std::vector<std::pair<uint32_t, bool>> AudioEngine::poll_talk_transitions() {
|
||||
std::vector<std::pair<uint32_t, bool>> edges;
|
||||
std::lock_guard lk(streams_mu_);
|
||||
int64_t now = now_ms();
|
||||
for (auto& [ssrc, stream] : streams_) {
|
||||
bool now_talking = (now - stream.last_voice_ms.load(std::memory_order_relaxed)) <
|
||||
kTalkHangoverMs;
|
||||
if (now_talking != stream.talking) {
|
||||
stream.talking = now_talking;
|
||||
edges.emplace_back(ssrc, now_talking);
|
||||
}
|
||||
}
|
||||
return edges;
|
||||
}
|
||||
|
||||
uint32_t AudioEngine::stream_packets_lost(uint32_t ssrc) const {
|
||||
std::lock_guard lk(streams_mu_);
|
||||
auto it = streams_.find(ssrc);
|
||||
@@ -205,7 +250,10 @@ void AudioEngine::capture_data_cb(ma_device* dev, void* /*out*/,
|
||||
}
|
||||
|
||||
void AudioEngine::on_capture(const int16_t* pcm, ma_uint32 frames) {
|
||||
if (capture_cb_) capture_cb_(pcm, static_cast<int>(frames));
|
||||
// The real hardware capture device is always the "primary" tap (kind 0 / MIC). A second
|
||||
// concurrent local stream (e.g. SCREEN_AUDIO) is fed via inject_capture() in M3 — there is
|
||||
// only one real capture device.
|
||||
if (capture_cb_) capture_cb_(0, pcm, static_cast<int>(frames));
|
||||
}
|
||||
|
||||
void AudioEngine::playback_data_cb(ma_device* dev, void* out,
|
||||
@@ -226,24 +274,43 @@ void AudioEngine::on_playback(int16_t* out, ma_uint32 frames) {
|
||||
for (auto& [ssrc, stream] : streams_) {
|
||||
if (stream.mute || !stream.decoder.valid()) continue;
|
||||
|
||||
// M3: a stream's Opus channel count (mono/stereo, per-channel AudioConfig) may differ
|
||||
// from the engine-wide playback channel count (always mono in M3 — see PROGRESS.md).
|
||||
// Decode into a buffer sized for the *decoder's* channel count (opus_decode's
|
||||
// frame_size parameter is samples-per-channel, not total samples — pass `frames`,
|
||||
// not `frames * channels`), then convert at this boundary.
|
||||
int dec_channels = std::max(1, stream.decoder.channels());
|
||||
auto maybe_frame = stream.jitter.pop(stream.playout_ts);
|
||||
std::vector<int16_t> pcm(frames * params_.channels);
|
||||
std::vector<int16_t> pcm(frames * static_cast<ma_uint32>(dec_channels));
|
||||
int n;
|
||||
|
||||
if (maybe_frame) {
|
||||
n = stream.decoder.decode(
|
||||
maybe_frame->payload.data(),
|
||||
static_cast<int>(maybe_frame->payload.size()),
|
||||
pcm.data(), static_cast<int>(pcm.size()));
|
||||
pcm.data(), static_cast<int>(frames));
|
||||
} else {
|
||||
n = stream.decoder.decode(nullptr, 0, pcm.data(),
|
||||
static_cast<int>(pcm.size()));
|
||||
n = stream.decoder.decode(nullptr, 0, pcm.data(), static_cast<int>(frames));
|
||||
}
|
||||
|
||||
if (n > 0) {
|
||||
// `n` is samples-per-channel (matches the frame_samples convention used by
|
||||
// OpusEncoder::encode elsewhere in the codebase).
|
||||
if (stream.recv_ns)
|
||||
stream.recv_ns->process_capture(pcm.data(), n,
|
||||
static_cast<int>(params_.sample_rate));
|
||||
|
||||
float g = stream.gain;
|
||||
for (int i = 0; i < n * static_cast<int>(params_.channels); ++i)
|
||||
mix[i] += static_cast<int32_t>(static_cast<float>(pcm[i]) * g);
|
||||
for (int i = 0; i < n; ++i) {
|
||||
// Downmix decoder output to the engine's mono accumulator if needed
|
||||
// (average L/R); upmix is unnecessary since the mix buffer is per-channel.
|
||||
int32_t sample = (dec_channels == 2)
|
||||
? (static_cast<int32_t>(pcm[i * 2]) +
|
||||
static_cast<int32_t>(pcm[i * 2 + 1])) / 2
|
||||
: static_cast<int32_t>(pcm[i]);
|
||||
for (uint32_t c = 0; c < params_.channels; ++c)
|
||||
mix[i * params_.channels + c] += static_cast<int32_t>(sample * g);
|
||||
}
|
||||
}
|
||||
stream.playout_ts += frames;
|
||||
}
|
||||
|
||||
@@ -31,6 +31,8 @@
|
||||
#include "codec/opus_codec.h"
|
||||
#endif
|
||||
|
||||
#include "audio/apm_processor.h"
|
||||
|
||||
namespace voicecat::audio {
|
||||
|
||||
// ── JitterBuffer ─────────────────────────────────────────────────────────────
|
||||
@@ -83,8 +85,12 @@ struct AudioParams {
|
||||
// Owns miniaudio capture/playback, per-ssrc jitter buffers + Opus decoders, and the mixer.
|
||||
class AudioEngine {
|
||||
public:
|
||||
// Callback type for encoded capture frames ready to be sent.
|
||||
using CaptureCallback = std::function<void(const int16_t* pcm, int samples)>;
|
||||
// Callback type for encoded capture frames ready to be sent. `kind` identifies which
|
||||
// local stream this PCM belongs to (a vc_stream_kind value; 0 = MIC for the real capture
|
||||
// device, which is always the "primary" tap). M3: multiple concurrent local streams are
|
||||
// possible (e.g. MIC + SCREEN_AUDIO), each fed via its own injection tap (see
|
||||
// inject_capture) since there is only one real hardware capture device.
|
||||
using CaptureCallback = std::function<void(int kind, const int16_t* pcm, int samples)>;
|
||||
|
||||
AudioEngine();
|
||||
~AudioEngine();
|
||||
@@ -100,8 +106,11 @@ class AudioEngine {
|
||||
bool running() const { return running_.load(std::memory_order_acquire); }
|
||||
|
||||
// Inject synthetic PCM directly into the capture pipeline (bypasses real device).
|
||||
// Thread-safe; can be called from any thread including tests.
|
||||
void inject_capture(const int16_t* pcm, size_t n);
|
||||
// Thread-safe; can be called from any thread including tests. `kind` selects which local
|
||||
// stream's injection tap to feed (each gets its own ring buffer); the 2-arg overload
|
||||
// targets kind 0 (MIC) for source compatibility with existing callers.
|
||||
void inject_capture(int kind, const int16_t* pcm, size_t n);
|
||||
void inject_capture(const int16_t* pcm, size_t n) { inject_capture(0, pcm, n); }
|
||||
|
||||
// Called by the net thread when a decoded voice frame arrives for a remote stream.
|
||||
void push_recv_frame(uint32_t ssrc, JitterBuffer::Frame f);
|
||||
@@ -109,8 +118,17 @@ class AudioEngine {
|
||||
// Per-stream receive-side controls (safe from any thread).
|
||||
void set_stream_gain(uint32_t ssrc, float gain); // 0.0–2.0, default 1.0
|
||||
void set_stream_mute(uint32_t ssrc, bool mute);
|
||||
// Listener-chosen, local-only noise reduction on a specific remote stream (docs/voice.md
|
||||
// §10) — lazily instantiates an ApmProcessor on first enable, frees it on disable.
|
||||
void set_stream_noise_reduction(uint32_t ssrc, bool enable);
|
||||
void remove_stream(uint32_t ssrc);
|
||||
|
||||
// Edge-triggered talk-state transitions since the last call (docs/voice.md §7: talk state
|
||||
// is derived from recent frame arrival, no protocol message). Call from a lightweight
|
||||
// poller, not the audio callback thread. Returns {ssrc, now_talking} for each stream whose
|
||||
// state flipped.
|
||||
std::vector<std::pair<uint32_t, bool>> poll_talk_transitions();
|
||||
|
||||
// Get stats for a remote stream's jitter buffer.
|
||||
uint32_t stream_packets_lost(uint32_t ssrc) const;
|
||||
uint32_t stream_target_depth_ms(uint32_t ssrc) const;
|
||||
@@ -138,12 +156,16 @@ class AudioEngine {
|
||||
CaptureCallback capture_cb_;
|
||||
std::atomic<bool> running_{false};
|
||||
|
||||
// Inject ring: stores raw int16 PCM written by inject_capture().
|
||||
// The encode thread reads from this (no real capture device needed in tests).
|
||||
std::mutex inject_mu_;
|
||||
std::vector<int16_t> inject_ring_; // circular, size = frame_samples_
|
||||
std::atomic<size_t> inject_write_{0};
|
||||
std::atomic<size_t> inject_read_{0};
|
||||
// Inject ring(s): stores raw int16 PCM written by inject_capture(), one ring per local
|
||||
// stream kind so e.g. MIC and SCREEN_AUDIO can each be fed independently in tests.
|
||||
// The encode thread reads from these (no real capture device needed in tests).
|
||||
struct InjectTap {
|
||||
std::vector<int16_t> ring; // circular, size = kInjectCapSamples
|
||||
std::atomic<size_t> write{0};
|
||||
std::atomic<size_t> read{0};
|
||||
};
|
||||
std::mutex inject_mu_;
|
||||
std::unordered_map<int, std::unique_ptr<InjectTap>> inject_taps_;
|
||||
static constexpr size_t kInjectCapSamples = 48000 * 2; // 2 s @48 kHz mono
|
||||
|
||||
// Per remote stream (protected by streams_mu_).
|
||||
@@ -155,11 +177,24 @@ class AudioEngine {
|
||||
float gain = 1.0f;
|
||||
bool mute = false;
|
||||
uint32_t playout_ts = 0;
|
||||
|
||||
// M3: listener-chosen, local-only noise reduction (docs/voice.md §10). Lazily
|
||||
// created only when enabled — bounded by how many remote streams this listener
|
||||
// subscribes to, so no separate instance cap is needed.
|
||||
bool noise_reduction_enabled = false;
|
||||
std::unique_ptr<ApmProcessor> recv_ns;
|
||||
|
||||
// M3: talk-indicator edge detection (docs/voice.md §7) — updated by push_recv_frame
|
||||
// (already off the real-time audio thread), polled by poll_talk_transitions().
|
||||
std::atomic<int64_t> last_voice_ms{0};
|
||||
bool talking = false;
|
||||
};
|
||||
mutable std::mutex streams_mu_;
|
||||
std::unordered_map<uint32_t, RemoteStream> streams_;
|
||||
|
||||
int frame_samples_ = 960; // 20 ms @48 kHz
|
||||
|
||||
static constexpr int64_t kTalkHangoverMs = 300;
|
||||
};
|
||||
|
||||
} // namespace voicecat::audio
|
||||
|
||||
Reference in New Issue
Block a user