Bump to v2.2.0: Opus native binding, efficiency tidy-up, diag self-meter

Single biggest change: added the Concentus.Native NuGet package. Concentus
2.0+ auto-detects native libopus at runtime and routes encode calls
through it; encoder state lives on the C side and is reused across calls
rather than `new`ing ~15 working buffers per call (Concentus issue #22,
open since 2018). Measured on the desktop test at 15:36:55 — Opus 10 ms
allocation rate dropped from 4,625 KB/s to 108 KB/s, a 97.7% reduction.
Process CPU dropped from 4.7% to 1.6% in the same config. Audio is bit-
for-bit identical (it's literally the same encoder, just better
packaged). `OpusEncoderState.cs` itself unchanged on the call site.

Diagnostic / measurement layer (gated on Enable-logs, zero cost when off):
* ProcessSelfMeter: CPU%, managed heap MB, working set MB, allocation
  rate per second, GC counts per generation
* Per-thread work-time counters: captureMs / sendMs / recvMs / renderMs
  expressed as milliseconds of CPU consumed by each audio thread per
  second
* Inter-packet arrival gap measured at the user-space UDP socket
  (rxNetGapMs) — pinpoints whether arrival jitter is in the network or
  our own dispatch path

Small efficiency wins (each one was small but cumulative):
* deviceRefreshTimer interval 1s -> 3s (item 4)
* WaitHandle array allocations eliminated in MixingEngine.MixLoop and
  MultiOutputPlayout.ProduceLoop (item 6)
* MultiOutputPlayout caches its output-buffer snapshot and only rebuilds
  on SetOutputDevices, instead of rebuilding every 10 ms (item 7)
* HeartbeatService reuses an outbound ping byte[] instead of allocating
  per send (item 14)
* PeerDiscoveryService caches broadcast addresses and invalidates on
  Windows' NetworkChange event instead of walking all NICs every 1.5 s
  (item 16)

Legacy / dead-code removal:
* KeepAlive packet's implementation (struct, enums, writer, reader, size
  constant) — all dead since HeartbeatService landed 2026-05-06. Kept
  the RemPacketType.KeepAlive enum value and silent-drop dispatch for
  wire compat with any pre-2026-05-06 build still in the wild (item 30)
* driftDropFramesTotal / driftRepeatFramesTotal fields and accessors —
  Phase-2 splice corrector relics, never incremented since Phase-4
  resampler design landed; backed five always-zero diag log columns
  (items 34 + 35)
* DriftAccumulator (always returned 0) — same shape, removed alongside
  the driftAcc= column (item 35)
* TakeMaxFanOutCacheBytes / Ms + fanCacheMs column — FanOutSource was
  retired in May (item 36)

Project documentation:
* RemSoundefficiency.md added as the canonical record of the efficiency
  analysis, every item's status, and the measured wins from this round
* Honest item-by-item review of the original 50-item list — several
  items I had sized optimistically in the original analysis turned out
  to be already-done (item 20), already-optimal (item 22), or below
  the meter floor (items 9, 15, 17, 25). Recorded so future passes
  don't re-investigate.

Wire format and audio pipeline unchanged from v1.5 onward — v1.5 through
v2.2 peers interoperate.
This commit is contained in:
Ednunp
2026-05-23 15:56:03 +01:00
parent 79b28b6c02
commit 6d6d6897e4
22 changed files with 847 additions and 232 deletions
+20
View File
@@ -110,11 +110,17 @@ public sealed class AudioSender : IDisposable
// Both are reset on each Take() so the SNAP gets per-second peaks.
private long maxEmitTicks;
private long maxSendCallTicks;
// Cumulative counters mirroring the max ones above. The diag log samples these once
// a second to report "milliseconds-of-CPU-per-second" for the send-side audio thread —
// i.e. per-thread CPU usage from item 2 of RemSoundefficiency.md. Drain-on-read so the
// value reads naturally as "this last second's load". 2026-05-22.
private long cumulativeEmitTicks;
internal void RecordEmitTicks(long ticks)
{
long current;
do { current = Volatile.Read(ref maxEmitTicks); }
while (ticks > current && Interlocked.CompareExchange(ref maxEmitTicks, ticks, current) != current);
Interlocked.Add(ref cumulativeEmitTicks, ticks);
}
internal void RecordSendCallTicks(long ticks)
{
@@ -124,6 +130,20 @@ public sealed class AudioSender : IDisposable
}
public int TakeMaxEmitMs() => (int)(Interlocked.Exchange(ref maxEmitTicks, 0) * 1000 / Stopwatch.Frequency);
public int TakeMaxSendCallMs() => (int)(Interlocked.Exchange(ref maxSendCallTicks, 0) * 1000 / Stopwatch.Frequency);
/// <summary>Cumulative milliseconds the send-side audio thread spent inside
/// <see cref="SenderLane.OnMixedSamples"/> (encode + sendto + per-packet bookkeeping)
/// since the last call. Resets on read. Diag log emits this as sendMs per second
/// — direct measurement of "how busy is the send thread". 2026-05-22.</summary>
public double TakeSendWorkMs() =>
Interlocked.Exchange(ref cumulativeEmitTicks, 0) * 1000.0 / Stopwatch.Frequency;
/// <summary>Cumulative milliseconds the capture-side threads spent doing per-callback
/// work (ASIO buffer copy + mix loop; WASAPI capture body; MixingEngine.MixLoop per
/// tick) since the last call. Resets on read. Diag log emits this as captureMs per
/// second. Sister metric to <see cref="TakeSendWorkMs"/> — the two together split
/// "what is the sender side spending its CPU on". 2026-05-22.</summary>
public double TakeCaptureWorkMs() =>
engine.TakeCumulativeCaptureTicks() * 1000.0 / Stopwatch.Frequency;
// Pre-encode discontinuity probe — per-lane (each <see cref="SenderLane"/> owns its own).
// The aggregate accessor returns the max across both lanes since the last read; per-lane