Skip to content

feat(diarization): add the model card's low-latency streaming modes - #73

Merged
mudler merged 2 commits into
masterfrom
feat/diar-low-latency
Sep 28, 2026
Merged

mudler merged 2 commits into
masterfrom
feat/diar-low-latency

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

What

Adds the Nemotron-3-Diarization model card's low-latency streaming modes, so speaker labels can keep up with live audio. Until now streaming diarization only ran the checkpoint's own configuration: 21.12 s chunks, no look-ahead.

Mode Input latency Chunk Look-ahead FIFO
PARAKEET_DIAR_LATENCY_MODEL 21.12 s 264 0 0
PARAKEET_DIAR_LATENCY_LOW 1.04 s 9 4 264
PARAKEET_DIAR_LATENCY_VERY_LOW 0.64 s 6 2 264
PARAKEET_DIAR_LATENCY_ULTRA_LOW 0.32 s 3 1 264

Sizes are 80 ms frames.

  • StreamingDiarization buffers mel frames and runs a chunk once its look-ahead arrives. It encodes each chunk with its context frames, as NeMo's streaming_feat_loader does, and only the chunk enters the FIFO and speaker cache.
  • C-API (additive, ABI 7):
    • diarize_stream_begin_latency and sas_stream_begin_latency pick a mode.
    • diarize_stream_active returns who is speaking now.
    • diarize_stream_time returns how far diarization has got.
  • Streaming SAS: it no longer drops words at low latency. It transcribes at least 4 s windows and resumes right after the last committed word.

Parity with NeMo (same mode)

Clip model 1.04 s 0.64 s 0.32 s
two_speakers.wav, 23.6 s 5/5 segments 5/5 6/6 7/7
68.5 s, 2 speakers (FIFO + cache compression) 26/26 26/26 not run not run

All matched segments are identical to the 10 ms frame. F16 probability max diff is 0.01 or less. Each mode is checked both whole-clip and live, fed in 100 ms PCM pieces through StreamingMel.

The 68.5 s clip was only baselined in the 1.04 s mode, because NeMo runs these modes very slowly on CPU.

Q8_0 keeps the segments on these clips. On another 31.5 s clip, a frame sitting at the 0.5 threshold splits a segment in Q8_0 (99.5% frame agreement). That's documented, and the Hugging Face README note is corrected.

Speed

12.3 min, 3 speakers, F16, CPU (Ryzen 9 9950X3D), model load included:

Mode Time x real time Agreement with NeMo's default output
model 5.7 s 130x 100%
1.04 s 107 s 6.9x 99.4%
0.64 s 166 s 4.4x 99.1%
0.32 s 295 s 2.5x 98.3%

Each low-latency chunk costs about 0.1 s of compute.

How to verify

python scripts/gen_diar_baseline.py --model nvidia/Nemotron-3-Diarization \
    --audio tests/fixtures/two_speakers.wav --output /tmp/diar_baseline.gguf   # all modes, slow in NeMo
PARAKEET_TEST_DIAR_GGUF=nemotron-3-diarization-f16.gguf PARAKEET_TEST_BASELINE_DIAR=/tmp/diar_baseline.gguf \
PARAKEET_TEST_GGUF=tdt_ctc-110m-f16.gguf ctest --test-dir build -R "diar|sas|combined"
./build/examples/cli/diarize nemotron-3-diarization-f16.gguf meeting.wav --stream low

This is the parakeet.cpp side of speaker labels on LocalAI's live transcription; the LocalAI side is in mudler/LocalAI#12323.

🤖 Generated with Claude Code

Streaming diarization only ran the checkpoint's own configuration:
21.12 s chunks with no look-ahead, too slow for live speaker labels.
The Nemotron-3-Diarization model card documents three low-latency
configurations for the same checkpoint (1.04, 0.64 and 0.32 s input
latency), which need a FIFO and per-chunk look-ahead.

StreamingDiarization now takes a DiarStreamConfig (speaker cache,
FIFO, chunk, left/right context, update period) with the presets as
DiarLatency. It buffers mel frames itself and runs a chunk once its
look-ahead has arrived, encoding the chunk with its context frames as
NeMo's streaming_feat_loader does; only the chunk is output and
enters the FIFO and speaker cache.

Every mode matches NeMo in the same mode: all four give NeMo's
segments exactly on the 23.6 s fixture, and the 1.04 s mode, which
fills the FIFO and compresses the cache, gives all 26 segments on a
68.5 s clip. On CPU the 1.04 s mode runs 6.9x faster than real time.

C-API, additive within ABI 7:
- diarize_stream_begin_latency and sas_stream_begin_latency pick a
  PARAKEET_DIAR_LATENCY_* mode.
- diarize_stream_active returns the segments still open ("who is
  speaking now"), diarize_stream_time the diarized time.
- diarize_stream_chunk_samples now reports the input latency.

Streaming SAS transcribed every time diarization advanced. At 1.04 s
that meant 1 to 2 s ASR windows, which drop words, and the commit
point skipped audio the ASR had missed. It now waits for 4 s of
uncommitted audio and resumes right after the last committed word.

gen_diar_baseline.py captures each mode (--modes) from one diarize()
run per mode, and the diarize example takes --stream <mode>.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Adds the latency modes, their parity with NeMo and their speed on a
12.3 min clip. Q8_0 was described as segment-exact; on a 31.5 s clip
a frame at the 0.5 threshold splits one segment (99.5% frame
agreement), so the note now says so and the suggested test tolerance
is 0.15.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
mudler added a commit to mudler/LocalAI that referenced this pull request Sep 28, 2026
With a diar_model attached, only unary transcription carried speakers.

- stream=true: the final segments carry the speaker who said most of
  each utterance, from the same diarization as a unary request.
- Live transcription: a low-latency diarization stream runs next to
  the ASR stream and gets each audio payload first. Every finalized
  word carries a speaker; an open speaker segment is assumed to
  continue through words diarization has not reached. The final
  result lists one segment per speaker turn, relabelled once the whole
  session is diarized. Live config param diarize=false opts out.
- diar_latency picks the model card's mode (low 1.04 s, default;
  very_low 0.64 s; ultra_low 0.32 s; model 21.12 s).
- Word-level timestamps on unary requests carry their speaker too.

The streaming diarization entry points are probed with Dlsym, so live
labels need parakeet.cpp with the latency modes
(mudler/parakeet.cpp#73); an older libparakeet.so still transcribes
live, without speakers.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
@mudler
mudler merged commit 3a1e15e into master Sep 28, 2026
7 checks passed
mudler added a commit to mudler/LocalAI that referenced this pull request Sep 29, 2026
The previous pin (238057c) had speaker diarization but not the
streaming latency modes that live speaker labels need
(mudler/parakeet.cpp#73). 6dea76a is current master: it includes #73,
plus sound-event detection (additive C-API, ABI 8) and a GCC 16 build
fix. It adds the third_party/ced.cpp submodule, which the recursive
submodule fetch picks up.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants