feat(diarization): add the model card's low-latency streaming modes - #73
Merged
Merged
Conversation
Streaming diarization only ran the checkpoint's own configuration:
21.12 s chunks with no look-ahead, too slow for live speaker labels.
The Nemotron-3-Diarization model card documents three low-latency
configurations for the same checkpoint (1.04, 0.64 and 0.32 s input
latency), which need a FIFO and per-chunk look-ahead.
StreamingDiarization now takes a DiarStreamConfig (speaker cache,
FIFO, chunk, left/right context, update period) with the presets as
DiarLatency. It buffers mel frames itself and runs a chunk once its
look-ahead has arrived, encoding the chunk with its context frames as
NeMo's streaming_feat_loader does; only the chunk is output and
enters the FIFO and speaker cache.
Every mode matches NeMo in the same mode: all four give NeMo's
segments exactly on the 23.6 s fixture, and the 1.04 s mode, which
fills the FIFO and compresses the cache, gives all 26 segments on a
68.5 s clip. On CPU the 1.04 s mode runs 6.9x faster than real time.
C-API, additive within ABI 7:
- diarize_stream_begin_latency and sas_stream_begin_latency pick a
PARAKEET_DIAR_LATENCY_* mode.
- diarize_stream_active returns the segments still open ("who is
speaking now"), diarize_stream_time the diarized time.
- diarize_stream_chunk_samples now reports the input latency.
Streaming SAS transcribed every time diarization advanced. At 1.04 s
that meant 1 to 2 s ASR windows, which drop words, and the commit
point skipped audio the ASR had missed. It now waits for 4 s of
uncommitted audio and resumes right after the last committed word.
gen_diar_baseline.py captures each mode (--modes) from one diarize()
run per mode, and the diarize example takes --stream <mode>.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Adds the latency modes, their parity with NeMo and their speed on a 12.3 min clip. Q8_0 was described as segment-exact; on a 31.5 s clip a frame at the 0.5 threshold splits one segment (99.5% frame agreement), so the note now says so and the suggested test tolerance is 0.15. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
mudler
added a commit
to mudler/LocalAI
that referenced
this pull request
Sep 28, 2026
With a diar_model attached, only unary transcription carried speakers. - stream=true: the final segments carry the speaker who said most of each utterance, from the same diarization as a unary request. - Live transcription: a low-latency diarization stream runs next to the ASR stream and gets each audio payload first. Every finalized word carries a speaker; an open speaker segment is assumed to continue through words diarization has not reached. The final result lists one segment per speaker turn, relabelled once the whole session is diarized. Live config param diarize=false opts out. - diar_latency picks the model card's mode (low 1.04 s, default; very_low 0.64 s; ultra_low 0.32 s; model 21.12 s). - Word-level timestamps on unary requests carry their speaker too. The streaming diarization entry points are probed with Dlsym, so live labels need parakeet.cpp with the latency modes (mudler/parakeet.cpp#73); an older libparakeet.so still transcribes live, without speakers. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
mudler
added a commit
to mudler/LocalAI
that referenced
this pull request
Sep 29, 2026
The previous pin (238057c) had speaker diarization but not the streaming latency modes that live speaker labels need (mudler/parakeet.cpp#73). 6dea76a is current master: it includes #73, plus sound-event detection (additive C-API, ABI 8) and a GCC 16 build fix. It adds the third_party/ced.cpp submodule, which the recursive submodule fetch picks up. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds the Nemotron-3-Diarization model card's low-latency streaming modes, so speaker labels can keep up with live audio. Until now streaming diarization only ran the checkpoint's own configuration: 21.12 s chunks, no look-ahead.
PARAKEET_DIAR_LATENCY_MODELPARAKEET_DIAR_LATENCY_LOWPARAKEET_DIAR_LATENCY_VERY_LOWPARAKEET_DIAR_LATENCY_ULTRA_LOWSizes are 80 ms frames.
StreamingDiarizationbuffers mel frames and runs a chunk once its look-ahead arrives. It encodes each chunk with its context frames, as NeMo'sstreaming_feat_loaderdoes, and only the chunk enters the FIFO and speaker cache.diarize_stream_begin_latencyandsas_stream_begin_latencypick a mode.diarize_stream_activereturns who is speaking now.diarize_stream_timereturns how far diarization has got.Parity with NeMo (same mode)
two_speakers.wav, 23.6 sAll matched segments are identical to the 10 ms frame. F16 probability max diff is 0.01 or less. Each mode is checked both whole-clip and live, fed in 100 ms PCM pieces through
StreamingMel.The 68.5 s clip was only baselined in the 1.04 s mode, because NeMo runs these modes very slowly on CPU.
Q8_0 keeps the segments on these clips. On another 31.5 s clip, a frame sitting at the 0.5 threshold splits a segment in Q8_0 (99.5% frame agreement). That's documented, and the Hugging Face README note is corrected.
Speed
12.3 min, 3 speakers, F16, CPU (Ryzen 9 9950X3D), model load included:
Each low-latency chunk costs about 0.1 s of compute.
How to verify
python scripts/gen_diar_baseline.py --model nvidia/Nemotron-3-Diarization \ --audio tests/fixtures/two_speakers.wav --output /tmp/diar_baseline.gguf # all modes, slow in NeMo PARAKEET_TEST_DIAR_GGUF=nemotron-3-diarization-f16.gguf PARAKEET_TEST_BASELINE_DIAR=/tmp/diar_baseline.gguf \ PARAKEET_TEST_GGUF=tdt_ctc-110m-f16.gguf ctest --test-dir build -R "diar|sas|combined" ./build/examples/cli/diarize nemotron-3-diarization-f16.gguf meeting.wav --stream lowThis is the parakeet.cpp side of speaker labels on LocalAI's live transcription; the LocalAI side is in mudler/LocalAI#12323.
🤖 Generated with Claude Code