feat: identify speakers in the scene stream (voice-detect.cpp, C-API v9) - #78
Merged
Merged
Conversation
Enrolled speakers are centroids of L2-normalized embeddings. identify() takes a threshold and a runner-up margin so a voice that is not enrolled comes out unknown instead of as the nearest name. The registry saves and loads as a small binary blob and rejects corrupt or mismatched input. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
deserialize() now validates that dim >= 1 when n > 0 (number of speakers > 0). A crafted blob with dim 0 and speakers was accepted, creating entries with empty sum vectors. A later enroll() would then cause an out-of-bounds heap write when trying to accumulate the new embedding. Empty registries (dim 0, n 0) still serialize and deserialize correctly. The fix adds validation to reject implausible headers and includes tests for: - A hand-built corrupt blob (dim 0, n 1) must throw - An empty registry round-trips and accepts subsequent enrollment - Failed enrollments on non-empty registries leave state unchanged - names() correctly lists remaining speakers after removal Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Collects each slot's clean audio (overlap with other speakers is skipped), embeds it once it has enough and again as it grows, and names the slot from the registry. A different name replaces the current one only after winning twice in a row, and an unknown match never removes a name. identify_offline reuses the same logic for finished recordings. Driven by an embedding callback, so it is tested with a fake encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…pen contract Adds tests for overlap with segments closed in the same call and in an earlier call, which exercise the history scan. Documents that the caller must list every started, unclosed segment in `open`, and what identify_offline does to max_voice_sec and ring_sec. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…e JSON Words and utterances get a name and a match score, SceneUpdate gets the current name of every slot, and the scene JSON and the text renderer print them. With no speaker model the JSON is unchanged, byte for byte. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
voice-detect.cpp is a submodule built as a static target that shares our ggml and dr_wav, like ced.cpp. pk::SpeakerEncoder is the only code that touches it, through voicedetect_capi.h, and PARAKEET_WITH_VOICEDETECT=OFF builds without it. The test checks that two clips of one voice score above two clips of different voices on two_speakers.wav, and that the folded encoder matches a standalone voice-detect build. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
SceneStream runs an optional SpeakerIdentifier after diarization, so words, utterances and the per-slot names in each update carry the enrolled name. It needs diarization and a registry and says so when they are missing. Names apply to words committed after the slot is identified; earlier words keep the label they were emitted with. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ndent naming The speaker checks now run before the generic "needs at least one model" check, so a speaker-only config gets the specific message. The test asserts each message, enrolls the voices in reverse arrival order so slot i cannot map to registry entry i, and checks that a voice missing from the registry stays unnamed. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
A voice-detect GGUF loads into a fourth context kind. A registry handle enrolls, saves and loads voices; the scene stream takes a speaker context and registry through a new begin function; speaker-attributed ASR has a named variant. Existing signatures are unchanged, and the scene options grow only at the end, read according to the caller's size. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
… a stream leak The speaker option fields were passed as arguments to a helper that checked the caller's size, so they were read before the check and a caller built against the v8 header read past its struct. Each field is now read by offset only after the size covers it. The scene wrapper is also allocated after the SceneStream is built, so a throwing constructor no longer leaks it. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
parakeet-cli enroll builds a registry from labeled clips, and scene --speakers --registry names diarized speakers in the transcript and the JSON. docs/speaker.md covers the models, the C-API, the timing rule and what has and has not been measured. The speaker test also streams ASR and checks that utterances carry the right name when PARAKEET_TEST_GGUF is set. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…sistent enroll now writes <registry>.tmp, checks every write and the close, then renames over the registry, so a failed write leaves the old file intact. docs/speaker.md explains why the sample scores differ from the measured ones and drops first-person wording. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
SpeakerIdentifier only took audio from closed segments, so a speaker stayed "Speaker N" for their whole first turn and got a name only after the first pause. Open segments are now consumed as they grow. Each slot keeps a cursor, and a call adds the clean audio from the cursor to the segment end, for closed and open segments alike, so audio taken while a segment was open is not added again when it closes. A short clean tail at the growing end is held back and joined to the audio that follows it, so small feeds do not lose it to the 0.2 s minimum piece. On tests/fixtures/two_speakers.wav with the low latency preset, slot 0 is now named at 3.4 s of stream time instead of 6.2 s. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The scene JSON printed "names" and the per-item "name"/"name_score"
fields only once some slot had been seen, so the first documents of a
stream with a speaker part had a different shape than the later ones.
SceneUpdate now has a `named` flag, set on every update when the stream
has a speaker part. The writer then prints "names" (possibly {}) and the
name fields (empty name, score 0.0000) from the first document on.
Without a speaker part the output is byte for byte what it was, and a
golden test made with the old serializer guards that.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…able one A new helper, pk::write_file_atomic, writes <path>.tmp in the same directory, checks every write and the close, and moves the tmp file over the target: MoveFileExA with MOVEFILE_REPLACE_EXISTING on Windows, where rename() fails when the target exists, and rename() elsewhere. On any failure the tmp file is removed and the old file is left as it was. parakeet-cli enroll and parakeet_capi_speaker_registry_save both use it. The C-API save used to truncate the existing file before writing, and enroll could not add to an existing registry on Windows. enroll also treated any failure to open the registry as "no registry yet" and replaced it with a new one. Now only a missing file starts a new registry. Any other failure exits 1 with the path and the reason, and a directory given as the registry says so instead of reporting a truncated file. Also: test_capi_speaker uses std::filesystem for its temp file so it builds on Windows, and the named SAS entry point keeps the message of a std::exception instead of "unknown error". Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ce-detect OFF build docs/speaker.md gets a starting threshold per encoder, from the fixture numbers only: 0.5 for WeSpeaker ResNet34 and CAM++, 0.7 for ECAPA (an impostor reached 0.566 there), ERes2Net not measured. It recommends WeSpeaker to start with and says the threshold needs checking on your own audio. The code default stays 0.5; the scene help text points ECAPA users to 0.7. The Enroll section now says what happens to one voice under two names and to near-duplicate names, the clip-to-clip table says which clips it uses, and the timing section says a slot is named while it is still talking. CI: the CED-OFF job now also builds with PARAKEET_WITH_VOICEDETECT=OFF and checks that `parakeet-cli enroll` exits 2 with "built without speaker identification". SpeakerRegistry refuses a NaN or Inf embedding on enroll and treats a NaN or Inf probe as unknown, like an all-zero one. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
voice-detect.cpp now lives under localai-org, like ced.cpp. The old mudler URL redirects, but the canonical one is what a fresh clone and the docs should use. The pinned commit is unchanged and is on the pushed feat/embedding branch. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
localai-org/voice-detect.cpp#1 is merged. Move the submodule from the feature branch commit b44c586 to the merge commit b74a896 on master. The tree is identical, so nothing else changes. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Names the speakers the diarizer finds. Sortformer gives "speaker 0" and "speaker 1" in arrival order; this matches each slot's voice against enrolled speakers, so a transcript reads
Ada: ...andBen: .... It runs in the same one-pass scene stream as ASR, diarization and sound events, and it also names speakers in offline speaker-attributed ASR. Nothing changes for callers that do not pass a speaker model.Docs:
docs/speaker.md.How it works
third_party/voice-detect.cpp) behindPARAKEET_WITH_VOICEDETECT(default ON), the same way ced.cpp is behindPARAKEET_WITH_CED.pk::SpeakerEncoderis the only file that touches it, throughvoicedetect_capi.h. A voice-detect GGUF loads throughparakeet_capi_loadas a fourth context kind. The voice-detect side is build: allow embedding voice-detect.cpp in another ggml project localai-org/voice-detect.cpp#1 (embeddable build,VOICEDETECT_EXTERNAL_DR_WAV,voicedetect_capi_embedding_dim). This PR pins that commit.pk::SpeakerRegistry: enrolled speakers are the L2-normalized mean of their clip embeddings.identifytakes a threshold and a runner-up margin, so a voice that is not enrolled comes out unknown instead of as the nearest name. It saves and loads as a small binary blob (version 1, not a stable interchange format yet).pk::SpeakerIdentifier: collects each slot's clean audio (overlap with another speaker is skipped, not resolved), embeds it once it has 2 s and again as it grows, and names the slot. A different name replaces the current one only after winning twice in a row, and an unknown match never removes a name. It takes audio from open segments as well as closed ones, so a speaker is named during a long turn, not only after their first pause.SceneStreamruns the identifier after diarization. Words and utterances carrynameandname_score, every update carries the current name of each slot innames. Words committed before a slot is identified keep their earlier label. Without a speaker part the JSON is byte for byte what it was.parakeet_capi_scene_stream_begin_speaker,parakeet_capi_transcribe_and_diarize_named_json. Existing symbols and signatures are unchanged, and the scene options struct grows only at the end and is read only within the caller'ssize. Not used by LocalAI yet.What was measured, and what was not
Everything below is one fixture,
tests/fixtures/two_speakers.wav: two read-speech voices, alternating.The default
accept_threshold(0.5) is a starting point and is encoder specific: ECAPA admitted an impostor near 0.57 with it. WeSpeaker ResNet34 is the encoder to start with. The genuine scores in the tests (0.92 to 0.98) come from enrolling with clips cut from the same recording, so they are optimistic; expect lower ones across sessions and check the threshold on your own audio.Not measured yet: a third voice, noisy or overlapping speech, telephone audio, cross-session enrollment, ERes2Net.
scene --soundwas not run here (see below).Tests
test_speaker_registry,test_speaker_identifier(fake encoder; overlap, refresh, hysteresis, unknown voice, open-segment consumption),test_write_atomic, plus extensions totest_sas_mergeandtest_scene_render.test_speaker_encoder(also checks the folded encoder matches a standalone voice-detect build, cosine 1.000000),test_speaker_identify(real diarization plus the speaker encoder, enrollment in reverse order so arrival-order naming cannot pass, an unenrolled voice must stay unknown, ASR plus speaker utterance naming),test_capi_speaker(also under AddressSanitizer for the versioned scene options).test_diarization_accuracystill gives 100% frame agreement and the closed-loop TDT transcript onspeech.wavis unchanged. I built with-DPARAKEET_WITH_CED=OFFbecausethird_party/ced.cppis not checked out on my machine, sotest_scene_streamand the sound tests were skipped and a build with CED and voice-detect both ON has not been run outside CI. Please look at thebuildjob for that one.PARAKEET_WITH_VOICEDETECT=OFFand checks thatenrollexits 2 with "built without speaker identification".Notes for review
submodules: recursivealso fetches voice-detect's ownthird_party/ggml, which is never built (our ggml is reused). It only costs a clone.libparakeet.sonow also exports the static dependency'svoicedetect_capi_*symbols, the same pattern asced_capi_*. Symbol visibility for the static deps would be a fine follow-up.VOICEDETECT_DEVICE,VOICEDETECT_THREADS), likeCED_DEVICE. An embedded build does not apply voice-detect's CUDA / cuDNN patch, which does not matter on CPU.AGENTS.md: the ggml sentence said there are no local patches; the CMake applies the in-tree ones at configure time, so it now says that.fsyncbefore the registry rename.🤖 Generated with Claude Code