Skip to content

feat: identify speakers in the scene stream (voice-detect.cpp, C-API v9) - #78

Merged
mudler merged 19 commits into
masterfrom
feat/voice-identification
Sep 30, 2026
Merged

mudler merged 19 commits into
masterfrom
feat/voice-identification

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

Names the speakers the diarizer finds. Sortformer gives "speaker 0" and "speaker 1" in arrival order; this matches each slot's voice against enrolled speakers, so a transcript reads Ada: ... and Ben: .... It runs in the same one-pass scene stream as ASR, diarization and sound events, and it also names speakers in offline speaker-attributed ASR. Nothing changes for callers that do not pass a speaker model.

parakeet-cli enroll --model wespeaker_resnet34_f32.gguf --name Ada --input ada.wav --registry voices.bin
parakeet-cli enroll --model wespeaker_resnet34_f32.gguf --name Ben --input ben.wav --registry voices.bin
parakeet-cli scene --model asr.gguf --diar diar.gguf --speakers wespeaker_resnet34_f32.gguf --registry voices.bin --input meeting.wav

Docs: docs/speaker.md.

How it works

  • voice-detect.cpp is folded in as a static submodule (third_party/voice-detect.cpp) behind PARAKEET_WITH_VOICEDETECT (default ON), the same way ced.cpp is behind PARAKEET_WITH_CED. pk::SpeakerEncoder is the only file that touches it, through voicedetect_capi.h. A voice-detect GGUF loads through parakeet_capi_load as a fourth context kind. The voice-detect side is build: allow embedding voice-detect.cpp in another ggml project localai-org/voice-detect.cpp#1 (embeddable build, VOICEDETECT_EXTERNAL_DR_WAV, voicedetect_capi_embedding_dim). This PR pins that commit.
  • pk::SpeakerRegistry: enrolled speakers are the L2-normalized mean of their clip embeddings. identify takes a threshold and a runner-up margin, so a voice that is not enrolled comes out unknown instead of as the nearest name. It saves and loads as a small binary blob (version 1, not a stable interchange format yet).
  • pk::SpeakerIdentifier: collects each slot's clean audio (overlap with another speaker is skipped, not resolved), embeds it once it has 2 s and again as it grows, and names the slot. A different name replaces the current one only after winning twice in a row, and an unknown match never removes a name. It takes audio from open segments as well as closed ones, so a speaker is named during a long turn, not only after their first pause.
  • SceneStream runs the identifier after diarization. Words and utterances carry name and name_score, every update carries the current name of each slot in names. Words committed before a slot is identified keep their earlier label. Without a speaker part the JSON is byte for byte what it was.
  • C-API v9, additive only: a speaker context kind, a registry handle (new, free, enroll, save, load, identify), parakeet_capi_scene_stream_begin_speaker, parakeet_capi_transcribe_and_diarize_named_json. Existing symbols and signatures are unchanged, and the scene options struct grows only at the end and is read only within the caller's size. Not used by LocalAI yet.

What was measured, and what was not

Everything below is one fixture, tests/fixtures/two_speakers.wav: two read-speech voices, alternating.

encoder same voice / different voice (clip to clip cosine) impostor against a registry with the other voice only starting threshold
WeSpeaker ResNet34 0.869 / -0.012 -0.019 0.5
CAM++ 0.894 / 0.383 0.374 0.5
ECAPA 0.934 / 0.558 0.546 to 0.566 0.7

The default accept_threshold (0.5) is a starting point and is encoder specific: ECAPA admitted an impostor near 0.57 with it. WeSpeaker ResNet34 is the encoder to start with. The genuine scores in the tests (0.92 to 0.98) come from enrolling with clips cut from the same recording, so they are optimistic; expect lower ones across sessions and check the threshold on your own audio.

Not measured yet: a third voice, noisy or overlapping speech, telephone audio, cross-session enrollment, ERes2Net. scene --sound was not run here (see below).

Tests

  • Model independent: test_speaker_registry, test_speaker_identifier (fake encoder; overlap, refresh, hysteresis, unknown voice, open-segment consumption), test_write_atomic, plus extensions to test_sas_merge and test_scene_render.
  • Model dependent (skip with 77 when the env vars are unset): test_speaker_encoder (also checks the folded encoder matches a standalone voice-detect build, cosine 1.000000), test_speaker_identify (real diarization plus the speaker encoder, enrollment in reverse order so arrival-order naming cannot pass, an unenrolled voice must stay unknown, ASR plus speaker utterance naming), test_capi_speaker (also under AddressSanitizer for the versioned scene options).
  • Run locally with WeSpeaker, CAM++ and ECAPA: all pass, and test_diarization_accuracy still gives 100% frame agreement and the closed-loop TDT transcript on speech.wav is unchanged. I built with -DPARAKEET_WITH_CED=OFF because third_party/ced.cpp is not checked out on my machine, so test_scene_stream and the sound tests were skipped and a build with CED and voice-detect both ON has not been run outside CI. Please look at the build job for that one.
  • CI: the CED-off gate now also builds with PARAKEET_WITH_VOICEDETECT=OFF and checks that enroll exits 2 with "built without speaker identification".

Notes for review

  • Adding a submodule means submodules: recursive also fetches voice-detect's own third_party/ggml, which is never built (our ggml is reused). It only costs a clone.
  • libparakeet.so now also exports the static dependency's voicedetect_capi_* symbols, the same pattern as ced_capi_*. Symbol visibility for the static deps would be a fine follow-up.
  • voice-detect keeps its own process-global backend and device selection (VOICEDETECT_DEVICE, VOICEDETECT_THREADS), like CED_DEVICE. An embedded build does not apply voice-detect's CUDA / cuDNN patch, which does not matter on CPU.
  • AGENTS.md: the ggml sentence said there are no local patches; the CMake applies the in-tree ones at configure time, so it now says that.
  • Possible follow-ups: per-encoder default threshold from the GGUF architecture key, a threshold sweep on a second fixture, fsync before the registry rename.

🤖 Generated with Claude Code

Enrolled speakers are centroids of L2-normalized embeddings. identify()
takes a threshold and a runner-up margin so a voice that is not enrolled
comes out unknown instead of as the nearest name. The registry saves and
loads as a small binary blob and rejects corrupt or mismatched input.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
deserialize() now validates that dim >= 1 when n > 0 (number of speakers > 0).
A crafted blob with dim 0 and speakers was accepted, creating entries with
empty sum vectors. A later enroll() would then cause an out-of-bounds heap
write when trying to accumulate the new embedding. Empty registries (dim 0,
n 0) still serialize and deserialize correctly.

The fix adds validation to reject implausible headers and includes tests for:
- A hand-built corrupt blob (dim 0, n 1) must throw
- An empty registry round-trips and accepts subsequent enrollment
- Failed enrollments on non-empty registries leave state unchanged
- names() correctly lists remaining speakers after removal

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Collects each slot's clean audio (overlap with other speakers is
skipped), embeds it once it has enough and again as it grows, and names
the slot from the registry. A different name replaces the current one
only after winning twice in a row, and an unknown match never removes a
name. identify_offline reuses the same logic for finished recordings.
Driven by an embedding callback, so it is tested with a fake encoder.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…pen contract

Adds tests for overlap with segments closed in the same call and in an
earlier call, which exercise the history scan. Documents that the caller
must list every started, unclosed segment in `open`, and what
identify_offline does to max_voice_sec and ring_sec.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…e JSON

Words and utterances get a name and a match score, SceneUpdate gets the
current name of every slot, and the scene JSON and the text renderer
print them. With no speaker model the JSON is unchanged, byte for byte.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
voice-detect.cpp is a submodule built as a static target that shares our
ggml and dr_wav, like ced.cpp. pk::SpeakerEncoder is the only code that
touches it, through voicedetect_capi.h, and PARAKEET_WITH_VOICEDETECT=OFF
builds without it. The test checks that two clips of one voice score
above two clips of different voices on two_speakers.wav, and that the
folded encoder matches a standalone voice-detect build.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
SceneStream runs an optional SpeakerIdentifier after diarization, so
words, utterances and the per-slot names in each update carry the
enrolled name. It needs diarization and a registry and says so when
they are missing. Names apply to words committed after the slot is
identified; earlier words keep the label they were emitted with.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ndent naming

The speaker checks now run before the generic "needs at least one
model" check, so a speaker-only config gets the specific message. The
test asserts each message, enrolls the voices in reverse arrival order
so slot i cannot map to registry entry i, and checks that a voice
missing from the registry stays unnamed.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
A voice-detect GGUF loads into a fourth context kind. A registry handle
enrolls, saves and loads voices; the scene stream takes a speaker
context and registry through a new begin function; speaker-attributed
ASR has a named variant. Existing signatures are unchanged, and the
scene options grow only at the end, read according to the caller's size.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
… a stream leak

The speaker option fields were passed as arguments to a helper that
checked the caller's size, so they were read before the check and a
caller built against the v8 header read past its struct. Each field is
now read by offset only after the size covers it. The scene wrapper is
also allocated after the SceneStream is built, so a throwing constructor
no longer leaks it.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
parakeet-cli enroll builds a registry from labeled clips, and scene
--speakers --registry names diarized speakers in the transcript and the
JSON. docs/speaker.md covers the models, the C-API, the timing rule and
what has and has not been measured. The speaker test also streams ASR and
checks that utterances carry the right name when PARAKEET_TEST_GGUF is set.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…sistent

enroll now writes <registry>.tmp, checks every write and the close, then
renames over the registry, so a failed write leaves the old file intact.
docs/speaker.md explains why the sample scores differ from the measured
ones and drops first-person wording.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
SpeakerIdentifier only took audio from closed segments, so a speaker
stayed "Speaker N" for their whole first turn and got a name only after
the first pause.

Open segments are now consumed as they grow. Each slot keeps a cursor,
and a call adds the clean audio from the cursor to the segment end, for
closed and open segments alike, so audio taken while a segment was open
is not added again when it closes. A short clean tail at the growing end
is held back and joined to the audio that follows it, so small feeds do
not lose it to the 0.2 s minimum piece.

On tests/fixtures/two_speakers.wav with the low latency preset, slot 0
is now named at 3.4 s of stream time instead of 6.2 s.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The scene JSON printed "names" and the per-item "name"/"name_score"
fields only once some slot had been seen, so the first documents of a
stream with a speaker part had a different shape than the later ones.

SceneUpdate now has a `named` flag, set on every update when the stream
has a speaker part. The writer then prints "names" (possibly {}) and the
name fields (empty name, score 0.0000) from the first document on.
Without a speaker part the output is byte for byte what it was, and a
golden test made with the old serializer guards that.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…able one

A new helper, pk::write_file_atomic, writes <path>.tmp in the same
directory, checks every write and the close, and moves the tmp file over
the target: MoveFileExA with MOVEFILE_REPLACE_EXISTING on Windows, where
rename() fails when the target exists, and rename() elsewhere. On any
failure the tmp file is removed and the old file is left as it was.

parakeet-cli enroll and parakeet_capi_speaker_registry_save both use it.
The C-API save used to truncate the existing file before writing, and
enroll could not add to an existing registry on Windows.

enroll also treated any failure to open the registry as "no registry
yet" and replaced it with a new one. Now only a missing file starts a new
registry. Any other failure exits 1 with the path and the reason, and a
directory given as the registry says so instead of reporting a truncated
file.

Also: test_capi_speaker uses std::filesystem for its temp file so it
builds on Windows, and the named SAS entry point keeps the message of a
std::exception instead of "unknown error".

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ce-detect OFF build

docs/speaker.md gets a starting threshold per encoder, from the fixture
numbers only: 0.5 for WeSpeaker ResNet34 and CAM++, 0.7 for ECAPA (an
impostor reached 0.566 there), ERes2Net not measured. It recommends
WeSpeaker to start with and says the threshold needs checking on your
own audio. The code default stays 0.5; the scene help text points ECAPA
users to 0.7. The Enroll section now says what happens to one voice
under two names and to near-duplicate names, the clip-to-clip table says
which clips it uses, and the timing section says a slot is named while
it is still talking.

CI: the CED-OFF job now also builds with PARAKEET_WITH_VOICEDETECT=OFF
and checks that `parakeet-cli enroll` exits 2 with "built without
speaker identification".

SpeakerRegistry refuses a NaN or Inf embedding on enroll and treats a
NaN or Inf probe as unknown, like an all-zero one.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
voice-detect.cpp now lives under localai-org, like ced.cpp. The old
mudler URL redirects, but the canonical one is what a fresh clone and the
docs should use. The pinned commit is unchanged and is on the pushed
feat/embedding branch.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
localai-org/voice-detect.cpp#1 is merged. Move the submodule from the
feature branch commit b44c586 to the merge commit b74a896 on master. The
tree is identical, so nothing else changes.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 2c3bb73 into master Sep 30, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants