Skip to content

feat(parakeet-cpp): name speakers from the shared voice registry - #12382

Merged
mudler merged 13 commits into
masterfrom
feat/parakeet-speaker-names
Oct 1, 2026
Merged

mudler merged 13 commits into
masterfrom
feat/parakeet-speaker-names

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

Register a voice once with /v1/voice/register and the parakeet-cpp backend uses it to name the speakers it finds: in /v1/audio/diarization results and in live transcription. Speakers that are not registered keep their SPEAKER_NN label.

It is opt-in per model. Nothing changes for models without speaker_model:, for other diarization backends, or for clients that ignore the new fields.

What changes

  • Voice registry. Registry.List, and Metadata.Model: the speaker encoder that produced an embedding (the voice-detect backend's model name, by default the GGUF file name). /v1/voice/register now stores it. A voice's encoder and a model's speaker_model: file have to match, so an encoder switch cannot turn into wrong names.
  • Known voices. Core reads the model config's speaker_model:, takes the registered voices made by that encoder (plus untagged legacy voices of the same size), and sends them with the diarize request and the live session config (KnownVoice in the proto).
  • Diarization. Segments and the speakers summary gain name, segments also name_score. speaker stays SPEAKER_NN; RTTM is unchanged. A result without names marshals byte for byte as before (golden test).
  • Live transcription. Speaker segments carry the name; the realtime transcription segment event gains speaker_name.
  • parakeet-cpp backend. New options speaker_model:, speaker_threshold: (a distance, 1 minus cosine, the unit of /v1/voice/identify, default 0.5) and speaker_margin: (default 0.05). It builds a per-request parakeet.cpp registry from the known voices and calls the named diarization and scene-stream functions. A voice of the wrong embedding size is skipped with a warning that does not name it.
  • Gallery and docs. Three entries (parakeet-cpp-nemotron-3-diarization-speakers, parakeet-cpp-nemotron-3-diarization-asr-speakers, parakeet-cpp-realtime-scene-speakers), all using the WeSpeaker ResNet34 file the voice-detect-wespeaker-resnet34 entry already installs, so no new files on the Hub. Docs in voice-recognition.md, audio-diarization.md, audio-to-text.md and openai-realtime.md.
  • Pin. PARAKEET_VERSION moves from 623a968 to 8c8cec0, which contains the speaker code (feat: identify speakers in the scene stream (voice-detect.cpp, C-API v9) parakeet.cpp#78) and the C-API v10 functions this needs (Welcome to LocalAI Discussions! #79). Those two commits are the whole range, both additive.

Checked

  • Go unit tests for the touched packages pass, including specs that stub the C library and count every registry and stream free.
  • Real library: specs that load libparakeet.so built from the pinned tree and diarize two_speakers.wav with two known voices passed in reverse order: 5 segments, slot 0 named for voice A and slot 1 for voice B, scores about 0.92 and 0.97. With only voice B registered and speaker_threshold:0.3, slot 0 stays unnamed. speaker_threshold:0.01 names nobody, which shows the float32 threshold really reaches C through purego. The live path names all 5 segments correctly. These specs are gated on env vars and skip in default CI.
  • I did not run a full LocalAI end to end (install the gallery model, register through the API, call the endpoint). Worth doing before merge.

Things to know

  • Accuracy is measured on one fixture (two read-speech voices, in parakeet.cpp's docs). The default threshold is a starting point: ECAPA needs a stricter one. Check it on your own audio.
  • Registry. It is in memory and global, so names disappear on a restart and all users share it. That was already true of the registry. Because of it, anyone allowed to call a model with speaker_model: can learn which registered names match their audio; the docs say to restrict such models with the model allowlist. A possible follow-up is to skip the known voices when the caller lacks the voice-recognition feature.
  • model_name:. A model_name: option on the voice-detect model config overrides the tag and the voices then won't match. Documented.
  • With include_text the names use parakeet.cpp's default threshold and margin, because that C function takes neither. In a live stream a segment that closes before its speaker is identified has no name.
  • The swagger files are edited by hand for the two schema definitions (swag isn't installed here).

🤖 Generated with Claude Code

The voice registry could register, identify and forget but not list, and
it did not remember which speaker encoder produced an embedding. Add
Metadata.Model and Registry.List, answered from the index the store
registry already keeps for Forget. Needed so a backend can be given the
registered voices that match its own speaker encoder.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
… messages

Assisted-by: Claude:claude-haiku-4-5 [Claude Code]
When a diarization model has a speaker_model option, the endpoint sends
the registered voices made by that encoder to the backend. The backend's
name and name_score come back as extra fields next to the normalized
SPEAKER_NN speaker, and the speakers summary carries the first name seen
for each speaker. RTTM output and results without names are unchanged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ker names

Live sessions now send the registered voices that match the model's
speaker_model to the backend, and each speaker segment carries the name
the backend matched. The realtime segment event gains an optional
speaker_name field.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…registries

Adds the speaker bindings (ABI v9 and v10, probed separately), the
speaker_model, speaker_threshold and speaker_margin options, and a
per-request registry builder over the known voices.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Diarize builds a per-request speaker registry from the known voices when a
speaker model is loaded, calls the named C functions, and puts each slot's
registered name and score on the segments. The registry is freed on every
path. A library without ABI 10 reports Unimplemented instead of dropping
the names.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The live scene stream now begins with a known-voice registry when a
speaker model is loaded and the live config carries voices, and each
closed speaker segment takes its slot's current name from the feed's
names map. A segment that closes before its slot is identified has an
empty name. The registry is freed after the stream, on every path.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Add three gallery entries that load the WeSpeaker ResNet34 speaker model
next to the diarization or realtime scene models, and document speaker
names in the voice recognition, diarization, audio to text and realtime
pages.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ng the request

A registered voice with the wrong embedding size, or one the C side
refused, failed the whole diarization request, so one legacy voice broke
the model for every user. Skip such voices with a warning that does not
carry the voice name, and take the plain path when none is left.

Also map an exact 0 speaker threshold or margin to a tiny positive value,
since the C side reads 0 as "use the default", and fix a stale comment
about which contexts Free() walks.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…er; document the privacy limit

The different-encoder warning fired on every request. Log it once per
feature and speaker model, then at debug level. Document that the global
voice registry lets any caller of a speaker_model model learn matching
names, and that skipped wrong-sized voices are logged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…ck speaker naming against the real library

The pin moves from 623a968 to 8c8cec0, which brings in everything merged
in parakeet.cpp since: the voice identification change (C-API v9, #78) and
raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79).

New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL,
_DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two
speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings,
with the voices passed in reversed order. They also check that the float32
threshold reaches C through purego. The shared test loader now registers
the v9/v10 and scene symbols as main.go does.

The rebase onto origin/master had no conflicts.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 2ae6cae into master Oct 1, 2026
78 of 79 checks passed
@mudler
mudler deleted the feat/parakeet-speaker-names branch October 1, 2026 06:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants