Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
e4261a6
feat(speaker): add SpeakerRegistry for enrolled voices
mudler Sep 30, 2026
85af742
fix(speaker): reject a registry blob with dim 0 and speakers
mudler Sep 30, 2026
e22e5e5
feat(speaker): add SpeakerIdentifier for diarization slots
mudler Sep 30, 2026
f0bd407
fix(speaker): test overlap against closed segments and document the o…
mudler Sep 30, 2026
5837b4a
feat(speaker): carry speaker names through words, utterances and scen…
mudler Sep 30, 2026
b6e42cf
feat(speaker): fold voice-detect.cpp in behind SpeakerEncoder
mudler Sep 30, 2026
b1a9544
feat(speaker): identify speakers inside the scene stream
mudler Sep 30, 2026
07a9984
fix(speaker): check speaker preconditions first and test order-indepe…
mudler Sep 30, 2026
5c7ef6c
feat(capi): speaker identification, ABI v9
mudler Sep 30, 2026
1f639c5
fix(capi): read scene options only within the caller's size and avoid…
mudler Sep 30, 2026
b9bafb4
feat(cli): enroll speakers and name them in scene output
mudler Sep 30, 2026
7260f7e
fix(cli): write the speaker registry atomically and keep the docs con…
mudler Sep 30, 2026
88500cb
fix(speaker): name a slot while its segment is still open
mudler Sep 30, 2026
c0d5896
fix(scene): keep one JSON shape for a stream with a speaker part
mudler Sep 30, 2026
ca4b586
fix(speaker): save registries atomically and do not replace an unread…
mudler Sep 30, 2026
9b2b5a2
docs(speaker): give a starting threshold per encoder and gate the voi…
mudler Sep 30, 2026
29053a9
test(speaker): word a comment without an arrow
mudler Sep 30, 2026
4c742ac
build: point the voice-detect submodule at localai-org
mudler Sep 30, 2026
4cff7d5
build: pin voice-detect.cpp to its merge commit
mudler Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 11 additions & 4 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,11 +32,12 @@ jobs:
# job once a models bundle is published (Phase 4).
run: ctest --test-dir build --output-on-failure -LE model

- name: build without ced (PARAKEET_WITH_CED=OFF)
# PARAKEET_WITH_CED is on by default, so this is the gate that catches
# anything that quietly starts depending on ced.cpp being present.
- name: build without ced and voice-detect (PARAKEET_WITH_CED=OFF, PARAKEET_WITH_VOICEDETECT=OFF)
# PARAKEET_WITH_CED and PARAKEET_WITH_VOICEDETECT are on by default, so
# this is the gate that catches anything that quietly starts depending
# on ced.cpp or voice-detect.cpp being present.
run: |
cmake -B build-noced -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=OFF -DPARAKEET_WITH_CED=OFF
cmake -B build-noced -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=OFF -DPARAKEET_WITH_CED=OFF -DPARAKEET_WITH_VOICEDETECT=OFF
cmake --build build-noced -j
ctest --test-dir build-noced --output-on-failure -LE model
# scene --sound must fail cleanly (exit 2) with a clear message.
Expand All @@ -46,6 +47,12 @@ jobs:
echo "$out"
test "$rc" -eq 2
grep -q "built without sound tagging" <<< "$out"
# enroll must fail the same way without voice-detect.cpp.
rc=0
out=$(build-noced/examples/cli/parakeet-cli enroll --model x.gguf --name a --input x.wav --registry /tmp/r.bin 2>&1) || rc=$?
echo "$out"
test "$rc" -eq 2
grep -q "built without speaker identification" <<< "$out"

# -------------------------------------------------------------------------
# server-e2e: drive the real parakeet-server over HTTP.
Expand Down
3 changes: 3 additions & 0 deletions .gitmodules
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,6 @@
[submodule "third_party/ced.cpp"]
path = third_party/ced.cpp
url = https://github.com/localai-org/ced.cpp
[submodule "third_party/voice-detect.cpp"]
path = third_party/voice-detect.cpp
url = https://github.com/localai-org/voice-detect.cpp
37 changes: 35 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,9 @@ src/ libparakeet implementation
sound_stream.hpp/cpp, pk::SoundStream: sliding-window sound-event detection over live PCM
scene_stream.hpp/cpp, pk::SceneStream: combined ASR + diarization + sound-event stream
scene_render.hpp/cpp, pk::SceneRenderer + format_span/is_speech_label: `parakeet-cli scene` text rendering
speaker_registry.hpp/cpp, pk::SpeakerRegistry: enrolled voices (centroid per name), match, binary save/load
speaker_identifier.hpp/cpp, pk::SpeakerIdentifier: names diarization slots from their clean audio; identify_offline
speaker_encoder.hpp/cpp, pk::SpeakerEncoder: the only code that talks to voice-detect.cpp (voicedetect_capi.h)
examples/cli/ parakeet-cli binary
subcommands: info, transcribe (+ --stream), quantize, scene (ASR + diar + sound, one time-ordered feed)
sound-window-eval: measures CED short-window accuracy vs whole-clip top-1
Expand Down Expand Up @@ -123,6 +126,11 @@ tests/ ctest targets
test_sound_capi.cpp , sound_stream_* C-API (PARAKEET_TEST_CED_GGUF)
test_scene_stream.cpp , pk::SceneStream / scene_stream_* C-API, all three models together (PARAKEET_TEST_GGUF + PARAKEET_TEST_DIAR_GGUF + PARAKEET_TEST_CED_GGUF)
test_scene_render.cpp , SceneRenderer / format_span / is_speech_label (model-independent)
test_speaker_registry.cpp, SpeakerRegistry enroll/match/serialize (model-independent)
test_speaker_identifier.cpp, SpeakerIdentifier with a fake embedder (model-independent)
test_speaker_encoder.cpp, SpeakerEncoder vs voice-detect reference embedding (PARAKEET_TEST_VD_GGUF + PARAKEET_TEST_VD_REF_WAV + PARAKEET_TEST_VD_REF_JSON)
test_speaker_identify.cpp, scene stream names both fixture voices (PARAKEET_TEST_DIAR_GGUF + PARAKEET_TEST_VD_GGUF; PARAKEET_TEST_GGUF adds the named-utterance block)
test_capi_speaker.cpp , speaker C-API v9 (PARAKEET_TEST_DIAR_GGUF + PARAKEET_TEST_VD_GGUF; PARAKEET_TEST_GGUF optional; PARAKEET_TEST_VD_GGUF_ALT for the size-mismatch check)
python/check_convert.py , converter round-trip (model-dependent)
python/check_baseline.py, baseline dumper (model-dependent)
fixtures/clip.wav , 2 s 16 kHz mono WAV for stage parity tests
Expand All @@ -134,6 +142,9 @@ third_party/ vendored deps
built as a static `ced` target linked into libparakeet, not a separate
process; dr_wav is shared via CED_EXTERNAL_DR_WAV so there is one
DR_WAV_IMPLEMENTATION in the whole build
voice-detect.cpp/, submodule, speaker encoders (PARAKEET_WITH_VOICEDETECT, on by default);
static `voicedetect` target linked into libparakeet, dr_wav shared via
VOICEDETECT_EXTERNAL_DR_WAV
dr_wav.h , vendored single header
models/ output dir for converted GGUFs (gitignored;
MANIFEST.md tracks the expected published set)
Expand All @@ -142,6 +153,8 @@ docs/
quantization.md , quantization allowlist, policy, measured size + WER per type
parity.md , full model coverage matrix + per-stage tensor parity
diarization.md , speaker diarization + speaker-attributed ASR: parity, C-API, speed
sound.md , sound-event detection (CED) and the combined scene stream
speaker.md , speaker identification: enroll, scene naming, C-API v9, measured numbers
.github/workflows/
ci.yml , build job (per-push) + closed-loop job (pull_request + dispatch)
```
Expand All @@ -164,6 +177,7 @@ cmake -B build -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=ON && cmake --build build
| `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule |
| `PARAKEET_GGML_HIPBLAS` | OFF | Forward GGML_HIPBLAS to the submodule |
| `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp |
| `PARAKEET_WITH_VOICEDETECT` | ON | Speaker identification through voice-detect.cpp |

Use `-DGGML_NATIVE=OFF` when building for CI or portable binaries.

Expand Down Expand Up @@ -256,6 +270,8 @@ parakeet-cli info <model.gguf>
parakeet-cli transcribe --model <model.gguf> --input <audio.wav> [--decoder ctc|tdt] [--stream] [--timestamps] [--json]
parakeet-cli quantize <in.gguf> <out.gguf> <type>
parakeet-cli scene [--model <asr.gguf>] [--diar <diar.gguf>] [--sound <ced.gguf>] --input <audio.wav> [--latency model|low|very_low|ultra_low] [--chunk-ms N] [--show-speech] [--json]
parakeet-cli scene ... --speakers <speaker.gguf> --registry <file> [--speaker-threshold F] # names diarized speakers
parakeet-cli enroll --model <speaker.gguf> --name <name> --input <wav> [--input <wav> ...] --registry <file>
```

`--timestamps` prints one `<start>-<end> <word> (<conf>)` line per word (also
Expand Down Expand Up @@ -310,7 +326,22 @@ parakeet_capi_sound_stream_begin / _feed / _active / _drain_scores_json / _free
parakeet_capi_free_sound_segments
parakeet_capi_num_classes
parakeet_capi_class_label
parakeet_capi_model_kind # which kind of ctx (NONE/ASR/DIARIZATION/SOUND)
parakeet_capi_model_kind # which kind of ctx (NONE/ASR/DIARIZATION/SOUND/SPEAKER)
```

Speaker identification (ABI v9, additive; not used by LocalAI yet). A
voice-detect.cpp speaker GGUF loads into a fourth `parakeet_ctx`
kind (`PARAKEET_MODEL_KIND_SPEAKER`, 4) through the same `parakeet_capi_load`;
see `docs/speaker.md`:

```
parakeet_capi_speaker_dim
parakeet_capi_speaker_registry_new / _free / _size / _last_error
parakeet_capi_speaker_enroll
parakeet_capi_speaker_registry_save / _load
parakeet_capi_speaker_identify_pcm_json
parakeet_capi_scene_stream_begin_speaker
parakeet_capi_transcribe_and_diarize_named_json
```

Combined scene stream (ABI v8, additive; not used by LocalAI yet). One stream
Expand Down Expand Up @@ -454,7 +485,9 @@ See `docs/conversion.md` for the authoritative schema. Quick summary:

## ggml submodule

Pinned at v0.13.0 in `third_party/ggml`. No local patches. To bump:
Pinned at v0.13.0 in `third_party/ggml`. CMake applies the patches in
`third_party/ggml-patches` in-tree at configure time (`scripts/apply_ggml_patches.sh`),
so the submodule shows as modified. To bump:
1. Update the submodule SHA.
2. Run `ctest --test-dir build --output-on-failure`.
3. Fix any API breakage in `src/model_loader.cpp`.
Expand Down
22 changes: 22 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ option(PARAKEET_GGML_METAL "Forward GGML_METAL" OFF)
option(PARAKEET_GGML_VULKAN "Forward GGML_VULKAN" OFF)
option(PARAKEET_GGML_HIP "Forward GGML_HIP (ROCm)" OFF)
option(PARAKEET_WITH_CED "Sound-event detection through ced.cpp" ON)
option(PARAKEET_WITH_VOICEDETECT "Speaker identification through voice-detect.cpp" ON)

set(GGML_CUDA ${PARAKEET_GGML_CUDA} CACHE BOOL "" FORCE)
set(GGML_METAL ${PARAKEET_GGML_METAL} CACHE BOOL "" FORCE)
Expand Down Expand Up @@ -91,6 +92,19 @@ if(PARAKEET_WITH_CED)
target_link_libraries(ced PRIVATE dr_wav_impl)
endif()

if(PARAKEET_WITH_VOICEDETECT)
if(NOT EXISTS "${CMAKE_CURRENT_SOURCE_DIR}/third_party/voice-detect.cpp/CMakeLists.txt")
message(FATAL_ERROR "third_party/voice-detect.cpp is missing: run `git submodule update --init third_party/voice-detect.cpp`, or configure with -DPARAKEET_WITH_VOICEDETECT=OFF")
endif()
set(VOICEDETECT_BUILD_CLI OFF CACHE BOOL "" FORCE)
set(VOICEDETECT_BUILD_TESTS OFF CACHE BOOL "" FORCE)
set(VOICEDETECT_SHARED OFF CACHE BOOL "" FORCE)
set(VOICEDETECT_EXTERNAL_DR_WAV ON CACHE BOOL "" FORCE) # dr_wav_impl provides it
add_subdirectory(third_party/voice-detect.cpp EXCLUDE_FROM_ALL)
set_target_properties(voicedetect PROPERTIES POSITION_INDEPENDENT_CODE ON)
target_link_libraries(voicedetect PRIVATE dr_wav_impl)
endif()

set(PARAKEET_SRC
src/parakeet.cpp
src/model.cpp
Expand Down Expand Up @@ -131,6 +145,9 @@ set(PARAKEET_SRC
src/diarization_streaming.cpp
src/ced_tagger.cpp
src/sound_stream.cpp
src/speaker_registry.cpp
src/speaker_identifier.cpp
src/speaker_encoder.cpp
src/scene_render.cpp)

if(PARAKEET_SHARED)
Expand All @@ -154,6 +171,11 @@ if(PARAKEET_WITH_CED)
target_compile_definitions(parakeet PRIVATE PARAKEET_WITH_CED=1)
endif()

if(PARAKEET_WITH_VOICEDETECT)
target_link_libraries(parakeet PRIVATE voicedetect)
target_compile_definitions(parakeet PRIVATE PARAKEET_WITH_VOICEDETECT=1)
endif()

if(PARAKEET_BUILD_CLI)
add_subdirectory(examples/cli)
endif()
Expand Down
10 changes: 10 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,7 @@ cmake --build build-shared -j
| `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule |
| `PARAKEET_GGML_HIP` | OFF | Forward GGML_HIP (ROCm) to the submodule |
| `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp |
| `PARAKEET_WITH_VOICEDETECT` | ON | Speaker identification through voice-detect.cpp |

To build for a GPU backend, forward its flag, e.g. Apple Metal:

Expand Down Expand Up @@ -337,6 +338,15 @@ parakeet-cli scene --model asr.gguf --diar diar.gguf --sound ced-base-q8_0.gguf
See [`docs/sound.md`](docs/sound.md) for the CED GGUFs, the sound and scene
stream C-API (ABI v8), and the `--sound-model` server option.

### Naming speakers

With a voice-detect.cpp speaker encoder (`PARAKEET_WITH_VOICEDETECT`, on by
default) the scene stream can say who is talking instead of `Speaker 0`.
Enroll each person from a short clip with `parakeet-cli enroll`, then pass
`--speakers <speaker.gguf> --registry <file>` to `scene`. Only one two-voice
fixture has been measured so far. See [`docs/speaker.md`](docs/speaker.md) for
the models, the commands, the C-API (ABI v9) and what is still untested.

---

## C-API (`libparakeet.so`)
Expand Down
Loading
Loading