Add sound-event detection and a combined scene stream - #75
Merged
Merged
Conversation
Add third_party/ced.cpp (localai-org/ced.cpp, feat/embedding @ e1a3cfa) as a submodule and wire it into the build behind PARAKEET_WITH_CED (default ON). CED GGUFs (general.architecture "ced") now load through parakeet_capi_load into a third parakeet_ctx kind, a "tagger", exactly like ASR and diarization models do today; pk::CedTagger in the new src/ced_tagger.hpp/.cpp is the only parakeet code that talks to ced_capi.h, so the rest of the codebase never needs to know ced.cpp exists. Bump PARAKEET_CAPI_ABI_VERSION to 8 (additive) and add parakeet_capi_num_classes / parakeet_capi_class_label for tagger introspection. ced.cpp is built with CED_EXTERNAL_DR_WAV so it reuses parakeet's own dr_wav implementation instead of linking a second copy; that makes the two static libraries genuinely circular at link time (parakeet needs ced's classify symbols, ced needs parakeet's dr_wav symbols), so both targets declare a link dependency on each other and let CMake repeat the archives in the generated link line. tests/test_ced_parity.cpp checks ced.cpp built against parakeet's own patched ggml reproduces ced.cpp's own PyTorch baseline: max|d| = 1.714e-07 on CPU f32, matching standalone ced.cpp's reported 1.7e-7. It reads the model and baseline paths from PARAKEET_TEST_CED_GGUF / PARAKEET_TEST_CED_BASELINE and skips (77) when either is unset, since ced.cpp's baseline fixtures are gitignored there. Confirmed the PARAKEET_WITH_CED=OFF build compiles clean and still skips the test. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
…s own TU
The reciprocal target_link_libraries(ced PRIVATE parakeet) added to make
CED_EXTERNAL_DR_WAV link only works because CMake allows cycles among
STATIC libraries. It breaks any -DPARAKEET_SHARED=ON configure ("Cyclic
dependencies are allowed only among static libraries"), including the
libparakeet.so build LocalAI dlopens and the release workflow's shared
jobs, since PARAKEET_WITH_CED defaults ON.
Move the single dr_wav implementation out of src/audio_io.cpp (which now
only declares dr_wav's functions, like ced's own audio_io.cpp does) into
a new source file, src/dr_wav_impl.cpp, built as an OBJECT library.
libparakeet and, when PARAKEET_WITH_CED is on, ced both link against
that object library directly instead of against each other, so there is
no longer any cycle: parakeet depends on ced and dr_wav_impl, ced
depends on dr_wav_impl, and dr_wav_impl depends on nothing.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Adds pk::SoundStream on top of the Task 2 SoundScorer: windows end every hop, hysteresis segments open at on_threshold and close at off_threshold with one-hop start/end timing, is_last scores an unscored tail and force-closes everything open, and a per-window score queue (drain_windows) keeps the top-k classes for the C-API. Includes model-independent tests with a fake scorer that mimics CED's mean pooling, exercising single/overlapping sounds, hysteresis, piece-size invariance, min-duration filtering, open segments, validation, and scorer failure. Fixes a buffer-trim defect: the naive keep_from = next_end_ - win_n_ looks one hop ahead of what has actually streamed, so an is_last tail window ending between two hops could read before the buffer start. keep_from now clamps to min(next_end_, samples_in_) - win_n_, and score_window guards its precondition so a future regression throws instead of reading out of bounds. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
While a SoundStream is still open, safe_until() promised no future segment starts earlier than the last window actually scored, but it missed a case: is_last can score a shorter tail window and open a class at that window's newest hop, which starts earlier than the last regularly scored window whenever the stream ends mid-hop. Widen the bound to min(scored_end_, max(0, samples_in_ - hop_n_)), since a tail window is always at least one hop wide, plus any already-open class's start. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add parakeet_capi_sound_stream_begin/feed/active/drain_scores_json/free and parakeet_capi_free_sound_segments, wrapping pk::SoundStream over a loaded CED tagger context. parakeet_sound_opts carries a size field so callers built against an older header still work (only the fields their struct covers are read). Segment output is a malloc'd C array of parakeet_sound_segment (index, borrowed label, start/end/peak); the per-window top-k scores are JSON only, since a variable-length tag list per window does not fit a fixed C struct and LocalAI consumes JSON anyway. Tested end to end against a real ced-base GGUF on ced.cpp's rooster, thunder and guitar demo clips fed in small pieces: each clip's label closes inside its own 6 s slot, window scores in the drain match a direct scorer call on the same samples, and an ASR context is rejected with a clear message. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
open_segments()/to_c_sound_segments can allocate and throw across the C boundary; wrap parakeet_capi_sound_stream_active in the same try/catch its siblings use (last_error on failure, cleared on success), and add a streaming-loop check that still-open segments around 3.5 s in have end == stream time and start <= end. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add sound-window-eval, a one-off tool that scores each WAV's whole clip (up to 10 s) as the top-1 reference, then slides 1/2/3/5 s windows (hop = window / 2) over the same clip and reports how often each window size agrees with that reference. Run it with ced-tiny-q8_0 on all 2000 ESC-50 clips plus the three ced.cpp demo clips, and with ced-base-q8_0 on ESC-50 folds 1 and 2 (800 clips) plus the same three demo clips. ced-base's 3 s windows agree with the clip top-1 75.1% of the time, above the 70% floor for switching to a 5 s window, so the SoundOpts defaults (window 3 s, hop 1 s) stay unchanged. Add docs/sound.md: what CED is and where the GGUFs are, loading a tagger through parakeet_capi_load, the windowing and timing rules (one-hop boundary resolution), the measurement and its tables, the sound stream C-API with a short C example, and the CED_DEVICE note. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
The streaming SAS path lived inside the C-API. Move it into C++ classes so the scene stream can reuse it: - pk::DiarPcmStream: PCM to incremental mel to StreamingDiarization, the old parakeet_diar_stream internals and diar_stream_advance. - pk::AsrCommitter: the text commit rules (4 s minimum window, 1 s right context, resume after the last committed word, drop a word heard twice) over an injected transcriber, unit tested with a fake. - pk::SceneStream: runs diarization and ASR over one PCM stream and merges words with speakers. The sound part is not wired yet. parakeet_capi_diarize_stream_* and parakeet_capi_sas_stream_* now call through these classes. Errors still go to the context that failed. The SAS output is unchanged. test_combined_offline and test_streaming_diarization print the same segments and transcripts before and after, and a scratch harness that feeds two_speakers.wav in 0.5 s pieces through sas_stream_begin_latency in all four latency modes prints the same utterances (speaker, text, times, conf) with the old and the new library. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Before the SAS logic moved into SceneStream, the stream counted as finished as soon as the diarizer took the is_last chunk, before transcription. SceneStream set its flag only after a feed completed, so an is_last feed that threw in the diarizer or the transcriber could be fed again: the mel tail was finalized twice, the diarizer got a second last chunk and the ASR buffer got the audio again. SceneStream::finished() now also reports the diarizer's state, so the next feed fails with "stream already finished" as before. Checked by injecting a throw into each part on the last feed. Also note in comments that a successful sas_stream_feed clears the ASR context's last error, and that a scene with only a tagger does nothing yet. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Wire the CED sound part into pk::SceneStream alongside the existing ASR and diarization parts, following the same "at least one" part contract the constructor already enforced. The sound step runs after diarization and ASR, without touching their commit ordering, and adds its closed and still-open segments to the SceneUpdate. Add scene_update_to_json to serialize the shared document (t, utterances, words, speakers, sounds, active.speakers, active.sounds), matching the design spec's key order and field shapes exactly. safe_until now takes the minimum of the ASR commit bound and the sound stream's own safe_until() when both parts run, so a renderer is never told a window is final when a sound segment could still open earlier inside it. With only a sound part it is the sound bound; after finish it is always t. Expose all of this across the C-API as parakeet_scene_stream (ABI v8, additive): parakeet_capi_scene_opts_default, _scene_stream_begin, _feed_json, _drain_scores_json, _last_error and _free. Any of the three contexts may be NULL as long as one is given; a context of the wrong kind is rejected with a message on that context, and an internal failure is attributed to the part that was running via SceneStream::failed_part(), extending the same pattern the sas_stream wrapper already used for ASR/diarization. test_scene_stream proves composition changes nothing: with all three parts, the utterances match sas_stream and the sounds match sound_stream on the same PCM and feed schedule; a sound-only scene returns sounds with empty word arrays; asr+tagger without diarization carries speaker -1 throughout; and a tagger passed where ASR is expected is rejected with a message naming the CED sound model. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Narrow the safe_until doc comment: it only promises that no utterance, word or sound segment starts before it. An already-open speaker segment can still close later with an earlier start, since diarization gives no such bound and speaker segments are not what a renderer commits to the screen. The computed value is unchanged. Document SceneStream::feed's error paths: a part that throws loses the rest of that call (nothing after it in the diar/asr/sound order runs), and with diarization present, an is_last feed that throws ends the stream without flushing whatever the throwing part would otherwise have flushed. Add try/catch around parakeet_capi_scene_stream_drain_scores_json and parakeet_capi_sound_stream_drain_scores_json, which could previously let a C++ exception (for example bad_alloc) cross the C boundary. parakeet_capi_scene_stream_begin now also catches (...) and clears last_error on every context it was given once the stream is built successfully, matching sound_stream_begin's convention. Point parakeet_capi.h's scene document comment at docs/sound.md instead of the untracked docs/superpowers/specs. Strengthen test_scene_stream: check the reference sas_stream and sound_stream feed calls actually succeed, assert the references are non-empty so the later equality checks cannot pass vacuously, compare scene utterances against the sas reference by speaker and start/end (not text alone), and look for the rooster label only inside the closed "sounds" arrays rather than the whole document (which also contains a growing, still-open "active.sounds" entry for the same class). Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add pk::SceneRenderer (src/scene_render.hpp/cpp) to turn a pk::SceneUpdate stream into a time-ordered transcript, and wire a new `parakeet-cli scene` subcommand around it. The subcommand drives pk::SceneStream directly over ASR, diarization and CED sound tagging (each optional, at least one required) and prints speaker-attributed utterances alongside sound annotations such as "(Knock 0.81)", ordered by their start time as each chunk's safe_until promise makes them final. --json instead prints scene_update_to_json per update, the same document the C-API's scene_stream_feed_json returns. format_span renders "[mm:ss.s - mm:ss.s]" with the tenths truncated, rounding to hundredths first to absorb float32 timestamp noise before truncating so genuinely intended tenths survive. is_speech_label hides CED's plain speech classes from the annotation stream unless --show-speech is passed, since ASR/diarization already carry that text. Add tests/test_scene_render.cpp (model-independent, no label) as the TDD spec for format_span/is_speech_label/SceneRenderer. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
"Speech synthesizer" is a child of Speech in the AudioSet ontology and CED emits it over clean narration, so the scene renderer now hides it by default the same way it hides the other speech labels (visible again with --show-speech). Also guard format_span against non-finite input (NaN/inf now clamp to 0 before the long long cast, avoiding UB) and cap --chunk-ms at 60000 so chunk_ms * 16 cannot overflow int. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add --sound-model <path> to parakeet-server, loading a ced.cpp sound
tagger at startup (local path only, not resolved through the ASR
model's alias/URL fetcher). When set, verbose_json responses gain a
sound_events array of {label,start,end,score}; json and text
responses are unchanged and never pay for the sound pass.
The sound pass runs under the same infer_mu lock as ASR inference,
since CedTagger is not thread-safe. A failure in the sound pass is
logged to stderr and the transcription is still returned, without
the field.
format_transcription gains an optional SoundEventOut vector param,
covered by an extended test_server_format.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Document the CED sound-event tagger, pk::SoundStream, pk::SceneStream and the ABI v8 C-API surface added in earlier commits: the scene stream C example and JSON shape, the parakeet-cli scene subcommand with real demo output, and the parakeet-server --sound-model field in docs/sound.md; a short Sound events section with a CLI example in README.md; and the new files, tests, CMake option and ABI v8 symbols in AGENTS.md. Add a build-without-ced job to CI (PARAKEET_WITH_CED=OFF) so the default-on option keeps working when it is turned off. Validate on three GPU backends (Vulkan on strix, CUDA on dgx, Metal on an M4 Mac): parakeet-cli scene reproduces the CPU transcript on all three. Two non-regressions came out of the runs, recorded in docs/sound.md: a staging artifact (CIFS rsync drops the executable bit on tests/server_e2e.sh) and a pre-existing CUDA/Metal multi-context teardown crash in ggml (not Vulkan) that test_combined_offline, test_streaming_diarization and test_scene_stream now exercise by holding more than one GPU-backed context alive at once; the offline CLI/server paths did not hit it. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
AsrCommitter only moved its commit point past the last committed word. When a window held no word with right context (silence, music, long non-speech), nothing was dropped, so the uncommitted buffer grew with the stream. In an ASR-only scene, which commits on every feed, each feed then transcribed the whole growing buffer: quadratic cost and unbounded memory. When no word is kept and the stream is not ending, move the commit point to the start of the first word heard, or to the start of the right context when no word was heard. Behavior with at least one kept word is unchanged; the SAS stream output on two_speakers.wav is identical in every latency mode. Add model-free tests for 20 s of non-speech (the commit point advances and the buffer stays bounded) and for speech after a silent stretch (every word committed once, at absolute times), plus a buffered_samples() accessor. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
`parakeet-cli scene --diar x.gguf` without `--model` printed nothing: the renderer only turned utterances and sounds into lines, and without ASR there are no utterances. SceneRenderer now takes a has_asr flag. With diarization and no ASR it prints each closed speaker segment as `[mm:ss.s - mm:ss.s] Speaker N`, ordered with the sound lines. The scene's safe_until does not cover speaker segments, so flush() also caps its bound at the earliest start a later segment can have: the start of the earliest open segment, or the diarized time when none is open. With ASR present the output is unchanged. The speech-label filter for sound lines works as before in this mode. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
parakeet_capi_scene_stream_begin now rejects an unknown diar_latency the same way sas_stream_begin_latency does, with "unknown diarization latency mode" on the diar ctx. Without a diar ctx the value is unused and not checked. The sound feed error appended the tagger's last error even when it was empty or left over from an earlier call. CedTagger's scorer now clears it on every call, and the C-API appends it only when it is set. Document that the per-window score queue grows until drained (drain regularly, or set top_k = 0), and that after a scene feed error later timestamps may be misaligned, so callers should end the stream. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
test_sound_capi and test_scene_stream now skip (77) when the build has no CED support, like test_ced_parity. test_ced_parity checks label(0) for NULL before comparing it. Rewrite comments in the sound stream and its tests as plain technical rationale, say that test_sound_capi compares window scores with CedTagger::scorer(), and drop a stale note on where dr_wav's implementation lives. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Configuring with PARAKEET_WITH_CED=ON and no third_party/ced.cpp now stops with a message that says how to fix it, instead of a confusing add_subdirectory error. The PARAKEET_WITH_CED=OFF CI step also checks that `parakeet-cli scene --sound` exits 2 with "built without sound tagging". The output is captured before grep because steps run with -e -o pipefail. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Segment boundaries follow the one-hop grid, but a slowly rising or falling score can move them by more than one hop, so drop "never more". The GPU teardown abort also hits test_combined_offline and test_streaming_diarization on the base branch; it comes from the process-global backend's allocator being freed in a static destructor after the GPU context is gone. Remove the staging note, which belongs in a run log. Update the scene demo output for the ASR change that releases non-speech audio, point the ced.cpp links at localai-org, and say that test_ced_parity checks against the PyTorch baseline. Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Parakeet's word start times come from the frame where the first token is emitted, often one or two 80 ms encoder frames after the sound actually starts. When a window commits no word, the cut used to land exactly on the first uncommitted word's start (or the right-context limit), so up to that lag could be lost from the word's onset for good, or from a word the ASR had not emitted yet just before the cut. Back off by a fixed 0.3 s onset margin in that path only, clamped so the commit point never moves backwards; paths that keep at least one word are unchanged. Added a unit test where a word straddles the cut and checked it survives to the next commit with its correct absolute times. Verified the identity check: parakeet_capi_sas_stream_begin_latency over tests/fixtures/two_speakers.wav in 0.5 s pieces, for every PARAKEET_DIAR_LATENCY_* mode, produces byte-for-byte identical output before and after this change (continuous speech never lands in the keep == 0 path). The scene demo clip, which has a silent gap, shows two merged utterance lines instead of three at one boundary; updated the docs/sound.md example accordingly. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
The GPU runs note counted the transcript lines that changed after the ASR release fix. The onset margin changed that count again, so describe the effect instead of a number. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
LocalAI's parakeet-cpp backend is about to load ASR, diarization and sound-event (CED) models through the same parakeet_capi_load and needs to tell them apart without probing individual entry points. Add parakeet_capi_model_kind(ctx), returning PARAKEET_MODEL_KIND_NONE on a NULL context, or ASR/DIARIZATION/SOUND depending on which of ctx->model, ctx->diar, ctx->tagger is set. Declared in the ABI v8 section next to the tagger introspection calls; purely additive, ABI stays 8. Covered by a check in test_capi.cpp (ASR ctx + NULL), an extension to test_combined_offline.cpp (ASR + diarization ctx loaded through the C-API), and test_sound_capi.cpp (tagger ctx, plus its optional ASR ctx). Documented in docs/sound.md's C-API section and AGENTS.md's ABI v8 symbol list. Assisted-by: Claude:claude-sonnet-5 [Claude Code]
mudler
added a commit
to mudler/LocalAI
that referenced
this pull request
Sep 29, 2026
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and #74 (the missing <algorithm> include that broke the image builds) are on master now. Pin 6dea76a instead of the #75 PR head, and update the header comment the bump bot reads. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
mudler
added a commit
to mudler/LocalAI
that referenced
this pull request
Sep 29, 2026
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and #74 (the missing <algorithm> include that broke the image builds) are on master now. Pin 6dea76a instead of the #75 PR head, and update the header comment the bump bot reads. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
mudler
added a commit
to mudler/LocalAI
that referenced
this pull request
Sep 29, 2026
…ne events (#12335) * feat(parakeet-cpp): load diarization and CED models and companions Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound event and combined scene stream C symbols through the same purego.Dlsym probe pattern already used for the batched JSON entry point, so the backend still loads against an older libparakeet.so. Load now classifies the loaded GGUF by role (ASR, diarization or sound) via parakeet_capi_model_kind and can load up to two companion models from Options[] (asr_model:, diarization_model:, sound_model:, paths resolved against opts.ModelPath), verifying each companion's kind and freeing every context opened so far on any failure. Free releases the primary and every companion. AudioTranscription now names the loaded role when it is not ASR instead of a generic model not loaded error. The dynamic batcher starts only when an ASR context ends up loaded, primary or companion. This is groundwork only: the Diarize and SoundDetection RPCs and the live scene stream that actually use these new roles land in later commits. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): reset role fields on a failed companion load loadRoles' freeLoaded only released the C contexts it had opened; it left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed contexts, so a later Free() on the same instance would double-free. Zero all four alongside the CppFree calls. Also route AudioTranscriptionStream and AudioTranscriptionLive through notASRError when ctxPtr is unset but a diarization or sound model is loaded, matching AudioTranscription: both used to return the generic model-not-loaded error instead of naming the loaded role. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet-cpp): add speaker diarization Implement the Diarize RPC for the parakeet-cpp Go backend, wired to Nemotron-3-Diarization through libparakeet.so's diarization C-API. Plain diarization uses parakeet_capi_diarize_pcm; when include_text is set and an ASR companion is loaded, parakeet_capi_transcribe_and_ diarize_json fills each segment's text instead. Speaker labels are the decimal index, or "unknown" for -1 (no diarized speaker overlaps). min_duration_off merges same-speaker segments across a short gap before min_duration_on drops the segments still too short, then ids are renumbered. num_speakers/min_speakers/max_speakers/clustering_ threshold have no Sortformer equivalent and are logged at debug instead of rejected. Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc- 110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B speaker segmentation and matching speaker-attributed transcripts. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet-cpp): add sound event detection Wire the SoundDetection RPC to the CED tagger context (p.tagCtx) loaded by Task 1's role classification. It runs the whole clip through a one-shot parakeet_capi_sound_stream_* session (window 10s, hop 10s, top_k set to the tagger's class count so every drained window carries a full score list), averages each class's score across the drained windows, sorts descending, then applies the request's threshold and top_k (0 keeps every class). No tagCtx returns FailedPrecondition; a libparakeet.so missing the sound_stream symbols returns Unimplemented. Every C call runs under engineMu, and the stream is always freed, even when a feed or drain call fails partway through. Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo clip: "Chicken, rooster" tops the list at score 0.91. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock SoundDetection now checks ctx before each 10 s feed slice (mirroring driver.go's feedSlices) and returns Canceled if the caller gave up, so a long clip can be interrupted instead of feeding to completion regardless. The stream is still freed on every path, cancellation included. Also narrow engineMu to the C calls: the drained JSON document is now decoded after the lock is released, splitting soundStreamScores into a locked soundStreamDrain (opts, begin, feed, drain, free) and an unlocked json.Unmarshal. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet-cpp): stream speaker and sound events during live transcription Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent, repeated on TranscriptLiveResponse. When a diarization or sound companion model is loaded, AudioTranscriptionLive now runs a no-ASR scene stream (parakeet_capi_scene_stream_begin) beside the ASR streaming session, feeding it the same PCM slices and forwarding any closed speaker or sound events alongside the matching ASR delta, or on their own when a slice has no ASR output. The scene stream is freed and reopened on a mid-stream Config reset, flushed with is_last before the closing FinalResult, and degrades gracefully (a warning, not an error) when begin or a later feed call fails, so live transcription keeps working ASR-only. Existing live behavior is unchanged when no companion is configured, and no scene C call is made in that case. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): keep scene events off the ASR critical path in live Emit each slice's ASR result right after the ASR feed, before the scene feed for that slice runs, so a companion diarization/sound model never adds scene compute latency in front of the delta or <EOU> that drives realtime turn detection. Closed speakers/sounds go out afterward as their own response, so a slice with both now produces two responses, ASR first. The live feed log line now reports ASR and scene wall time separately. Re-check the diarization/sound contexts a scene stream was begun with against the live contexts before every feed, under the same lock: Free() can race between an ASR feed and the matching scene feed and free the model the stream borrows. A mismatch now returns without touching the C side. Freeing the stream itself stays unconditional; the scene stream's destructor only releases its own buffers and never touches the borrowed contexts. Also recover a panicking stub inside the live test goroutine instead of crashing the test binary, and reset the live decode-lag tracker on a mid-stream config reset, matching what its own comment already promised. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(realtime): surface live speaker and sound events Carry the backend's closed speaker segments and sound events (TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to seconds), and forward them from the semantic_vad live path. Each speaker segment emits conversation.item.input_audio_transcription.segment with speaker, start, end and empty text under the turn's item id. Each sound event emits conversation.item.sound_detection with one tag (label, score = peak, index) and the event's new optional start/end seconds fields, omitted when unset so the existing unary/windowed sound-detection path is unaffected. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(realtime): keep start/end on a zero-second transcription segment ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used omitempty, so a speaker segment starting at 0.0s dropped its "start" key. Nothing emitted this event before the live scene-event path, so drop omitempty: the segment always carries real times. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(gallery): add parakeet-cpp diarization, CED and realtime scene models Add gallery entries for the new parakeet-cpp capabilities: standalone Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC 110M ASR model for speaker-attributed text, CED-Tiny and CED-Base sound classifiers, and a realtime scene bundle combining the streaming EOU ASR model with diarization and sound companions. SHA256 taken from the Hub API; licenses from each model card (openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED, cc-by-4.0 for the Parakeet ASR models). Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: document parakeet-cpp diarization, sound detection and live scene events Cover the new parakeet-cpp capabilities across the feature pages: Nemotron-3-Diarization as a diarization backend (with and without speaker text, the ignored speaker-count hints, the Sortformer voice-like-sound quirk), CED as a sound classification backend, the asr_model/diarization_model/sound_model/diarization_latency companion options, and the realtime live speaker/sound events (event shapes, the speech-turn-only limitation, and using this or pipeline.sound_detection but not both). Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(gallery): correct the realtime-scene license and wording nits parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA open model license instead. Switch to the gallery's usual spelling for that license and keep the diarization/CED licenses called out in the description. Also: audio-diarization.md now says getting per-segment text needs both an asr_model companion and include_text=true on the request, and audio-to-text.md's option table reads "Use on" (a pairing the loader does not enforce) instead of "Allowed on". Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): reject a companion role that duplicates the primary's loadRoles let a companion option (asr_model:/diarization_model:/ sound_model:) assign into a role field the primary already occupied, for example asr_model: on an already-ASR primary. The companion's context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only walks those three fields, so the original primary context was never freed again. Reject a companion whose role the primary already holds before its GGUF is even loaded, freeing everything loadRoles opened so far, the same way a wrong-kind companion is already rejected. Also warn, rather than silently fall through, when parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a successfully loaded primary; the primary is still treated as ASR, matching today's behavior. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): cap live scene sound score retention sceneBegin started the live diarization/sound companion stream with the C API's default sound options, whose top_k keeps 5 scores per window forever until drained. The live scene path never drains sound scores (only the offline SoundDetection RPC does, with its own fresh stream), so this window queue on the C side grew for the whole session's lifetime. Set opts.Sound.TopK = 0 before starting the scene stream: this disables score retention while leaving sound event detection (onset/ offset), which the live path actually consumes, unaffected. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize mergeCloseSegments only compared neighbors in the single start-sorted segment list, so two same-speaker segments never merged once another speaker's turn fell between them (A, B, A): the short B segment broke the adjacency the merge relied on. Group segments by speaker first, merge within each speaker's own start-ordered run, then re-sort the result by start so interleaved speakers come back out in timeline order. Also harden Diarize's entry points the same way streamFeedDoc/ sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the include_text path, p.ctxPtr) under engineMu right before the C call, so a Free() racing between Diarize's own checks and the lock can no longer reach the C side with a freed context. When the include_text call returns NULL, last_error is now read from both contexts and whichever came back non-empty is reported, since either side of the pairing can be the one that failed. A WAV decode failure is reported as InvalidArgument instead of an unwrapped/untyped error. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): harden SoundDetection's engine checks soundStreamDrain ran every C call under engineMu but never re-checked p.tagCtx there, so a Free() racing between SoundDetection's own tagCtx==0 check and this lock could still reach the C side with a freed context. Re-check p.tagCtx under the lock and return ModelNotLoaded when it was cleared, mirroring diarizeCall's own re-check. A WAV decode failure is now reported as InvalidArgument instead of an unwrapped/untyped error. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(parakeet-cpp): cover a mid-session scene feed failure feedSlicesScene already degrades gracefully when a scene feed call fails mid-session: it frees the broken stream and carries the ASR-only session forward. Add a spec covering that path end to end: the scene stream is freed exactly once, later audio slices still produce ASR responses, and no speaker/sound events appear before or after the failure. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: fix the parakeet-cpp companion role table and realtime scene docs audio-to-text.md's companion option table read "Use on" with a note that the loader did not enforce the pairing; it now rejects a companion whose role duplicates the primary's, so restore the "Allowed on" wording and describe the real enforcement. openai-realtime.md's live speaker/sound section claimed a mid-stream session.update resets the companion stream and that it flushes on session close; neither happens, since the realtime core opens one live stream (and so one scene stream) per speech turn and closes it at that turn's commit, with no mid-stream Config in between. Document that lifecycle instead, state precisely that start/end are seconds from the start of the turn's own audio, and note that the diarization model starts a fresh session every turn, so a speaker index is only meaningful within one turn. The example sound tag ("Rooster", index 17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf is "Baby laughter". Replaced with "Chicken, rooster" at its real index, 99. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(parakeet-cpp): use CED's real index for Chicken, rooster The scene feed comment and the live test's canned document gave "Chicken, rooster" index 365. In CED's AudioSet label list it is 99, which is also what the realtime docs show. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(realtime): call the test event accessor The scene-event tests range over a method instead of its returned slice. Call the synchronized accessor so the OpenAI test package compiles. Assisted-by: Codex:gpt-6 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet-cpp): pin parakeet.cpp master with sound events mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and #74 (the missing <algorithm> include that broke the image builds) are on master now. Pin 6dea76a instead of the #75 PR head, and update the header comment the bump bot reads. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main parakeet.cpp #76 moved its ced.cpp submodule from the head of localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where #3 landed with an identical tree. Pin 623a968 so the image builds no longer depend on that branch. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(transcription): carry speaker labels on words and streamed segments A diarizing backend could label transcript segments, but two paths dropped the label: TranscriptWord had no speaker field, so live transcription words and word-level timestamps could not carry one, and the stream=true transcript.text.done event left the speaker out of its segments. TranscriptWord gains an optional speaker (proto field 4, additive). It flows through the live event and result mapping, the JSON word output of the endpoint and the CLI, and transcript.text.done now includes a segment's speaker when there is one. Empty labels are omitted, so responses without diarization are unchanged. Assisted-by: Claude:claude-opus-5-5 [Claude Code] (cherry picked from commit 2f0049f) Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(importers): detect the parakeet.cpp diarization GGUF The Nemotron-3-Diarization GGUFs are published in mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The parakeet-cpp importer did not recognise that name, so a direct `local-ai models import` of the file fell through to another importer. A direct URL to the file now imports with the diarization usecase. A repo import still picks ASR weights when the repo also ships the diarization model, and falls back to the diarization weights only when there are no others. Ported from #12323. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(config): advertise diarization and sound detection for parakeet-cpp The capability table listed parakeet-cpp as transcription only, though the backend now answers Diarize (Nemotron-3-Diarization) and SoundDetection (CED) depending on the model kind it loads. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet-cpp): label transcript segments with the diarization companion A diarization_model companion only fed live speaker events and Diarize; /v1/audio/transcriptions ignored it. With the companion attached and diarize=true (the OpenAI endpoint's default), unary transcription now labels each segment with its speaker and splits segments at speaker turns; with word timestamps each word carries its speaker. The stream=true final result labels each utterance with the speaker who said most of it. Both use the checkpoint's own diarization over the whole clip, as NeMo's diarize() does. Words take the speaker whose segments overlap them most, or the nearest segment within 0.5 s, the same rule as parakeet.cpp's speaker-attributed ASR. Docs: the diarization_model row and a paragraph on transcript speakers; Nemotron-3-Diarization handles up to 8 speakers. Ported from #12323. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(realtime): speaker segments from committed-turn transcription Speaker events reached a realtime session only from the live semantic_vad path, which needs a cache-aware streaming transcription model. Committed-turn transcription (server_vad, or any offline model) always asked the backend for diarize=false and dropped the segments' speakers. pipeline.diarization (off by default) asks the transcription model for speaker labels on each committed turn and emits every labelled segment as a conversation.item.input_audio_transcription.segment event, with its text, before the turn's completed event. It is opt-in because some backends fail a diarization request they cannot serve. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(gallery): add parakeet-cpp-realtime-scene-tdt parakeet-cpp-realtime-scene pairs the streaming EOU model with the diarization and CED companions; its speaker and sound events need a cache-aware streaming model. This entry does the same with Parakeet TDT 0.6B v3 (multilingual, offline) for realtime under server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments and sound tags from one parakeet-cpp backend. Files and sha256 match the Hub and are shared with the existing TDT v3, diarization and CED-Tiny entries. A real-model spec checks the combination on a clip with two speakers and a rooster: A-B-A-B speaker turns, and "Chicken, rooster" among the sound tags. The test loader now binds the sound entry points like main.go. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(gallery): add CED-Base variants of the parakeet-cpp scene models parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for Tiny). Measured on CPU over a 37 s clip: the live diarization + sound stream runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny, because diarization dominates; sound detection per committed turn costs 0.031 against 0.005. The realtime docs list both and note that any CED size works as sound_model. Files and sha256 match the Hub and are shared with the existing parakeet-cpp-ced-base entry. The TDT variant passes the real-model scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found). Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(config): register pipeline.diarization in the config metadata TestAllFieldsHaveRegistryEntries fails on the branch because the new pipeline.diarization field has no registry entry. Add one so the model editor shows it as a toggle next to the sound detection options. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Add sound-event detection to parakeet.cpp through ced.cpp (Xiaomi's CED AudioSet tagger, 527 classes), and a combined "scene" stream that returns words, speakers and sound events on one timeline from a single feed call.
third_party/ced.cpp), reached only throughced_capi.hviapk::CedTagger. A CED GGUF loads withparakeet_capi_loadinto a third context kind.pk::SoundStream: CED on a sliding window (default 3 s every 1 s), per-class hysteresis segments ({label, start, end, peak}) plus a queue of per-window top-k scores.pk::DiarPcmStream+pk::AsrCommitter+pk::SceneStream. The SAS output is identical before and after (checked with the existing tests and a harness over all four latency modes).parakeet_capi_sound_stream_*,parakeet_capi_scene_stream_*(any of ASR / diarization / tagger may be NULL;feed_jsonreturns one document per call),parakeet_capi_num_classes,parakeet_capi_class_label. Options structs are size-versioned;parakeet_scene_opts.flagsis reserved for speech gating later.parakeet-cli scene [--model] [--diar] [--sound] --input x.wavprints a time-ordered transcript with sound annotations (--jsonprints the C-API documents).parakeet-server --sound-model ced.ggufadds"sound_events"toverbose_jsonresponses.PARAKEET_WITH_CED(default ON). With OFF the new functions stay exported and fail with "built without sound tagging"; CI builds and tests that configuration too.docs/sound.md, a README section, AGENTS.md.Why
Sound detection so far lived only in LocalAI's separate
cedbackend. Having it in the same stream as ASR and diarization gives one clock for words, speakers and events, and it is the place speech gating can go later. A tagging head on parakeet's own encoder was measured first and rejected: a frozen probe on the best parakeet layers reaches 80 to 86% on ESC-50 against 96 to 98% for CED, so CED runs as its own model (ced-tiny costs about 10 ms per 5 s clip).Example
parakeet-cli scene --model tdt_ctc-110m --diar nemotron-3-diarization --sound ced-base-q8_0 --latency lowon two speakers, a rooster and more speech:How to verify
Local results: 22/22 model-independent tests (ON and OFF builds), all sound, scene and SAS model tests pass, and
test_ced_paritygives max|d| = 1.7e-7 on CPU (the same as standalone ced.cpp).GPU (see
docs/sound.md): Vulkan on a Radeon 8060S, CUDA on a GB10 and Metal on an M4 all reproduce the CPUscenetranscript. On CUDA and Metal,test_combined_offline,test_streaming_diarizationandtest_scene_streamprint PASS and then abort at process exit. The same abort happens on master without this work: the process-global backend's allocator is freed in a static destructor after the GPU context is gone. That needs its own fix.Before merge
e1a3cfa, the head of Allow embedding ced.cpp in another ggml project localai-org/ced.cpp#3 (stacked on Metal backend crashes on macOS arm64 with tdt-0.6b-v3-q8_0.gguf #2, the GPU backend). Once those merge, the pin should move to a commit on ced.cpp main.AsrCommitternow backs off by a 0.3 s onset margin from the first uncommitted word's start (or the right-context limit) before releasing audio, instead of cutting exactly there. Verified the SAS stream output is byte-for-byte unchanged ontwo_speakers.wavacross all latency modes, since continuous speech never hits this path.Follow-ups
flagsfield is reserved for it).StreamingSessionfor cache-aware models in the scene stream.🤖 Generated with Claude Code