Skip to content

Add sound-event detection and a combined scene stream - #75

Merged
mudler merged 24 commits into
masterfrom
feat/sound-events
Sep 29, 2026
Merged

mudler merged 24 commits into
masterfrom
feat/sound-events

Conversation

@localai-org-maint-bot

@localai-org-maint-bot localai-org-maint-bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What

Add sound-event detection to parakeet.cpp through ced.cpp (Xiaomi's CED AudioSet tagger, 527 classes), and a combined "scene" stream that returns words, speakers and sound events on one timeline from a single feed call.

  • ced.cpp is a submodule (third_party/ced.cpp), reached only through ced_capi.h via pk::CedTagger. A CED GGUF loads with parakeet_capi_load into a third context kind.
  • pk::SoundStream: CED on a sliding window (default 3 s every 1 s), per-class hysteresis segments ({label, start, end, peak}) plus a queue of per-window top-k scores.
  • The speaker-attributed ASR stream now runs on pk::DiarPcmStream + pk::AsrCommitter + pk::SceneStream. The SAS output is identical before and after (checked with the existing tests and a harness over all four latency modes).
  • C-API, ABI 8, all additive: parakeet_capi_sound_stream_*, parakeet_capi_scene_stream_* (any of ASR / diarization / tagger may be NULL; feed_json returns one document per call), parakeet_capi_num_classes, parakeet_capi_class_label. Options structs are size-versioned; parakeet_scene_opts.flags is reserved for speech gating later.
  • parakeet-cli scene [--model] [--diar] [--sound] --input x.wav prints a time-ordered transcript with sound annotations (--json prints the C-API documents).
  • parakeet-server --sound-model ced.gguf adds "sound_events" to verbose_json responses.
  • PARAKEET_WITH_CED (default ON). With OFF the new functions stay exported and fail with "built without sound tagging"; CI builds and tests that configuration too.
  • Docs: docs/sound.md, a README section, AGENTS.md.

Why

Sound detection so far lived only in LocalAI's separate ced backend. Having it in the same stream as ASR and diarization gives one clock for words, speakers and events, and it is the place speech gating can go later. A tagging head on parakeet's own encoder was measured first and rejected: a frozen probe on the best parakeet layers reaches 80 to 86% on ESC-50 against 96 to 98% for CED, so CED runs as its own model (ced-tiny costs about 10 ms per 5 s clip).

Example

parakeet-cli scene --model tdt_ctc-110m --diar nemotron-3-diarization --sound ced-base-q8_0 --latency low on two speakers, a rooster and more speech:

[00:00.4 - 00:03.2]  Speaker 0: mister Quilter is the apostle of the middle classes, and
[00:03.6 - 00:05.4]  Speaker 0: we're glad to welcome his gospel.
...
[00:24.0 - 00:30.0]  (Chicken, rooster 0.86)
[00:26.0 - 00:30.0]  (Crowing, cock-a-doodle-doo 0.65)
...

How to verify

git submodule update --init --recursive
cmake -B build -DPARAKEET_BUILD_TESTS=ON && cmake --build build -j
ctest --test-dir build -LE model --output-on-failure

# model tests
export PARAKEET_TEST_GGUF=<parakeet ASR gguf> PARAKEET_TEST_DIAR_GGUF=<nemotron-3-diarization gguf> \
       PARAKEET_TEST_BASELINE_DIAR=<diar baseline> PARAKEET_TEST_CED_GGUF=<ced-base-q8_0.gguf>
ctest --test-dir build -R "test_asr_committer|test_sound|test_scene|test_combined_offline|test_streaming_diarization" --output-on-failure
CED_DEVICE=cpu PARAKEET_TEST_CED_GGUF=<ced-base-f32.gguf> PARAKEET_TEST_CED_BASELINE=<ced tests/fixtures/ced-base.baseline.gguf> \
  ctest --test-dir build -R test_ced_parity --output-on-failure

Local results: 22/22 model-independent tests (ON and OFF builds), all sound, scene and SAS model tests pass, and test_ced_parity gives max|d| = 1.7e-7 on CPU (the same as standalone ced.cpp).

GPU (see docs/sound.md): Vulkan on a Radeon 8060S, CUDA on a GB10 and Metal on an M4 all reproduce the CPU scene transcript. On CUDA and Metal, test_combined_offline, test_streaming_diarization and test_scene_stream print PASS and then abort at process exit. The same abort happens on master without this work: the process-global backend's allocator is freed in a static destructor after the GPU context is gone. That needs its own fix.

Before merge

  • Repin ced.cpp to main. The submodule points at localai-org/ced.cpp e1a3cfa, the head of Allow embedding ced.cpp in another ggml project localai-org/ced.cpp#3 (stacked on Metal backend crashes on macOS arm64 with tdt-0.6b-v3-q8_0.gguf #2, the GPU backend). Once those merge, the pin should move to a commit on ced.cpp main.
  • Applied: when a window commits no word (after a long pause), AsrCommitter now backs off by a 0.3 s onset margin from the first uncommitted word's start (or the right-context limit) before releasing audio, instead of cutting exactly there. Verified the SAS stream output is byte-for-byte unchanged on two_speakers.wav across all latency modes, since continuous speech never hits this path.

Follow-ups

  • Speech gating on CED's Speech score (the flags field is reserved for it).
  • Streaming ASR through StreamingSession for cache-aware models in the scene stream.
  • Using the scene stream from LocalAI's parakeet backend.
  • Release the backend before static teardown (the CUDA/Metal exit abort above).

🤖 Generated with Claude Code

Add third_party/ced.cpp (localai-org/ced.cpp, feat/embedding @ e1a3cfa)
as a submodule and wire it into the build behind PARAKEET_WITH_CED
(default ON). CED GGUFs (general.architecture "ced") now load through
parakeet_capi_load into a third parakeet_ctx kind, a "tagger", exactly
like ASR and diarization models do today; pk::CedTagger in the new
src/ced_tagger.hpp/.cpp is the only parakeet code that talks to
ced_capi.h, so the rest of the codebase never needs to know ced.cpp
exists. Bump PARAKEET_CAPI_ABI_VERSION to 8 (additive) and add
parakeet_capi_num_classes / parakeet_capi_class_label for tagger
introspection.

ced.cpp is built with CED_EXTERNAL_DR_WAV so it reuses parakeet's own
dr_wav implementation instead of linking a second copy; that makes the
two static libraries genuinely circular at link time (parakeet needs
ced's classify symbols, ced needs parakeet's dr_wav symbols), so both
targets declare a link dependency on each other and let CMake repeat
the archives in the generated link line.

tests/test_ced_parity.cpp checks ced.cpp built against parakeet's own
patched ggml reproduces ced.cpp's own PyTorch baseline: max|d| =
1.714e-07 on CPU f32, matching standalone ced.cpp's reported 1.7e-7.
It reads the model and baseline paths from PARAKEET_TEST_CED_GGUF /
PARAKEET_TEST_CED_BASELINE and skips (77) when either is unset, since
ced.cpp's baseline fixtures are gitignored there. Confirmed the
PARAKEET_WITH_CED=OFF build compiles clean and still skips the test.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
…s own TU

The reciprocal target_link_libraries(ced PRIVATE parakeet) added to make
CED_EXTERNAL_DR_WAV link only works because CMake allows cycles among
STATIC libraries. It breaks any -DPARAKEET_SHARED=ON configure ("Cyclic
dependencies are allowed only among static libraries"), including the
libparakeet.so build LocalAI dlopens and the release workflow's shared
jobs, since PARAKEET_WITH_CED defaults ON.

Move the single dr_wav implementation out of src/audio_io.cpp (which now
only declares dr_wav's functions, like ced's own audio_io.cpp does) into
a new source file, src/dr_wav_impl.cpp, built as an OBJECT library.
libparakeet and, when PARAKEET_WITH_CED is on, ced both link against
that object library directly instead of against each other, so there is
no longer any cycle: parakeet depends on ced and dr_wav_impl, ced
depends on dr_wav_impl, and dr_wav_impl depends on nothing.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Adds pk::SoundStream on top of the Task 2 SoundScorer: windows end
every hop, hysteresis segments open at on_threshold and close at
off_threshold with one-hop start/end timing, is_last scores an
unscored tail and force-closes everything open, and a per-window
score queue (drain_windows) keeps the top-k classes for the C-API.
Includes model-independent tests with a fake scorer that mimics
CED's mean pooling, exercising single/overlapping sounds, hysteresis,
piece-size invariance, min-duration filtering, open segments,
validation, and scorer failure.

Fixes a buffer-trim defect: the naive keep_from = next_end_ - win_n_
looks one hop ahead of what has actually streamed, so an is_last tail
window ending between two hops could read before the buffer start.
keep_from now clamps to min(next_end_, samples_in_) - win_n_, and
score_window guards its precondition so a future regression throws
instead of reading out of bounds.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
While a SoundStream is still open, safe_until() promised no future
segment starts earlier than the last window actually scored, but it
missed a case: is_last can score a shorter tail window and open a
class at that window's newest hop, which starts earlier than the
last regularly scored window whenever the stream ends mid-hop. Widen
the bound to min(scored_end_, max(0, samples_in_ - hop_n_)), since a
tail window is always at least one hop wide, plus any already-open
class's start.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add parakeet_capi_sound_stream_begin/feed/active/drain_scores_json/free
and parakeet_capi_free_sound_segments, wrapping pk::SoundStream over a
loaded CED tagger context. parakeet_sound_opts carries a size field so
callers built against an older header still work (only the fields
their struct covers are read). Segment output is a malloc'd C array of
parakeet_sound_segment (index, borrowed label, start/end/peak); the
per-window top-k scores are JSON only, since a variable-length tag
list per window does not fit a fixed C struct and LocalAI consumes
JSON anyway.

Tested end to end against a real ced-base GGUF on ced.cpp's rooster,
thunder and guitar demo clips fed in small pieces: each clip's label
closes inside its own 6 s slot, window scores in the drain match a
direct scorer call on the same samples, and an ASR context is
rejected with a clear message.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
open_segments()/to_c_sound_segments can allocate and throw across the
C boundary; wrap parakeet_capi_sound_stream_active in the same
try/catch its siblings use (last_error on failure, cleared on
success), and add a streaming-loop check that still-open segments
around 3.5 s in have end == stream time and start <= end.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add sound-window-eval, a one-off tool that scores each WAV's whole
clip (up to 10 s) as the top-1 reference, then slides 1/2/3/5 s
windows (hop = window / 2) over the same clip and reports how often
each window size agrees with that reference.

Run it with ced-tiny-q8_0 on all 2000 ESC-50 clips plus the three
ced.cpp demo clips, and with ced-base-q8_0 on ESC-50 folds 1 and 2
(800 clips) plus the same three demo clips. ced-base's 3 s windows
agree with the clip top-1 75.1% of the time, above the 70% floor for
switching to a 5 s window, so the SoundOpts defaults (window 3 s,
hop 1 s) stay unchanged.

Add docs/sound.md: what CED is and where the GGUFs are, loading a
tagger through parakeet_capi_load, the windowing and timing rules
(one-hop boundary resolution), the measurement and its tables, the
sound stream C-API with a short C example, and the CED_DEVICE note.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
The streaming SAS path lived inside the C-API. Move it into C++ classes
so the scene stream can reuse it:

- pk::DiarPcmStream: PCM to incremental mel to StreamingDiarization,
  the old parakeet_diar_stream internals and diar_stream_advance.
- pk::AsrCommitter: the text commit rules (4 s minimum window, 1 s
  right context, resume after the last committed word, drop a word
  heard twice) over an injected transcriber, unit tested with a fake.
- pk::SceneStream: runs diarization and ASR over one PCM stream and
  merges words with speakers. The sound part is not wired yet.

parakeet_capi_diarize_stream_* and parakeet_capi_sas_stream_* now call
through these classes. Errors still go to the context that failed.

The SAS output is unchanged. test_combined_offline and
test_streaming_diarization print the same segments and transcripts
before and after, and a scratch harness that feeds two_speakers.wav in
0.5 s pieces through sas_stream_begin_latency in all four latency
modes prints the same utterances (speaker, text, times, conf) with the
old and the new library.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Before the SAS logic moved into SceneStream, the stream counted as
finished as soon as the diarizer took the is_last chunk, before
transcription. SceneStream set its flag only after a feed completed,
so an is_last feed that threw in the diarizer or the transcriber could
be fed again: the mel tail was finalized twice, the diarizer got a
second last chunk and the ASR buffer got the audio again.

SceneStream::finished() now also reports the diarizer's state, so the
next feed fails with "stream already finished" as before. Checked by
injecting a throw into each part on the last feed.

Also note in comments that a successful sas_stream_feed clears the ASR
context's last error, and that a scene with only a tagger does nothing
yet.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Wire the CED sound part into pk::SceneStream alongside the existing
ASR and diarization parts, following the same "at least one" part
contract the constructor already enforced. The sound step runs after
diarization and ASR, without touching their commit ordering, and adds
its closed and still-open segments to the SceneUpdate.

Add scene_update_to_json to serialize the shared document (t,
utterances, words, speakers, sounds, active.speakers, active.sounds),
matching the design spec's key order and field shapes exactly.

safe_until now takes the minimum of the ASR commit bound and the sound
stream's own safe_until() when both parts run, so a renderer is never
told a window is final when a sound segment could still open earlier
inside it. With only a sound part it is the sound bound; after finish
it is always t.

Expose all of this across the C-API as parakeet_scene_stream (ABI v8,
additive): parakeet_capi_scene_opts_default, _scene_stream_begin,
_feed_json, _drain_scores_json, _last_error and _free. Any of the
three contexts may be NULL as long as one is given; a context of the
wrong kind is rejected with a message on that context, and an
internal failure is attributed to the part that was running via
SceneStream::failed_part(), extending the same pattern the sas_stream
wrapper already used for ASR/diarization.

test_scene_stream proves composition changes nothing: with all three
parts, the utterances match sas_stream and the sounds match
sound_stream on the same PCM and feed schedule; a sound-only scene
returns sounds with empty word arrays; asr+tagger without diarization
carries speaker -1 throughout; and a tagger passed where ASR is
expected is rejected with a message naming the CED sound model.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Narrow the safe_until doc comment: it only promises that no utterance,
word or sound segment starts before it. An already-open speaker
segment can still close later with an earlier start, since
diarization gives no such bound and speaker segments are not what a
renderer commits to the screen. The computed value is unchanged.

Document SceneStream::feed's error paths: a part that throws loses
the rest of that call (nothing after it in the diar/asr/sound order
runs), and with diarization present, an is_last feed that throws ends
the stream without flushing whatever the throwing part would
otherwise have flushed.

Add try/catch around parakeet_capi_scene_stream_drain_scores_json and
parakeet_capi_sound_stream_drain_scores_json, which could previously
let a C++ exception (for example bad_alloc) cross the C boundary.
parakeet_capi_scene_stream_begin now also catches (...) and clears
last_error on every context it was given once the stream is built
successfully, matching sound_stream_begin's convention.

Point parakeet_capi.h's scene document comment at docs/sound.md
instead of the untracked docs/superpowers/specs.

Strengthen test_scene_stream: check the reference sas_stream and
sound_stream feed calls actually succeed, assert the references are
non-empty so the later equality checks cannot pass vacuously, compare
scene utterances against the sas reference by speaker and start/end
(not text alone), and look for the rooster label only inside the
closed "sounds" arrays rather than the whole document (which also
contains a growing, still-open "active.sounds" entry for the same
class).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add pk::SceneRenderer (src/scene_render.hpp/cpp) to turn a
pk::SceneUpdate stream into a time-ordered transcript, and wire a new
`parakeet-cli scene` subcommand around it. The subcommand drives
pk::SceneStream directly over ASR, diarization and CED sound tagging
(each optional, at least one required) and prints speaker-attributed
utterances alongside sound annotations such as "(Knock 0.81)",
ordered by their start time as each chunk's safe_until promise makes
them final. --json instead prints scene_update_to_json per update,
the same document the C-API's scene_stream_feed_json returns.

format_span renders "[mm:ss.s - mm:ss.s]" with the tenths truncated,
rounding to hundredths first to absorb float32 timestamp noise before
truncating so genuinely intended tenths survive. is_speech_label
hides CED's plain speech classes from the annotation stream unless
--show-speech is passed, since ASR/diarization already carry that
text.

Add tests/test_scene_render.cpp (model-independent, no label) as the
TDD spec for format_span/is_speech_label/SceneRenderer.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
"Speech synthesizer" is a child of Speech in the AudioSet ontology and
CED emits it over clean narration, so the scene renderer now hides it
by default the same way it hides the other speech labels (visible
again with --show-speech).

Also guard format_span against non-finite input (NaN/inf now clamp to
0 before the long long cast, avoiding UB) and cap --chunk-ms at 60000
so chunk_ms * 16 cannot overflow int.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Add --sound-model <path> to parakeet-server, loading a ced.cpp sound
tagger at startup (local path only, not resolved through the ASR
model's alias/URL fetcher). When set, verbose_json responses gain a
sound_events array of {label,start,end,score}; json and text
responses are unchanged and never pay for the sound pass.

The sound pass runs under the same infer_mu lock as ASR inference,
since CedTagger is not thread-safe. A failure in the sound pass is
logged to stderr and the transcription is still returned, without
the field.

format_transcription gains an optional SoundEventOut vector param,
covered by an extended test_server_format.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Document the CED sound-event tagger, pk::SoundStream, pk::SceneStream
and the ABI v8 C-API surface added in earlier commits: the scene
stream C example and JSON shape, the parakeet-cli scene subcommand
with real demo output, and the parakeet-server --sound-model field in
docs/sound.md; a short Sound events section with a CLI example in
README.md; and the new files, tests, CMake option and ABI v8 symbols
in AGENTS.md.

Add a build-without-ced job to CI (PARAKEET_WITH_CED=OFF) so the
default-on option keeps working when it is turned off.

Validate on three GPU backends (Vulkan on strix, CUDA on dgx, Metal on
an M4 Mac): parakeet-cli scene reproduces the CPU transcript on all
three. Two non-regressions came out of the runs, recorded in
docs/sound.md: a staging artifact (CIFS rsync drops the executable
bit on tests/server_e2e.sh) and a pre-existing CUDA/Metal multi-context
teardown crash in ggml (not Vulkan) that test_combined_offline,
test_streaming_diarization and test_scene_stream now exercise by
holding more than one GPU-backed context alive at once; the offline
CLI/server paths did not hit it.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
AsrCommitter only moved its commit point past the last committed word.
When a window held no word with right context (silence, music, long
non-speech), nothing was dropped, so the uncommitted buffer grew with
the stream. In an ASR-only scene, which commits on every feed, each
feed then transcribed the whole growing buffer: quadratic cost and
unbounded memory.

When no word is kept and the stream is not ending, move the commit
point to the start of the first word heard, or to the start of the
right context when no word was heard. Behavior with at least one kept
word is unchanged; the SAS stream output on two_speakers.wav is
identical in every latency mode.

Add model-free tests for 20 s of non-speech (the commit point advances
and the buffer stays bounded) and for speech after a silent stretch
(every word committed once, at absolute times), plus a
buffered_samples() accessor.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
`parakeet-cli scene --diar x.gguf` without `--model` printed nothing:
the renderer only turned utterances and sounds into lines, and without
ASR there are no utterances.

SceneRenderer now takes a has_asr flag. With diarization and no ASR it
prints each closed speaker segment as `[mm:ss.s - mm:ss.s]  Speaker N`,
ordered with the sound lines. The scene's safe_until does not cover
speaker segments, so flush() also caps its bound at the earliest start
a later segment can have: the start of the earliest open segment, or
the diarized time when none is open. With ASR present the output is
unchanged.

The speech-label filter for sound lines works as before in this mode.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
parakeet_capi_scene_stream_begin now rejects an unknown diar_latency
the same way sas_stream_begin_latency does, with "unknown diarization
latency mode" on the diar ctx. Without a diar ctx the value is unused
and not checked.

The sound feed error appended the tagger's last error even when it was
empty or left over from an earlier call. CedTagger's scorer now clears
it on every call, and the C-API appends it only when it is set.

Document that the per-window score queue grows until drained (drain
regularly, or set top_k = 0), and that after a scene feed error later
timestamps may be misaligned, so callers should end the stream.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
test_sound_capi and test_scene_stream now skip (77) when the build has
no CED support, like test_ced_parity. test_ced_parity checks label(0)
for NULL before comparing it.

Rewrite comments in the sound stream and its tests as plain technical
rationale, say that test_sound_capi compares window scores with
CedTagger::scorer(), and drop a stale note on where dr_wav's
implementation lives.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Configuring with PARAKEET_WITH_CED=ON and no third_party/ced.cpp now
stops with a message that says how to fix it, instead of a confusing
add_subdirectory error.

The PARAKEET_WITH_CED=OFF CI step also checks that `parakeet-cli scene
--sound` exits 2 with "built without sound tagging". The output is
captured before grep because steps run with -e -o pipefail.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Segment boundaries follow the one-hop grid, but a slowly rising or
falling score can move them by more than one hop, so drop "never
more". The GPU teardown abort also hits test_combined_offline and
test_streaming_diarization on the base branch; it comes from the
process-global backend's allocator being freed in a static destructor
after the GPU context is gone. Remove the staging note, which belongs
in a run log.

Update the scene demo output for the ASR change that releases
non-speech audio, point the ced.cpp links at localai-org, and say that
test_ced_parity checks against the PyTorch baseline.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Parakeet's word start times come from the frame where the first token
is emitted, often one or two 80 ms encoder frames after the sound
actually starts. When a window commits no word, the cut used to land
exactly on the first uncommitted word's start (or the right-context
limit), so up to that lag could be lost from the word's onset for
good, or from a word the ASR had not emitted yet just before the cut.

Back off by a fixed 0.3 s onset margin in that path only, clamped so
the commit point never moves backwards; paths that keep at least one
word are unchanged. Added a unit test where a word straddles the cut
and checked it survives to the next commit with its correct absolute
times.

Verified the identity check: parakeet_capi_sas_stream_begin_latency
over tests/fixtures/two_speakers.wav in 0.5 s pieces, for every
PARAKEET_DIAR_LATENCY_* mode, produces byte-for-byte identical output
before and after this change (continuous speech never lands in the
keep == 0 path). The scene demo clip, which has a silent gap, shows
two merged utterance lines instead of three at one boundary; updated
the docs/sound.md example accordingly.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
The GPU runs note counted the transcript lines that changed after the
ASR release fix. The onset margin changed that count again, so describe
the effect instead of a number.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
LocalAI's parakeet-cpp backend is about to load ASR, diarization and
sound-event (CED) models through the same parakeet_capi_load and
needs to tell them apart without probing individual entry points.

Add parakeet_capi_model_kind(ctx), returning PARAKEET_MODEL_KIND_NONE
on a NULL context, or ASR/DIARIZATION/SOUND depending on which of
ctx->model, ctx->diar, ctx->tagger is set. Declared in the ABI v8
section next to the tagger introspection calls; purely additive, ABI
stays 8.

Covered by a check in test_capi.cpp (ASR ctx + NULL), an extension to
test_combined_offline.cpp (ASR + diarization ctx loaded through the
C-API), and test_sound_capi.cpp (tagger ctx, plus its optional ASR
ctx). Documented in docs/sound.md's C-API section and AGENTS.md's ABI
v8 symbol list.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
@mudler
mudler merged commit 2c67ae2 into master Sep 29, 2026
10 checks passed
mudler added a commit to mudler/LocalAI that referenced this pull request Sep 29, 2026
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
mudler added a commit to mudler/LocalAI that referenced this pull request Sep 29, 2026
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
mudler added a commit to mudler/LocalAI that referenced this pull request Sep 29, 2026
…ne events (#12335)

* feat(parakeet-cpp): load diarization and CED models and companions

Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.

Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.

This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reset role fields on a failed companion load

loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.

Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add speaker diarization

Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.

Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.

Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add sound event detection

Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).

No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.

Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock

SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.

Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): stream speaker and sound events during live transcription

Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.

The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): keep scene events off the ASR critical path in live

Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.

Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.

Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): surface live speaker and sound events

Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.

Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): keep start/end on a zero-second transcription segment

ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models

Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.

SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: document parakeet-cpp diarization, sound detection and live scene events

Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): correct the realtime-scene license and wording nits

parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.

Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reject a companion role that duplicates the primary's

loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.

Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.

Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cap live scene sound score retention

sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.

Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize

mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.

Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): harden SoundDetection's engine checks

soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(parakeet-cpp): cover a mid-session scene feed failure

feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fix the parakeet-cpp companion role table and realtime scene docs

audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.

openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): use CED's real index for Chicken, rooster

The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): call the test event accessor

The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.

Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp master with sound events

mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main

parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(transcription): carry speaker labels on words and streamed segments

A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.

TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit 2f0049f)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(importers): detect the parakeet.cpp diarization GGUF

The Nemotron-3-Diarization GGUFs are published in
mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The
parakeet-cpp importer did not recognise that name, so a direct
`local-ai models import` of the file fell through to another importer.

A direct URL to the file now imports with the diarization usecase. A
repo import still picks ASR weights when the repo also ships the
diarization model, and falls back to the diarization weights only
when there are no others.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): advertise diarization and sound detection for parakeet-cpp

The capability table listed parakeet-cpp as transcription only, though
the backend now answers Diarize (Nemotron-3-Diarization) and
SoundDetection (CED) depending on the model kind it loads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): label transcript segments with the diarization companion

A diarization_model companion only fed live speaker events and
Diarize; /v1/audio/transcriptions ignored it.

With the companion attached and diarize=true (the OpenAI endpoint's
default), unary transcription now labels each segment with its
speaker and splits segments at speaker turns; with word timestamps
each word carries its speaker. The stream=true final result labels
each utterance with the speaker who said most of it. Both use the
checkpoint's own diarization over the whole clip, as NeMo's diarize()
does. Words take the speaker whose segments overlap them most, or the
nearest segment within 0.5 s, the same rule as parakeet.cpp's
speaker-attributed ASR.

Docs: the diarization_model row and a paragraph on transcript
speakers; Nemotron-3-Diarization handles up to 8 speakers.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): speaker segments from committed-turn transcription

Speaker events reached a realtime session only from the live
semantic_vad path, which needs a cache-aware streaming transcription
model. Committed-turn transcription (server_vad, or any offline
model) always asked the backend for diarize=false and dropped the
segments' speakers.

pipeline.diarization (off by default) asks the transcription model for
speaker labels on each committed turn and emits every labelled segment
as a conversation.item.input_audio_transcription.segment event, with
its text, before the turn's completed event. It is opt-in because some
backends fail a diarization request they cannot serve.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add parakeet-cpp-realtime-scene-tdt

parakeet-cpp-realtime-scene pairs the streaming EOU model with the
diarization and CED companions; its speaker and sound events need a
cache-aware streaming model. This entry does the same with Parakeet
TDT 0.6B v3 (multilingual, offline) for realtime under server_vad:
set it as both transcription and sound_detection and turn on
pipeline.diarization, and each committed turn gets speaker segments
and sound tags from one parakeet-cpp backend.

Files and sha256 match the Hub and are shared with the existing TDT v3,
diarization and CED-Tiny entries. A real-model spec checks the
combination on a clip with two speakers and a rooster: A-B-A-B speaker
turns, and "Chicken, rooster" among the sound tags. The test loader
now binds the sound entry points like main.go.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add CED-Base variants of the parakeet-cpp scene models

parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship
with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds
more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for
Tiny).

Measured on CPU over a 37 s clip: the live diarization + sound stream
runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny,
because diarization dominates; sound detection per committed turn costs
0.031 against 0.005. The realtime docs list both and note that any CED
size works as sound_model.

Files and sha256 match the Hub and are shared with the existing
parakeet-cpp-ced-base entry. The TDT variant passes the real-model
scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found).

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): register pipeline.diarization in the config metadata

TestAllFieldsHaveRegistryEntries fails on the branch because the new
pipeline.diarization field has no registry entry. Add one so the model
editor shows it as a toggle next to the sound detection options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants