feat(parakeet-cpp): speaker diarization, sound detection and live scene events - #12335
Merged
Merged
Conversation
localai-org-maint-bot
force-pushed
the
feat/parakeet-diarization-sound
branch
from
September 29, 2026 03:04
a6bfba1 to
e972a1b
Compare
Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound event and combined scene stream C symbols through the same purego.Dlsym probe pattern already used for the batched JSON entry point, so the backend still loads against an older libparakeet.so. Load now classifies the loaded GGUF by role (ASR, diarization or sound) via parakeet_capi_model_kind and can load up to two companion models from Options[] (asr_model:, diarization_model:, sound_model:, paths resolved against opts.ModelPath), verifying each companion's kind and freeing every context opened so far on any failure. Free releases the primary and every companion. AudioTranscription now names the loaded role when it is not ASR instead of a generic model not loaded error. The dynamic batcher starts only when an ASR context ends up loaded, primary or companion. This is groundwork only: the Diarize and SoundDetection RPCs and the live scene stream that actually use these new roles land in later commits. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
loadRoles' freeLoaded only released the C contexts it had opened; it left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed contexts, so a later Free() on the same instance would double-free. Zero all four alongside the CppFree calls. Also route AudioTranscriptionStream and AudioTranscriptionLive through notASRError when ctxPtr is unset but a diarization or sound model is loaded, matching AudioTranscription: both used to return the generic model-not-loaded error instead of naming the loaded role. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Implement the Diarize RPC for the parakeet-cpp Go backend, wired to Nemotron-3-Diarization through libparakeet.so's diarization C-API. Plain diarization uses parakeet_capi_diarize_pcm; when include_text is set and an ASR companion is loaded, parakeet_capi_transcribe_and_ diarize_json fills each segment's text instead. Speaker labels are the decimal index, or "unknown" for -1 (no diarized speaker overlaps). min_duration_off merges same-speaker segments across a short gap before min_duration_on drops the segments still too short, then ids are renumbered. num_speakers/min_speakers/max_speakers/clustering_ threshold have no Sortformer equivalent and are logged at debug instead of rejected. Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc- 110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B speaker segmentation and matching speaker-attributed transcripts. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Wire the SoundDetection RPC to the CED tagger context (p.tagCtx) loaded by Task 1's role classification. It runs the whole clip through a one-shot parakeet_capi_sound_stream_* session (window 10s, hop 10s, top_k set to the tagger's class count so every drained window carries a full score list), averages each class's score across the drained windows, sorts descending, then applies the request's threshold and top_k (0 keeps every class). No tagCtx returns FailedPrecondition; a libparakeet.so missing the sound_stream symbols returns Unimplemented. Every C call runs under engineMu, and the stream is always freed, even when a feed or drain call fails partway through. Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo clip: "Chicken, rooster" tops the list at score 0.91. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
SoundDetection now checks ctx before each 10 s feed slice (mirroring driver.go's feedSlices) and returns Canceled if the caller gave up, so a long clip can be interrupted instead of feeding to completion regardless. The stream is still freed on every path, cancellation included. Also narrow engineMu to the C calls: the drained JSON document is now decoded after the lock is released, splitting soundStreamScores into a locked soundStreamDrain (opts, begin, feed, drain, free) and an unlocked json.Unmarshal. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…cription Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent, repeated on TranscriptLiveResponse. When a diarization or sound companion model is loaded, AudioTranscriptionLive now runs a no-ASR scene stream (parakeet_capi_scene_stream_begin) beside the ASR streaming session, feeding it the same PCM slices and forwarding any closed speaker or sound events alongside the matching ASR delta, or on their own when a slice has no ASR output. The scene stream is freed and reopened on a mid-stream Config reset, flushed with is_last before the closing FinalResult, and degrades gracefully (a warning, not an error) when begin or a later feed call fails, so live transcription keeps working ASR-only. Existing live behavior is unchanged when no companion is configured, and no scene C call is made in that case. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Emit each slice's ASR result right after the ASR feed, before the scene feed for that slice runs, so a companion diarization/sound model never adds scene compute latency in front of the delta or <EOU> that drives realtime turn detection. Closed speakers/sounds go out afterward as their own response, so a slice with both now produces two responses, ASR first. The live feed log line now reports ASR and scene wall time separately. Re-check the diarization/sound contexts a scene stream was begun with against the live contexts before every feed, under the same lock: Free() can race between an ASR feed and the matching scene feed and free the model the stream borrows. A mismatch now returns without touching the C side. Freeing the stream itself stays unconditional; the scene stream's destructor only releases its own buffers and never touches the borrowed contexts. Also recover a panicking stub inside the live test goroutine instead of crashing the test binary, and reset the live decode-lag tracker on a mid-stream config reset, matching what its own comment already promised. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Carry the backend's closed speaker segments and sound events (TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to seconds), and forward them from the semantic_vad live path. Each speaker segment emits conversation.item.input_audio_transcription.segment with speaker, start, end and empty text under the turn's item id. Each sound event emits conversation.item.sound_detection with one tag (label, score = peak, index) and the event's new optional start/end seconds fields, omitted when unset so the existing unary/windowed sound-detection path is unaffected. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used omitempty, so a speaker segment starting at 0.0s dropped its "start" key. Nothing emitted this event before the live scene-event path, so drop omitempty: the segment always carries real times. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…models Add gallery entries for the new parakeet-cpp capabilities: standalone Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC 110M ASR model for speaker-attributed text, CED-Tiny and CED-Base sound classifiers, and a realtime scene bundle combining the streaming EOU ASR model with diarization and sound companions. SHA256 taken from the Hub API; licenses from each model card (openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED, cc-by-4.0 for the Parakeet ASR models). Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…ne events Cover the new parakeet-cpp capabilities across the feature pages: Nemotron-3-Diarization as a diarization backend (with and without speaker text, the ignored speaker-count hints, the Sortformer voice-like-sound quirk), CED as a sound classification backend, the asr_model/diarization_model/sound_model/diarization_latency companion options, and the realtime live speaker/sound events (event shapes, the speech-turn-only limitation, and using this or pipeline.sound_detection but not both). Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA open model license instead. Switch to the gallery's usual spelling for that license and keep the diarization/CED licenses called out in the description. Also: audio-diarization.md now says getting per-segment text needs both an asr_model companion and include_text=true on the request, and audio-to-text.md's option table reads "Use on" (a pairing the loader does not enforce) instead of "Allowed on". Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
loadRoles let a companion option (asr_model:/diarization_model:/ sound_model:) assign into a role field the primary already occupied, for example asr_model: on an already-ASR primary. The companion's context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only walks those three fields, so the original primary context was never freed again. Reject a companion whose role the primary already holds before its GGUF is even loaded, freeing everything loadRoles opened so far, the same way a wrong-kind companion is already rejected. Also warn, rather than silently fall through, when parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a successfully loaded primary; the primary is still treated as ASR, matching today's behavior. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
sceneBegin started the live diarization/sound companion stream with the C API's default sound options, whose top_k keeps 5 scores per window forever until drained. The live scene path never drains sound scores (only the offline SoundDetection RPC does, with its own fresh stream), so this window queue on the C side grew for the whole session's lifetime. Set opts.Sound.TopK = 0 before starting the scene stream: this disables score retention while leaving sound event detection (onset/ offset), which the live path actually consumes, unaffected. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…rize mergeCloseSegments only compared neighbors in the single start-sorted segment list, so two same-speaker segments never merged once another speaker's turn fell between them (A, B, A): the short B segment broke the adjacency the merge relied on. Group segments by speaker first, merge within each speaker's own start-ordered run, then re-sort the result by start so interleaved speakers come back out in timeline order. Also harden Diarize's entry points the same way streamFeedDoc/ sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the include_text path, p.ctxPtr) under engineMu right before the C call, so a Free() racing between Diarize's own checks and the lock can no longer reach the C side with a freed context. When the include_text call returns NULL, last_error is now read from both contexts and whichever came back non-empty is reported, since either side of the pairing can be the one that failed. A WAV decode failure is reported as InvalidArgument instead of an unwrapped/untyped error. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
soundStreamDrain ran every C call under engineMu but never re-checked p.tagCtx there, so a Free() racing between SoundDetection's own tagCtx==0 check and this lock could still reach the C side with a freed context. Re-check p.tagCtx under the lock and return ModelNotLoaded when it was cleared, mirroring diarizeCall's own re-check. A WAV decode failure is now reported as InvalidArgument instead of an unwrapped/untyped error. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
feedSlicesScene already degrades gracefully when a scene feed call fails mid-session: it frees the broken stream and carries the ASR-only session forward. Add a spec covering that path end to end: the scene stream is freed exactly once, later audio slices still produce ASR responses, and no speaker/sound events appear before or after the failure. Assisted-by: Claude:claude-sonnet-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.
openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The scene feed comment and the live test's canned document gave "Chicken, rooster" index 365. In CED's AudioSet label list it is 99, which is also what the realtime docs show. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The scene-event tests range over a method instead of its returned slice. Call the synchronized accessor so the OpenAI test package compiles. Assisted-by: Codex:gpt-6 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and #74 (the missing <algorithm> include that broke the image builds) are on master now. Pin 6dea76a instead of the #75 PR head, and update the header comment the bump bot reads. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
parakeet.cpp #76 moved its ced.cpp submodule from the head of localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where #3 landed with an identical tree. Pin 623a968 so the image builds no longer depend on that branch. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A diarizing backend could label transcript segments, but two paths dropped the label: TranscriptWord had no speaker field, so live transcription words and word-level timestamps could not carry one, and the stream=true transcript.text.done event left the speaker out of its segments. TranscriptWord gains an optional speaker (proto field 4, additive). It flows through the live event and result mapping, the JSON word output of the endpoint and the CLI, and transcript.text.done now includes a segment's speaker when there is one. Empty labels are omitted, so responses without diarization are unchanged. Assisted-by: Claude:claude-opus-5-5 [Claude Code] (cherry picked from commit 2f0049f) Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The Nemotron-3-Diarization GGUFs are published in mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The parakeet-cpp importer did not recognise that name, so a direct `local-ai models import` of the file fell through to another importer. A direct URL to the file now imports with the diarization usecase. A repo import still picks ASR weights when the repo also ships the diarization model, and falls back to the diarization weights only when there are no others. Ported from #12323. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The capability table listed parakeet-cpp as transcription only, though the backend now answers Diarize (Nemotron-3-Diarization) and SoundDetection (CED) depending on the model kind it loads. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…mpanion A diarization_model companion only fed live speaker events and Diarize; /v1/audio/transcriptions ignored it. With the companion attached and diarize=true (the OpenAI endpoint's default), unary transcription now labels each segment with its speaker and splits segments at speaker turns; with word timestamps each word carries its speaker. The stream=true final result labels each utterance with the speaker who said most of it. Both use the checkpoint's own diarization over the whole clip, as NeMo's diarize() does. Words take the speaker whose segments overlap them most, or the nearest segment within 0.5 s, the same rule as parakeet.cpp's speaker-attributed ASR. Docs: the diarization_model row and a paragraph on transcript speakers; Nemotron-3-Diarization handles up to 8 speakers. Ported from #12323. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Speaker events reached a realtime session only from the live semantic_vad path, which needs a cache-aware streaming transcription model. Committed-turn transcription (server_vad, or any offline model) always asked the backend for diarize=false and dropped the segments' speakers. pipeline.diarization (off by default) asks the transcription model for speaker labels on each committed turn and emits every labelled segment as a conversation.item.input_audio_transcription.segment event, with its text, before the turn's completed event. It is opt-in because some backends fail a diarization request they cannot serve. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
parakeet-cpp-realtime-scene pairs the streaming EOU model with the diarization and CED companions; its speaker and sound events need a cache-aware streaming model. This entry does the same with Parakeet TDT 0.6B v3 (multilingual, offline) for realtime under server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments and sound tags from one parakeet-cpp backend. Files and sha256 match the Hub and are shared with the existing TDT v3, diarization and CED-Tiny entries. A real-model spec checks the combination on a clip with two speakers and a rooster: A-B-A-B speaker turns, and "Chicken, rooster" among the sound tags. The test loader now binds the sound entry points like main.go. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for Tiny). Measured on CPU over a 37 s clip: the live diarization + sound stream runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny, because diarization dominates; sound detection per committed turn costs 0.031 against 0.005. The realtime docs list both and note that any CED size works as sound_model. Files and sha256 match the Hub and are shared with the existing parakeet-cpp-ced-base entry. The TDT variant passes the real-model scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found). Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
TestAllFieldsHaveRegistryEntries fails on the branch because the new pipeline.diarization field has no registry entry. Add one so the model editor shows it as a toggle next to the sound detection options. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
mudler
force-pushed
the
feat/parakeet-diarization-sound
branch
from
September 29, 2026 20:50
6a4ac25 to
7b91c7f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The parakeet-cpp backend now serves the speaker diarization, sound-event detection and live "scene" events that parakeet.cpp gained in mudler/parakeet.cpp#75 (on top of the Nemotron-3-Diarization work already on master).
parakeet_capi_model_kind. Companion models come from options, paths relative to the models dir:asr_model:on a diarization model, for speaker-attributed text;diarization_model:andsound_model:on an ASR model, for live speaker and sound events;diarization_latency:<model|low|very_low|ultra_low>(defaultlow).Diarize(Sortformer, up to 8 speakers). Withinclude_text=trueand anasr_modelcompanion, segments carry text.min_duration_on/min_duration_offapply; speaker-count hints are ignored (Sortformer has no clustering stage).SoundDetectionwith a CED GGUF: 10 s windows averaged per class,top_k0 = all (as the proto documents),thresholddrops lower scores.AudioTranscriptionLive, a no-ASR scene stream runs beside the ASR stream when companions are configured. Closed speaker segments and sound events go out in two new, additiveTranscriptLiveResponsefields (7 and 8). ASR deltas and EOU are sent first and are unchanged.conversation.item.input_audio_transcription.segment(the event existed but was never sent) and sound events emitconversation.item.sound_detectionwith one tag and new optionalstart/end. Both are per speech turn: times are relative to the turn's audio, and speaker labels are only consistent within a turn.diarization_model:companion,/v1/audio/transcriptionslabels each segment with itsspeakerand splits segments at speaker turns. With word timestamps, each word carries its speaker;TranscriptWordgains an optionalspeaker(proto field 4, additive). Withstream=true, thetranscript.text.donesegments carry it.diarize=falseskips it per request. Words take the speaker whose segments overlap them most, or the nearest segment within 0.5 s, which is parakeet.cpp's own merge rule.nemotron-3-diarization-*.ggufimports as a parakeet-cpp diarization model. Repo imports still pick ASR weights.DiarizeandSoundDetection(usecasesdiarization,sound_classification) next to transcription.pipeline.diarization, off by default). Committed-turn transcription (server_vad, any offline model) now asks for speakers. Each labelled segment is emitted asconversation.item.input_audio_transcription.segment, with text, before the turn'scompletedevent. Previously only the livesemantic_vadpath emitted speaker events, and that needs a streaming model.parakeet-cpp-realtime-scene-tdt: Parakeet TDT 0.6B v3 (multilingual) plus Nemotron-3-Diarization plus CED-Tiny in one parakeet-cpp model. Set it as bothtranscriptionandsound_detectionwithdiarization: true, and each committed turn gets speaker segments and sound tags. On a real clip with two speakers and a rooster it gives A-B-A-B speaker turns and "Chicken, rooster" among the sound tags; that check is a real-model spec.parakeet-cpp-realtime-scene-baseandparakeet-cpp-realtime-scene-tdt-baseare the scene models with CED-Base (86M) in place of CED-Tiny. Measured on CPU over a 37 s clip:parakeet-cpp-nemotron-3-diarization,parakeet-cpp-nemotron-3-diarization-asr,parakeet-cpp-ced-tiny,parakeet-cpp-ced-base,parakeet-cpp-realtime-scene.Every new C symbol is Dlsym-probed, so the backend still loads an older
libparakeet.so(the new RPCs then return Unimplemented and companion options are rejected with a clear message).Why
This lets one parakeet-cpp backend do speech-to-text, "who spoke when" and "what sound was that" from the same models LocalAI already ships, and puts speakers and sounds on the realtime websocket next to the live transcript.
How it was verified
go test ./... -count=1inbackend/go/parakeet-cpp(122 specs, stubbed C calls),go test ./core/backend/... ./core/http/endpoints/openai/...,go vet(only the existingunsafe.Pointernote atgoparakeetcpp.go:928).libparakeet.sofrom parakeet.cpp#75:Diarizeon a two-speaker LibriSpeech clip gives the right A-B-A-B turns, with text when paired with tdt_ctc-110m.SoundDetectionwith ced-tiny on a rooster clip gives "Chicken, rooster" 0.91 first.low) and ced-tiny the session took 13.7 s, the scene part about 4.7 s (roughly 0.13 of real time).Docker e2e, not run here:
make test-extra-backend-parakeet-cpp-transcription. The e2e harness has no diarization or sound capability yet.libparakeet.sobuilt at the pin (623a968, fetched as the Makefile does) and real models: all 129 backend specs pass with none skipped. That includes a new real-model spec, where unary transcription with adiarization_modelcompanion gives A-B-A-B speakers on the two-speaker fixture and none withoutdiarize.core/backend,core/schema,core/config,core/http/endpoints/openaiandcore/gallery/importerstests pass, and golangci-lint reports 0 issues.Before merge
PARAKEET_VERSIONis parakeet.cpp master623a968: Dependency Dashboard #75 (sound events), fix(deps): update github.com/go-skynet/go-gpt4all-j.cpp digest to 1f7bff5 #74 (the missing<algorithm>include that broke every image build on the old pin) and feat: automatic updates with renovate, docs updates #76 (ced.cpp submodule on ced.cpp main).make test-extra-backend-parakeet-cpp-transcription(the failing CI job) passes locally: the Docker image builds and the gRPC specs pass (4 passed, 22 skipped, 0 failed).core/backend,core/configand the realtime package tests pass.Signed-off-by(AI commits carryAssisted-byonly).Known nits and follow-ups
parakeet-cpp-realtime_eou_120m-v1gallery entry sayscc-by-4.0, but the model card uses the NVIDIA Open Model License (the new entry uses the right one).🤖 Generated with Claude Code