Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
c17810f
feat(sound): load CED models through a ced.cpp submodule
mudler Sep 28, 2026
4eedac2
fix(sound): drop the ced/parakeet circular link, isolate dr_wav in it…
mudler Sep 28, 2026
1282193
feat(sound): add a sliding-window sound-event stream
mudler Sep 28, 2026
5a6d5a2
fix(sound): make safe_until() account for the is_last tail window
mudler Sep 28, 2026
1d26da6
feat(sound): expose the sound stream in the C-API (ABI v8)
mudler Sep 28, 2026
304b2f6
fix(sound): guard sound_stream_active against exceptions
mudler Sep 28, 2026
48e01be
docs(sound): measure short CED windows and document the sound stream
mudler Sep 28, 2026
f77cf51
refactor(sas): move the commit logic into AsrCommitter and SceneStream
mudler Sep 28, 2026
c272566
fix(sas): end the stream once the diarizer takes the last chunk
mudler Sep 28, 2026
dc6ee4c
feat(sound): add the combined scene stream to the C-API
mudler Sep 28, 2026
6a332bb
fix(sound): harden the scene stream's docs, tests and C boundary
mudler Sep 28, 2026
d1c5e4b
feat(cli): add a scene subcommand with sound events
mudler Sep 28, 2026
fffcdd5
fix(cli): hide CED's Speech synthesizer label and harden scene args
mudler Sep 28, 2026
10c92ad
feat(server): add sound events to verbose_json with --sound-model
mudler Sep 28, 2026
568f626
docs(sound): document sound events, scene stream and ABI v8
mudler Sep 28, 2026
f6102a2
fix(sas): release audio that holds no committable word
mudler Sep 28, 2026
6211543
fix(cli): print speaker segments in diarization-only scenes
mudler Sep 28, 2026
92a213c
fix(capi): validate scene latency and tidy sound feed errors
mudler Sep 28, 2026
6219643
test(sound): skip without CED and rewrite comments as rationale
mudler Sep 28, 2026
605b5b7
build: fail early without the ced.cpp submodule, check noced CLI
mudler Sep 28, 2026
65d170c
docs(sound): correct boundary, GPU teardown and link details
mudler Sep 28, 2026
4ac35ad
fix(sas): back off from the first word start when nothing commits
mudler Sep 28, 2026
3f41a13
docs(sound): fix a stale note on the GPU runs
mudler Sep 28, 2026
c018537
feat(capi): report which kind of model a context holds
mudler Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,21 @@ jobs:
# job once a models bundle is published (Phase 4).
run: ctest --test-dir build --output-on-failure -LE model

- name: build without ced (PARAKEET_WITH_CED=OFF)
# PARAKEET_WITH_CED is on by default, so this is the gate that catches
# anything that quietly starts depending on ced.cpp being present.
run: |
cmake -B build-noced -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=OFF -DPARAKEET_WITH_CED=OFF
cmake --build build-noced -j
ctest --test-dir build-noced --output-on-failure -LE model
# scene --sound must fail cleanly (exit 2) with a clear message.
# Capture first: steps run with -e -o pipefail.
rc=0
out=$(build-noced/examples/cli/parakeet-cli scene --sound x --input y 2>&1) || rc=$?
echo "$out"
test "$rc" -eq 2
grep -q "built without sound tagging" <<< "$out"

# -------------------------------------------------------------------------
# server-e2e: drive the real parakeet-server over HTTP.
#
Expand Down
3 changes: 3 additions & 0 deletions .gitmodules
Original file line number Diff line number Diff line change
@@ -1,3 +1,6 @@
[submodule "third_party/ggml"]
path = third_party/ggml
url = https://github.com/ggml-org/ggml
[submodule "third_party/ced.cpp"]
path = third_party/ced.cpp
url = https://github.com/localai-org/ced.cpp
47 changes: 46 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,14 @@ src/ libparakeet implementation
diarization_encoder/head.*, RoPE Transformer encoder + speaker head
diarization_streaming.*, NeMo cache-aware streaming diarization (speaker cache + FIFO)
sas_merge.hpp/cpp , ASR words x speaker segments -> speaker-attributed utterances
asr_committer.hpp/cpp, shared "finalize a word/utterance once" logic used by SAS and SceneStream
ced_tagger.hpp/cpp , pk::CedTagger: loads a CED GGUF (ced.cpp) into a tagger context, pk::SoundScorer
sound_stream.hpp/cpp, pk::SoundStream: sliding-window sound-event detection over live PCM
scene_stream.hpp/cpp, pk::SceneStream: combined ASR + diarization + sound-event stream
scene_render.hpp/cpp, pk::SceneRenderer + format_span/is_speech_label: `parakeet-cli scene` text rendering
examples/cli/ parakeet-cli binary
subcommands: info, transcribe (+ --stream), quantize
subcommands: info, transcribe (+ --stream), quantize, scene (ASR + diar + sound, one time-ordered feed)
sound-window-eval: measures CED short-window accuracy vs whole-clip top-1
diarize binary: diarize <diar.gguf> <wav> [--stream]
scripts/ Python tooling
convert_parakeet_to_gguf.py, .nemo/.hf -> GGUF (--dtype f32|f16|q8_0)
Expand Down Expand Up @@ -111,13 +117,23 @@ tests/ ctest targets
test_streaming_diarization.cpp, streaming diarization == NeMo streaming, every latency mode (same baseline)
test_combined_offline.cpp, SAS + streaming diarization/SAS through the C-API
test_sas_merge.cpp , SAS merge/grouping (model-independent)
test_asr_committer.cpp , shared word/utterance finalize logic (model-independent)
test_ced_parity.cpp , CedTagger scores == ced.cpp PyTorch baseline (PARAKEET_TEST_CED_GGUF f32 + PARAKEET_TEST_CED_BASELINE)
test_sound_stream.cpp , pk::SoundStream windowing/on-off-min_duration logic (model-independent)
test_sound_capi.cpp , sound_stream_* C-API (PARAKEET_TEST_CED_GGUF)
test_scene_stream.cpp , pk::SceneStream / scene_stream_* C-API, all three models together (PARAKEET_TEST_GGUF + PARAKEET_TEST_DIAR_GGUF + PARAKEET_TEST_CED_GGUF)
test_scene_render.cpp , SceneRenderer / format_span / is_speech_label (model-independent)
python/check_convert.py , converter round-trip (model-dependent)
python/check_baseline.py, baseline dumper (model-dependent)
fixtures/clip.wav , 2 s 16 kHz mono WAV for stage parity tests
fixtures/speech.wav , LibriSpeech 2086-149220-0033, ~7.4 s
fixtures/two_speakers.wav, LibriSpeech 1272 + 2086 alternating A-B-A-B, 23.6 s
third_party/ vendored deps
ggml/ , submodule pinned at v0.13.0
ced.cpp/ , submodule, CED sound-event tagger (PARAKEET_WITH_CED, on by default);
built as a static `ced` target linked into libparakeet, not a separate
process; dr_wav is shared via CED_EXTERNAL_DR_WAV so there is one
DR_WAV_IMPLEMENTATION in the whole build
dr_wav.h , vendored single header
models/ output dir for converted GGUFs (gitignored;
MANIFEST.md tracks the expected published set)
Expand Down Expand Up @@ -147,6 +163,7 @@ cmake -B build -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=ON && cmake --build build
| `PARAKEET_GGML_METAL` | OFF | Forward GGML_METAL to the submodule |
| `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule |
| `PARAKEET_GGML_HIPBLAS` | OFF | Forward GGML_HIPBLAS to the submodule |
| `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp |

Use `-DGGML_NATIVE=OFF` when building for CI or portable binaries.

Expand Down Expand Up @@ -238,6 +255,7 @@ The binary is at `build/examples/cli/parakeet-cli`.
parakeet-cli info <model.gguf>
parakeet-cli transcribe --model <model.gguf> --input <audio.wav> [--decoder ctc|tdt] [--stream] [--timestamps] [--json]
parakeet-cli quantize <in.gguf> <out.gguf> <type>
parakeet-cli scene [--model <asr.gguf>] [--diar <diar.gguf>] [--sound <ced.gguf>] --input <audio.wav> [--latency model|low|very_low|ultra_low] [--chunk-ms N] [--show-speech] [--json]
```

`--timestamps` prints one `<start>-<end> <word> (<conf>)` line per word (also
Expand Down Expand Up @@ -282,6 +300,33 @@ parakeet_capi_free_diar_segments
parakeet_capi_sas_stream_begin / _begin_latency / _feed / _free
```

Sound-event detection (ABI v8, additive; not used by LocalAI yet). A CED GGUF
(ced.cpp) loads into its own `parakeet_ctx` kind (a "tagger") through the same
`parakeet_capi_load`; see `docs/sound.md`:

```
parakeet_capi_sound_opts_default
parakeet_capi_sound_stream_begin / _feed / _active / _drain_scores_json / _free
parakeet_capi_free_sound_segments
parakeet_capi_num_classes
parakeet_capi_class_label
parakeet_capi_model_kind # which kind of ctx (NONE/ASR/DIARIZATION/SOUND)
```

Combined scene stream (ABI v8, additive; not used by LocalAI yet). One stream
that carries any mix of an ASR context, a diarization context, and a tagger
context, and emits speaker-attributed words/utterances plus sound-event
segments in one time-ordered JSON document per feed; see `docs/sound.md`:

```
parakeet_capi_scene_opts_default
parakeet_capi_scene_stream_begin
parakeet_capi_scene_stream_feed_json
parakeet_capi_scene_stream_drain_scores_json
parakeet_capi_scene_stream_last_error
parakeet_capi_scene_stream_free
```

`parakeet_capi_transcribe_path_json(ctx, wav, decoder)` returns malloc'd UTF-8
JSON `{"text":..,"words":[{"w","start","end","conf"}],"tokens":[{"id","t","conf"}]}`
(times in seconds, conf in `(0,1]`), built from
Expand Down
42 changes: 41 additions & 1 deletion CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ option(PARAKEET_GGML_CUDA "Forward GGML_CUDA" OFF)
option(PARAKEET_GGML_METAL "Forward GGML_METAL" OFF)
option(PARAKEET_GGML_VULKAN "Forward GGML_VULKAN" OFF)
option(PARAKEET_GGML_HIP "Forward GGML_HIP (ROCm)" OFF)
option(PARAKEET_WITH_CED "Sound-event detection through ced.cpp" ON)

set(GGML_CUDA ${PARAKEET_GGML_CUDA} CACHE BOOL "" FORCE)
set(GGML_METAL ${PARAKEET_GGML_METAL} CACHE BOOL "" FORCE)
Expand Down Expand Up @@ -63,6 +64,33 @@ endif()

add_subdirectory(third_party/ggml)

# Single dr_wav implementation for the whole build. src/audio_io.cpp only
# declares dr_wav's functions (#include "dr_wav.h", no DR_WAV_IMPLEMENTATION);
# this object library is the one translation unit that defines them, so both
# libparakeet and, when PARAKEET_WITH_CED is on, ced.cpp (built with
# CED_EXTERNAL_DR_WAV, which likewise only declares them) link against the
# same symbols instead of each pulling in its own copy (a second
# DR_WAV_IMPLEMENTATION define would be a multiple-definition link error).
# An OBJECT library's objects are linked directly into each consumer rather
# than referenced as a separate archive dependency, so this does not create a
# link-time cycle between the parakeet and ced targets.
add_library(dr_wav_impl OBJECT src/dr_wav_impl.cpp)
target_include_directories(dr_wav_impl PRIVATE ${CMAKE_SOURCE_DIR}/third_party)
set_target_properties(dr_wav_impl PROPERTIES POSITION_INDEPENDENT_CODE ON)

if(PARAKEET_WITH_CED)
if(NOT EXISTS "${CMAKE_CURRENT_SOURCE_DIR}/third_party/ced.cpp/CMakeLists.txt")
message(FATAL_ERROR "third_party/ced.cpp is missing: run `git submodule update --init --recursive`, or configure with -DPARAKEET_WITH_CED=OFF")
endif()
set(CED_BUILD_CLI OFF CACHE BOOL "" FORCE)
set(CED_BUILD_TESTS OFF CACHE BOOL "" FORCE)
set(CED_SHARED OFF CACHE BOOL "" FORCE)
set(CED_EXTERNAL_DR_WAV ON CACHE BOOL "" FORCE) # dr_wav_impl provides it
add_subdirectory(third_party/ced.cpp EXCLUDE_FROM_ALL)
set_target_properties(ced PROPERTIES POSITION_INDEPENDENT_CODE ON)
target_link_libraries(ced PRIVATE dr_wav_impl)
endif()

set(PARAKEET_SRC
src/parakeet.cpp
src/model.cpp
Expand Down Expand Up @@ -97,7 +125,13 @@ set(PARAKEET_SRC
src/diarization_encoder.cpp
src/diarization_head.cpp
src/sas_merge.cpp
src/diarization_streaming.cpp)
src/asr_committer.cpp
src/diar_pcm_stream.cpp
src/scene_stream.cpp
src/diarization_streaming.cpp
src/ced_tagger.cpp
src/sound_stream.cpp
src/scene_render.cpp)

if(PARAKEET_SHARED)
add_library(parakeet SHARED ${PARAKEET_SRC})
Expand All @@ -113,6 +147,12 @@ target_include_directories(parakeet PUBLIC include PRIVATE src ${CMAKE_SOURCE_DI
target_compile_definitions(parakeet PUBLIC $<$<CXX_COMPILER_ID:MSVC>:_USE_MATH_DEFINES>)
target_compile_definitions(parakeet PRIVATE PARAKEET_VERSION="${PARAKEET_VERSION}")
target_link_libraries(parakeet PUBLIC ggml)
target_link_libraries(parakeet PRIVATE dr_wav_impl)

if(PARAKEET_WITH_CED)
target_link_libraries(parakeet PRIVATE ced)
target_compile_definitions(parakeet PRIVATE PARAKEET_WITH_CED=1)
endif()

if(PARAKEET_BUILD_CLI)
add_subdirectory(examples/cli)
Expand Down
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,7 @@ cmake --build build-shared -j
| `PARAKEET_GGML_METAL` | OFF | Forward GGML_METAL to the submodule |
| `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule |
| `PARAKEET_GGML_HIP` | OFF | Forward GGML_HIP (ROCm) to the submodule |
| `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp |

To build for a GPU backend, forward its flag, e.g. Apple Metal:

Expand Down Expand Up @@ -317,6 +318,27 @@ To batch from code, use the batched entry points (single-clip B=1 is just N=1):

---

## Sound events

parakeet.cpp can also tag everyday sounds (dog bark, glass breaking, applause,
alarms, music, and the rest of the 527-class AudioSet ontology) with
[CED](https://github.com/RicherMans/CED), through the
[ced.cpp](https://github.com/localai-org/ced.cpp) submodule (`PARAKEET_WITH_CED`,
on by default). `parakeet-cli scene` combines it with ASR and diarization into
one time-ordered feed:

```sh
parakeet-cli scene --model asr.gguf --diar diar.gguf --sound ced-base-q8_0.gguf \
--latency low --input audio.wav
[00:00.4 - 00:03.2] Speaker 0: mister Quilter is the apostle of the middle classes, and
[00:24.0 - 00:30.0] (Chicken, rooster 0.86)
```

See [`docs/sound.md`](docs/sound.md) for the CED GGUFs, the sound and scene
stream C-API (ABI v8), and the `--sound-model` server option.

---

## C-API (`libparakeet.so`)

`include/parakeet_capi.h` defines a flat, exception-free C-API meant for `dlopen` / FFI / LocalAI integration. Build the shared library with `-DPARAKEET_SHARED=ON`:
Expand Down
Loading
Loading