Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
ae599e1
feat(convert): import HF safetensors Parakeet TDT checkpoints
mudler Sep 29, 2026
b9105f7
feat(convert): keep ternary weights packed and carry the VAD head
mudler Sep 29, 2026
c70d419
feat(loader): read the ternary and VAD flags from the GGUF
mudler Sep 29, 2026
bd1ff06
feat(ternary): repack, int8 activation quantization and scalar reference
mudler Sep 29, 2026
e54e0c5
feat(ternary): run packed ternary encoder linears through a ggml op pair
mudler Sep 29, 2026
bc8775e
fix(ternary): validate and repack packed weights at load
mudler Sep 29, 2026
4bf3e96
feat(ternary): AVX-512 VNNI and AVX2 kernels with runtime dispatch
mudler Sep 29, 2026
f2eb381
feat(ternary): NEON dot-product kernel
mudler Sep 29, 2026
c8499e6
fix(ternary): move NEON -ffp-contract=off flag inside NOT MSVC guard
mudler Sep 29, 2026
9d151fb
docs: ternary Redux support, measurements and converter
mudler Sep 29, 2026
f7c9e8c
bench: compare ternary kernels with ggml Q8_0 and F16 mul_mat
mudler Sep 29, 2026
b4164da
docs: fix ternary numbers and wording after review
mudler Sep 29, 2026
7c8cc48
ternary: one output row per vector lane in the kernels
mudler Sep 29, 2026
93028b6
bench: run the scalar ternary reference with fewer repetitions
mudler Sep 29, 2026
6136e9b
ternary: tune the x86 tile shapes
mudler Sep 29, 2026
f228376
ternary: vectorize the activation quantization on x86
mudler Sep 29, 2026
8b66c63
ternary: build the AVX-512 quantizer for avx512f only
mudler Sep 29, 2026
8b18b08
docs: ternary speed after the kernel rework
mudler Sep 29, 2026
2a370e6
feat(vad): VAD head and a vad-probe tool
mudler Sep 30, 2026
ff2b86b
fix(vad): read head weights through the backend, asymmetric tests
mudler Sep 30, 2026
dcc7666
feat(vad): cut long audio at pauses
mudler Sep 30, 2026
fed1141
fix(vad_segmenter_test): correct test expectations and add rule tests
mudler Sep 30, 2026
ec0d983
feat(vad): transcribe long audio in VAD-cut segments
mudler Sep 30, 2026
31a0058
test(vad): self-contained long-clip checks, harden slicing and CLI flags
mudler Sep 30, 2026
53e7f5a
scripts: long-form clip builder and VAD WER evaluation
mudler Sep 30, 2026
daae921
docs: long-form VAD validation for Ultra and Redux
mudler Sep 30, 2026
6a7ae55
docs: correct Ultra and Redux claims, sharpen the VAD reading
mudler Sep 30, 2026
d038c7e
loader: validate packed ternary tensors before reading them
mudler Sep 30, 2026
82788d2
vad: harden the config, the segmenter and the frame offsets
mudler Sep 30, 2026
a8682ce
docs: scalar-kernel fallback, benchmark provenance, batch and two-spe…
mudler Sep 30, 2026
f545dc4
harden: scan all tensors for packed ternary, guard VAD options, check…
mudler Sep 30, 2026
ff22b4a
docs: multilingual and long-form validation for Ultra and Redux
mudler Sep 30, 2026
2a91486
fix: wire packed ternary weights into the local and chunked attention…
mudler Sep 30, 2026
29a1dae
test: packed Redux through local and chunked attention; docs for the fix
mudler Sep 30, 2026
b4aa8fe
docs: restore the Ultra parity and Commands sections dropped by the l…
mudler Sep 30, 2026
3a543ed
docs: fill packed Redux plain TED-LIUM column after the long-audio fix
mudler Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 40 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,9 @@ before, so do not change them without an A/B benchmark that proves parity.
(so the unsupported op can run on CPU); when every op is supported, the fast
gallocr path runs. If you think gallocr can go, you are about to reintroduce
that regression.
- **Ternary weights stay packed in the GGUF and are repacked once per loader; do not dequantize per call.**
The packed Redux form is what makes it 6.8x smaller than F16 and what the
CPU kernels in `src/ternary*.cpp` read.
- **Zero-copy weights.** `clone_weight` returns loader tensors directly so the
same device buffer is reused every utterance; do not copy weights per call.

Expand Down Expand Up @@ -84,13 +87,18 @@ src/ libparakeet implementation
ced_tagger.hpp/cpp , pk::CedTagger: loads a CED GGUF (ced.cpp) into a tagger context, pk::SoundScorer
sound_stream.hpp/cpp, pk::SoundStream: sliding-window sound-event detection over live PCM
scene_stream.hpp/cpp, pk::SceneStream: combined ASR + diarization + sound-event stream
ternary.hpp/cpp , packed ternary (moondream/parakeet-redux) linears: repack once, int8 activations, scalar ref + two-op ggml custom op
ternary_kernels.hpp, ternary_kernels_x86.cpp (AVX-512 VNNI / AVX2), ternary_kernels_neon.cpp
vad_head.hpp/cpp , voice-activity head of Ultra/Redux (plain C++ loops), Model::vad_probabilities
vad_segmenter.hpp/cpp, pk::segment_by_vad: cut long audio at VAD pauses (SegmenterOpts), used by `transcribe --vad`
scene_render.hpp/cpp, pk::SceneRenderer + format_span/is_speech_label: `parakeet-cli scene` text rendering
examples/cli/ parakeet-cli binary
subcommands: info, transcribe (+ --stream), quantize, scene (ASR + diar + sound, one time-ordered feed)
sound-window-eval: measures CED short-window accuracy vs whole-clip top-1
diarize binary: diarize <diar.gguf> <wav> [--stream]
scripts/ Python tooling
convert_parakeet_to_gguf.py, .nemo/.hf -> GGUF (--dtype f32|f16|q8_0)
convert_hf_parakeet_to_gguf.py, HF safetensors (moondream/parakeet-ultra, -redux) to GGUF (--template, --ternary keep|dequant, --vad keep|drop)
gen_nemo_baseline.py , NeMo intermediates -> baseline.gguf
gen_stream_baseline.py , NeMo cache-aware streaming encode+decode -> stream baseline.gguf
gen_diar_baseline.py , NeMo offline + streaming diarization -> diar baseline.gguf
Expand All @@ -117,6 +125,13 @@ tests/ ctest targets
test_streaming_diarization.cpp, streaming diarization == NeMo streaming, every latency mode (same baseline)
test_combined_offline.cpp, SAS + streaming diarization/SAS through the C-API
test_sas_merge.cpp , SAS merge/grouping (model-independent)
test_ternary.cpp , ternary repack, int8 quant, every kernel == scalar (model-independent)
test_ternary_model.cpp , packed Redux == dequantized Redux transcript (PARAKEET_TEST_GGUF_REDUX_KEEP + _DEQ)
test_model_loader_ternary.cpp, ternary + VAD flags from GGUF KVs
bench_ternary.cpp , single-thread throughput of each ternary kernel (not a ctest)
test_vad_head.cpp , VAD head probabilities (PARAKEET_TEST_GGUF_ULTRA)
test_vad_segmenter.cpp , segmenter cut rules (model-independent)
test_transcribe_vad.cpp , --vad path vs plain pass on long audio (PARAKEET_TEST_GGUF_ULTRA, PARAKEET_TEST_GGUF)
test_asr_committer.cpp , shared word/utterance finalize logic (model-independent)
test_ced_parity.cpp , CedTagger scores == ced.cpp PyTorch baseline (PARAKEET_TEST_CED_GGUF f32 + PARAKEET_TEST_CED_BASELINE)
test_sound_stream.cpp , pk::SoundStream windowing/on-off-min_duration logic (model-independent)
Expand All @@ -141,6 +156,7 @@ docs/
conversion.md , GGUF schema reference
quantization.md , quantization allowlist, policy, measured size + WER per type
parity.md , full model coverage matrix + per-stage tensor parity
ternary.md , packed ternary Redux: GGUF form, kernels, limits, measured speed
diarization.md , speaker diarization + speaker-attributed ASR: parity, C-API, speed
.github/workflows/
ci.yml , build job (per-push) + closed-loop job (pull_request + dispatch)
Expand Down Expand Up @@ -189,6 +205,10 @@ ctest --test-dir build --output-on-failure
Tests return exit code 77 (ctest SKIP) when the venv or checkpoint is absent,
so they never break a CI environment that lacks them.

Ultra/Redux tests read `PARAKEET_TEST_GGUF_ULTRA` (F16 Ultra),
`PARAKEET_TEST_GGUF_REDUX_KEEP` (packed ternary) and `PARAKEET_TEST_GGUF_REDUX_DEQ`
(dequantized Redux); they skip (77) when unset.

### Test labels

| Label | Tests | Needs |
Expand Down Expand Up @@ -221,6 +241,17 @@ Convert (HuggingFace id or local `.nemo`):
Featurizer window and filterbank are lifted from the checkpoint at runtime;
mel/fft parameters do not need to be specified manually.

## Ternary GGUF flags

Two optional GGUF flags, both read into `ParakeetConfig`:

- `parakeet.ternary.present` (with `parakeet.ternary.group_size` = 128): the
encoder linears are stored as `<name>.qweight` (I8) + `<name>.scales` (F16).
CPU only, offline only (no streaming). `PARAKEET_TERNARY_KERNEL=scalar|avx2|vnni|neon`
forces a kernel. See `docs/ternary.md`.
- `parakeet.vad.present` (with `parakeet.vad.d_in/hidden/kernel/frame_sec`): the
file carries `vad_head.*` tensors.

## Quantization policy

See `docs/quantization.md` for the full policy. Summary:
Expand Down Expand Up @@ -255,6 +286,8 @@ The binary is at `build/examples/cli/parakeet-cli`.
parakeet-cli info <model.gguf>
parakeet-cli transcribe --model <model.gguf> --input <audio.wav> [--decoder ctc|tdt] [--stream] [--timestamps] [--json]
parakeet-cli quantize <in.gguf> <out.gguf> <type>
parakeet-cli transcribe --model <ultra-or-redux.gguf> --input <long.wav> --vad [--vad-threshold F] [--vad-min-pause SEC] [--vad-max-seg SEC]
parakeet-cli vad-probe --model <m.gguf> --input <wav|-> [--variant N] # dump VAD head probabilities as t_sec,p
parakeet-cli scene [--model <asr.gguf>] [--diar <diar.gguf>] [--sound <ced.gguf>] --input <audio.wav> [--latency model|low|very_low|ultra_low] [--chunk-ms N] [--show-speech] [--json]
```

Expand Down Expand Up @@ -288,6 +321,13 @@ parakeet_capi_stream_finalize # flush the end-of-stream tail
parakeet_capi_stream_free
```

VAD segmentation (additive, ABI unchanged; not used by LocalAI yet). Needs a GGUF
with a VAD head (Ultra/Redux); see `docs/ternary.md`:

```
parakeet_capi_transcribe_path_json_vad # same JSON as _json, long audio cut at VAD pauses
```

Speaker diarization (ABI v7, additive; not used by LocalAI yet). A
diarization GGUF loads into its own `parakeet_ctx`; see `docs/diarization.md`:

Expand Down
12 changes: 12 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,8 @@ endif()
set(PARAKEET_SRC
src/parakeet.cpp
src/model.cpp
src/vad_head.cpp
src/vad_segmenter.cpp
src/parakeet_capi.cpp
src/common.cpp
src/audio_io.cpp
Expand All @@ -115,6 +117,9 @@ set(PARAKEET_SRC
src/prediction.cpp
src/joint.cpp
src/prompt_kernel.cpp
src/ternary.cpp
src/ternary_kernels_x86.cpp
src/ternary_kernels_neon.cpp
src/tdt.cpp
src/rnnt.cpp
src/transducer_batch.cpp
Expand Down Expand Up @@ -143,6 +148,13 @@ if(PARAKEET_SHARED)
else()
add_library(parakeet STATIC ${PARAKEET_SRC})
endif()
if(NOT MSVC)
set_source_files_properties(src/ternary.cpp src/ternary_kernels_x86.cpp src/ternary_kernels_neon.cpp
PROPERTIES COMPILE_OPTIONS "-ffp-contract=off")
endif()
if(CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64|arm64" AND NOT APPLE AND NOT MSVC)
set_property(SOURCE src/ternary_kernels_neon.cpp APPEND PROPERTY COMPILE_OPTIONS "-march=armv8.2-a+dotprod")
endif()
target_include_directories(parakeet PUBLIC include PRIVATE src ${CMAKE_SOURCE_DIR}/third_party)
target_compile_definitions(parakeet PUBLIC $<$<CXX_COMPILER_ID:MSVC>:_USE_MATH_DEFINES>)
target_compile_definitions(parakeet PRIVATE PARAKEET_VERSION="${PARAKEET_VERSION}")
Expand Down
24 changes: 24 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,26 @@ Every model below is validated at WER 0 against NeMo and published as GGUF (f16,
| [parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1) | RNNT, streaming | 120M | cache-aware streaming with end-of-utterance detection (`--stream`) | NVIDIA |
| [nemotron-3.5-asr-streaming-0.6b](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b) | RNNT, streaming | 0.6B | multilingual (40+ locales), prompt-conditioned, offline and cache-aware streaming, pick a language with `--lang` (default `auto`). OpenMDW-1.1 | NVIDIA |


### Moondream Ultra and Redux (not yet published)

[moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and
[moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) are Moondream's
post-trained (Ultra, F16) and ternary-encoder (Redux) derivatives of parakeet-tdt-0.6b-v3. They are
HF safetensors, converted with `scripts/convert_hf_parakeet_to_gguf.py`. They are not part of the
NeMo-validated set above: there is no NeMo baseline for them, so parity is transcript-level against
our own v3 path (see [`docs/parity.md`](docs/parity.md)), and no GGUFs are published yet.

- Redux packs the encoder as ternary weights: a 213 MB GGUF, 6.8x smaller than F16. It runs on CPU
only and offline only. On x86 with AVX-512 VNNI it reaches median RTF 75.6 per utterance on
LibriSpeech-100 (8 threads) against 46.1 for the same model in F16; on a single 180 s clip the
gain is about 10 percent; WER on the 100 LibriSpeech utterances is 1.96 percent. See
[`docs/ternary.md`](docs/ternary.md). SIMD kernels exist for x86-64 with AVX2 or AVX-512 VNNI and
aarch64 with dotprod; MSVC builds, Windows on ARM and aarch64 without dotprod use a slow scalar
kernel (about 1 GMAC/s), and the load logs a warning. The packed file also stays resident next to
the repacked planes, so memory use is more than the file size.
- Both carry a voice-activity head, used by `transcribe --vad` to cut long audio at pauses. On
synthetic long-form clips it does not change WER meaningfully.
---

## Performance
Expand Down Expand Up @@ -242,6 +262,10 @@ parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt \
# Read WAV bytes from stdin (useful with ffmpeg/curl pipelines)
ffmpeg -i input.mp3 -f wav - | parakeet-cli transcribe --model m.gguf --input -

# Long audio on Ultra/Redux: cut at VAD pauses, transcribe each piece (offline only).
# Tune with --vad-threshold F, --vad-min-pause SEC, --vad-max-seg SEC
parakeet-cli transcribe --model ultra.gguf --input long.wav --vad

# Print model metadata (arch, dims, mel params, vocab size, TDT durations)
parakeet-cli info m.gguf

Expand Down
30 changes: 30 additions & 0 deletions docs/conversion.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,36 @@ reversed relative to numpy, so the GGUF tensor shape reads `[257, 80, 1]`. The
Hann `window` buffer (`preprocessor.featurizer.window`, shape `(win_length,)`) is
exported the same way.

## HF safetensors checkpoints (`convert_hf_parakeet_to_gguf.py`)

`scripts/convert_hf_parakeet_to_gguf.py` imports checkpoints in the transformers
`ParakeetForTDT` layout that derive from `nvidia/parakeet-tdt-0.6b-v3`
(`moondream/parakeet-ultra`, `moondream/parakeet-redux`). It needs `numpy` and
`gguf` but not NeMo. The HF repos carry no mel filterbank, window or
SentencePiece vocab, so those and all KV metadata come from a v3 GGUF made by
`convert_parakeet_to_gguf.py`. The HF `config.json` is checked against that
template and every tensor shape is checked too; a mismatch aborts the run.
HF tensor names are mapped back to the verbatim NeMo names described below.

| Argument | Meaning |
|---|---|
| `--hf DIR` | local directory with `model.safetensors`, `config.json` (and `ternary.json` for Redux) |
| `--template GGUF` | v3 GGUF that supplies KVs, vocab, filterbank and window |
| `--output GGUF` | output path |
| `--dtype f32\|f16\|q8_0` | storage type of ordinary linear weights (default `f32`) |
| `--name NAME` | `general.name` (default: the `--hf` directory name) |
| `--ternary dequant\|keep` | `dequant` (default) expands ternary weights to ordinary ones. `keep` stores them packed, see `docs/ternary.md` |
| `--vad keep\|drop` | keep (default) or drop the `vad_head.*` tensors and `parakeet.vad.*` KVs when the checkpoint has them |

Extra GGUF content written by this converter:

| Key or tensor | Written when | Meaning |
|---|---|---|
| `parakeet.ternary.present` (bool), `parakeet.ternary.group_size` (u32, 128) | `--ternary keep` | the file holds packed ternary linears |
| `<base>.qweight` (I8), `<base>.scales` (F16) | `--ternary keep` | replace `<base>.weight` of each ternary linear |
| `parakeet.vad.present` (bool), `parakeet.vad.d_in`, `parakeet.vad.hidden`, `parakeet.vad.kernel` (u32), `parakeet.vad.frame_sec` (f32) | the checkpoint has a VAD head and `--vad keep` | shape of the head; the loader reads them into `ParakeetConfig::vad` |
| `vad_head.proj.*`, `vad_head.ctx.*`, `vad_head.out.*` (weight and bias) | same | the VAD head tensors, written as stored |

## Worked example — `parakeet-tdt_ctc-110m`

```
Expand Down
57 changes: 57 additions & 0 deletions docs/parity.md
Original file line number Diff line number Diff line change
Expand Up @@ -389,6 +389,63 @@ heads — and the C++ port reproduces each head exactly, including the second he
# -> MODEL nvidia/parakeet-tdt_ctc-1.1b HEAD rnnt arch=hybrid_tdt_ctc xscaling=false WER 0.0000 ... PASS
```

## Ultra and Redux (moondream, HF safetensors)

`moondream/parakeet-ultra` (F16) and `moondream/parakeet-redux` (ternary
encoder) share the v3 architecture and are converted with
`scripts/convert_hf_parakeet_to_gguf.py` (see `docs/conversion.md`). There is no
NeMo baseline for these checkpoints, so parity means the transcript on
`tests/fixtures/speech.wav` equals the reference transcript in `AGENTS.md`,
and the packed ternary kernels agree with the dequantized model.

| Model | GGUF form | Kernel | Transcript equals the reference transcript in AGENTS.md |
|---|---|---|---|
| parakeet-ultra | F16 | n/a | yes |
| parakeet-redux | packed ternary (`--ternary keep`) | scalar | yes |
| parakeet-redux | packed ternary (`--ternary keep`) | avx2 | yes |
| parakeet-redux | packed ternary (`--ternary keep`) | vnni | yes |
| parakeet-redux | dequantized F16 | n/a | yes |

WER on long audio. Measured on three synthetic long-form clips per set, built from
LibriSpeech utterances of `benchmarks/librispeech_manifest.tsv` (30 utterances,
218 to 354 s each, gaps of low-level noise), not on TED-LIUM or other real long
recordings. Reference = the joined manifest texts; WER from `scripts/asr_metrics.py`
(case and punctuation normalized); `parakeet-cli transcribe --decoder tdt`, plain
single pass and `--vad`. Mean over three clips, percent. Details, per-clip values and the
parameter sweep are in `docs/ternary.md`.

| Model | Gap between utterances | Plain WER | `--vad` WER |
|---|---|---:|---:|
| parakeet-ultra F16 | 0.45 s | 1.71 | 1.69 |
| parakeet-ultra F16 | 0.16 s | 1.69 | 1.78 |
| parakeet-ultra F16 | none | 1.63 | 1.84 |
| parakeet-redux packed ternary | 0.45 s | 1.97 | 1.92 |
| parakeet-redux packed ternary | 0.16 s | 1.95 | 1.69 |
| parakeet-redux packed ternary | none | 1.83 | 1.71 |

Ultra Q8_0 and dequantized Redux were not measured on the long-form sets. On the 100 LibriSpeech
utterances (`docs/ternary.md`) Ultra Q8_0 has 1.71 percent and Redux packed 1.96 percent.

Speed numbers are in `docs/ternary.md`.

Multilingual and real long-form results (measured, no NeMo baseline; details, commands and caveats in
`docs/ternary.md`, section "Multilingual and long-form validation"). Scoring uses the plain `normalize` from
`scripts/asr_metrics.py`, not the Open ASR Leaderboard normalizer, so compare models with each other and not with
upstream cards.

| Check | v3 F16 | Ultra F16 | Redux packed | Redux dequantized F16 |
|---|---:|---:|---:|---:|
| FLEURS, 25 languages x first 50 test utterances, mean WER percent | 12.77 | 10.71 | 12.09 | 11.99 |
| TED-LIUM long-form, 11 talks, plain single pass, mean WER percent | 4.40 | 3.73 | 4.32 | 5.07 |
| TED-LIUM long-form, `--vad`, mean WER percent | no VAD head | 3.64 | 4.38 | 5.17 |

The 5.5 s clip among the 11 talks skews the Redux dequantized means; without it the packed and dequantized `--vad`
means are 4.82 and 4.85. Without the short clip the plain means are 4.75 packed and 4.74 dequantized. Ultra against the transformers `ParakeetForTDT` reference (fp32, CPU) on 60 FLEURS utterances
(en_us, de_de, fr_fr): 56 of 60 transcripts identical after normalization, WER of ours against HF 0.33 percent.
Packed Redux single-pass used to segfault above 8192 encoder frames (about 11 minutes, local attention paths without the
packed branch); fixed, and a 714 s clip gives the same transcript as the dequantized GGUF. On the seven talks that crashed, packed plain
WER is 4.92 against 4.91 dequantized, and the packed transcripts differ from the dequantized ones by 0.28 percent of words.

## Test suite status

`ctest --test-dir build --output-on-failure` (with `PARAKEET_TEST_GGUF`,
Expand Down
6 changes: 6 additions & 0 deletions docs/quantization.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,6 +133,12 @@ remain in place as a safety net — those tensors are never quantized.

---

### Ternary tensors

Packed ternary GGUFs (`--ternary keep`, see `docs/ternary.md`) store their
linears as `.qweight` (I8) and `.scales` (F16). `parakeet-cli quantize` never
re-quantizes them; it copies both tensors verbatim.

## Measured size + WER

WER is word-level vs NeMo (`scripts/validate_vs_nemo.py` on
Expand Down
Loading
Loading