Skip to content

feat: support Moondream Parakeet Ultra and Redux - #77

Open
localai-org-maint-bot wants to merge 30 commits into
masterfrom
feat/hf-ternary-vad
Open

localai-org-maint-bot wants to merge 30 commits into
masterfrom
feat/hf-ternary-vad

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

What

Adds support for moondream/parakeet-ultra and moondream/parakeet-redux, two derivatives of parakeet-tdt-0.6b-v3. Both ship as HF safetensors that only Moondream's Photon runtime reads today.

  • Converter. scripts/convert_hf_parakeet_to_gguf.py reads the HF layout and maps tensor names back to the NeMo names the loader expects. Config, vocab and the mel buffers come from a v3 GGUF (--template), and every shape is checked against it. --ternary keep stores Redux's ternary weights packed. --vad keep carries the VAD head.
  • Native ternary kernel for Redux. The packed weights are repacked once at load. Activations are quantized to int8 per row, and a ggml custom op pair runs the matmul. There are AVX-512 VNNI and AVX2 kernels with runtime dispatch (no global -march) and a NEON kernel. Every kernel is bit-identical to a scalar reference. The path is taken only when a tensor <name>.qweight exists, so other models build the same graph as before.
  • VAD segmentation. parakeet-cli transcribe --vad cuts long audio at pauses into segments of at most 30 s, using the model's own VAD head, and offsets word and token timestamps back. parakeet-cli vad-probe prints the per-frame probabilities. The C-API gains parakeet_capi_transcribe_path_json_vad (additive, no ABI bump).
  • Docs and tests. docs/ternary.md has the measurements and the evidence behind the VAD wiring. docs/conversion.md, docs/parity.md, models/MANIFEST.md, AGENTS.md and the README are updated. New tests cover the converter, the loader flags, the kernels, the VAD head, the segmenter and the end to end paths.

Results

All measured on one Ryzen 9 9950X3D. Details and raw numbers are in docs/ternary.md.

Redux packed ternary Ultra Q8_0 (stands in for v3 Q8_0) Redux dequantized F16
File size 213 MB 941 MB 1.44 GB
LibriSpeech-100, per utterance, 8 threads, median RTF 75.6x 43.1x 46.1x
WER on those 100 utterances 1.96% 1.71% (Ultra) 1.92%
  • Single core kernel throughput: 547 GMAC/s (VNNI) against about 87 for ggml's own Q8_0 matmul on the same shape.
  • On one 180 s clip the end to end gain is only about 10 percent (single pass, load included).
  • Upstream reports 113x for its own engine on a different machine. That number is theirs, and I did not reproduce or compare it.
  • VAD segmentation does not change WER meaningfully on synthetic long-form clips built from LibriSpeech (differences of about 0.2 points in both directions, three clips per set).

Limits and known gaps

  • The VAD wiring is inferred. Moondream publishes no reference for the head. The chosen wiring (subsampler tap, ReLU after proj and ctx, no residual) had the best mean AUC in a 108 variant search on real speech (0.93 for Ultra, 0.95 for Redux). Its pause detection is weak, and the segmenter falls back to hard cuts when it finds no pause. The evidence and caveats are in docs/ternary.md.
  • Packed ternary is CPU only and offline only. A packed GGUF is refused at load on GPU backends and in streaming, with a message. Convert with --ternary dequant for those.
  • Some builds fall back to the scalar kernel (about 1 GMAC/s): MSVC, Windows on ARM, and aarch64 without the dot product extension. The loader logs a warning when that happens.
  • NEON was verified only under qemu-user (bit-identical to the reference). It has not run on real ARM hardware, and its speed is unmeasured.
  • Known follow-up. The check that refuses a GGUF that has packed tensors but no parakeet.ternary.present flag looks at two layer 0 tensor names. A hand-crafted file with a packed tensor in a later layer gets past it. On CPU the type and size checks still prevent an out-of-bounds read. The fix is to scan every tensor name and add a negative test.
  • The numbers come from this repository's standard build, which applies the in-tree ggml patches at configure time. The ggml pin itself is unchanged.

How to verify

cmake -B build -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=ON && cmake --build build -j
ctest --test-dir build --output-on-failure -LE model      # 25 of 25 pass

To exercise the model dependent tests, convert the two checkpoints and export the paths (steps in docs/conversion.md and docs/ternary.md):

python3 scripts/convert_hf_parakeet_to_gguf.py --hf <ultra dir> --template <v3 f16 gguf> --output ultra-f16.gguf --dtype f16
python3 scripts/convert_hf_parakeet_to_gguf.py --hf <redux dir> --template <v3 f16 gguf> --output redux-keep.gguf --dtype f16 --ternary keep
PARAKEET_TEST_GGUF_ULTRA=ultra-f16.gguf PARAKEET_TEST_GGUF_REDUX_KEEP=redux-keep.gguf ctest --test-dir build --output-on-failure

With every model variable set, 90 of 92 tests pass. The two failures, test_relpos_attention_local_chunked and test_capi_timestamps, also fail on master's own build. They expect the 110M model's baselines, and I ran them with the 0.6B v3 model.

The ced.cpp submodule was not available where this was built, so it was built with -DPARAKEET_WITH_CED=OFF and the CED tests did not run.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

🤖 Generated with Claude Code

Adds convert_hf_parakeet_to_gguf.py for the transformers ParakeetForTDT
layout (moondream/parakeet-ultra and parakeet-redux). Tensor names are
mapped back to the NeMo names the loader expects; config, vocab and the
mel buffers come from a v3 GGUF template and every shape is checked
against it. Ternary weights are dequantized for now.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
--ternary keep writes <name>.qweight (upstream bytes as I8) and
<name>.scales (F16) instead of a dequantized weight, plus
parakeet.ternary.* KVs. The vad_head.* tensors and parakeet.vad.* KVs are
written when present. Default output is unchanged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Both default to absent, so existing models load exactly as before.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The reference defines correctness for the SIMD kernels that follow.
Silence quantizes to a zero scale and gives an exact zero output.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
ternary_linear inserts an int8 activation quantization op and a ternary
matmul op wherever the encoder would call ggml_mul_mat, but only when
<base>.qweight exists, so other models take the same path as before.
Packed GGUFs are refused on GPU backends and in streaming with a clear
message.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Graph building must not throw: a malformed packed GGUF escaped
Backend::compute and leaked state. ternary_prepare now checks and
repacks every encoder linear in Model::load, which returns null with a
logged message on failure. The shape check in ternary_linear throws
instead of aborting.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Kernels are compiled with per-function target attributes, so portable
builds keep working, and are selected at run time. They equal the scalar
reference bit for bit (integer group sums, fp-contract off).
PARAKEET_TERNARY_KERNEL forces one for testing and benchmarks.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed code-1 times int8 activations with vdotq_s32; the integer result
equals the reference, so no group-sum correction is needed.

Cross-compiled with aarch64-linux-gnu-g++ for syntax verification.
NEON kernel unverified on hardware (no ARM machine available in this session).

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The -ffp-contract=off flag is unknown to MSVC and causes D9002 warning.
Move src/ternary_kernels_neon.cpp into the existing if(NOT MSVC) block
that guards the x86 ternary files. Keep the aarch64 -march flag in its
separate conditional block with APPEND to preserve the base property.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Adds single-thread ggml_mul_mat timings to bench_ternary and rewrites the
measured-speed section of docs/ternary.md with a Q8_0 baseline (via Ultra),
per-utterance LibriSpeech RTF and the kernel ceiling table.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Recompute the microbench table from a fresh raw run, correct the derived
percentages, add the measurement commits, reword the WER status in parity and
the manifest, and update the converter docstring for --ternary keep and the
VAD head.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The kernels reduced a whole vector to one int32 for every (row, group,
activation row) triple, and used 256-bit vectors. Repack the weights in
blocks of 16 output rows so that one vector lane holds one row: a 64-byte
step holds 16 rows by 16 elements as four 2-bit planes, and each plane is
multiplied with one broadcast group of 4 activations. The 16 group sums
then come out in one vector, and the float update runs on 16 rows at once,
with the same multiply and add per element as the scalar reference.

The AVX-512 VNNI kernel uses 512-bit vectors, 2 row blocks by 4
activation rows. The AVX2 kernel accumulates a whole group in int16 (at
most 16256 in magnitude) and widens once per group. The NEON kernel uses
sdot by lane on the same layout.

Rows are padded to a multiple of 16. The kernels compute whole blocks and
store only the rows of [r0, r1), so any row split stays correct.

Single thread on a Ryzen 9 9950X3D (bench_ternary, N=4096 K=1024 T=200):
vnni 84 to 444 GMAC/s, avx2 79 to 155 GMAC/s.

Tests: check that the repacked layout holds the input codes and scales,
thread splits with more threads than rows, untouched outputs outside the
row range, row and column remainders, and a case with large group sums
where a fused multiply-add would round differently from the reference.
NEON verified bit-identical under qemu-aarch64.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The scalar reference is about 50x slower than the vector kernels and
dominated the run time of bench_ternary.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
AVX-512 VNNI: 3 row blocks by 4 activation rows instead of 2 by 4, with
a 2-block tail so that a thread chunk of 8 blocks runs as 3 + 3 + 2. AVX2:
2 activation rows instead of 4 (4 spill with 16 ymm registers).

Single thread on a Ryzen 9 9950X3D, bench_ternary N=4096 K=1024 T=200:
vnni 444 to 541 GMAC/s, avx2 155 to 184 GMAC/s.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
After the kernel rework, ternary_quant_rows (scalar, one lrintf call per
element) took about as long as the matmul it feeds: 1.41 ms against
1.53 ms for K=4096 T=200 on one core. Add AVX-512 and AVX2 versions that
write the same bytes as the scalar version, which stays as
ternary_quant_rows_ref and defines them. ternary_quant_rows picks the
best one for the CPU.

The reference now clamps to +-127 before rounding instead of after. The
result is the same for every non-NaN input, since +-127 are integers; a
NaN input now becomes -127 on every platform, where before it depended on
what lrintf returns for NaN.

The test compares every quantizer byte for byte with the reference, for
K in 128, 1024, 4096, on random rows, all-zero rows, rounding ties,
denormals, values near FLT_MAX, negative zero, NaN and infinity.
bench_ternary now also times the quantization.

Single thread on a Ryzen 9 9950X3D, K=4096 T=200: 1.405 ms to 0.047 ms.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
It is selected when the CPU has avx512f, so it must not be compiled with
the avx512vl and avx512vnni target of the matmul kernel.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Replace the speed section with measurements of the new kernels: single
thread GMAC/s, per-utterance LibriSpeech RTF (with the previous kernel
built from b4164da in the same session) and the long clip at 8 and 16
threads. Every table now comes with the command line that produced it,
instead of a pointer to a report outside the repository. Describe the new
weight layout and the vectorized activation quantizer in the runtime path.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The head reads the subsampler output and gives one speech probability per
80 ms frame. The three undocumented wiring choices (ReLU after proj,
residual ctx, ReLU after ctx) were fixed from labeled probe clips on both
models: subsampler tap, ReLU, ReLU, no residual. Pause detection is weak,
see the VAD head wiring section of docs/ternary.md.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The head weights were read from tensor->data, which is only valid for
host memory. Read them with weight_to_host_f32 so device backends work,
after checking presence, type and size ourselves so failures stay
runtime_errors. The analytic tests now use asymmetric weights (non
identity proj, o != i and right-hand ctx taps) with hand-computed
expectations, so a transposed indexing bug fails them.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Pure function from per-frame speech probabilities to contiguous segments
of at most 30 s, with a hard cut when no pause is found.

Assisted-by: Claude:claude-haiku-4-5 [Claude Code]
Fix test_cuts_land_in_pauses pauses to actually exercise the rule (second
pause now in second window's last third). Add four rule-branch tests:
fallback_window (pause outside last third), min_seg_respected (pause before
min_seg), tie_goes_to_later_run (equal pauses), exact_totals (boundary cases).
Improve random_property to generate bursty silence runs. Add 3-line rule
comment to vad_segmenter.hpp. Adjust fallback_window total to 42.4s to get
2 segments per the rule.

Assisted-by: Claude:claude-haiku-4-5 [Claude Code]
Model::transcribe_pcm_vad(_with_timestamps), transcribe --vad and an
additive C-API function. Audio up to 30 s takes the plain path, so short
clips are unchanged. Word and token times are offset by the segment start.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The VAD test now builds a 62 s clip from tracked fixtures; the 180 s
clip is optional. Slices share exact boundary samples, an empty
segmentation falls back to the plain path, and the --vad-* flags reject
bad values.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
make_longform.py joins LibriSpeech utterances into long clips with a known
reference, filling gaps with low-level noise instead of digital zeros.
eval_vad_longform.py compares plain and --vad transcription WER.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Synthetic long clips with known references, WER with and without VAD
segmentation, and the parameter sweep behind the defaults (unchanged).
Also fixes the kernel speedup range, drops internal task labels, and
lists the VAD C-API symbol, CLI flags, sources and test env vars.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Move Ultra and Redux out of the NeMo-validated README table, restate the
VAD WER result as no meaningful change, document the caller-side pinning,
and make eval_vad_longform.py report CLI errors and empty matches.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
A doctored GGUF could make the repack read past the tensor data. weight_for
now checks types (I8 qweight, F16 scales), 2-D shapes, K % 128, overflow-safe
byte sizes, and ternary_prepare checks all 11 linears against d_model and
ff_dim. group_size other than 128 fails the load, a GGUF with .qweight tensors
but no parakeet.ternary.present flag is refused, and the GPU refusal message
now also suggests PARAKEET_DEVICE=cpu. ternary_prepare logs once when the
scalar kernel is selected. A negative-load test doctors copies of the Redux
GGUF and checks that Model::load returns nullptr.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The loader rejects a VAD config with a non-finite or non-positive frame_sec,
zero sizes or an even kernel. segment_by_vad returns one segment for degenerate
options instead of dividing by zero. Token frame offsets of VAD slices are now
in encoder frames (hop * subsampling / rate), and a VAD frame that is not a
whole number of encoder frames throws. The --vad-* options reject non-finite
values.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…aker checks

Document which builds and CPUs select the slow scalar ternary kernel, that
packed models keep both the packed tensors and the repacked planes resident,
and how the numbers were measured (standard build, in-tree ggml patches).
State the RTF basis in the README, qualify the v3 F16 row of the long-clip
table, and extend test_ternary_model with the two-speaker clip and a
batched-equals-per-item check on the packed model.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants