feat: support Moondream Parakeet Ultra and Redux - #77
Open
localai-org-maint-bot wants to merge 30 commits into
Open
localai-org-maint-bot wants to merge 30 commits into
localai-org-maint-bot wants to merge 30 commits into
Conversation
Adds convert_hf_parakeet_to_gguf.py for the transformers ParakeetForTDT layout (moondream/parakeet-ultra and parakeet-redux). Tensor names are mapped back to the NeMo names the loader expects; config, vocab and the mel buffers come from a v3 GGUF template and every shape is checked against it. Ternary weights are dequantized for now. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
--ternary keep writes <name>.qweight (upstream bytes as I8) and <name>.scales (F16) instead of a dequantized weight, plus parakeet.ternary.* KVs. The vad_head.* tensors and parakeet.vad.* KVs are written when present. Default output is unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Both default to absent, so existing models load exactly as before. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The reference defines correctness for the SIMD kernels that follow. Silence quantizes to a zero scale and gives an exact zero output. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
ternary_linear inserts an int8 activation quantization op and a ternary matmul op wherever the encoder would call ggml_mul_mat, but only when <base>.qweight exists, so other models take the same path as before. Packed GGUFs are refused on GPU backends and in streaming with a clear message. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Graph building must not throw: a malformed packed GGUF escaped Backend::compute and leaked state. ternary_prepare now checks and repacks every encoder linear in Model::load, which returns null with a logged message on failure. The shape check in ternary_linear throws instead of aborting. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Kernels are compiled with per-function target attributes, so portable builds keep working, and are selected at run time. They equal the scalar reference bit for bit (integer group sums, fp-contract off). PARAKEET_TERNARY_KERNEL forces one for testing and benchmarks. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed code-1 times int8 activations with vdotq_s32; the integer result equals the reference, so no group-sum correction is needed. Cross-compiled with aarch64-linux-gnu-g++ for syntax verification. NEON kernel unverified on hardware (no ARM machine available in this session). Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The -ffp-contract=off flag is unknown to MSVC and causes D9002 warning. Move src/ternary_kernels_neon.cpp into the existing if(NOT MSVC) block that guards the x86 ternary files. Keep the aarch64 -march flag in its separate conditional block with APPEND to preserve the base property. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Adds single-thread ggml_mul_mat timings to bench_ternary and rewrites the measured-speed section of docs/ternary.md with a Q8_0 baseline (via Ultra), per-utterance LibriSpeech RTF and the kernel ceiling table. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Recompute the microbench table from a fresh raw run, correct the derived percentages, add the measurement commits, reword the WER status in parity and the manifest, and update the converter docstring for --ternary keep and the VAD head. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The kernels reduced a whole vector to one int32 for every (row, group, activation row) triple, and used 256-bit vectors. Repack the weights in blocks of 16 output rows so that one vector lane holds one row: a 64-byte step holds 16 rows by 16 elements as four 2-bit planes, and each plane is multiplied with one broadcast group of 4 activations. The 16 group sums then come out in one vector, and the float update runs on 16 rows at once, with the same multiply and add per element as the scalar reference. The AVX-512 VNNI kernel uses 512-bit vectors, 2 row blocks by 4 activation rows. The AVX2 kernel accumulates a whole group in int16 (at most 16256 in magnitude) and widens once per group. The NEON kernel uses sdot by lane on the same layout. Rows are padded to a multiple of 16. The kernels compute whole blocks and store only the rows of [r0, r1), so any row split stays correct. Single thread on a Ryzen 9 9950X3D (bench_ternary, N=4096 K=1024 T=200): vnni 84 to 444 GMAC/s, avx2 79 to 155 GMAC/s. Tests: check that the repacked layout holds the input codes and scales, thread splits with more threads than rows, untouched outputs outside the row range, row and column remainders, and a case with large group sums where a fused multiply-add would round differently from the reference. NEON verified bit-identical under qemu-aarch64. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The scalar reference is about 50x slower than the vector kernels and dominated the run time of bench_ternary. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
AVX-512 VNNI: 3 row blocks by 4 activation rows instead of 2 by 4, with a 2-block tail so that a thread chunk of 8 blocks runs as 3 + 3 + 2. AVX2: 2 activation rows instead of 4 (4 spill with 16 ymm registers). Single thread on a Ryzen 9 9950X3D, bench_ternary N=4096 K=1024 T=200: vnni 444 to 541 GMAC/s, avx2 155 to 184 GMAC/s. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
After the kernel rework, ternary_quant_rows (scalar, one lrintf call per element) took about as long as the matmul it feeds: 1.41 ms against 1.53 ms for K=4096 T=200 on one core. Add AVX-512 and AVX2 versions that write the same bytes as the scalar version, which stays as ternary_quant_rows_ref and defines them. ternary_quant_rows picks the best one for the CPU. The reference now clamps to +-127 before rounding instead of after. The result is the same for every non-NaN input, since +-127 are integers; a NaN input now becomes -127 on every platform, where before it depended on what lrintf returns for NaN. The test compares every quantizer byte for byte with the reference, for K in 128, 1024, 4096, on random rows, all-zero rows, rounding ties, denormals, values near FLT_MAX, negative zero, NaN and infinity. bench_ternary now also times the quantization. Single thread on a Ryzen 9 9950X3D, K=4096 T=200: 1.405 ms to 0.047 ms. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
It is selected when the CPU has avx512f, so it must not be compiled with the avx512vl and avx512vnni target of the matmul kernel. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Replace the speed section with measurements of the new kernels: single thread GMAC/s, per-utterance LibriSpeech RTF (with the previous kernel built from b4164da in the same session) and the long clip at 8 and 16 threads. Every table now comes with the command line that produced it, instead of a pointer to a report outside the repository. Describe the new weight layout and the vectorized activation quantizer in the runtime path. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The head reads the subsampler output and gives one speech probability per 80 ms frame. The three undocumented wiring choices (ReLU after proj, residual ctx, ReLU after ctx) were fixed from labeled probe clips on both models: subsampler tap, ReLU, ReLU, no residual. Pause detection is weak, see the VAD head wiring section of docs/ternary.md. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The head weights were read from tensor->data, which is only valid for host memory. Read them with weight_to_host_f32 so device backends work, after checking presence, type and size ourselves so failures stay runtime_errors. The analytic tests now use asymmetric weights (non identity proj, o != i and right-hand ctx taps) with hand-computed expectations, so a transposed indexing bug fails them. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Pure function from per-frame speech probabilities to contiguous segments of at most 30 s, with a hard cut when no pause is found. Assisted-by: Claude:claude-haiku-4-5 [Claude Code]
Fix test_cuts_land_in_pauses pauses to actually exercise the rule (second pause now in second window's last third). Add four rule-branch tests: fallback_window (pause outside last third), min_seg_respected (pause before min_seg), tie_goes_to_later_run (equal pauses), exact_totals (boundary cases). Improve random_property to generate bursty silence runs. Add 3-line rule comment to vad_segmenter.hpp. Adjust fallback_window total to 42.4s to get 2 segments per the rule. Assisted-by: Claude:claude-haiku-4-5 [Claude Code]
Model::transcribe_pcm_vad(_with_timestamps), transcribe --vad and an additive C-API function. Audio up to 30 s takes the plain path, so short clips are unchanged. Word and token times are offset by the segment start. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The VAD test now builds a 62 s clip from tracked fixtures; the 180 s clip is optional. Slices share exact boundary samples, an empty segmentation falls back to the plain path, and the --vad-* flags reject bad values. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
make_longform.py joins LibriSpeech utterances into long clips with a known reference, filling gaps with low-level noise instead of digital zeros. eval_vad_longform.py compares plain and --vad transcription WER. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Synthetic long clips with known references, WER with and without VAD segmentation, and the parameter sweep behind the defaults (unchanged). Also fixes the kernel speedup range, drops internal task labels, and lists the VAD C-API symbol, CLI flags, sources and test env vars. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Move Ultra and Redux out of the NeMo-validated README table, restate the VAD WER result as no meaningful change, document the caller-side pinning, and make eval_vad_longform.py report CLI errors and empty matches. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
A doctored GGUF could make the repack read past the tensor data. weight_for now checks types (I8 qweight, F16 scales), 2-D shapes, K % 128, overflow-safe byte sizes, and ternary_prepare checks all 11 linears against d_model and ff_dim. group_size other than 128 fails the load, a GGUF with .qweight tensors but no parakeet.ternary.present flag is refused, and the GPU refusal message now also suggests PARAKEET_DEVICE=cpu. ternary_prepare logs once when the scalar kernel is selected. A negative-load test doctors copies of the Redux GGUF and checks that Model::load returns nullptr. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
The loader rejects a VAD config with a non-finite or non-positive frame_sec, zero sizes or an even kernel. segment_by_vad returns one segment for degenerate options instead of dividing by zero. Token frame offsets of VAD slices are now in encoder frames (hop * subsampling / rate), and a VAD frame that is not a whole number of encoder frames throws. The --vad-* options reject non-finite values. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…aker checks Document which builds and CPUs select the slow scalar ternary kernel, that packed models keep both the packed tensors and the repacked planes resident, and how the numbers were measured (standard build, in-tree ggml patches). State the RTF basis in the README, qualify the v3 F16 row of the long-clip table, and extend test_ternary_model with the two-speaker clip and a batched-equals-per-item check on the packed model. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds support for
moondream/parakeet-ultraandmoondream/parakeet-redux, two derivatives ofparakeet-tdt-0.6b-v3. Both ship as HF safetensors that only Moondream's Photon runtime reads today.scripts/convert_hf_parakeet_to_gguf.pyreads the HF layout and maps tensor names back to the NeMo names the loader expects. Config, vocab and the mel buffers come from a v3 GGUF (--template), and every shape is checked against it.--ternary keepstores Redux's ternary weights packed.--vad keepcarries the VAD head.-march) and a NEON kernel. Every kernel is bit-identical to a scalar reference. The path is taken only when a tensor<name>.qweightexists, so other models build the same graph as before.parakeet-cli transcribe --vadcuts long audio at pauses into segments of at most 30 s, using the model's own VAD head, and offsets word and token timestamps back.parakeet-cli vad-probeprints the per-frame probabilities. The C-API gainsparakeet_capi_transcribe_path_json_vad(additive, no ABI bump).docs/ternary.mdhas the measurements and the evidence behind the VAD wiring.docs/conversion.md,docs/parity.md,models/MANIFEST.md,AGENTS.mdand the README are updated. New tests cover the converter, the loader flags, the kernels, the VAD head, the segmenter and the end to end paths.Results
All measured on one Ryzen 9 9950X3D. Details and raw numbers are in
docs/ternary.md.Limits and known gaps
projandctx, no residual) had the best mean AUC in a 108 variant search on real speech (0.93 for Ultra, 0.95 for Redux). Its pause detection is weak, and the segmenter falls back to hard cuts when it finds no pause. The evidence and caveats are indocs/ternary.md.--ternary dequantfor those.parakeet.ternary.presentflag looks at two layer 0 tensor names. A hand-crafted file with a packed tensor in a later layer gets past it. On CPU the type and size checks still prevent an out-of-bounds read. The fix is to scan every tensor name and add a negative test.How to verify
To exercise the model dependent tests, convert the two checkpoints and export the paths (steps in
docs/conversion.mdanddocs/ternary.md):With every model variable set, 90 of 92 tests pass. The two failures,
test_relpos_attention_local_chunkedandtest_capi_timestamps, also fail onmaster's own build. They expect the 110M model's baselines, and I ran them with the 0.6B v3 model.The ced.cpp submodule was not available where this was built, so it was built with
-DPARAKEET_WITH_CED=OFFand the CED tests did not run.Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
🤖 Generated with Claude Code