From a8cdfb625eadc05df506f4426fbdd4b118543f09 Mon Sep 17 00:00:00 2001 From: Alex-Wengg Date: Thu, 24 Sep 2026 13:45:34 -0400 Subject: [PATCH] docs(asr): refresh Redux/Ultra numbers on the v3-family long-form path #955's review fix moved Redux and Ultra onto v3's no-mel, silence-aligned chunking default. Re-measured on full sets with v3 back to back (ANE default): test-clean WER v3 2.27 / redux 2.71 / ultra 2.13, test-other 4.12 / 5.12 / 3.81 (FLEURS unchanged: 14.81 / 13.06 / 11.67). RTFx test-clean 128.6x / 83.9x / 126.7x, test-other 114.7x / 76.0x / 110.1x. Replaces the earlier speed tables, which had Redux/Ultra on the mel-context path. --- Documentation/ASR/ParakeetRedux.md | 25 +++++++++++----------- Documentation/ASR/ParakeetUltra.md | 34 ++++++++++++------------------ Documentation/Models.md | 4 ++-- 3 files changed, 29 insertions(+), 34 deletions(-) diff --git a/Documentation/ASR/ParakeetRedux.md b/Documentation/ASR/ParakeetRedux.md index 603cca2ec..e12e3378d 100644 --- a/Documentation/ASR/ParakeetRedux.md +++ b/Documentation/ASR/ParakeetRedux.md @@ -52,20 +52,21 @@ there: Full LibriSpeech, `asr-benchmark`, M-series Mac. WER is corpus-level (total edit distance over total reference words, which is what the published leaderboards report); RTFx is total audio divided by total processing time. -| Set | Encoder | v3 WER | redux WER | v3 RTFx | redux RTFx | -|---|---|------:|----------:|--------:|-----------:| -| test-clean (2620 files) | GPU | **2.30 %** | 2.67 % | **118×** | 105× | -| test-other (2939 files) | GPU | **4.10 %** | 5.15 % | **106×** | 99× | -| test-clean (2620 files) | ANE | 2.27 % | 2.68 % | 101× | 80× | - -**v3 is the more accurate model on English**, by 0.37 points on test-clean and 1.05 on test-other. That reproduces -the upstream model card's own deltas (+0.44 and +1.21) almost exactly; the absolute values are higher here because +| Set | v3 WER | redux WER | v3 RTFx | redux RTFx | +|---|------:|----------:|--------:|-----------:| +| test-clean (2620 files) | **2.27 %** | 2.71 % | **128.6×** | 83.9× | +| test-other (2939 files) | **4.12 %** | 5.12 % | **114.7×** | 76.0× | + +Default compute units (ANE), v3 and redux run back to back on the same machine, both on the v3-family long-form path +(no mel context, silence-aligned window starts). + +**v3 is the more accurate model on English**, by 0.44 points on test-clean and 1.00 on test-other, which reproduces +the upstream model card's own deltas (+0.44 and +1.21). The absolute values are higher than the card's because FluidAudio decodes in 15 s windows and scores with a simpler normalizer than the Open ASR Leaderboard, which -penalises both models equally. Compute placement is WER-neutral (redux 2.67 % on GPU vs 2.68 % on ANE), as it is -for v3. +penalises both models equally. Compute placement is WER-neutral for both models. -Redux is also slower end to end: 11 % behind v3 on GPU (its encoder is 21.3 ms/window vs 18.3 ms), and 21 % behind -v3 on ANE (80× vs 101×). Redux on ANE is 23 % slower than redux on GPU; it still defaults to the ANE for iOS background execution. +Redux is also slower end to end: ~34 % behind v3 on the ANE. On GPU (`.cpuAndGPU`) the gap closes to a few percent +(its encoder is 21.3 ms/window vs 18.3 ms), but redux keeps the ANE default for iOS background execution. Conversion fidelity is not the issue — the Core ML redux transcripts match a PyTorch fp32 decode of the redux checkpoint to 0.19 % WER. Everything above is the checkpoint's own behaviour. **Choose redux for download size, and diff --git a/Documentation/ASR/ParakeetUltra.md b/Documentation/ASR/ParakeetUltra.md index ca65d01e3..7412e3b34 100644 --- a/Documentation/ASR/ParakeetUltra.md +++ b/Documentation/ASR/ParakeetUltra.md @@ -35,8 +35,8 @@ Full corpora, corpus-level WER (total edits over total reference words), same bu | Set | v3 | redux | ultra | |---|---:|---:|---:| -| LibriSpeech test-clean (2620 files) | 2.27 % | 2.67 % | **2.12 %** | -| LibriSpeech test-other (2939 files) | 4.12 % | 5.15 % | **3.79 %** | +| LibriSpeech test-clean (2620 files) | 2.27 % | 2.71 % | **2.13 %** | +| LibriSpeech test-other (2939 files) | 4.12 % | 5.12 % | **3.81 %** | | FLEURS, 24 languages × 100, mean | 14.81 % | 13.06 % | **11.67 %** | | FLEURS, duration-weighted | 14.65 % | 12.89 % | **11.51 %** | @@ -46,31 +46,25 @@ direction agrees with the upstream card in 24/24. Unlike redux, it does not give high-resource languages (French −0.8, Russian −1.3, English −0.2). The int8 encoder is WER-identical to an fp16 export (test-clean 2.12 vs 2.13 %, test-other 3.79 vs 3.79 %), and -compute placement is WER-neutral (ANE 2.12 %, GPU 2.13 %). +compute placement is WER-neutral (ANE 2.12 %, GPU 2.13 %). All numbers in this page use the v3-family long-form path +(no mel context, silence-aligned window starts), the library default for v3, redux and ultra. ## Speed -RTFx = total audio / total processing time. v3, redux and ultra run back to back per row on the same machine (load -3–6), full test sets: +RTFx = total audio / total processing time. v3, redux and ultra run back to back per row on the same machine, default +compute units (ANE), full test sets: -| Set | Encoder | v3 | redux | ultra | -|---|---|---:|---:|---:| -| test-clean | ANE | 88.3× | 67.8× | **89.5×** | -| test-clean | GPU | 93.1× | 86.6× | **94.2×** | -| test-other | ANE | 92.2× | 68.0× | **96.5×** | -| test-other | GPU | 93.4× | 89.2× | **100.7×** | - -FLEURS (24 languages × 100, `fleurs-benchmark --encoder-compute-units`), same protocol: - -| Encoder | v3 | redux | ultra | +| Set | v3 | redux | ultra | |---|---:|---:|---:| -| ANE | 92.8× | 70.8× | **93.0×** | -| GPU | 89.8× | 87.0× | **122.5×** | +| test-clean | **128.6×** | 83.9× | 126.7× | +| test-other | **114.7×** | 76.0× | 110.1× | -FLEURS clips are short, so its RTFx swings more with machine load than LibriSpeech; compare within a row. +Ultra is at parity with v3 (within 1–4 %); redux is ~34 % slower on the ANE. On GPU (`encoderComputeUnits: +.cpuAndGPU`) ultra was at or above v3 in every paired run. FLEURS clips are short and single-window; there ultra ran +at 135× vs v3 137× on ANE. -The shipped iOS 17 encoder was re-checked against v3 on test-clean: ANE 96.7× vs 94.9×, GPU 102.6× vs 103.4×, WER -unchanged (2.12 %). In isolation it is within 3–5 % of the iOS 18 export on GPU (16.6–17.4 vs 16.1–16.5 ms/window). +The shipped iOS 17 encoder matches an iOS 18 export of the same weights on test-other in WER (3.79 / 3.80 %) and speed +(ANE 96.4× vs 93.3×, GPU 96.7× vs 96.5×). ## Which to ship diff --git a/Documentation/Models.md b/Documentation/Models.md index b17163167..8358ab24f 100644 --- a/Documentation/Models.md +++ b/Documentation/Models.md @@ -12,8 +12,8 @@ Long-form audio processed via `SlidingWindowAsrManager` — chunked, overlapped, |-------|-------------|---------| | **Parakeet TDT v2** | Batch speech-to-text, English only (0.6B params). TDT architecture. | First ASR model added. | | **Parakeet TDT v3** | Batch speech-to-text, 25 European languages (0.6B params). Default ASR model. | Released after v2 to add multilingual support. | -| **Parakeet Ultra** ([docs](ASR/ParakeetUltra.md)) | moondream's post-training of v3: same architecture, languages and API (`AsrModelVersion.ultra`). More accurate than v3 on every benchmark (LibriSpeech test-clean 2.12 vs 2.27 %, test-other 3.79 vs 4.12 %, FLEURS 24-language mean 11.67 vs 14.81 %) at the same speed. int8 encoder (595 MB), ANE, iOS 17+. | **Recommended** for new integrations. | -| **Parakeet Redux** ([docs](ASR/ParakeetRedux.md)) | moondream's ternary re-training of v3 (`AsrModelVersion.redux`). 183 MB 2-bit encoder (~220 MB model dir). Better than v3 on FLEURS (13.06 vs 14.81 %), worse on English (2.67 vs 2.27 %). iOS 18+ / macOS 15+ only (use Ultra on iOS 17); ANE by default, several-minute first compile. | For size-constrained apps on iOS 18+. | +| **Parakeet Ultra** ([docs](ASR/ParakeetUltra.md)) | moondream's post-training of v3: same architecture, languages and API (`AsrModelVersion.ultra`). More accurate than v3 on every benchmark (LibriSpeech test-clean 2.13 vs 2.27 %, test-other 3.81 vs 4.12 %, FLEURS 24-language mean 11.67 vs 14.81 %) at the same speed. int8 encoder (595 MB), ANE, iOS 17+. | **Recommended** for new integrations. | +| **Parakeet Redux** ([docs](ASR/ParakeetRedux.md)) | moondream's ternary re-training of v3 (`AsrModelVersion.redux`). 183 MB 2-bit encoder (~220 MB model dir). Better than v3 on FLEURS (13.06 vs 14.81 %), worse on English (2.71 vs 2.27 %). iOS 18+ / macOS 15+ only (use Ultra on iOS 17); ANE by default, several-minute first compile. | For size-constrained apps on iOS 18+. | | **Parakeet TDT-CTC-110M** | Hybrid TDT-CTC batch model (110M params). 3.01% WER on LibriSpeech test-clean. 96.5x RTFx on M2 Mac. Fused preprocessor+encoder for reduced memory footprint. iOS compatible. | Smaller, faster alternative to v3 with competitive accuracy. | | **Parakeet TDT Japanese** | Batch speech-to-text, Japanese only (0.6B params). Hybrid model: INT8 CTC-trained preprocessor + encoder paired with a TDT decoder + joint. 6.85% CER on JSUT, 10.8x RTFx on M2. | CTC-only Japanese inference was removed in 846924a1d; only the preprocessor + encoder from the original CTC repo are reused. | | **Cohere Transcribe** ([FluidAudio#487](https://github.com/FluidInference/FluidAudio/pull/487), [#537](https://github.com/FluidInference/FluidAudio/pull/537)) | Batch encoder-decoder speech-to-text, 14 languages (en/fr/de/es/it/pt/nl/pl/el/ar/ja/zh/ko/vi). 48-layer Conformer encoder + 8-layer transformer decoder with external KV cache. Mixed precision: INT8 encoder (1.8 GB, iOS 18+) + FP32 ANE-resident static-shape decoder (v2, ~1.6× faster on Apple Silicon than the dynamic FP16 v1 decoder). Hard 35 s per-call audio cap (`max_audio_clip_s` from upstream config), 16 384-token SentencePiece vocab. Language must be passed explicitly via the conditioned prompt. | First Cohere Transcribe port; ANE-optimized v2 decoder (#537) lands fixed `[1, 1, 1, 108]` `attention_mask` so the decoder stays on the Neural Engine. |