Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 13 additions & 12 deletions Documentation/ASR/ParakeetRedux.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,20 +52,21 @@ there:
Full LibriSpeech, `asr-benchmark`, M-series Mac. WER is corpus-level (total edit distance over total reference
words, which is what the published leaderboards report); RTFx is total audio divided by total processing time.

| Set | Encoder | v3 WER | redux WER | v3 RTFx | redux RTFx |
|---|---|------:|----------:|--------:|-----------:|
| test-clean (2620 files) | GPU | **2.30 %** | 2.67 % | **118×** | 105× |
| test-other (2939 files) | GPU | **4.10 %** | 5.15 % | **106×** | 99× |
| test-clean (2620 files) | ANE | 2.27 % | 2.68 % | 101× | 80× |

**v3 is the more accurate model on English**, by 0.37 points on test-clean and 1.05 on test-other. That reproduces
the upstream model card's own deltas (+0.44 and +1.21) almost exactly; the absolute values are higher here because
| Set | v3 WER | redux WER | v3 RTFx | redux RTFx |
|---|------:|----------:|--------:|-----------:|
| test-clean (2620 files) | **2.27 %** | 2.71 % | **128.6×** | 83.9× |
| test-other (2939 files) | **4.12 %** | 5.12 % | **114.7×** | 76.0× |

Default compute units (ANE), v3 and redux run back to back on the same machine, both on the v3-family long-form path
(no mel context, silence-aligned window starts).

**v3 is the more accurate model on English**, by 0.44 points on test-clean and 1.00 on test-other, which reproduces
the upstream model card's own deltas (+0.44 and +1.21). The absolute values are higher than the card's because
FluidAudio decodes in 15 s windows and scores with a simpler normalizer than the Open ASR Leaderboard, which
penalises both models equally. Compute placement is WER-neutral (redux 2.67 % on GPU vs 2.68 % on ANE), as it is
for v3.
penalises both models equally. Compute placement is WER-neutral for both models.

Redux is also slower end to end: 11 % behind v3 on GPU (its encoder is 21.3 ms/window vs 18.3 ms), and 21 % behind
v3 on ANE (80× vs 101×). Redux on ANE is 23 % slower than redux on GPU; it still defaults to the ANE for iOS background execution.
Redux is also slower end to end: ~34 % behind v3 on the ANE. On GPU (`.cpuAndGPU`) the gap closes to a few percent
(its encoder is 21.3 ms/window vs 18.3 ms), but redux keeps the ANE default for iOS background execution.

Conversion fidelity is not the issue — the Core ML redux transcripts match a PyTorch fp32 decode of the redux
checkpoint to 0.19 % WER. Everything above is the checkpoint's own behaviour. **Choose redux for download size, and
Expand Down
34 changes: 14 additions & 20 deletions Documentation/ASR/ParakeetUltra.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,8 +35,8 @@ Full corpora, corpus-level WER (total edits over total reference words), same bu

| Set | v3 | redux | ultra |
|---|---:|---:|---:|
| LibriSpeech test-clean (2620 files) | 2.27 % | 2.67 % | **2.12 %** |
| LibriSpeech test-other (2939 files) | 4.12 % | 5.15 % | **3.79 %** |
| LibriSpeech test-clean (2620 files) | 2.27 % | 2.71 % | **2.13 %** |
| LibriSpeech test-other (2939 files) | 4.12 % | 5.12 % | **3.81 %** |
| FLEURS, 24 languages × 100, mean | 14.81 % | 13.06 % | **11.67 %** |
| FLEURS, duration-weighted | 14.65 % | 12.89 % | **11.51 %** |

Expand All @@ -46,31 +46,25 @@ direction agrees with the upstream card in 24/24. Unlike redux, it does not give
high-resource languages (French −0.8, Russian −1.3, English −0.2).

The int8 encoder is WER-identical to an fp16 export (test-clean 2.12 vs 2.13 %, test-other 3.79 vs 3.79 %), and
compute placement is WER-neutral (ANE 2.12 %, GPU 2.13 %).
compute placement is WER-neutral (ANE 2.12 %, GPU 2.13 %). All numbers in this page use the v3-family long-form path
(no mel context, silence-aligned window starts), the library default for v3, redux and ultra.

## Speed

RTFx = total audio / total processing time. v3, redux and ultra run back to back per row on the same machine (load
3–6), full test sets:
RTFx = total audio / total processing time. v3, redux and ultra run back to back per row on the same machine, default
compute units (ANE), full test sets:

| Set | Encoder | v3 | redux | ultra |
|---|---|---:|---:|---:|
| test-clean | ANE | 88.3× | 67.8× | **89.5×** |
| test-clean | GPU | 93.1× | 86.6× | **94.2×** |
| test-other | ANE | 92.2× | 68.0× | **96.5×** |
| test-other | GPU | 93.4× | 89.2× | **100.7×** |

FLEURS (24 languages × 100, `fleurs-benchmark --encoder-compute-units`), same protocol:

| Encoder | v3 | redux | ultra |
| Set | v3 | redux | ultra |
|---|---:|---:|---:|
| ANE | 92.8× | 70.8× | **93.0×** |
| GPU | 89.8× | 87.0× | **122.5×** |
| test-clean | **128.6×** | 83.9× | 126.7× |
| test-other | **114.7×** | 76.0× | 110.1× |

FLEURS clips are short, so its RTFx swings more with machine load than LibriSpeech; compare within a row.
Ultra is at parity with v3 (within 1–4 %); redux is ~34 % slower on the ANE. On GPU (`encoderComputeUnits:
.cpuAndGPU`) ultra was at or above v3 in every paired run. FLEURS clips are short and single-window; there ultra ran
at 135× vs v3 137× on ANE.

The shipped iOS 17 encoder was re-checked against v3 on test-clean: ANE 96.7× vs 94.9×, GPU 102.6× vs 103.4×, WER
unchanged (2.12 %). In isolation it is within 3–5 % of the iOS 18 export on GPU (16.6–17.4 vs 16.1–16.5 ms/window).
The shipped iOS 17 encoder matches an iOS 18 export of the same weights on test-other in WER (3.79 / 3.80 %) and speed
(ANE 96.4× vs 93.3×, GPU 96.7× vs 96.5×).

## Which to ship

Expand Down
4 changes: 2 additions & 2 deletions Documentation/Models.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@ Long-form audio processed via `SlidingWindowAsrManager` — chunked, overlapped,
|-------|-------------|---------|
| **Parakeet TDT v2** | Batch speech-to-text, English only (0.6B params). TDT architecture. | First ASR model added. |
| **Parakeet TDT v3** | Batch speech-to-text, 25 European languages (0.6B params). Default ASR model. | Released after v2 to add multilingual support. |
| **Parakeet Ultra** ([docs](ASR/ParakeetUltra.md)) | moondream's post-training of v3: same architecture, languages and API (`AsrModelVersion.ultra`). More accurate than v3 on every benchmark (LibriSpeech test-clean 2.12 vs 2.27 %, test-other 3.79 vs 4.12 %, FLEURS 24-language mean 11.67 vs 14.81 %) at the same speed. int8 encoder (595 MB), ANE, iOS 17+. | **Recommended** for new integrations. |
| **Parakeet Redux** ([docs](ASR/ParakeetRedux.md)) | moondream's ternary re-training of v3 (`AsrModelVersion.redux`). 183 MB 2-bit encoder (~220 MB model dir). Better than v3 on FLEURS (13.06 vs 14.81 %), worse on English (2.67 vs 2.27 %). iOS 18+ / macOS 15+ only (use Ultra on iOS 17); ANE by default, several-minute first compile. | For size-constrained apps on iOS 18+. |
| **Parakeet Ultra** ([docs](ASR/ParakeetUltra.md)) | moondream's post-training of v3: same architecture, languages and API (`AsrModelVersion.ultra`). More accurate than v3 on every benchmark (LibriSpeech test-clean 2.13 vs 2.27 %, test-other 3.81 vs 4.12 %, FLEURS 24-language mean 11.67 vs 14.81 %) at the same speed. int8 encoder (595 MB), ANE, iOS 17+. | **Recommended** for new integrations. |
| **Parakeet Redux** ([docs](ASR/ParakeetRedux.md)) | moondream's ternary re-training of v3 (`AsrModelVersion.redux`). 183 MB 2-bit encoder (~220 MB model dir). Better than v3 on FLEURS (13.06 vs 14.81 %), worse on English (2.71 vs 2.27 %). iOS 18+ / macOS 15+ only (use Ultra on iOS 17); ANE by default, several-minute first compile. | For size-constrained apps on iOS 18+. |
| **Parakeet TDT-CTC-110M** | Hybrid TDT-CTC batch model (110M params). 3.01% WER on LibriSpeech test-clean. 96.5x RTFx on M2 Mac. Fused preprocessor+encoder for reduced memory footprint. iOS compatible. | Smaller, faster alternative to v3 with competitive accuracy. |
| **Parakeet TDT Japanese** | Batch speech-to-text, Japanese only (0.6B params). Hybrid model: INT8 CTC-trained preprocessor + encoder paired with a TDT decoder + joint. 6.85% CER on JSUT, 10.8x RTFx on M2. | CTC-only Japanese inference was removed in 846924a1d; only the preprocessor + encoder from the original CTC repo are reused. |
| **Cohere Transcribe** ([FluidAudio#487](https://github.com/FluidInference/FluidAudio/pull/487), [#537](https://github.com/FluidInference/FluidAudio/pull/537)) | Batch encoder-decoder speech-to-text, 14 languages (en/fr/de/es/it/pt/nl/pl/el/ar/ja/zh/ko/vi). 48-layer Conformer encoder + 8-layer transformer decoder with external KV cache. Mixed precision: INT8 encoder (1.8 GB, iOS 18+) + FP32 ANE-resident static-shape decoder (v2, ~1.6× faster on Apple Silicon than the dynamic FP16 v1 decoder). Hard 35 s per-call audio cap (`max_audio_clip_s` from upstream config), 16 384-token SentencePiece vocab. Language must be passed explicitly via the conditioned prompt. | First Cohere Transcribe port; ANE-optimized v2 decoder (#537) lands fixed `[1, 1, 1, 108]` `attention_mask` so the decoder stays on the Neural Engine. |
Expand Down
Loading