fix(kokoro): export Prosody with fp32 compute (FluidAudio #947) - #108
Open
Alex-Wengg wants to merge 1 commit into
Open
Alex-Wengg wants to merge 1 commit into
Alex-Wengg wants to merge 1 commit into
Conversation
Core ML CPU/ANE fp16 corrupts Prosody F0/N at the utterance onset for many T_a >= 400. Keep fp16 I/O + int8 palettization, compute in fp32. zh script inputs aligned to fp16 to match the shipped interface. Published as KokoroProsody_v2.mlmodelc. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Alex-Wengg
added a commit
to FluidInference/FluidAudio
that referenced
this pull request
Sep 25, 2026
…Prosody_v2 (#947) (#963) Fixes #947. The shipped fp16 `KokoroProsody` miscomputes F0/N over the first 1-3 s of the utterance on the Core ML CPU and ANE paths once `T_a` reaches about 400 frames (~10 s of audio). The opening words then come out ~12-15 dB quiet. Not iOS-specific: it reproduces on macOS too. All variants are affected (en and ja share weights; zh too). - Prosody stage is now `KokoroProsody_v2.mlmodelc` (fp32 compute, same fp16 I/O). It's renamed rather than overwritten, so cached clients re-download. - New unit tests pin the stage bundle names to `ModelNames.KokoroAne.requiredCoreMLModels`. **Verification** - Reporter's sentence with `af_heart`: onset -38.2 dB → -26.5 dB (body -24 dB); ASR WER 7.7% → 3.8%. - Prosody vs PyTorch across T=20..1980: 0/99 broken lengths (v1 broken at 57/99). - Fresh HF download: sha of all 3 bundles matches the build; ja and zh synthesize. **Links** - Model conversion: FluidInference/mobius#108 - HF: [9db203c5](https://huggingface.co/FluidInference/kokoro-82m-coreml/commit/9db203c5bd9f40052ae06e9a64af6ee9b07a72ab) (`ANE/`, `ANE-ja/`, `ANE-zh/`) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Alex-Wengg
added a commit
to FluidInference/FluidAudio
that referenced
this pull request
Sep 25, 2026
Adds a short convention under **HUGGINGFACE UPLOADS** in `CLAUDE.md` so every model fix can be traced from the issue to the code to the HF commit, and back. - Fixed models ship as `_v2` next to the original. - HF commit messages name the issue and link the PRs. - mobius and FluidAudio PRs link the issue, each other, and the HF commit. - A `## Changelog` row goes on the HF model card. First applied to #947 / #963 / FluidInference/mobius#108. The [kokoro-82m-coreml card](https://huggingface.co/FluidInference/kokoro-82m-coreml) now has the table, backfilled for the #836, #852, #914 and #926 uploads. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the model side of FluidInference/FluidAudio#947 (opening words ~12-15 dB quiet on long Kokoro ANE utterances).
The fp16
KokoroProsodymiscomputes F0/N over the first 1-3 s of the utterance on the Core ML CPU and ANE paths, for manyT_a >= 400(57/99 lengths broken for en/ja, 56/99 for zh). GPU fp16 and fp32 are exact. The LSTM, palettization and each exposed sub-op are clean, which points to a fused fp16 kernel in the upsamplingAdainResBlk1d.Links
KokoroProsody_v2.mlmodelcinANE/,ANE-ja/,ANE-zh/, commit 9db203c5🤖 Generated with Claude Code