Skip to content

fix(kokoro): export Prosody with fp32 compute (FluidAudio #947) - #108

Open
Alex-Wengg wants to merge 1 commit into
mainfrom
fix/947-kokoro-prosody-fp32
Open

Alex-Wengg wants to merge 1 commit into
mainfrom
fix/947-kokoro-prosody-fp32

Conversation

@Alex-Wengg

@Alex-Wengg Alex-Wengg commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Fixes the model side of FluidInference/FluidAudio#947 (opening words ~12-15 dB quiet on long Kokoro ANE utterances).

The fp16 KokoroProsody miscomputes F0/N over the first 1-3 s of the utterance on the Core ML CPU and ANE paths, for many T_a >= 400 (57/99 lengths broken for en/ja, 56/99 for zh). GPU fp16 and fp32 are exact. The LSTM, palettization and each exposed sub-op are clean, which points to a fused fp16 kernel in the upsampling AdainResBlk1d.

  • Prosody now exports with fp32 compute. It keeps fp16 I/O and int8 palettization.
  • The zh script's Prosody inputs are fp16 now, matching the shipped interface.
  • Result: 0/99 broken lengths, worst F0 MAE 0.3 Hz, about +3.7 ms per call at T=450.

Links

🤖 Generated with Claude Code

Core ML CPU/ANE fp16 corrupts Prosody F0/N at the utterance onset for many
T_a >= 400. Keep fp16 I/O + int8 palettization, compute in fp32. zh script
inputs aligned to fp16 to match the shipped interface. Published as
KokoroProsody_v2.mlmodelc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Alex-Wengg added a commit to FluidInference/FluidAudio that referenced this pull request Sep 25, 2026
…Prosody_v2 (#947) (#963)

Fixes #947.

The shipped fp16 `KokoroProsody` miscomputes F0/N over the first 1-3 s
of the utterance on the Core ML CPU and ANE paths once `T_a` reaches
about 400 frames (~10 s of audio). The opening words then come out
~12-15 dB quiet. Not iOS-specific: it reproduces on macOS too. All
variants are affected (en and ja share weights; zh too).

- Prosody stage is now `KokoroProsody_v2.mlmodelc` (fp32 compute, same
fp16 I/O). It's renamed rather than overwritten, so cached clients
re-download.
- New unit tests pin the stage bundle names to
`ModelNames.KokoroAne.requiredCoreMLModels`.

**Verification**
- Reporter's sentence with `af_heart`: onset -38.2 dB → -26.5 dB (body
-24 dB); ASR WER 7.7% → 3.8%.
- Prosody vs PyTorch across T=20..1980: 0/99 broken lengths (v1 broken
at 57/99).
- Fresh HF download: sha of all 3 bundles matches the build; ja and zh
synthesize.

**Links**
- Model conversion: FluidInference/mobius#108
- HF:
[9db203c5](https://huggingface.co/FluidInference/kokoro-82m-coreml/commit/9db203c5bd9f40052ae06e9a64af6ee9b07a72ab)
(`ANE/`, `ANE-ja/`, `ANE-zh/`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Alex-Wengg added a commit to FluidInference/FluidAudio that referenced this pull request Sep 25, 2026
Adds a short convention under **HUGGINGFACE UPLOADS** in `CLAUDE.md` so
every model fix can be traced from the issue to the code to the HF
commit, and back.

- Fixed models ship as `_v2` next to the original.
- HF commit messages name the issue and link the PRs.
- mobius and FluidAudio PRs link the issue, each other, and the HF
commit.
- A `## Changelog` row goes on the HF model card.

First applied to #947 / #963 / FluidInference/mobius#108. The
[kokoro-82m-coreml
card](https://huggingface.co/FluidInference/kokoro-82m-coreml) now has
the table, backfilled for the #836, #852, #914 and #926 uploads.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant