Skip to content

fix: GCC 16 build, and TDT beam search failing on parakeet-tdt-0.6b-v3 - #74

Merged
mudler merged 2 commits into
mudler:masterfrom
EnacheB:fix/gcc16-tdt-beam
Sep 29, 2026
Merged

mudler merged 2 commits into
mudler:masterfrom
EnacheB:fix/gcc16-tdt-beam

Conversation

@EnacheB

@EnacheB EnacheB commented Sep 28, 2026

Copy link
Copy Markdown

Two small, independent fixes.

1. Build fails on GCC 16. src/diarization_encoder.cpp uses std::copy_n without including <algorithm>; GCC ≤ 15 pulled it in transitively, GCC 16 does not. One-line include.

2. TDT beam search throws on the v3 model. On master, parakeet-cli transcribe --beam-size 4 fails on the repo's own tests/fixtures/speech.wav with parakeet-tdt-0.6b-v3 (and on two_speakers.wav at beam 2 and 4):

tdt_beam_search: zero-duration expansion did not reduce score

The check required every zero-duration expansion to strictly lower the hypothesis score. A score is a sum of log-probs, so it can never rise, but for a near-certain label the log-prob (~-6e-8) is below float32 resolution at a running score of a few units (spacing ~2.4e-7), so score + logp == score. That is correct behaviour, and the pruning only relies on scores not increasing, so the check now rejects only an increase. The 110m anchor model used in CI never produces such a label, which is why this went unnoticed.

Test: test_transcribe_0_6b now also runs beam 4 on speech.wav and checks the top hypothesis equals the greedy reference. It fails on master with the error above and passes with the fix. (It skips unless PARAKEET_TEST_GGUF_06B is set, like the existing greedy check.)

Tested: full ctest with the 110m model; test_transcribe_0_6b with tdt-0.6b-v3-q8_0 from mudler/parakeet-cpp-gguf.

🤖 Generated with Claude Code

GCC 16 no longer pulls it in transitively, so the build failed.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
…o score change

tdt_beam_search required every zero-duration expansion to strictly lower
the hypothesis score. For a near-certain label the float32 log-prob
(~-6e-8) is below the resolution of a running score of a few units
(spacing ~2.4e-7), so score + logp == score and the search threw
"zero-duration expansion did not reduce score". parakeet-tdt-0.6b-v3 hits
this on tests/fixtures/speech.wav at beam 4 and on two_speakers.wav at
beam 2 and 4; the 110m anchor used in CI never does. Only an increase
breaks the search, so check for that.

test_transcribe_0_6b now also runs beam 4 on speech.wav and checks the top
hypothesis matches the greedy reference.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
@mudler
mudler merged commit 6dea76a into mudler:master Sep 29, 2026
@EnacheB
EnacheB deleted the fix/gcc16-tdt-beam branch September 29, 2026 18:49
EnacheB added a commit to EnacheB/hyprwhspr that referenced this pull request Sep 29, 2026
On 98 real dictation recordings (134 chunks), beam 4 changed 34 chunks
against greedy, mostly dropping fillers and the fragments greedy decoding
invents from noise ("Um range vanquisher title.", "managing your own
investment."). It costs ~0.2 s per 10 s of audio on a Radeon 780M.

parakeet.cpp v0.5.0's beam search throws "zero-duration expansion did not
reduce score" on 28 of those 134 chunks; mudler/parakeet.cpp#74 fixes it.
Until a release carries that fix, the pinned v0.5.0 runtime fails about
one chunk in five.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants