Run ced.cpp on any ggml GPU backend - #2
Merged
Merged
Conversation
Each triangular mel filter is non-zero over a narrow band of the 257 FFT bins, but the frontend multiplied the full row for every frame. The product now runs over each filter's non-zero band only. The skipped terms are exact zeros, so the output stays bit-identical. The mel frontend runs on the host. With the encoder on a GPU it was most of the latency: on a GB10 it took 101 of 117 ms for ced-base on a 36 s clip. With this change the full classify takes 62 ms. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The runner always used a static CPU backend and created and freed a
graph allocator for every graph. GPU builds compiled but never used
the GPU.
Each loaded model now owns a backend. It picks the first GPU or
integrated GPU, or the device named in CED_DEVICE ("cpu", "CUDA0",
"Vulkan0", "MTL0"), and falls back to the CPU. One ggml_gallocr is kept
for the model's lifetime. A graph goes through ggml_backend_sched with
a CPU fallback only when the device has no kernel for one of its ops.
On a GPU the weights are uploaded to one device buffer at load. The
tensors the host reads (mel window, filterbank, init_bn stats) keep a
host copy, and the init_bn scale and shift are folded once at load.
classify now runs each chunk as one graph instead of two with a host
round trip in between.
CPU output is bit-identical to the previous code (all 527 probs, four
models, two clips) and CPU speed is unchanged. On GPU the top-5 tags
match the CPU on CUDA (GB10), Vulkan (Radeon 8060S) and Metal (M4).
The GPU matmuls move intermediate activations by up to 1.15e-2, so the
per-stage parity tests use a 2e-2 tolerance off the CPU. The CPU
tolerances and all end-to-end probability checks are unchanged.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
mudler
approved these changes
Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
GPU builds of ced.cpp compiled but never used the GPU: the runner always ran on a static CPU backend and created and freed a graph allocator for every graph. This PR makes the model run on CUDA, Vulkan and Metal, and keeps the CPU path bit-identical.
What changed
ced::Backend). It picks the first GPU or integrated GPU, or the device named inCED_DEVICE(cpu,CUDA0,Vulkan0,MTL0), and falls back to the CPU. Oneggml_gallocrlives for the model's lifetime. A graph goes throughggml_backend_schedwith a CPU fallback only when the device has no kernel for one of its ops.classify(embed, blocks and head), instead of two graphs with a host round trip. The parity entry points (embed_from_input_values,forward_from_tokens) are unchanged.ced-cli infoandced-cli benchprint the device. README documentsCED_DEVICE, anddocs/BENCHMARKS.mdhas a GPU section.Results
ctest is 7/7 on every backend.
main; speed unchanged (±0.5%)Top-5 tags were compared for tiny and base at f32, f16 and q8_0 on a 6 s and a 36 s clip. Quantized models move more on GPU (up to about 1e-2 for tiny q8_0), because the CPU also quantizes the activations of q8_0 matmuls. For small models the host mel frontend is still most of the GPU latency. A GPU mel frontend is a possible follow-up.
How to verify
🤖 Generated with Claude Code