Skip to content

Run ced.cpp on any ggml GPU backend - #2

Merged
mudler merged 2 commits into
mainfrom
feat/gpu-backend
Sep 28, 2026
Merged

mudler merged 2 commits into
mainfrom
feat/gpu-backend

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Contributor

GPU builds of ced.cpp compiled but never used the GPU: the runner always ran on a static CPU backend and created and freed a graph allocator for every graph. This PR makes the model run on CUDA, Vulkan and Metal, and keeps the CPU path bit-identical.

What changed

  • One backend per loaded model (ced::Backend). It picks the first GPU or integrated GPU, or the device named in CED_DEVICE (cpu, CUDA0, Vulkan0, MTL0), and falls back to the CPU. One ggml_gallocr lives for the model's lifetime. A graph goes through ggml_backend_sched with a CPU fallback only when the device has no kernel for one of its ops.
  • Weights on the device. On a GPU the weights are uploaded to one device buffer at load. The mel window, filterbank and init_bn stats keep a host copy, and init_bn is folded once at load.
  • One graph per chunk in classify (embed, blocks and head), instead of two graphs with a host round trip. The parity entry points (embed_from_input_values, forward_from_tokens) are unchanged.
  • Faster mel frontend. The filterbank product skips the zero bins of each mel filter. Output is bit-identical. With the encoder on a GPU the host mel was most of the latency.
  • ced-cli info and ced-cli bench print the device. README documents CED_DEVICE, and docs/BENCHMARKS.md has a GPU section.
  • Test tolerance. GPU matmuls move intermediate activations by up to 1.15e-2. The per-stage parity checks use a 2e-2 tolerance when the device is not the CPU. The CPU tolerances and every end-to-end probability check are unchanged.

Results

ctest is 7/7 on every backend.

Backend Machine vs CPU ced-base f32, 36 s clip
CPU x86 devbox, M4 all 527 probs bit-identical to main; speed unchanged (±0.5%)
Metal Apple M4 same top-5 tags; f32 within 1.4e-4 99 ms (CPU: 1152 ms)
Vulkan (RADV) Radeon 8060S same top-5 tags; f32 within 1.7e-4 71 ms (CPU: 446 ms)
CUDA 13 NVIDIA GB10 same top-5 tags; f32 within 1.3e-4 62 ms

Top-5 tags were compared for tiny and base at f32, f16 and q8_0 on a 6 s and a 36 s clip. Quantized models move more on GPU (up to about 1e-2 for tiny q8_0), because the CPU also quantizes the activations of q8_0 matmuls. For small models the host mel frontend is still most of the GPU latency. A GPU mel frontend is a possible follow-up.

How to verify

cmake -B build -DCED_BUILD_TESTS=ON -DCED_GGML_VULKAN=ON   # or CUDA / METAL
cmake --build build -j
ctest --test-dir build --output-on-failure                 # 7/7, logs "device: ..."
build/examples/cli/ced-cli bench models/ced-base-f32.gguf clip.wav --iters 30
CED_DEVICE=cpu build/examples/cli/ced-cli bench models/ced-base-f32.gguf clip.wav --iters 5

🤖 Generated with Claude Code

Each triangular mel filter is non-zero over a narrow band of the 257
FFT bins, but the frontend multiplied the full row for every frame.
The product now runs over each filter's non-zero band only. The skipped
terms are exact zeros, so the output stays bit-identical.

The mel frontend runs on the host. With the encoder on a GPU it was
most of the latency: on a GB10 it took 101 of 117 ms for ced-base on a
36 s clip. With this change the full classify takes 62 ms.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The runner always used a static CPU backend and created and freed a
graph allocator for every graph. GPU builds compiled but never used
the GPU.

Each loaded model now owns a backend. It picks the first GPU or
integrated GPU, or the device named in CED_DEVICE ("cpu", "CUDA0",
"Vulkan0", "MTL0"), and falls back to the CPU. One ggml_gallocr is kept
for the model's lifetime. A graph goes through ggml_backend_sched with
a CPU fallback only when the device has no kernel for one of its ops.

On a GPU the weights are uploaded to one device buffer at load. The
tensors the host reads (mel window, filterbank, init_bn stats) keep a
host copy, and the init_bn scale and shift are folded once at load.
classify now runs each chunk as one graph instead of two with a host
round trip in between.

CPU output is bit-identical to the previous code (all 527 probs, four
models, two clips) and CPU speed is unchanged. On GPU the top-5 tags
match the CPU on CUDA (GB10), Vulkan (Radeon 8060S) and Metal (M4).

The GPU matmuls move intermediate activations by up to 1.15e-2, so the
per-stage parity tests use a 2e-2 tolerance off the CPU. The CPU
tolerances and all end-to-end probability checks are unchanged.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
@mudler
mudler merged commit b102376 into main Sep 28, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants