Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,8 @@ cmake --build build-shared -j

To build for a GPU backend, forward its flag, e.g. `cmake -B build -DCED_GGML_METAL=ON`.

At runtime a GPU build picks the first GPU it finds and falls back to CPU. Set `CED_DEVICE` to choose: `CED_DEVICE=cpu` forces the CPU, and a device name such as `CUDA0`, `Vulkan0` or `MTL0` selects that device (`ced-cli info` prints the one in use). The weights are uploaded to the device once at load. If the device has no kernel for an op, that graph runs through ggml's scheduler with a CPU fallback.

---

## Running inference
Expand Down
32 changes: 32 additions & 0 deletions docs/BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,38 @@ classified per wall-second).
before the first inference — a large practical gap for serving and CLI use.
- **Parity holds throughout**: identical top-5 tags across all variants.

## GPU

The same build runs on any ggml GPU backend (see `CED_DEVICE` in the README).
Numbers below are mean `ced-cli bench` latency, 30 iterations after 3 warmup,
next to the CPU of the same machine at 4 threads. The 36 s clip is split into
four ~10 s chunks, so it measures the multi-chunk path.

| Device | Model | 6 s clip | 36 s clip | same host CPU, 36 s |
|---|---|--:|--:|--:|
| Apple M4, Metal | base f32 | 17.5 ms | 99.5 ms | 1151.7 ms |
| Apple M4, Metal | base q8_0 | 16.9 ms | 97.6 ms | 419.5 ms |
| Apple M4, Metal | tiny q8_0 | 5.7 ms | 30.2 ms | 63.6 ms |
| Radeon 8060S, Vulkan (RADV) | base f32 | 14.7 ms | 70.6 ms | 446.4 ms |
| Radeon 8060S, Vulkan (RADV) | base q8_0 | 10.3 ms | 48.7 ms | 425.1 ms |
| Radeon 8060S, Vulkan (RADV) | tiny q8_0 | 7.0 ms | 33.2 ms | 73.7 ms |
| NVIDIA GB10, CUDA 13 | base f32 | 11.3 ms | 62.6 ms | 2393.7 ms |
| NVIDIA GB10, CUDA 13 | base q8_0 | 10.5 ms | 61.8 ms | 2332.7 ms |
| NVIDIA GB10, CUDA 13 | tiny q8_0 | 9.9 ms | 55.4 ms | 289.6 ms |

- **ced-base is 6x (Vulkan) to 12x (Metal) faster than the CPU of the same
machine.** The GB10 CPU column is from a container build that is slow on
CPU in general, so do not read the CUDA ratio as representative.
- **The GPU output keeps the CPU tags.** Over tiny and base at f32, f16 and
q8_0 on both clips, the top-5 tags are the same on every backend. f32
probabilities agree with the CPU to about 2e-4. Quantized models move more
(up to ~1e-2 for tiny q8_0), because the CPU also quantizes the activations
of q8_0 matmuls and the GPU backends do not.
- **Small models are bound by the host.** The mel frontend runs on the CPU,
and for ced-tiny it is a large part of the GPU latency (on the GB10 about
47 of the 55 ms on the 36 s clip). The GPU compute itself is a few
milliseconds per chunk.

## Reproduce

```sh
Expand Down
4 changes: 3 additions & 1 deletion examples/cli/main.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ static int cmd_info(const char* path) {
std::printf(" mel : n_mels=%u n_fft=%u hop=%u win=%u sr=%u\n", c.n_mels,
c.n_fft, c.hop_size, c.win_size, c.sample_rate);
std::printf(" target_length : %u frames\n", c.target_length);
std::printf(" device : %s\n", m.device_name().c_str());
return 0;
}

Expand Down Expand Up @@ -122,7 +123,8 @@ static int cmd_bench(int argc, char** argv) {
double median = ms[ms.size() / 2];
double minv = ms.front(), maxv = ms.back();

std::printf("model=%s clip=%.2fs threads=%d iters=%d\n", model, clip_s, threads, iters);
std::printf("model=%s device=%s clip=%.2fs threads=%d iters=%d\n", model,
m.device_name().c_str(), clip_s, threads, iters);
std::printf(" latency ms: min=%.2f median=%.2f mean=%.2f max=%.2f\n", minv, median, mean, maxv);
std::printf(" RTF (clip_s/mean_s): %.1fx realtime (%.1f clips/s)\n",
clip_s / (mean / 1000.0), 1000.0 / mean);
Expand Down
337 changes: 180 additions & 157 deletions src/ced.cpp

Large diffs are not rendered by default.

16 changes: 15 additions & 1 deletion src/ced.hpp
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
#pragma once
#include <memory>
#include <string>
#include <vector>

#include "ced_runner.hpp"
#include "model_loader.hpp"

namespace ced {
Expand All @@ -10,6 +12,8 @@ class Ced {
public:
bool load(const std::string& path);
const CedConfig& config() const { return loader_.config(); }
// Compute device the model runs on ("cpu", "CUDA0", "Vulkan0", ...).
const std::string& device_name() const { return backend_->device_name(); }

// Parity entry point: run the 12 ViT blocks + final norm + mean-pool head
// from precomputed patch tokens. `tokens` is n_tokens * embed_dim, token-
Expand All @@ -36,12 +40,22 @@ class Ced {
std::vector<float>& pos_out, std::vector<float>& tokens,
int& n_tokens, int n_threads = 4);

// End-to-end: waveform -> logits/probs over the 527 AudioSet classes.
// End-to-end: waveform -> logits/probs over the 527 AudioSet classes. Each
// target_length chunk runs as ONE graph (embed + blocks + head).
bool classify(const std::vector<float>& wav, std::vector<float>& logits,
std::vector<float>& probs, int n_threads = 4);

private:
struct Embed;
struct Head;
Embed build_embed(ggml_context* ctx, std::vector<GraphInput>& inputs,
const std::vector<float>& input_values, int T) const;
Head build_blocks(ggml_context* ctx, ggml_tensor* tokens, int n_tokens) const;

ModelLoader loader_;
std::unique_ptr<Backend> backend_;
// init_bn (BatchNorm2d, eval) folded into a per-mel scale/shift at load.
std::vector<float> bn_scale_, bn_shift_;
};

} // namespace ced
128 changes: 107 additions & 21 deletions src/ced_runner.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -5,21 +5,84 @@
#include "ggml-cpu.h"
#include "ggml.h"

#include <cctype>
#include <cstdio>
#include <cstdlib>

namespace ced {

bool run_graph(int n_threads, const BuildFn& build,
std::vector<std::vector<float>>& outs) {
static ggml_backend_t backend = ggml_backend_cpu_init();
if (!backend) {
std::fprintf(stderr, "ced: cpu backend init failed\n");
return false;
// Generous metadata context (no_alloc: data lives in allocator buffers). The
// fused classify graph (embed + 12 blocks + head) is well under 1k nodes.
static constexpr size_t kGraphSize = 8192;

struct Backend::Impl {
ggml_backend_t backend = nullptr; // selected device (GPU or CPU)
ggml_backend_t cpu_fallback = nullptr; // GPU path only, for unsupported ops
ggml_gallocr_t galloc = nullptr; // persistent, reused by every graph
ggml_backend_sched_t sched = nullptr; // created only if a graph needs it
};

static bool iequals(const std::string& a, const char* b) {
if (!b) return false;
size_t i = 0;
for (; i < a.size() && b[i]; ++i)
if (std::tolower((unsigned char)a[i]) != std::tolower((unsigned char)b[i]))
return false;
return i == a.size() && b[i] == '\0';
}

Backend::Backend() : impl_(new Impl()) {
const char* env = std::getenv("CED_DEVICE");
const std::string want = env ? env : "";
if (!iequals(want, "cpu")) {
for (size_t i = 0; i < ggml_backend_dev_count(); ++i) {
ggml_backend_dev_t dev = ggml_backend_dev_get(i);
const auto type = ggml_backend_dev_type(dev);
const char* name = ggml_backend_dev_name(dev);
const bool selected = want.empty()
? (type == GGML_BACKEND_DEVICE_TYPE_GPU ||
type == GGML_BACKEND_DEVICE_TYPE_IGPU)
: iequals(want, name);
if (!selected) continue;
impl_->backend = ggml_backend_dev_init(dev, nullptr);
if (impl_->backend) {
device_name_ = name ? name : "";
is_cpu_ = type == GGML_BACKEND_DEVICE_TYPE_CPU;
break;
}
}
if (!want.empty() && !impl_->backend)
std::fprintf(stderr, "ced: CED_DEVICE=%s not found, using CPU\n", want.c_str());
}
if (!impl_->backend) {
impl_->backend = ggml_backend_cpu_init();
device_name_ = "cpu";
is_cpu_ = true;
}
if (!impl_->backend) {
std::fprintf(stderr, "ced: backend init failed\n");
return;
}
if (!is_cpu_) impl_->cpu_fallback = ggml_backend_cpu_init();
}

Backend::~Backend() {
// Allocators before the backends they reference.
if (impl_->sched) ggml_backend_sched_free(impl_->sched);
if (impl_->galloc) ggml_gallocr_free(impl_->galloc);
if (impl_->cpu_fallback) ggml_backend_free(impl_->cpu_fallback);
if (impl_->backend) ggml_backend_free(impl_->backend);
delete impl_;
}

bool Backend::ok() const { return impl_->backend != nullptr; }

ggml_backend_t Backend::handle() const { return impl_->backend; }

bool Backend::compute(int n_threads, const BuildFn& build,
std::vector<std::vector<float>>& outs) {
if (!impl_->backend) return false;

// Generous metadata context (no_alloc: data lives in gallocr buffers). The
// 12-block ViT graph is a few hundred nodes; size for headroom.
constexpr size_t kGraphSize = 8192;
size_t mem = ggml_tensor_overhead() * (kGraphSize * 2) +
ggml_graph_overhead_custom(kGraphSize, false);
struct ggml_init_params ip { mem, nullptr, /*no_alloc=*/true };
Expand All @@ -34,31 +97,56 @@ bool run_graph(int n_threads, const BuildFn& build,
}

ggml_cgraph* gf = ggml_new_graph_custom(ctx, kGraphSize, false);
// Mark captured tensors as outputs so the gallocr does NOT reuse their
// Mark captured tensors as outputs so the allocator does NOT reuse their
// buffers for downstream nodes (intermediates like enc_norm/logits feed
// later ops and would otherwise read back freed/overwritten memory).
for (ggml_tensor* t : capture) {
ggml_set_output(t);
ggml_build_forward_expand(gf, t);
}

ggml_gallocr_t alloc =
ggml_gallocr_new(ggml_backend_get_default_buffer_type(backend));
if (!alloc || !ggml_gallocr_alloc_graph(alloc, gf)) {
std::fprintf(stderr, "ced: gallocr alloc failed\n");
if (alloc) ggml_gallocr_free(alloc);
// A GPU graph goes through the scheduler only when the device lacks a
// kernel for one of its ops; otherwise it takes the gallocr path.
bool use_sched = false;
if (impl_->cpu_fallback) {
const int n = ggml_graph_n_nodes(gf);
for (int i = 0; i < n && !use_sched; ++i)
use_sched = !ggml_backend_supports_op(impl_->backend, ggml_graph_node(gf, i));
}

const int nt = n_threads > 0 ? n_threads : 4;
bool alloc_ok = false;
if (use_sched) {
if (!impl_->sched) {
ggml_backend_t backs[2] = {impl_->backend, impl_->cpu_fallback};
impl_->sched = ggml_backend_sched_new(backs, nullptr, 2, kGraphSize,
/*parallel=*/false, /*op_offload=*/true);
}
if (impl_->sched) {
ggml_backend_cpu_set_n_threads(impl_->cpu_fallback, nt);
ggml_backend_sched_reset(impl_->sched);
alloc_ok = ggml_backend_sched_alloc_graph(impl_->sched, gf);
}
} else {
if (!impl_->galloc)
impl_->galloc =
ggml_gallocr_new(ggml_backend_get_default_buffer_type(impl_->backend));
alloc_ok = impl_->galloc && ggml_gallocr_alloc_graph(impl_->galloc, gf);
}
if (!alloc_ok) {
std::fprintf(stderr, "ced: graph alloc failed\n");
ggml_free(ctx);
return false;
}

// Inputs are gallocr-allocated; upload host data now.
for (const GraphInput& in : inputs)
ggml_backend_tensor_set(in.t, in.data, 0, in.nbytes);

ggml_backend_cpu_set_n_threads(backend, n_threads > 0 ? n_threads : 4);
if (ggml_backend_graph_compute(backend, gf) != GGML_STATUS_SUCCESS) {
if (is_cpu_) ggml_backend_cpu_set_n_threads(impl_->backend, nt);
const ggml_status st = use_sched ? ggml_backend_sched_graph_compute(impl_->sched, gf)
: ggml_backend_graph_compute(impl_->backend, gf);
if (st != GGML_STATUS_SUCCESS) {
std::fprintf(stderr, "ced: graph compute failed\n");
ggml_gallocr_free(alloc);
ggml_free(ctx);
return false;
}
Expand All @@ -69,8 +157,6 @@ bool run_graph(int n_threads, const BuildFn& build,
outs[i].resize(n);
ggml_backend_tensor_get(capture[i], outs[i].data(), 0, n * sizeof(float));
}

ggml_gallocr_free(alloc);
ggml_free(ctx);
return true;
}
Expand Down
47 changes: 40 additions & 7 deletions src/ced_runner.hpp
Original file line number Diff line number Diff line change
@@ -1,29 +1,62 @@
#pragma once
#include <cstddef>
#include <functional>
#include <string>
#include <vector>

struct ggml_context;
struct ggml_tensor;
struct ggml_backend;
typedef struct ggml_backend* ggml_backend_t;

namespace ced {

// An input leaf to be filled with host data AFTER the gallocr allocates the
// graph (gallocr decides the final address, so data is set post-alloc).
// An input leaf to be filled with host data AFTER the graph is allocated (the
// allocator decides the final address, so data is set post-alloc).
struct GraphInput {
ggml_tensor* t = nullptr;
const void* data = nullptr;
size_t nbytes = 0;
};

// build() creates input leaves (registering each in `inputs`), builds the graph,
// and returns the tensors to capture. run_graph allocates the graph on a CPU
// backend, uploads the inputs, computes, and reads each captured tensor's f32
// data into `outs` (parallel to the returned vector).
// and returns the tensors to capture. Backend::compute allocates the graph,
// uploads the inputs, computes, and reads each captured tensor's f32 data into
// `outs` (parallel to the returned vector).
using BuildFn =
std::function<std::vector<ggml_tensor*>(ggml_context*, std::vector<GraphInput>&)>;

bool run_graph(int n_threads, const BuildFn& build,
std::vector<std::vector<float>>& outs);
// Persistent compute backend + reusable graph allocator, one per loaded model.
//
// Device: CED_DEVICE names a registry device ("cpu", "CUDA0", "Vulkan0",
// "Metal", case-insensitive); unset picks the first GPU / integrated GPU and
// falls back to CPU.
//
// Every graph runs through ONE persistent ggml_gallocr, so the compute buffer
// is kept across calls instead of being allocated and freed per graph. On a GPU
// device, a graph that contains an op the device has no kernel for is routed
// through ggml_backend_sched with a CPU fallback instead; graphs the device
// fully supports stay on the gallocr path.
class Backend {
public:
Backend();
~Backend();
Backend(const Backend&) = delete;
Backend& operator=(const Backend&) = delete;

bool ok() const;
bool is_cpu() const { return is_cpu_; }
const std::string& device_name() const { return device_name_; }
ggml_backend_t handle() const;

bool compute(int n_threads, const BuildFn& build,
std::vector<std::vector<float>>& outs);

private:
struct Impl;
Impl* impl_;
bool is_cpu_ = true;
std::string device_name_ = "cpu";
};

} // namespace ced
15 changes: 14 additions & 1 deletion src/mel.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,19 @@ void mel_spectrogram(const CedConfig& cfg, const float* window, const float* fil
input_values.assign((size_t)n_mels * T, 0.0f);
std::vector<float> frame(n_fft), re, im, power(n_freqs);

// Each triangular mel filter is non-zero over a narrow band only; restrict
// the filterbank product to [lo, hi) per row. The skipped terms are exact
// zeros, so the sums (and their order) are unchanged.
std::vector<int> lo(n_mels, 0), hi(n_mels, 0);
for (int m = 0; m < n_mels; ++m) {
const float* fb = filterbank + (size_t)m * n_freqs;
int f0 = 0, f1 = n_freqs;
while (f0 < n_freqs && fb[f0] == 0.0f) ++f0;
while (f1 > f0 && fb[f1 - 1] == 0.0f) --f1;
lo[m] = f0;
hi[m] = f1;
}

for (int t = 0; t < T; ++t) {
const int start = t * hop;
for (int i = 0; i < n_fft; ++i) {
Expand All @@ -46,7 +59,7 @@ void mel_spectrogram(const CedConfig& cfg, const float* window, const float* fil
for (int m = 0; m < n_mels; ++m) {
const float* fb = filterbank + (size_t)m * n_freqs;
float acc = 0.0f;
for (int f = 0; f < n_freqs; ++f) acc += fb[f] * power[f];
for (int f = lo[m]; f < hi[m]; ++f) acc += fb[f] * power[f];
input_values[(size_t)m * T + t] = acc; // mel power, dB applied below
}
}
Expand Down
Loading
Loading