diff --git a/docs/kv-transport-pipelining.md b/docs/kv-transport-pipelining.md index dae92d143c2c..c39df7df6a66 100644 --- a/docs/kv-transport-pipelining.md +++ b/docs/kv-transport-pipelining.md @@ -304,7 +304,7 @@ A device-resident KV run is unaffected, and was measured to confirm it: 38.5612 ## Scope and limits - Only persistent host inputs marked with `GGML_TENSOR_FLAG_TRANSPORT` are candidates. The stable prefix remains a per-evaluation value. Unmarked inputs, weights, user inputs, transposed V, and copies with later readers stay ordered. -- CUDA is the only enabled backend. Meta, SYCL, WebGPU, and other backends stay ordered until their event behavior and transport path are validated. +- CUDA is the only enabled backend, on its own or as every simple device of a meta backend. SYCL, WebGPU, and other backends stay ordered until their event behavior and transport path are validated. - The ring costs `(depth + 2) x (largest staged split)` of device memory, and a staged split is both K and V of one attention layer over the window the graph reads. That is linear in context length, and it is what bounds the feature at depth rather than anything about the transfer itself. - **A cap is per graph, not per sequence.** `--kv-pipeline-budget` bounds the window one graph delivers, which is `n_kv * n_stream` over every sequence in the ubatch, so it cannot be applied to one sequence of a batch and not another. - **A multi-stream window is delivered one range per stream**, keyed on the last dimension, and the copy packs those ranges so a slot holds the window rather than the whole cache. A window whose streams are not on that dimension keeps the single flat range, which is correct but not accelerated. Each range carries its own stream's prefix; ranges that agree on it are issued as one strided copy, so streams at the same depth still cost a single call. @@ -312,22 +312,71 @@ A device-resident KV run is unaffected, and was measured to confirm it: 38.5612 - **One ring per accelerator.** A layer-split model pipelines on every device that qualifies; a device with no room within the budget falls back to the ordered path on its own without disabling the others. The exception is a graph that cannot be allocated next to the rings: there every device that was holding one gives it back for good, because the allocator does not say which of them it competed with. - **The producer of a staged input must be the CPU or the consumer itself.** Neither part of a staged delivery is ordered against a third device: the stable prefix goes on the transfer stream and the rest on the consumer's own stream, where the ordered path would have synchronized the producer first. An input a second accelerator writes keeps the ordered path. - **It turns graph-level pipeline parallelism off while it is delivering.** A graph that delivered has to block the host on its consumer before the next graph writes the host cache, because the host source of a delivery is read long after the call that issued it returned. That block is what `n_copies > 1` exists to avoid, so the two do not overlap: with `-sm layer` over several GPUs and `--kv-cpu-pinned`, `llama_context` enables both and the ring wins. Use `--kv-pipeline-depth 0` to keep the graph-level pipelining instead. -- **Tensor parallelism keeps the ordered path.** See [Tensor parallelism](#tensor-parallelism). +- **Tensor parallelism pipelines through the meta backend**, and each device allocates the whole ring. See [Tensor parallelism](#tensor-parallelism). - **A host write to the cache waits for the delivery.** `llama_memory_clear(mem, true)` waits for the scheduler before it clears the buffers, because a delivery the last decode issued can still be reading them. This was already needed without the transport: with a device-resident cache the same call cleared the buffers under the running graph, and `llama_decode` followed by that clear changed the logits of that decode on every trial. - The scheduler must be configured with the device's own default buffer type. A scheduler built on a split or host buffer type keeps the ordered path. - `GGML_KV_PIPELINE_DEPTH` and `GGML_KV_PIPELINE_BUDGET_MIB` set the defaults of a scheduler that nothing else configures. `llama_context` always configures its own from the context parameters, so under `llama-server` and `llama-bench` use `--kv-pipeline-depth` / `LLAMA_ARG_KV_PIPELINE_DEPTH` and `--kv-pipeline-budget` / `LLAMA_ARG_KV_PIPELINE_BUDGET` instead. ## Tensor parallelism -`-sm tensor` is not pipelined. The scheduler explicitly excludes meta devices. A host-resident cache needs a validated strided head-split write before this can be enabled. +`-sm tensor` pipelines the same way. The consumer is a meta backend, and the scheduler sees one ring, one transfer backend and one pair of events per slot; each of them fans out to the simple devices underneath. -Both sit behind a correctness problem that is not this feature's: **`-sm tensor` together with `--no-kv-offload` currently produces wrong output.** On one build and one prompt, `-sm layer --no-kv-offload` and `-sm tensor` with a device-resident cache agree exactly, while `-sm tensor --no-kv-offload` differs. It does not crash or warn; it generates fluent, different text. +This depends on the host-resident cache being split by head, which is #66: before it, `-sm tensor --no-kv-offload` mirrored the cache to every device and produced wrong output. The copy arrives permuted as `[head_dim, n_kv, n_head_kv, n_stream]`, so each device's heads are one run inside every cell rather than a contiguous block. -The cause is the GQA head mapping. Tensor parallelism splits attention by head, but a host-resident cache is one undivided tensor, so the scheduler's copy of it is classified `MIRRORED` and the whole window goes to every device. With 24 query heads split 12/12 and 4 KV heads mirrored, the kernel derives the GQA ratio from the tensors it is handed -- 12/4 = 3 rather than 6 -- and the second device's queries, renumbered from 0, read the first device's keys. With an uneven split the same fault surfaces as a crash instead: `GGML_ASSERT(Q->ne[2] % K->ne[2] == 0)`, because 24 heads split 13/11 is not divisible by 4. +What the meta backend adds: -Head-splitting the copy rather than mirroring it fixes it. That was prototyped and reproduced the layer-split output byte for byte, and needs four coordinated changes: classify the scheduler's copy at all (it is a leaf in a compute buffer, so it never reaches the device's split-state callback), use the head axis for the permuted `[head_dim, n_kv, n_head_kv, 1]` shape rather than the cache tensor's own axis, express the granularity in heads aligned to the query split divided by the GQA ratio, and add a strided write because the heads are interleaved within each row rather than laid out end to end. +- **Events.** A meta event is one event per simple device. Recording it records each part on that device's stream, and waiting on it makes each simple backend wait for its own device's part. `caps.events` still reports false, so nothing else starts using them. +- **A ranged head-split write.** `set_tensor_async` and `set_tensor_2d_async` accept a part of the window: whole cells from an offset, once per stream. Each device takes its run of heads from every cell with one 2d copy per stream, so the early and the late delivery are the same calls as on a single device, one per device. +- **A transfer backend without a communicator.** The transfer backend never computes, so `ggml_backend_meta_init_transfer` builds its streams without starting a second NCCL context. +- **The ring is a meta buffer.** A copy in it gets its per-device tensors from the same split-state callback as the copy the graph allocator would have made. The meta graph compute also rotates the compute containers of buffers that appear only as sources, or the ring's per-device tensors would accumulate one set per plan. + +**Each device allocates the whole slot.** A meta buffer places a tensor at the same offset in every device's buffer and sizes each of those buffers for the whole tensor, so a device that holds half the heads still allocates the full slot. The meta compute buffers already work this way. The budget and the headroom check are applied per device, against the device with the least free memory, so a ring that fits the budget costs that much on every device. + +### Measurements + +RTX 4070 (gen4 x16) + RTX 3060 (gen3 x4), CUDA with NCCL, `Qwen3.8-27B-UD-IQ2_M.gguf`, `-ngl 99 -sm tensor -t 3 -fa on -ctk q8_0 -ctv q8_0 -b 512 -ub 512 -nkvo --kv-cpu-pinned`, under `taskset -c 0,2,4`. The recurrent state stays on the devices, which `-sm tensor` forces. + +`LLAMA_KV_SM=tensor docs/repro/r4-kv-pipeline-ab.sh`, `--kv-pipeline-budget 512`, both passes shown: + +| depth | ordered | pipelined | gain | +|---:|---|---|---:| +| 4,096 | 15.8398, 15.8351 | 21.9574, 21.9368 | **+38.6%** | +| 16,384 | 7.0819, 7.0839 | 8.7811, 8.7829 | **+24.0%** | +| 32,768 | 4.0809, 4.0826 | 4.8851, 4.8852 | **+19.7%** | + +`llama-server`, the tasks of the exactness gate at `-c 32768`: + +| task | prompt | ordered | pipelined | gain | +|---|---:|---:|---:|---:| +| prose | 1,709 | 20.675 | 25.402 | **+22.9%** | +| code | 3,270 | 17.036 | 22.589 | **+32.6%** | +| prose | 14,821 | 7.582 | 9.451 | **+24.7%** | +| code | 29,670 | 4.428 | 5.317 | **+20.1%** | + +Per decode graph at 16,384, `GGML_SCHED_TRANSPORT_DEBUG=2`: + +| | ordered | pipelined | +|---|---:|---:| +| total | 138.85 ms | 111.78 ms | +| blocked in the ordered copy | 103.81 ms | 12.16 ms | +| blocked waiting for the consumer | 31.06 ms | 95.43 ms | +| bytes delivered early / late | 0 / 0 MiB | 549.3 / 3.2 MiB | +| ring | - | 3 slots x 35 MiB, per device | + +The copy is three times the compute here, where on the single RTX 4070 above the two were about equal. The 553 MiB cross in 104 ms, about 5.3 GB/s, and the 3060's half of them crosses a gen3 x4 link. The pipeline hides the compute behind the copy, and the consumer wait now contains the rest of the transfer, so the token is bounded by the slower link rather than by the order of the work. The ceiling is `max(copy, compute)` plus the work outside the split loop, the same as on one device, and at 16,384 the pipeline is within a few milliseconds of it. + +So the gain is a property of the link, as it is on one device, with one addition: the devices compute in lock step, so the slowest link sets the pace for all of them. Devices on equal links carry an equal share of the bytes each; that has not been measured here. + +### Validation under `-sm tensor` + +On the same two devices: + +- `docs/repro/r4-kv-pipeline-exact.sh` with `LLAMA_KV_SM=tensor`: all eight tasks identical at `N = 0`, `1` and `4`. +- `docs/repro/r4-kv-pipeline-parallel-exact.sh` with `LLAMA_KV_SM=tensor`: `125cb9c2082d36cf` at `N = 0`, `1` and `4`, 8 concurrent sequences over a cache split into streams. With `-sm none` and `-sm layer` the same gate still gives `17f946c340db110b` and `db661b7a08686b97` at `N = 0` and `1`. +- Greedy `llama-completion`, 64 tokens behind a 3k prompt: Qwen3.8-27B-UD-IQ2_M gives `64e86551f7ef1638` with a device-resident cache and at `N = 0`, `1` and `4` with a host one. gemma-4-26B-A4B gives `5525e3f5ac7337d7` at `N = 0` and `1` with `-ts 50,50`, and `4f8986fb3655a567` at both with `-ts 55,45`. +- `test-llama-archs` adds a `Meta -nkvo -np 2 -kvpd 1` configuration, which stages the cache on the meta ring and delivers it through the ranged head-split write. It passes on 2, 3 and 4 CUDA devices. It evaluates one ubatch, so it delivers only the late part. +- `test-alloc` passes. Its meta test now covers a device of the meta type that is not the ggml meta backend, which stays ordered. ## Future work -- Fix `-sm tensor` with `--no-kv-offload` (above). Until then it should not be used: it is wrong rather than slow. -- Add the strided head-split delivery above, validate it, and then measure it. +- Allocate each device's part of a meta ring at its own share of the heads, instead of the whole slot on every device. diff --git a/docs/multi-gpu.md b/docs/multi-gpu.md index 3d57bfbbd461..fcab4546adc6 100644 --- a/docs/multi-gpu.md +++ b/docs/multi-gpu.md @@ -87,6 +87,7 @@ llama-cli -m model.gguf -sm tensor -ctk f16 -ctv f16 - KV cache types must be non-quantized: `f32`, `f16`, or `bf16`. Support for quantized KV cache is not implemented and trying to use it will result in an error. - Mark this configuration as experimental in your tooling: validate output quality before deploying. - `--no-kv-offload` works in this mode: the host-resident cache is split by attention head like the rest. Two limits: a backend without a native 2d copy (CPU, Metal) pays one transfer per cache cell, and Gemma 4 is less accurate this way (perplexity 235.03 with a host cache versus 227.23 with a device cache), so keep the cache on the devices for that architecture. +- With `--no-kv-offload --kv-cpu-pinned`, `--kv-pipeline-depth 1` delivers the cache while the previous split computes, as with one device. See [kv-transport-pipelining.md](kv-transport-pipelining.md#tensor-parallelism). - A recurrent or hybrid model always keeps its recurrent state on the devices in this mode, even with `--no-recurrent-state-offload`. The linear-attention op writes the state back together with its output, so the two do not agree on a host-resident state. - `--split-mode tensor`is not implemented for all architectures. The following will fail with *"LLAMA_SPLIT_MODE_TENSOR not implemented for architecture '...'"*: diff --git a/docs/repro/r4-kv-pipeline-ab.sh b/docs/repro/r4-kv-pipeline-ab.sh index d0e50427ab74..04e039e90f4c 100755 --- a/docs/repro/r4-kv-pipeline-ab.sh +++ b/docs/repro/r4-kv-pipeline-ab.sh @@ -3,11 +3,13 @@ # The two arms are the same binary: --kv-pipeline-depth 0 is the ordered path. # # LLAMA_KV_MODEL=/path/model.gguf docs/repro/r4-kv-pipeline-ab.sh [depth ...] +# LLAMA_KV_SM=tensor splits the model and its cache by head over every device. set -euo pipefail MODEL="${LLAMA_KV_MODEL:?set LLAMA_KV_MODEL to a .gguf path}" BUILD="${LLAMA_KV_BUILD:-build}" PIN="${LLAMA_KV_TASKSET:-0,2,4}" BUDGET="${LLAMA_KV_BUDGET:-512}" +SM="${LLAMA_KV_SM:-none}" LOCK=/tmp/beellama-single-gpu.lock # An unpinned host cache and a host-resident recurrent state both cost more than the transport can win back, and without a budget the ring is declined at the larger contexts, so a build without these options does not measure what the doc reports. @@ -30,7 +32,7 @@ run () { # $1 label, $2 pipeline depth, $3 context depth, $4 reps err="$(mktemp)" rc=0 taskset -c "$PIN" "$BUILD/bin/llama-bench" -m "$MODEL" --kv-pipeline-depth "$2" \ - --kv-pipeline-budget "$BUDGET" -ngl 99 -sm none -mg 0 -t 3 -nkvo 1 -kvcp 1 -rso 1 \ + --kv-pipeline-budget "$BUDGET" -ngl 99 -sm "$SM" -mg 0 -t 3 -nkvo 1 -kvcp 1 -rso 1 \ -fa on -ctk q8_0 -ctv q8_0 -b 512 -ub 512 --no-warmup -p 0 -n 128 -d "$3" -r "$4" -o json \ > "$out" 2> "$err" || rc=$? if [ "$rc" -eq 0 ]; then @@ -55,7 +57,7 @@ for D in "${DEPTHS[@]}"; do R=5 fi echo "== context depth=$D reps=$R" - flock "$LOCK" bash -c "set -euo pipefail; $(declare -f run); BUILD='$BUILD'; MODEL='$MODEL'; PIN='$PIN'; BUDGET='$BUDGET' + flock "$LOCK" bash -c "set -euo pipefail; $(declare -f run); BUILD='$BUILD'; MODEL='$MODEL'; PIN='$PIN'; BUDGET='$BUDGET'; SM='$SM' run ordered 0 $D $R run pipelined 1 $D $R run ordered2 0 $D $R diff --git a/docs/repro/r4-kv-pipeline-exact.sh b/docs/repro/r4-kv-pipeline-exact.sh index 017926d29a0c..ab929e746ce1 100755 --- a/docs/repro/r4-kv-pipeline-exact.sh +++ b/docs/repro/r4-kv-pipeline-exact.sh @@ -4,6 +4,7 @@ # # LLAMA_KV_MODEL=/path/model.gguf docs/repro/r4-kv-pipeline-exact.sh [pipeline-depth ...] # LLAMA_KV_LENGTHS=2048,18432,65536 selects the prefill lengths (default 2048,18432). +# LLAMA_KV_SM=tensor splits the model and its cache by head over every device. set -u MODEL="${LLAMA_KV_MODEL:?set LLAMA_KV_MODEL to a .gguf path}" BUILD="${LLAMA_KV_BUILD:-build}" @@ -12,6 +13,7 @@ PORT="${LLAMA_KV_PORT:-18099}" LENGTHS="${LLAMA_KV_LENGTHS:-2048,18432}" CTX="${LLAMA_KV_CTX:-32768}" BUDGET="${LLAMA_KV_BUDGET:-512}" +SM="${LLAMA_KV_SM:-none}" HERE="$(cd "$(dirname "$0")" && pwd)" DEPTHS=(0 1 4); [ $# -gt 0 ] && DEPTHS=("$@") @@ -22,7 +24,7 @@ for I in "${!DEPTHS[@]}"; do echo "== pipeline depth=$D ctx=$CTX prefill lengths=$LENGTHS" LOG=$(mktemp /tmp/r4-kv-pipeline.XXXX.log) taskset -c "$PIN" "$BUILD/bin/llama-server" -m "$MODEL" --kv-pipeline-depth "$D" \ - --kv-pipeline-budget "$BUDGET" -ngl 99 -sm none -mg 0 -t 3 -nkvo --kv-cpu-pinned --recurrent-state-offload \ + --kv-pipeline-budget "$BUDGET" -ngl 99 -sm "$SM" -mg 0 -t 3 -nkvo --kv-cpu-pinned --recurrent-state-offload \ -fa on -ctk q8_0 -ctv q8_0 -b 512 -ub 512 -c "$CTX" --parallel 1 \ --host 127.0.0.1 --port "$PORT" --no-warmup > "$LOG" 2>&1 & SRV=$! diff --git a/ggml/src/ggml-backend-impl.h b/ggml/src/ggml-backend-impl.h index 36abb660535b..27c7340a605a 100644 --- a/ggml/src/ggml-backend-impl.h +++ b/ggml/src/ggml-backend-impl.h @@ -101,6 +101,9 @@ extern "C" { GGML_API size_t ggml_backend_meta_n_backends (ggml_backend_t meta_backend); GGML_API ggml_backend_t ggml_backend_meta_simple_backend(ggml_backend_t meta_backend, size_t index); + // a meta backend without a communicator, for moving data on streams of its own: a graph it computes reduces through copies + GGML_API ggml_backend_t ggml_backend_meta_init_transfer(ggml_backend_dev_t meta_dev); + // temporary workaround to statically allocate tensors from a context in a deduplicated way: GGML_API struct ggml_backend_buffer * ggml_backend_meta_alloc_ctx_tensors_from_buft(struct ggml_context * ctx, ggml_backend_buffer_type_t buft); diff --git a/ggml/src/ggml-backend-meta.cpp b/ggml/src/ggml-backend-meta.cpp index 23dc91dee3d4..c192c8a966cb 100644 --- a/ggml/src/ggml-backend-meta.cpp +++ b/ggml/src/ggml-backend-meta.cpp @@ -151,6 +151,12 @@ static ggml_backend_buffer_type_t ggml_backend_meta_device_get_buffer_type(ggml_ static ggml_backend_buffer_type_t ggml_backend_meta_device_get_host_buffer_type(ggml_backend_dev_t dev); +static ggml_backend_event_t ggml_backend_meta_device_event_new(ggml_backend_dev_t dev); + +static void ggml_backend_meta_device_event_free(ggml_backend_dev_t dev, ggml_backend_event_t event); + +static void ggml_backend_meta_device_event_synchronize(ggml_backend_dev_t dev, ggml_backend_event_t event); + static bool ggml_backend_meta_device_supports_op(ggml_backend_dev_t dev, const ggml_tensor * op) { GGML_ASSERT(ggml_backend_dev_is_meta(dev)); const ggml_backend_meta_device_context * meta_dev_ctx = (const ggml_backend_meta_device_context *) dev->context; @@ -190,9 +196,9 @@ static const ggml_backend_device_i ggml_backend_meta_device_iface = { /* .supports_op = */ ggml_backend_meta_device_supports_op, /* .supports_buft = */ ggml_backend_meta_device_supports_buft, /* .offload_op = */ nullptr, - /* .event_new = */ nullptr, - /* .event_free = */ nullptr, - /* .event_synchronize = */ nullptr, + /* .event_new = */ ggml_backend_meta_device_event_new, + /* .event_free = */ ggml_backend_meta_device_event_free, + /* .event_synchronize = */ ggml_backend_meta_device_event_synchronize, }; static bool ggml_backend_dev_is_meta(ggml_backend_dev_t dev) { @@ -212,6 +218,45 @@ static ggml_backend_dev_t ggml_backend_meta_dev_simple_dev(ggml_backend_dev_t me return meta_dev_ctx->simple_devs[index]; } +// a meta event is one event per simple device, in the order of the simple devices +static ggml_backend_event_t ggml_backend_meta_device_event_new(ggml_backend_dev_t dev) { + const size_t n_devs = ggml_backend_meta_dev_n_devs(dev); + auto * events = new std::vector(); + events->reserve(n_devs); + for (size_t i = 0; i < n_devs; i++) { + ggml_backend_event_t event = ggml_backend_event_new(ggml_backend_meta_dev_simple_dev(dev, i)); + if (event == nullptr) { + for (ggml_backend_event_t e : *events) { + ggml_backend_event_free(e); + } + delete events; + return nullptr; + } + events->push_back(event); + } + return new ggml_backend_event { + /* .device = */ dev, + /* .context = */ events, + }; +} + +static void ggml_backend_meta_device_event_free(ggml_backend_dev_t dev, ggml_backend_event_t event) { + GGML_UNUSED(dev); + auto * events = (std::vector *) event->context; + for (ggml_backend_event_t e : *events) { + ggml_backend_event_free(e); + } + delete events; + delete event; +} + +static void ggml_backend_meta_device_event_synchronize(ggml_backend_dev_t dev, ggml_backend_event_t event) { + GGML_UNUSED(dev); + for (ggml_backend_event_t e : *(std::vector *) event->context) { + ggml_backend_event_synchronize(e); + } +} + ggml_backend_dev_t ggml_backend_meta_device( ggml_backend_dev_t * devs, size_t n_devs, ggml_backend_meta_get_split_state_t get_split_state, void * get_split_state_ud) { GGML_ASSERT(n_devs <= GGML_BACKEND_META_MAX_DEVICES); @@ -1909,7 +1954,7 @@ struct ggml_backend_meta_context { void * comm_ctx = nullptr; ggml_backend_comm_allreduce_tensor_t comm_allreduce = nullptr; - ggml_backend_meta_context(ggml_backend_dev_t meta_dev, const char * params) { + ggml_backend_meta_context(ggml_backend_dev_t meta_dev, const char * params, bool init_comm) { const size_t n_devs = ggml_backend_meta_dev_n_devs(meta_dev); n_reduce_steps = std::ceil(std::log2(n_devs)); name = "Meta("; @@ -1927,7 +1972,7 @@ struct ggml_backend_meta_context { } name += ")"; - if (n_devs > 1) { + if (n_devs > 1 && init_comm) { ggml_backend_comm_init_t comm_init = (ggml_backend_comm_init_t) ggml_backend_reg_get_proc_address( ggml_backend_dev_backend_reg(ggml_backend_get_device(simple_backends[0])), "ggml_backend_comm_init"); if (comm_init != nullptr) { @@ -1968,12 +2013,68 @@ static void ggml_backend_meta_free(ggml_backend_t backend) { delete backend; } +// A host-resident attention cache copy, [head_dim, n_kv, n_head_kv, n_stream], split on its heads, which are interleaved per cell. +static bool ggml_backend_meta_is_strided_head_split(const ggml_tensor * tensor, const ggml_backend_meta_split_state & split_state) { + return !ggml_is_contiguous(tensor) && + split_state.axis == GGML_BACKEND_SPLIT_AXIS_2 && + split_state.n_segments == 1 && split_state.nr[0] == 1 && + tensor->nb[1] > tensor->nb[2]; +} + +// Write whole cells of a strided head split, n_copies times, from one stream of the copy to the next. +// Each device takes its own run of heads from every cell: one 2d copy per stream per device. +static void ggml_backend_meta_set_head_split_async(ggml_backend_t backend, ggml_tensor * tensor, + const ggml_backend_meta_split_state & split_state, const void * data, size_t offset, size_t size, + size_t n_copies, size_t stride_tensor, size_t stride_data) { + const size_t n_backends = ggml_backend_meta_n_backends(backend); + const size_t nb1 = tensor->nb[1]; + const size_t nb3 = tensor->nb[3]; + + GGML_ASSERT(n_copies <= 1 || stride_tensor == nb3); + GGML_ASSERT(nb3 > 0 && (offset % nb3) % nb1 == 0 && size % nb1 == 0); + + const size_t i3 = offset / nb3; + const size_t i1 = (offset % nb3) / nb1; + const size_t n_rows = size / nb1; + + size_t head_offset = 0; + for (size_t j = 0; j < n_backends; j++) { + const size_t nbytes = split_state.ne[j] * tensor->nb[2]; + if (nbytes == 0) { + continue; + } + ggml_backend_t simple_backend = ggml_backend_meta_simple_backend(backend, j); + ggml_tensor * simple_tensor = ggml_backend_meta_buffer_simple_tensor(tensor, j); + for (size_t k = 0; k < std::max(n_copies, 1); k++) { + ggml_backend_tensor_set_2d_async(simple_backend, simple_tensor, + (const char *) data + k*stride_data + head_offset, + (i3 + k)*simple_tensor->nb[3] + i1*simple_tensor->nb[1], nbytes, + n_rows, simple_tensor->nb[1], nb1); + } + head_offset += nbytes; + } + GGML_ASSERT(head_offset == (size_t) tensor->ne[2] * tensor->nb[2]); +} + static void ggml_backend_meta_set_tensor_async(ggml_backend_t backend, ggml_tensor * tensor, const void * data, size_t offset, size_t size) { const size_t n_backends = ggml_backend_meta_n_backends(backend); + const ggml_backend_meta_split_state split_state = ggml_backend_meta_get_split_state(tensor, /*assume_sync =*/ false); + + // a pipelined delivery of a host-resident cache writes part of the window + if (ggml_backend_meta_is_strided_head_split(tensor, split_state)) { + ggml_backend_meta_set_head_split_async(backend, tensor, split_state, data, offset, size, 1, 0, 0); + return; + } + if (split_state.axis == GGML_BACKEND_SPLIT_AXIS_MIRRORED) { + for (size_t j = 0; j < n_backends; j++) { + ggml_backend_tensor_set_async( + ggml_backend_meta_simple_backend(backend, j), ggml_backend_meta_buffer_simple_tensor(tensor, j), data, offset, size); + } + return; + } + GGML_ASSERT(offset == 0); GGML_ASSERT(ggml_is_contiguous(tensor)); - - const ggml_backend_meta_split_state split_state = ggml_backend_meta_get_split_state(tensor, /*assume_sync =*/ false); GGML_ASSERT(split_state.n_segments == 1); GGML_ASSERT(split_state.nr[0] == 1); @@ -2013,6 +2114,27 @@ static void ggml_backend_meta_set_tensor_async(ggml_backend_t backend, ggml_tens } } +static void ggml_backend_meta_set_tensor_2d_async(ggml_backend_t backend, ggml_tensor * tensor, const void * data, + size_t offset, size_t size, size_t n_copies, size_t stride_tensor, size_t stride_data) { + const size_t n_backends = ggml_backend_meta_n_backends(backend); + const ggml_backend_meta_split_state split_state = ggml_backend_meta_get_split_state(tensor, /*assume_sync =*/ false); + + if (ggml_backend_meta_is_strided_head_split(tensor, split_state)) { + ggml_backend_meta_set_head_split_async(backend, tensor, split_state, data, offset, size, n_copies, stride_tensor, stride_data); + return; + } + if (split_state.axis == GGML_BACKEND_SPLIT_AXIS_MIRRORED) { + for (size_t j = 0; j < n_backends; j++) { + ggml_backend_tensor_set_2d_async(ggml_backend_meta_simple_backend(backend, j), ggml_backend_meta_buffer_simple_tensor(tensor, j), + data, offset, size, n_copies, stride_tensor, stride_data); + } + return; + } + for (size_t i = 0; i < n_copies; i++) { + ggml_backend_meta_set_tensor_async(backend, tensor, (const char *) data + i*stride_data, offset + i*stride_tensor, size); + } +} + static void ggml_backend_meta_get_tensor_async(ggml_backend_t backend, const ggml_tensor * tensor, void * data, size_t offset, size_t size) { const size_t n_backends = ggml_backend_meta_n_backends(backend); GGML_ASSERT(offset == 0); @@ -2065,6 +2187,23 @@ static void ggml_backend_meta_synchronize(ggml_backend_t backend) { } } +static void ggml_backend_meta_event_record(ggml_backend_t backend, ggml_backend_event_t event) { + const auto & events = *(std::vector *) event->context; + GGML_ASSERT(events.size() == ggml_backend_meta_n_backends(backend)); + for (size_t j = 0; j < events.size(); j++) { + ggml_backend_event_record(events[j], ggml_backend_meta_simple_backend(backend, j)); + } +} + +// each simple backend waits for the event of its own device +static void ggml_backend_meta_event_wait(ggml_backend_t backend, ggml_backend_event_t event) { + const auto & events = *(std::vector *) event->context; + GGML_ASSERT(events.size() == ggml_backend_meta_n_backends(backend)); + for (size_t j = 0; j < events.size(); j++) { + ggml_backend_event_wait(ggml_backend_meta_simple_backend(backend, j), events[j]); + } +} + static enum ggml_status ggml_backend_meta_graph_compute(ggml_backend_t backend, struct ggml_cgraph * cgraph) { GGML_ASSERT(cgraph->grads == nullptr); const size_t n_backends = ggml_backend_meta_n_backends(backend); @@ -2096,6 +2235,13 @@ static enum ggml_status ggml_backend_meta_graph_compute(ggml_backend_t backend, if (ggml_backend_buffer_is_meta(cgraph->nodes[i]->buffer)) { used_buffers.emplace(cgraph->nodes[i]->buffer); } + // the copies in a transport ring are only sources, but their buffer holds compute tensors that must rotate too + for (int k = 0; k < GGML_MAX_SRC; k++) { + const ggml_tensor * src = cgraph->nodes[i]->src[k]; + if (src != nullptr && ggml_backend_buffer_is_meta(src->buffer)) { + used_buffers.emplace(src->buffer); + } + } } for (ggml_backend_buffer_t buf : used_buffers) { ggml_backend_meta_buffer_context * buf_ctx = (ggml_backend_meta_buffer_context *) buf->context; @@ -2577,7 +2723,7 @@ static const ggml_backend_i ggml_backend_meta_i = { /* .free = */ ggml_backend_meta_free, /* .set_tensor_async = */ ggml_backend_meta_set_tensor_async, /* .get_tensor_async = */ ggml_backend_meta_get_tensor_async, - /* .set_tensor_2d_async = */ nullptr, + /* .set_tensor_2d_async = */ ggml_backend_meta_set_tensor_2d_async, /* .get_tensor_2d_async = */ nullptr, /* .cpy_tensor_async = */ nullptr, /* .synchronize = */ ggml_backend_meta_synchronize, @@ -2586,8 +2732,8 @@ static const ggml_backend_i ggml_backend_meta_i = { /* .graph_plan_update = */ nullptr, /* .graph_plan_compute = */ nullptr, /* .graph_compute = */ ggml_backend_meta_graph_compute, - /* .event_record = */ nullptr, - /* .event_wait = */ nullptr, + /* .event_record = */ ggml_backend_meta_event_record, + /* .event_wait = */ ggml_backend_meta_event_wait, /* .graph_optimize = */ nullptr, }; @@ -2595,8 +2741,8 @@ bool ggml_backend_is_meta(ggml_backend_t backend) { return backend != nullptr && backend->iface.get_name == ggml_backend_meta_i.get_name; } -static ggml_backend_t ggml_backend_meta_device_init_backend(ggml_backend_dev_t dev, const char * params) { - ggml_backend_meta_context * backend_ctx = new ggml_backend_meta_context(dev, params); +static ggml_backend_t ggml_backend_meta_init_impl(ggml_backend_dev_t dev, const char * params, bool init_comm) { + ggml_backend_meta_context * backend_ctx = new ggml_backend_meta_context(dev, params, init_comm); ggml_backend_t backend = new struct ggml_backend; backend->guid = ggml_backend_meta_guid(); @@ -2606,6 +2752,15 @@ static ggml_backend_t ggml_backend_meta_device_init_backend(ggml_backend_dev_t d return backend; } +static ggml_backend_t ggml_backend_meta_device_init_backend(ggml_backend_dev_t dev, const char * params) { + return ggml_backend_meta_init_impl(dev, params, /*init_comm =*/ true); +} + +ggml_backend_t ggml_backend_meta_init_transfer(ggml_backend_dev_t dev) { + GGML_ASSERT(ggml_backend_dev_is_meta(dev)); + return ggml_backend_meta_init_impl(dev, nullptr, /*init_comm =*/ false); +} + size_t ggml_backend_meta_n_backends(ggml_backend_t meta_backend) { GGML_ASSERT(ggml_backend_is_meta(meta_backend)); const ggml_backend_meta_context * backend_ctx = (const ggml_backend_meta_context *) meta_backend->context; diff --git a/ggml/src/ggml-backend.cpp b/ggml/src/ggml-backend.cpp index 458c7e150f7b..acd1ab108dbe 100644 --- a/ggml/src/ggml-backend.cpp +++ b/ggml/src/ggml-backend.cpp @@ -2165,6 +2165,26 @@ static size_t ggml_backend_sched_transport_slot_alloc(size_t need, size_t limit, return std::max(size, need); } +// Free memory of the device a ring lands on, 0 when unknown. +// A meta buffer allocates the whole ring on every simple device, so the device with the least free memory decides; the meta device reports the sum. +static size_t ggml_backend_sched_transport_dev_free(ggml_backend_t backend) { + size_t dev_free = 0, dev_total = 0; + if (ggml_backend_is_meta(backend)) { + for (size_t j = 0; j < ggml_backend_meta_n_backends(backend); j++) { + ggml_backend_dev_t dev = ggml_backend_get_device(ggml_backend_meta_simple_backend(backend, j)); + size_t simple_free = 0, simple_total = 0; + ggml_backend_dev_memory(dev, &simple_free, &simple_total); + dev_free = j == 0 ? simple_free : std::min(dev_free, simple_free); + } + return dev_free; + } + ggml_backend_dev_t dev = ggml_backend_get_device(backend); + if (dev != NULL) { + ggml_backend_dev_memory(dev, &dev_free, &dev_total); + } + return dev_free; +} + // Created on demand, so a backend that never gets to stage anything does not carry a second device context for nothing. static bool ggml_backend_sched_transport_ensure_backend(ggml_backend_sched_t sched, int backend_id) { struct ggml_backend_sched_transport * tr = &sched->transport; @@ -2179,7 +2199,8 @@ static bool ggml_backend_sched_transport_ensure_backend(ggml_backend_sched_t sch return false; } - ggml_backend_t transfer = ggml_backend_dev_init(dev, NULL); + // a transfer backend never computes, so a meta one goes without the communicator a second set of streams would otherwise start + ggml_backend_t transfer = ggml_backend_is_meta(sched->backends[backend_id]) ? ggml_backend_meta_init_transfer(dev) : ggml_backend_dev_init(dev, NULL); if (transfer == NULL) { return false; } @@ -2484,11 +2505,7 @@ static void ggml_backend_sched_transport_plan(ggml_backend_sched_t sched) { ggml_backend_buffer_type_t buft = sched->bufts[bid]; // the graph allocator reserved before this, so leave it the room its buffers may still grow into - ggml_backend_dev_t dev = ggml_backend_get_device(sched->backends[bid]); - size_t dev_free = 0, dev_total = 0; - if (dev != NULL) { - ggml_backend_dev_memory(dev, &dev_free, &dev_total); - } + const size_t dev_free = ggml_backend_sched_transport_dev_free(sched->backends[bid]); if (dev_free > 0 && (dev_free <= GGML_SCHED_TRANSPORT_HEADROOM || ring_size > dev_free - GGML_SCHED_TRANSPORT_HEADROOM)) { if (!r->reported_no_room) { GGML_LOG_WARN("%s: transport ring on %s would need %zu MiB and leave less than " @@ -2949,8 +2966,7 @@ static enum ggml_status ggml_backend_sched_compute_splits(ggml_backend_sched_t s ggml_backend_buffer_t src_buf = input->view_src ? input->view_src->buffer : input->buffer; struct ggml_backend_sched_ranges rg; ggml_backend_sched_input_ranges(input, input_cpy, &rg); - // a meta backend writes a whole contiguous tensor, it cannot take one range per stream - const bool ranged = rg.n > 1 && src_buf != NULL && ggml_backend_buffer_is_host(src_buf) && !ggml_backend_is_meta(split_backend); + const bool ranged = rg.n > 1 && src_buf != NULL && ggml_backend_buffer_is_host(src_buf); // try async copy, but if not possible, we can still use a sync copy without synchronizing the dst backend, since we handle the synchronization here with multiple copies and events // TODO: add public function to facilitate this, since applications do not have direct access to the backend interface @@ -3227,6 +3243,23 @@ static void ggml_backend_sched_transport_teardown(ggml_backend_sched_t sched) { } +static bool ggml_backend_sched_transport_backend_supported(ggml_backend_t backend) { + ggml_backend_dev_t dev = ggml_backend_get_device(backend); + if (dev == NULL) { + return false; + } + + ggml_backend_reg_t reg = ggml_backend_dev_backend_reg(dev); + if (reg == NULL || strcmp(ggml_backend_reg_name(reg), "CUDA") != 0) { + return false; + } + + return backend->iface.set_tensor_async != NULL && + backend->iface.event_record != NULL && + backend->iface.event_wait != NULL && + dev->iface.event_new != NULL; +} + bool ggml_backend_sched_set_transport_pipeline_depth(ggml_backend_sched_t sched, int depth) { GGML_ASSERT(sched); @@ -3272,22 +3305,22 @@ bool ggml_backend_sched_set_transport_pipeline_depth(ggml_backend_sched_t sched, continue; } const enum ggml_backend_dev_type type = ggml_backend_dev_type(dev); - if (type == GGML_BACKEND_DEVICE_TYPE_META || type == GGML_BACKEND_DEVICE_TYPE_CPU) { + if (type == GGML_BACKEND_DEVICE_TYPE_CPU) { continue; } - ggml_backend_reg_t reg = ggml_backend_dev_backend_reg(dev); - if (reg == NULL || strcmp(ggml_backend_reg_name(reg), "CUDA") != 0) { - continue; - } - - if (backend->iface.set_tensor_async == NULL || - backend->iface.event_record == NULL || - backend->iface.event_wait == NULL) { - continue; + // a meta backend delivers through its simple backends, so each of them must qualify on its own + bool can_transport = true; + if (ggml_backend_is_meta(backend)) { + for (size_t j = 0; j < ggml_backend_meta_n_backends(backend) && can_transport; j++) { + can_transport = ggml_backend_sched_transport_backend_supported(ggml_backend_meta_simple_backend(backend, j)); + } + } else if (type == GGML_BACKEND_DEVICE_TYPE_META) { + can_transport = false; + } else { + can_transport = ggml_backend_sched_transport_backend_supported(backend); } - - if (dev->iface.event_new == NULL) { + if (!can_transport) { continue; } diff --git a/tests/test-alloc.cpp b/tests/test-alloc.cpp index 4b8e5e69556b..b49b25357322 100644 --- a/tests/test-alloc.cpp +++ b/tests/test-alloc.cpp @@ -2120,6 +2120,7 @@ static void test_transport_stops_after_backend_failure() { GGML_ASSERT(deliveries == 0); } +// only the ggml meta backend is reached through its simple backends, another device of the meta type stays ordered static void test_transport_excludes_meta() { dummy_backend meta = dummy_backend_init(SIZE_MAX, 8, true, GGML_BACKEND_DEVICE_TYPE_META, "CUDA", false); dummy_backend cpu = dummy_backend_init(SIZE_MAX, 8, true); diff --git a/tests/test-llama-archs.cpp b/tests/test-llama-archs.cpp index cf193bf168e6..20f41f005524 100644 --- a/tests/test-llama-archs.cpp +++ b/tests/test-llama-archs.cpp @@ -402,8 +402,9 @@ static bool silent_model_load_progress(float /*progress*/, void * /*user_data*/) // with offload_kqv=false the cache lives in host memory // n_seq_max > 1 gives the cache one stream per sequence struct kv_config { - bool offload_kqv = true; - uint32_t n_seq_max = 1; + bool offload_kqv = true; + uint32_t n_seq_max = 1; + uint32_t kv_pipeline_depth = 0; }; static std::pair get_model_and_ctx( @@ -423,6 +424,7 @@ static std::pair get_model_and_ctx( ctx_params.n_threads_batch = 4; ctx_params.offload_kqv = kvc.offload_kqv; ctx_params.n_seq_max = kvc.n_seq_max; + ctx_params.kv_pipeline_depth = kvc.kv_pipeline_depth; if (!encode) { ctx_params.n_ubatch = 64; } @@ -1459,6 +1461,11 @@ static int test_backends(const llm_arch target_arch, const size_t seed, const in kvc_host_streams.n_seq_max = 2; dev_configs.emplace_back(devices_meta, "Meta -nkvo -np 2", LLAMA_SPLIT_MODE_TENSOR, kvc_host_streams); + // the same copy, delivered ahead of the split that reads it + kv_config kvc_host_pipelined = kvc_host_streams; + kvc_host_pipelined.kv_pipeline_depth = 1; + dev_configs.emplace_back(devices_meta, "Meta -nkvo -np 2 -kvpd 1", LLAMA_SPLIT_MODE_TENSOR, kvc_host_pipelined); + for (const device_config & dc : dev_configs) { max_device_label_length = std::max(max_device_label_length, dc.label.length()); }