LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
This is an experimental fork built for one specific purpose: serving a Qwen3.6-27B
dense hybrid SSM/attention model ("f711") and a Qwen3.8-27B vision model in production
on an AMD Radeon AI PRO R9700 (32 GB, gfx1201/RDNA4) via ROCm/HIP, driving
Claude Code through the /v1/messages
endpoint.
It combines two upstream sources:
- ggml-org/llama.cpp — upstream, tracked
close to
master. - stew675/llama.cpp:rdna-boosts —
a set of RDNA-specific kernel fusions (
MUL_MAT,FLASH_ATTN_EXT,RMS_NORM,MUL) ported on top of (1).
...plus a single-target CI workflow (.github/workflows/f711-rocm.yml, ROCm 10.0.0
by default — 7.14 until Sep 2026, when AMD's whl-next index stopped serving it —
AMDGPU_TARGETS=gfx1201 only) so a build finishes in minutes instead of covering
every GPU target upstream CI builds for.
Upstream llama.cpp on this GPU/architecture combination (gfx1201/RDNA4, a dense
Qwen3.5/3.6-family hybrid SSM+attention model) left a lot of prefill throughput on
the table. rdna-boosts closes most of that gap. Before deploying it we ran two
independent correctness gates (not just a speed benchmark): test-backend-ops on
every op the branch touches, and a perplexity/code-quality comparison against a
control build differing only by those commits — see Results below.
Measured on production hardware, 15–16 Aug 2026. Hardware: AMD Radeon AI PRO
R9700 (32 GB VRAM, gfx1201/RDNA4), ROCm 7.14. Model: f711-AMD-Q6_K
(Qwen3.6-27B, dense, n_expert = 0, hybrid SSM/gated-delta-net + MTP head).
Speed — real 100k-token prompt, production -c 131072:
| ctx | arm | prefill | TTFT | gen | shared VRAM (spill) |
|---|---|---|---|---|---|
| 131072 | base | 312.4 t/s | 5m22s | 18.4 t/s | 242 MiB |
| 131072 | rdna-boosts | 573.9 t/s (+84%) | 2m55s | 18.5 t/s | 244 MiB |
| 147456 | base | collapses (>10 min) | — | — | 768 MiB (over the spill threshold) |
| 147456 | rdna-boosts | 519.6 t/s | 3m13s | 18.6 t/s | 542 MiB (still under threshold) |
rdna-boosts is also more VRAM-efficient at high context — it raises the
spill-collapse threshold instead of only being faster at the same one.
Correctness / quality:
test-backend-opson the ops the branch changes: 1194/1194 (MUL_MAT) and 4552/4552 (FLASH_ATTN_EXT) passing on gfx1201.- Perplexity, production config (
-fa on, fusions on): base 2.6898 → rdna-boosts 2.6992 (+0.35%). Bisection isolated the cause to one commit,6e478a115("fuse IMRoPE + set-rows for BF16 KV cache"), which mis-fires on plain f16 KV cache where it shouldn't apply — reported upstream, kept deployed anyway since it doesn't regress generated-code quality and the speed win is large. As an independent cross-check, the same model on the Vulkan backend (which shares none of these kernels) measures PPL 2.7124 — i.e. the base→rdna-boosts gap is smaller than the ordinary HIP↔Vulkan backend gap. - Code-generation quality benchmark (21 runs: generate C#/.NET code,
dotnet build- regex assertions against a pre-validated skeleton): base 16/21 build-ok / 98.3% asserts vs. rdna-boosts 15/21 / 99.2% asserts — no measurable regression (the one-build difference is within run-to-run noise at temperature 0.7).
Status: deployed in production since 16 Aug 2026, serving both models above.
The branch currently in production, and the default branch, is
f711-rdna-b10665-stack0911-nossm-maskskipv2-cachefix (still based on upstream b10665; see Changes since the b10665 rebase).
Earlier stages: f711-rdna-b10665-chatfix (rebased onto b10665 on 28 Aug 2026)
and the original deployment f711-rdna / build tag rdna-20260815.
Two later upstream commits from the same rdna-boosts branch (1b009339e,
7955770b2) were evaluated on 16 Aug and not deployed — correct, but no
measurable speed gain (−0.33%, within noise) on this model/GPU; not worth the
extra rebase-conflict surface.
What the stack gained on top of f711-rdna-b10665-chatfix (30 Aug – 22 Sep 2026):
Flash attention on RDNA4 WMMA (head size 256)
a88da45de/0fe0e31eb: K/V loads skip the LDS staging whenDKQ > 128. The follow-up fixes atile_maskwrite-after-read race that this exposed (reported in the #26419 review).FLASH_ATTN_EXTpassed 2920/2920 on gfx1201 in 5 consecutive runs.af43ef7ef: prefer whole-tile FA grids over stream-k on AMD WMMA. At 1 block/SM, stream-k collapses to 32 blocks on the R9700, measured at 16.4 vs 27 TFLOPS (DKQ=DV=256,-ub 512). NVIDIA behaviour is unchanged.efbb900c0: head-256 tuning from upstream #28102, adapted to our stack. Our wider head-size gates (up to 576) are kept, and so is our own(256, 256, 64)tuning.f8372c0c8: packed mask classes in WMMA FA, a V2 port of upstream #28943 (not merged upstream). A helper kernel classifies every (query tile, KV block) as fully masked, all-zero or mixed. Fully masked runs are skipped, and all-zero runs skip the mask load. Unlike the upstreamKV_maxhelper, this also fires on our production shape (-ub 512 --kv-unified). It replaces the earlier V1 port (e17236d92, reverted in404cf5282), andb11c1c90fadds the V2 patch's mask-pattern cases totest-backend-ops.
Correctness fixes
9c9d65afd: fixes the 7 Sep 2026 production abortmmvq.cu: GGML_ASSERT(ids || dst->ne[1] == 1). Themul_mat + addfusion through a view fired when the reshape moved tokens across dimensions. That happened under-np 2 --kv-unifiedwhenever both slots contributed prefill to one batch. The fix was located withe2bea2ccf/2c53c00ae, which add diagnostic-only fusion-site IDs, aGGML_CUDA_FUSION_MASKenv var that disables individual sites, and a diagnostic written unbuffered to stderr.814a00830: reverts the fused SSM gate/beta projection kernel. It fired only at batch width 1, so decode and speculative verify ran different math. It caused 80 vs 14 top-1 flips per 512 positions against the #28768 harness and gave no measurable speed gain (the fused arm was 0.79% slower).0a124702e(upstream #28068): the GDN q/k normalization now uses FLA'sx * rsqrt(sum(x²) + eps)instead ofx / max(‖x‖, eps).
Server
2acdede5c(#27624): clear stale prompt, checkpoint and KV/recurrent state when an LRU-selected slot is reused without a valid restored state.d54c21732(port of #28992): consult the RAM prompt cache even when the outgoing slot state is not worth saving. The server now looks up a returning, evicted conversation in the cache instead of re-prefilling it.e1cf8d36b(#28302): checkpoint min-step eviction runs only when the checkpoint list is full, so short prompts keep their resume checkpoint on hybrid models.f5130b07a(#28715): pass the draft model the correct position after an image when speculating.
CI
7f47acfa8/f5b6a8e3a: ROCm wheels now come fromwhl-next, and the default isrocm_version = 10.0.0, which the production binaries are built against.
A rebase onto upstream b10930 is in progress on f711-rdna-b10930-r0 and is not
yet the default branch.
The stack is re-validated on every rebase rather than assumed neutral. The
b10488 → b10665 rebase (177 upstream commits, 28 Aug 2026) is the reference
example — the full measurement record is in
docs/f711-r9700/rebase-b10665-2026-08-28.md
(in Slovak):
- Three of our 28 commits were dropped: two had landed upstream (
#27679,#27404), and ourattn_gatetensor-parallel granularity fix was superseded by an upstream branch forpattern_attn_gate_weightthat does the same thing. - One conflict needed a real merge:
ggml/src/ggml-cuda/rope.cu, where upstream's newn_offs/inplace(ggml_rope_set_offset) meets ourD-cast output plus ROPE+VIEW+SET_ROWS fusion inrope_multi. Resolved as a union of both, mirroringrope_neoxin the same file, which merged cleanly and already carries both. - Because that merge was hand-written, compiling was not treated as evidence. Three gates, all on production hardware: perplexity unchanged on both models (f711 2.6992, Qwen3.8 1.8577 — identical to the pre-rebase build), vision 3/3 assertions on both test images, and an interleaved 12-iteration speed A/B.
- On the speed A/B: raw
gen_tpsis not readable on a build with an MTP draft head. Draft acceptance swung 37.3–74.8% within a single arm, which is larger than the difference between arms. Normalising to actual forward passes per second —gen_tps × (gen_tok − accepted) / gen_tok— gives 9.78–9.89 across all 12 iterations with the arms overlapping, i.e. no change.
f711-rocm.yml packs build/bin only. That is not a runnable ROCm tree: it
lacks rocblas.dll, hipblas.dll, rocsolver.dll, libhipblaslt.dll and the
Tensile kernel directories (rocblas/library/, hipblaslt/library/gfx1201/),
without which --list-devices reports no device at all. We fill those in from a
fixed reference tree after unpacking, which also keeps rocBLAS constant across
measured arms.
Upstream's release-job step from
#26973 (bundling
amdhip64_7.dll, rocm_kpack.dll, amd_comgr.dll) does not substitute for
this — it addresses a different problem (the driver's HIP runtime in System32
winning the loader search). It was ported here, measured, and reverted: with those
DLLs bundled the artifact still enumerated no device.
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

