Project status: Active development will be limited for the near future, and significant updates are unlikely in the short term. The cost of the AI models used for development—particularly DeepSeek—has increased, while my subscription-based model access is currently being prioritized for other projects. Maintaining several AI-heavy projects in parallel is therefore not financially practical at the moment.
The repository is not abandoned. I will continue to maintain it when necessary and may still make smaller fixes or improvements, but larger optimization and research work will resume when resources allow.
Hardware-focused fork of ggml-org/llama.cpp
for local AI on Windows with two AMD RDNA4 GPUs. It combines the llama.cpp
runtime with a local web GUI (GUI 2.0, gui2/), reproducible benchmark
tooling, long-context work, and AMD-specific Vulkan and ROCm/HIP
optimizations. The desktop application is branded RDNA LLM Studio.
The primary workload is agentic coding with Qwen3.8-27B: large cold prompts, single-user requests, long contexts, tool use, vision, and speculative decode. The main performance priority is prompt evaluation. MTP is kept only when its decode gain does not impose an unacceptable prefill cost.
Specialized research and production fork, not a drop-in replacement for every upstream platform. Results and defaults are tuned for the reference dual-RX 9070 XT machine. This file is the presentation and headline results; details live in the linked documents below.
- Performance — current peer tables, lane contracts, and model quality (PPL)
- Benchmarking — canonical benchmark methodology and history
- Fork Details — fork-only features, backend fixes, runtime profiles
- MTP — MTP behavior and practical rules
- Backends & Models — supported backends, model/format matrix, vision
- RPC Backend — remote GPU stack and measured results
- Active Branches — branch status and resume points
- Build Guide — CPU / Vulkan / ROCm requirements and commands
- Supported backends — backend policy
- Contributing, Development, Upstream sync
- License & Security — MIT, with SECURITY.md
| Area | Current focus |
|---|---|
| Host platform | Windows 11 on AMD AM4 (Linux ROCm 10 also used for research lanes) |
| Accelerators | 2x Radeon RX 9070 XT 16 GB (gfx1201) |
| Backends | ROCm/HIP, Vulkan, and CPU (+ remote GPU via RPC) |
| Primary model | Qwen3.8-27B Q4_K_M with MTP; MXFP4 for speed lanes |
| Main objective | Maximum cold prompt evaluation without sacrificing useful decode speed |
| Serving | OpenAI-compatible llama-server plus the local web GUI (GUI 2.0) |
Supported backends, model/format matrix and vision: docs/BACKENDS_AND_MODELS.md.
- Maximize Qwen3.8 prompt-evaluation throughput for agent workloads.
- Use both GPUs without moving the active working set into system RAM.
- Make MTP improve decode while keeping long-prompt prefill close to baseline.
- Provide a practical GUI for building, launching, monitoring, and autotuning.
- Keep performance claims reproducible through cold, lane-locked benchmarks.
- Keep the fork maintainable by carrying only useful backends and upstream changes.
- Windows 11, AMD Ryzen 7 5800X3D, 64 GB RAM
- 2x AMD Radeon RX 9070 XT, 16 GB VRAM each, RDNA4
gfx1201 - AMD ROCm/HIP SDK 7.1 (reference) and 7.2 (current builds) for Windows; AMD proprietary Vulkan driver
- Main model:
Qwen3.8-27B-Q4_K_M.gguf(UD variantQwen3.8-27B-UD-Q4_K_M.ggufin the tables below); Vision projector:mmproj-F16.gguf
The two GPUs are normally used with layer split, not tensor split. GPU1 is the preferred output device because GPU0 also drives the desktop. Device order is backend- and workload-sensitive; exact routes are recorded with each benchmark. See Active Branches for branch status.
Full tables, lane contracts and evidence links live in PERFORMANCE.md.
All rows use one server slot, FlashAttention, cold prompt processing, no
prompt-cache reuse, no prime pass, batch 8192 / ubatch 1024, KV
f8_e4m3 / f8_e4m3, device route ROCm1,ROCm0 -sm layer -ts 1,1, -ngl 999,
seed 42, temperature 0.2, top-p 0.9, --no-warmup. Three-tier agent
workload: L1/L2 use a repository snapshot, L3 uses a deterministic synthetic
context (repo-snapshot is capped at ~53K tokens).
| Lane | Context | Actual prompt | Output | Context source |
|---|---|---|---|---|
| L1 | 16,384 | ~8.4K | 128 | repo-snapshot |
| L2 | 49,152 | ~33.9K | 256 | repo-snapshot |
| L3 | 98,304 | ~64.3K | 256 | synthetic |
Fresh Windows re-runs on the current binary family: the ROCm rows on
fea1c3180 plus the working-tree flash-attention change (build-rocm72,
HIP 7.2), the Vulkan rows on 7c49e6212 (build-vulkan-gcc16, not re-run
in the 2026-09-26 pass). Same lane contract as the Linux table below:
batch 8192 / ubatch 1024, KV f8_e4m3 / f8_e4m3,
FlashAttention, -ngl 999, one slot, cold prompt, seed 42, temp 0.2,
top-p 0.9, --no-warmup, repo-snapshot L1/L2, synthetic L3. Model
Qwen3.8-27B-UD-Q4_K_M.gguf (UD), single run per cell; actual prompt
geometry L1 8386 tok, L2 32996 tok, L3 64287 tok (Linux L1/L2 differ by
<3%: 8450/33865).
| Backend | Lane | Spec | Prompt TPS | Decode TPS | Aggregate TPS | Acceptance |
|---|---|---|---|---|---|---|
| ROCm 7.2 | L1 | none | 2061.11 | 24.70 | 14.198 | - |
| ROCm 7.2 | L1 | MTP n3 | 1817.81 | 36.73 | 16.344 | 46.8% |
| ROCm 7.2 | L2 | none | 1879.06 | 25.74 | 9.758 | - |
| ROCm 7.2 | L2 | MTP n3 | 1799.29 | 43.75 | 11.197 | 68.1% |
| ROCm 7.2 | L3 | none | 1555.55 | 24.85 | 4.959 | - |
| ROCm 7.2 | L3 | MTP n3 | 1495.92 | 39.09 | 5.169 | 67.5% |
| Vulkan | L1 | none | 1569.17 | 25.42 | 12.332 | - |
| Vulkan | L1 | MTP n3 | 1606.98 | 39.00 | 15.059 | 53.4% |
| Vulkan | L2 | none | 1470.80 | 23.61 | 7.693 | - |
| Vulkan | L2 | MTP n3 | 1580.02 | 35.36 | 9.103 | 48.6% |
| Vulkan | L3 | none | 1243.99 | 22.19 | 4.050 | - |
| Vulkan | L3 | MTP n3 | 1357.37 | 35.14 | 4.685 | 58.3% |
ROCm stays ahead on prompt processing (L1 +31%, L2 +28%, L3 +25% over
Vulkan) and now leads decode on five of six rows. Vulkan still leads the L1
rows (25.42 vs 24.70, 39.00 vs 36.73); ROCm leads L2/L3 on both spec modes
(none 25.74 vs 23.61 and 24.85 vs 22.19, MTP 43.75 vs 35.36 and 39.09 vs
35.14). The f8_e4m3
KV policy is used on both backends; MTP rows carry the automatic Vulkan
last-8-f16 tail and the HIP n_ubatch=256 draft cap from the same commit.
ROCm decode re-measure (2026-09-26). The ROCm rows were re-run with the
identical lane contract and decode improved on five of six rows: L1 none
23.82 -> 24.70 (+3.7%), L2 none 24.00 -> 25.74 (+7.2%), L3 none
20.92 -> 24.85 (+18.8%), L2 MTP 36.10 -> 43.75 (+21.2%), L3 MTP
36.37 -> 39.09 (+7.5%). L1 MTP moved 38.22 -> 36.73 on a lower acceptance
(51.0% -> 46.8%), while the L2/L3 MTP rows gained acceptance
(50.0% -> 68.1%, 60.3% -> 67.5%); MTP decode tracks acceptance, so the
none rows are the clean signal. Vulkan was not re-run in this pass.
| Format | Lane | Spec | Prompt TPS | Decode TPS | Aggregate TPS | Acceptance |
|---|---|---|---|---|---|---|
| MXFP4-requant (dense) | L1 | none | 2253.06 | 29.38 | 15.787 | - |
| Q4_K_M (UD) | L1 | MTP n3 | 1836.97 | 50.13 | 18.674 | 78.1% |
| MXFP4-requant (UD) | L1 | MTP n3 | 1825.13 | 50.99 | 18.716 | 57.6% |
| Q4_K_M (UD) | L2 | none | 1765.82 | 23.53 | 8.517 | - |
| MXFP4-requant (UD) | L2 | none | 2076.46 | 25.92 | 9.777 | - |
| Q4_K_M (UD) | L2 | MTP n3 | 1712.48 | 39.94 | 10.542 | 62.9% |
| MXFP4-requant (UD) | L2 | MTP n3 | 1666.09 | 49.195 | 10.028 | 66.0% |
| MXFP4-hybrid-attnQ6 | L2 | none | 1769.91 | 24.26 | 8.623 | - |
| NVFP4-native | L2 | none | 488.95 | 24.86 | 3.218 | - |
| Q4_K_M (UD) | L3 | none | 1539.60 | 20.80 | 4.735 | - |
| MXFP4-requant (UD) | L3 | none | 1765.49 | 22.60 | 5.362 | - |
| Q4_K_M (UD) | L3 | MTP n3 | 1329.83 | 33.27 | 4.568 | 64.5% |
| MXFP4-requant (UD) | L3 | MTP n3 | 1502.55 | 33.68 | 5.081 | 52.4% |
| MXFP4-hybrid-attnQ6 | L3 | none | 1718.59 | 21.63 | 5.199 | - |
| NVFP4-native | L3 | none | 473.85 | 21.75 | 1.736 | - |
Current binary includes W24 (nwarps=8 for MXFP4 decode) and W26 (MXFP4
prefill routed to MMQ): MXFP4 prefill +17-21% and decode +3-5% vs before.
MTP n3 is the decode optimum; acceptance profile 0.800/0.600/0.432 (L2),
and all p_min/n_max/draft-up-quantization attempts are negative. MTP
acceptance falls at L3 (52-64%) and NVFP4-native is not usable for
prompt-heavy lanes; hybrid is quality-first. L1/L2 MTP rows were recorded on
the W20/W24 binary (pre-W26 prefill routing); details and exact artifacts in
PERFORMANCE.md.
Windows re-runs (2026-09-23) repeat every Linux row on both Windows backends
with the same lane contract. Cells are prompt / decode in tok/s.
| Format | Lane | Spec | Linux ROCm 10 | Windows ROCm 7.2 | Windows Vulkan |
|---|---|---|---|---|---|
| MXFP4-requant (dense) | L1 | none | 2253.06 / 29.38 | 1999.23 / 27.05 | 1746.42 / 28.06 |
| Q4_K_M (UD) | L1 | MTP n3 | 1836.97 / 50.13 | 1810.81 / 38.22 | 1606.98 / 39.00 |
| MXFP4-requant (UD) | L1 | MTP n3 | 1825.13 / 50.99 | 1825.71 / 44.77 | 1739.19 / 46.35 |
| Q4_K_M (UD) | L2 | none | 1765.82 / 23.53 | 1819.46 / 24.00 | 1470.80 / 23.61 |
| MXFP4-requant (UD) | L2 | none | 2076.46 / 25.92 | 1813.11 / 25.25 | 1577.90 / 25.95 |
| Q4_K_M (UD) | L2 | MTP n3 | 1712.48 / 39.94 | 1744.44 / 36.10 | 1580.02 / 35.36 |
| MXFP4-requant (UD) | L2 | MTP n3 | 1666.09 / 49.20 | 1751.81 / 41.49 | 1703.11 / 44.99 |
| MXFP4-hybrid-attnQ6 | L2 | none | 1769.91 / 24.26 | 1756.82 / 24.44 | 1542.28 / 24.67 |
| NVFP4-native | L2 | none | 488.95 / 24.86 | 79.05 / 24.65 | 1560.12 / 25.81 |
| Q4_K_M (UD) | L3 | none | 1539.60 / 20.80 | 1525.66 / 20.92 | 1243.99 / 22.19 |
| MXFP4-requant (UD) | L3 | none | 1765.49 / 22.60 | 1528.46 / 22.01 | 1322.88 / 23.48 |
| Q4_K_M (UD) | L3 | MTP n3 | 1329.83 / 33.27 | 1472.01 / 36.37 | 1357.37 / 35.14 |
| MXFP4-requant (UD) | L3 | MTP n3 | 1502.55 / 33.68 | 1482.71 / 36.78 | 1447.82 / 41.64 |
| MXFP4-hybrid-attnQ6 | L3 | none | 1718.59 / 21.63 | 1484.54 / 21.35 | 1295.55 / 22.61 |
| NVFP4-native | L3 | none | 473.85 / 21.75 | 79.71 / 22.35 | 1307.73 / 23.17 |
Read this table together with the caveats:
- Linux ROCm rows used GPU peer-to-peer (P2P); ROCm on Windows does not provide P2P, so the Windows ROCm column is a non-P2P baseline. The Windows-vs-Linux delta mixes the OS/driver change with the loss of P2P for cross-device layer traffic.
- NVFP4-native is broken on Windows ROCm: prompt throughput collapses to 79 tok/s against 489 on Linux (-84%) while decode stays at parity. Windows Vulkan runs the same model normally (1560/1307), so the regression is in the Windows HIP path, not in the model.
- MXFP4 prompt throughput is 12-14% below Linux on Windows ROCm
(
2076 -> 1813at L2,1765 -> 1528at L3,2253 -> 1999at L1 dense); Q4_K_M and hybrid are at parity (±3%). MXFP4 decode is within 3%. - MTP acceptance is lower on Windows (47-60% against 52-78% on Linux),
which costs MTP decode (
Q4_K_M L1: 50.13 -> 38.22); the L3 MTP rows are the exception and land slightly above Linux. - Vulkan gives up prompt throughput (L3 1244 vs ROCm 1526) but wins decode on several rows and handles NVFP4-native far better than Windows ROCm.
- Single run per cell (r1): treat deltas below ~3% as noise.
| Format | bpw | PPL | Δ vs Q4_K_M |
|---|---|---|---|
| Q4_K_M | 5.01 | 6.7802 ± 0.050 | — baseline |
| MXFP4-native | 4.25 | 7.1391 ± 0.054 | +5.29% |
| MXFP4-requant | 4.25 | 7.1937 ± 0.055 | +6.10% |
| MXFP4-hybrid-attnQ6 | 4.85 | 6.9925 ± 0.054 | +3.13% |
| NVFP4-native | 4.50 | 6.9840 ± 0.052 | +3.01% |
Q4_K_M stays the production quality baseline. MXFP4/NVFP4 trade quality for
speed — compare speed only together with this table. imatrix for FP4 formats
is not implemented (quantize_mxfp4 ignores quant_weights).
Install GUI dependencies and launch from the repository root:
python -m pip install -r gui/requirements-gui.txt
python run.pyIn the GUI: open Build & Setup and configure a backend, build or select
llama-server, open Launch Server and pick a GGUF model. Start with
Spec: None for the baseline; for an MTP-enabled GGUF use depth 3 as the
starting point. Use the recommended explicit device order in
Benchmark / Autotune — Auto is discovery only, not a benchmark
contract. Full build commands and requirements: docs/build.md.
Canonical methodology, lane contracts and history:
BENCHMARKS.md, build_logs/agent-workload/BENCH_RUNS.csv,
BENCH_RECENT.md, BENCH_LANES.md, docs/research/RESULTS_LOG.md,
and the per-experiment notes in docs/research/. Always
compare neighboring controls (same model, backend, device order, split,
context, prompt/output length, batch/ubatch, KV, spec, cache policy and
background load); record an explicit -dev route for every dual-GPU result.
| Path | Purpose |
|---|---|
gui2/ |
GUI 2.0: local web UI (FastHTML + HTMX) |
src/, common/, include/ |
llama runtime and speculative pipeline |
ggml/src/ggml-vulkan/, ggml/src/ggml-hip/, ggml/src/ggml-cuda/, ggml/src/ggml-cpu/ |
Backends (CUDA layer is the HIP-compatible kernel source) |
scripts/ |
Benchmark and autotune runners (bench2.py, agent workload) |
PERFORMANCE.md |
Current benchmark tables and lane contracts |
docs/ |
Backends, build, RPC, research notes (accepted/rejected experiments) |
Read AGENTS.md, CONTRIBUTING.md and UPSTREAM_SYNC.md before changing the fork. When reporting performance, include the model, backend, device order, split, context, actual prompt/output tokens, batch/ubatch, KV types, speculative mode, cache policy and background load.
Derived from ggml-org/llama.cpp,
MIT LICENSE; bundled third-party components retain their own
notices. Security issues are handled privately per
SECURITY.md.