Name and Version
llama-cli version: b10709 (commit 9a9394a89)
Release: prism-b10709-9a9394a
Windows prebuilt: llama-prism-b10709-9a9394a-bin-win-hip-radeon-x64.zip
ROCm userspace: 7.1
HIP runtime loaded: C:\WINDOWS\SYSTEM32\amdhip64_7.dll
A source build from the same 9a9394a commit, compiled specifically for gfx1151, reproduced the failure as well.
Operating system
Windows 11 Pro, build 10.0.26200, x64.
GGML backend
HIP/ROCm (ROCm0). CPU is a useful passing control when repacking is disabled.
Hardware
- AMD Ryzen AI MAX+ 395
- Radeon 8060S,
gfx1151 / Strix Halo
- 128 GB unified system memory
Model
prism-ml/Ternary-Bonsai-2-27B-gguf
- File:
Ternary-Bonsai-2-27B-PQ2_0.gguf
- Size:
7,206,168,928 bytes
- SHA-256:
3907dc1658db1f78a9826bf8d5bcb8dc65db0d466388937af57f2294fae62ec1
- Reported type:
PQ2_0 - 2.13 bpw (group 128)
The hash matches the current Hugging Face artifact.
Problem description
The exact Prism release required by Bonsai 2 loads the model and runs quickly on HIP, but even a tiny direct llama-cli prompt produces corrupt repetitive punctuation/digits. This is independent of llama-server, chat templates, vision/mmproj, long context, and sampling.
The same binary and GGUF produce coherent tokens on CPU when --no-repack is used. HIP remains corrupt with --no-repack, including when only the 64 transformer blocks are offloaded and the output tensor remains on CPU. This localizes the remaining problem to GPU-executed model layers rather than the output head or GGUF contents.
Minimal HIP reproduction
$env:GGML_CUDA_NO_PINNED = '1'
.\llama-cli.exe `
-m Ternary-Bonsai-2-27B-PQ2_0.gguf `
--device ROCm0 -ngl 99 -fa on -c 32768 `
--no-repack --temp 0 `
-p 'What is 2 plus 2? Answer with the number only.' `
-n 32 -st
Observed generation:
[Start thinking]
33333333333333333333333333333333
Prompt: 445.0 t/s | Generation: 178.0 t/s
Removing --no-repack also fails, with punctuation/digit gibberish instead.
CPU passing control
.\llama-cli.exe `
-m Ternary-Bonsai-2-27B-PQ2_0.gguf `
-ngl 0 -c 4096 --no-repack --temp 0 `
-p 'What is 2 plus 2? Answer with the number only.' `
-n 16 -st
Observed generation begins coherently:
[Start thinking]
We need answer user's request: "What is 2 plus 2?
Prompt: 17.3 t/s | Generation: 3.1 t/s
GPU layers with output tensor left on CPU
-ngl 64 --no-repack still produces the identical repeated-3 corruption, at 46.9 t/s generation. Therefore the failure is not limited to the vocabulary-sized output tensor.
Additional isolation already performed
All of the following still produced corrupt HIP output:
- direct
llama-cli rather than llama-server;
- tiny prompt well below
n_ubatch;
--no-repack;
--no-mmap;
HIP_LAUNCH_BLOCKING=1;
GGML_CUDA_NO_PINNED=1;
- output tensor left on CPU via
-ngl 64;
- native Windows build from
9a9394a with:
-DGGML_HIP=ON
-DGPU_TARGETS=gfx1151
-DGGML_HIP_NO_VMM=ON
-DGGML_CUDA_FORCE_CUBLAS=ON
-DGGML_NATIVE=ON
The FORCE_CUBLAS build rules out the PQ2 MMQ kernel as the sole cause. GGML_CUDA_ENABLE_UNIFIED_MEMORY and experimental LLAMA_MMB_*/LLAMA_HC_* flags are not set.
The Vulkan prebuilt detects the 8060S but crashes while loading this PQ2_0 model, so it is not a usable comparison backend here.
Related evidence
Expected result
HIP output should agree semantically with the coherent CPU --no-repack control for the same model, prompt, and greedy sampling.
AI usage disclosure
OpenAI Codex assisted with running the controlled A/Bs, organizing the observed logs, searching for related reports, and drafting this issue. The user explicitly requested publication. No generated source-code patch is included.
Name and Version
A source build from the same
9a9394acommit, compiled specifically forgfx1151, reproduced the failure as well.Operating system
Windows 11 Pro, build 10.0.26200, x64.
GGML backend
HIP/ROCm (
ROCm0). CPU is a useful passing control when repacking is disabled.Hardware
gfx1151/ Strix HaloModel
prism-ml/Ternary-Bonsai-2-27B-ggufTernary-Bonsai-2-27B-PQ2_0.gguf7,206,168,928bytes3907dc1658db1f78a9826bf8d5bcb8dc65db0d466388937af57f2294fae62ec1PQ2_0 - 2.13 bpw (group 128)The hash matches the current Hugging Face artifact.
Problem description
The exact Prism release required by Bonsai 2 loads the model and runs quickly on HIP, but even a tiny direct
llama-cliprompt produces corrupt repetitive punctuation/digits. This is independent of llama-server, chat templates, vision/mmproj, long context, and sampling.The same binary and GGUF produce coherent tokens on CPU when
--no-repackis used. HIP remains corrupt with--no-repack, including when only the 64 transformer blocks are offloaded and the output tensor remains on CPU. This localizes the remaining problem to GPU-executed model layers rather than the output head or GGUF contents.Minimal HIP reproduction
Observed generation:
Removing
--no-repackalso fails, with punctuation/digit gibberish instead.CPU passing control
Observed generation begins coherently:
GPU layers with output tensor left on CPU
-ngl 64 --no-repackstill produces the identical repeated-3corruption, at 46.9 t/s generation. Therefore the failure is not limited to the vocabulary-sized output tensor.Additional isolation already performed
All of the following still produced corrupt HIP output:
llama-clirather than llama-server;n_ubatch;--no-repack;--no-mmap;HIP_LAUNCH_BLOCKING=1;GGML_CUDA_NO_PINNED=1;-ngl 64;9a9394awith:-DGGML_HIP=ON-DGPU_TARGETS=gfx1151-DGGML_HIP_NO_VMM=ON-DGGML_CUDA_FORCE_CUBLAS=ON-DGGML_NATIVE=ONThe FORCE_CUBLAS build rules out the PQ2 MMQ kernel as the sole cause.
GGML_CUDA_ENABLE_UNIFIED_MEMORYand experimentalLLAMA_MMB_*/LLAMA_HC_*flags are not set.The Vulkan prebuilt detects the 8060S but crashes while loading this PQ2_0 model, so it is not a usable comparison backend here.
Related evidence
--no-repackas a CPU workaround. That workaround makes CPU coherent here but does not repair HIP.prism-b10709-9a9394aworking on ROCmgfx1101, suggesting this failure isgfx1151-specific.gfx1151across other architectures, although its known trigger is prompts longer thann_ubatch; this reproduction fails on a tiny prompt.Expected result
HIP output should agree semantically with the coherent CPU
--no-repackcontrol for the same model, prompt, and greedy sampling.AI usage disclosure
OpenAI Codex assisted with running the controlled A/Bs, organizing the observed logs, searching for related reports, and drafting this issue. The user explicitly requested publication. No generated source-code patch is included.