You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Track investigation of the CUDA Q4 KV-cache prompt-processing slowdown reported in Bonsai-demo #145, and validation/integration of an appropriate fix in this fork.
The original measurements used the previous-generationTernary-Bonsai-27B-Q2_g64.gguf on upstream llama.cpp 22dc605c4ead20e36f447cc67b55ef87e523bd55 (b10257), RTX 4090, CUDA 12.6, full GPU offload, Flash Attention, one slot, batch 2048 / ubatch 512. They are reporter-provided results, not a fresh reproduction on current prism or Bonsai 2.
KV K / V
PP512 (t/s)
PP4096 (t/s)
F16 / F16
3427.7
3459.5
Q8_0 / Q8_0
3104.6
3320.5
Q4_0 / Q8_0
921.8
71.8
Q8_0 / Q4_0
647.4
Aborted due to slowdown
The reporter also observed approximately 49 t/s versus 3,283 t/s prefill on roughly 14K-token server prompts with Q4/Q8 versus Q8/Q8. See the original report for full settings and results.
Related work and scope
Upstream candidate fix: ggml-org/llama.cpp#27140, still open at the time of filing; targets slow prefill with small KV quants. Its applicability and correctness in our fork need validation.
Mean-centered Q4 KV addresses quantization quality; it is not an established fix for this performance issue.
Follow-up
Reproduce on current prism, recording exact commit, GPU, model, cache types, batch/ubatch, and prompt lengths. Check Bonsai 2 separately from the original model.
Compare F16/F16, Q8/Q8, Q4/Q8, Q8/Q4, and Q4/Q4 with otherwise matched settings, including longer prompts and decode timings.
Evaluate the upstream candidate fix or an equivalent correction, with numerical/quality checks as well as performance; include our mean-centering path where supported.
Record affected configurations and the first verified fixed release; update the demo issue/documentation accordingly.
No new benchmark or confirmed current-fork regression is claimed by this tracking issue.
馃 Tracking on behalf of the maintainers
Problem
Track investigation of the CUDA Q4 KV-cache prompt-processing slowdown reported in Bonsai-demo #145, and validation/integration of an appropriate fix in this fork.
The original measurements used the previous-generation
Ternary-Bonsai-27B-Q2_g64.ggufon upstream llama.cpp22dc605c4ead20e36f447cc67b55ef87e523bd55(b10257), RTX 4090, CUDA 12.6, full GPU offload, Flash Attention, one slot, batch 2048 / ubatch 512. They are reporter-provided results, not a fresh reproduction on currentprismor Bonsai 2.The reporter also observed approximately 49 t/s versus 3,283 t/s prefill on roughly 14K-token server prompts with Q4/Q8 versus Q8/Q8. See the original report for full settings and results.
Related work and scope
Follow-up
prism, recording exact commit, GPU, model, cache types, batch/ubatch, and prompt lengths. Check Bonsai 2 separately from the original model.No new benchmark or confirmed current-fork regression is claimed by this tracking issue.