Skip to content

CUDA: track Q4 KV-cache prefill slowdown and upstream fix validation聽#266

Description

@khosravipasha

馃 Tracking on behalf of the maintainers

Problem

Track investigation of the CUDA Q4 KV-cache prompt-processing slowdown reported in Bonsai-demo #145, and validation/integration of an appropriate fix in this fork.

The original measurements used the previous-generation Ternary-Bonsai-27B-Q2_g64.gguf on upstream llama.cpp 22dc605c4ead20e36f447cc67b55ef87e523bd55 (b10257), RTX 4090, CUDA 12.6, full GPU offload, Flash Attention, one slot, batch 2048 / ubatch 512. They are reporter-provided results, not a fresh reproduction on current prism or Bonsai 2.

KV K / V PP512 (t/s) PP4096 (t/s)
F16 / F16 3427.7 3459.5
Q8_0 / Q8_0 3104.6 3320.5
Q4_0 / Q8_0 921.8 71.8
Q8_0 / Q4_0 647.4 Aborted due to slowdown

The reporter also observed approximately 49 t/s versus 3,283 t/s prefill on roughly 14K-token server prompts with Q4/Q8 versus Q8/Q8. See the original report for full settings and results.

Related work and scope

Follow-up

  • Reproduce on current prism, recording exact commit, GPU, model, cache types, batch/ubatch, and prompt lengths. Check Bonsai 2 separately from the original model.
  • Compare F16/F16, Q8/Q8, Q4/Q8, Q8/Q4, and Q4/Q4 with otherwise matched settings, including longer prompts and decode timings.
  • Evaluate the upstream candidate fix or an equivalent correction, with numerical/quality checks as well as performance; include our mean-centering path where supported.
  • Record affected configurations and the first verified fixed release; update the demo issue/documentation accordingly.

No new benchmark or confirmed current-fork regression is claimed by this tracking issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions