Skip to content

CUDA: track Q4 KV-cache prefill slowdown and upstream fix validation聽#266

Description

@khosravipasha

馃 Tracking on behalf of the maintainers

Problem

Track investigation of the CUDA Q4 KV-cache prompt-processing slowdown reported in Bonsai-demo #145, and validation/integration of an appropriate fix in this fork.

The original measurements used the previous-generation Ternary-Bonsai-27B-Q2_g64.gguf on upstream llama.cpp 22dc605c4ead20e36f447cc67b55ef87e523bd55 (b10257), RTX 4090, CUDA 12.6, full GPU offload, Flash Attention, one slot, batch 2048 / ubatch 512. They are reporter-provided results, not a fresh reproduction on current prism or Bonsai 2.

KV K / V PP512 (t/s) PP4096 (t/s)
F16 / F16 3427.7 3459.5
Q8_0 / Q8_0 3104.6 3320.5
Q4_0 / Q8_0 921.8 71.8
Q8_0 / Q4_0 647.4 Aborted due to slowdown

The reporter also observed approximately 49 t/s versus 3,283 t/s prefill on roughly 14K-token server prompts with Q4/Q8 versus Q8/Q8. See the original report for full settings and results.

Related work and scope

Follow-up

  • Reproduce on current prism, recording exact commit, GPU, model, cache types, batch/ubatch, and prompt lengths. Check Bonsai 2 separately from the original model.
  • Compare F16/F16, Q8/Q8, Q4/Q8, Q8/Q4, and Q4/Q4 with otherwise matched settings, including longer prompts and decode timings.
  • Evaluate the upstream candidate fix or an equivalent correction, with numerical/quality checks as well as performance; include our mean-centering path where supported.
  • Record affected configurations and the first verified fixed release; update the demo issue/documentation accordingly.

No new benchmark or confirmed current-fork regression is claimed by this tracking issue.

Activity

  1. cheese-cakee commented on Oct 6, 2026

    @cheese-cakee

    Fresh check on current prism. On my setup the slowdown only shows up with mixed K/V types, and it is the CPU fallback from #267: in a default CUDA build, FLASH_ATTN_EXT with K type != V type is rejected by CUDA and runs on the CPU, so the cost grows with context. Same-type Q4_0/Q4_0 and Q8_0/Q8_0 have no pathological slowdown. #317 (backport of ggml-org#28079) fixes the mixed case.

    Setup: RTX 4050 Laptop (Ada, cc 8.9, 6 GB), CUDA 12.6, Linux (WSL2), -DCMAKE_CUDA_ARCHITECTURES=89. Three builds: prism 2459f68 default, #317 (4105522) default, and prism with GGML_CUDA_FA_ALL_QUANTS=ON. No file under ggml/src/ggml-cuda/fattn* or template-instances changed between 2459f68 and the current prism head 6bfcd79.

    Kernel level, Bonsai 2 27B attention shape (hd 256, 4 KV heads, GQA 6, 512 queries), test-backend-ops perf -o FLASH_ATTN_EXT, median us/run of 9 rotating rounds, lower is better:

    K / V kv 512: prism kv 512: #317 kv 4096: prism kv 4096: #317
    F16 / F16 433 446 2766 2734
    Q8_0 / Q8_0 511 501 3369 3353
    Q4_0 / Q4_0 470 468 3022 3050
    Q4_0 / Q8_0 not supported 451 not supported 2951
    Q8_0 / Q4_0 not supported 459 not supported 3005

    Q8_0 KV is about 20% and Q4_0 about 10% slower than F16 in the kernel, nothing like the 921 -> 72 t/s drop in the report. The GGML_CUDA_FA_ALL_QUANTS=ON build is within 1-4% of #317 in every cell. "not supported" means the default build cannot run that pair on CUDA at all.

    Model level, Qwen3-0.6B Q8_0 (not Bonsai: shows the mechanism, not Bonsai numbers), llama-bench -ngl 99 -fa 1 -p 512 -n 128 -d 0,4096 -r 3, median t/s of 5 rotating rounds:

    K / V test prism default #317 ALL_QUANTS=ON
    Q8_0 / Q4_0 pp512 888 13937 14378
    Q8_0 / Q4_0 pp512 @ d4096 47 8631 8881
    Q8_0 / Q4_0 tg128 @ d4096 17.3 99.6 143.3
    Q8_0 / Q8_0 pp512 @ d4096 8478 8375 8155
    Q8_0 / Q8_0 tg128 @ d4096 143.3 147.3 139.8

    With GGML_SCHED_DEBUG=2 on the default build, all FLASH_ATTN_EXT nodes run on the CPU for the mixed pair, with K/V copies to the CPU every layer, and nothing in the log says so. With #317 the mixed pair prefills at same-type speed. Decode at depth stays below the native kernel (99.6 vs 143.3 t/s) because #317 converts K/V to F16 for pairs without a compiled kernel; adding the pair to GGML_CUDA_FA_QUANTS removes that.

    Not done:

    • Bonsai 2 27B at model level. On 6 GB it only fits with the output head on the CPU, and the Windows desktop shares the GPU, so long prompts spill into shared memory and every KV type drops to the same ~57 t/s. I discarded those runs. This needs a GPU with VRAM headroom.
    • The previous-generation Ternary-Bonsai-27B-Q2_g64 model from the original report.
    • Upstream Slow prefill on small KV quants fixed聽ggml-org/llama.cpp#27140.
    • Quality checks (only speed here).

    If someone with a bigger card can confirm Q4_0/Q8_0 and Q8_0/Q4_0 prefill on Bonsai 2 27B with #317, I think this can close together with #267.

    AI was used to help write the benchmark scripts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions