Repository navigation
CUDA: track Q4 KV-cache prefill slowdown and upstream fix validation聽#266
Description
Activity
Fresh check on current
prism. On my setup the slowdown only shows up with mixed K/V types, and it is the CPU fallback from #267: in a default CUDA build, FLASH_ATTN_EXT with K type != V type is rejected by CUDA and runs on the CPU, so the cost grows with context. Same-type Q4_0/Q4_0 and Q8_0/Q8_0 have no pathological slowdown. #317 (backport of ggml-org#28079) fixes the mixed case.Setup: RTX 4050 Laptop (Ada, cc 8.9, 6 GB), CUDA 12.6, Linux (WSL2),
-DCMAKE_CUDA_ARCHITECTURES=89. Three builds:prism2459f68 default, #317 (4105522) default, andprismwithGGML_CUDA_FA_ALL_QUANTS=ON. No file underggml/src/ggml-cuda/fattn*ortemplate-instanceschanged between 2459f68 and the currentprismhead 6bfcd79.Kernel level, Bonsai 2 27B attention shape (hd 256, 4 KV heads, GQA 6, 512 queries),
test-backend-ops perf -o FLASH_ATTN_EXT, median us/run of 9 rotating rounds, lower is better:K / V kv 512: prismkv 512: #317 kv 4096: prismkv 4096: #317 F16 / F16 433 446 2766 2734 Q8_0 / Q8_0 511 501 3369 3353 Q4_0 / Q4_0 470 468 3022 3050 Q4_0 / Q8_0 not supported 451 not supported 2951 Q8_0 / Q4_0 not supported 459 not supported 3005 Q8_0 KV is about 20% and Q4_0 about 10% slower than F16 in the kernel, nothing like the 921 -> 72 t/s drop in the report. The
GGML_CUDA_FA_ALL_QUANTS=ONbuild is within 1-4% of #317 in every cell. "not supported" means the default build cannot run that pair on CUDA at all.Model level, Qwen3-0.6B Q8_0 (not Bonsai: shows the mechanism, not Bonsai numbers),
llama-bench -ngl 99 -fa 1 -p 512 -n 128 -d 0,4096 -r 3, median t/s of 5 rotating rounds:K / V test prismdefault#317 ALL_QUANTS=ONQ8_0 / Q4_0 pp512 888 13937 14378 Q8_0 / Q4_0 pp512 @ d4096 47 8631 8881 Q8_0 / Q4_0 tg128 @ d4096 17.3 99.6 143.3 Q8_0 / Q8_0 pp512 @ d4096 8478 8375 8155 Q8_0 / Q8_0 tg128 @ d4096 143.3 147.3 139.8 With
GGML_SCHED_DEBUG=2on the default build, all FLASH_ATTN_EXT nodes run on the CPU for the mixed pair, with K/V copies to the CPU every layer, and nothing in the log says so. With #317 the mixed pair prefills at same-type speed. Decode at depth stays below the native kernel (99.6 vs 143.3 t/s) because #317 converts K/V to F16 for pairs without a compiled kernel; adding the pair toGGML_CUDA_FA_QUANTSremoves that.Not done:
- Bonsai 2 27B at model level. On 6 GB it only fits with the output head on the CPU, and the Windows desktop shares the GPU, so long prompts spill into shared memory and every KV type drops to the same ~57 t/s. I discarded those runs. This needs a GPU with VRAM headroom.
- The previous-generation
Ternary-Bonsai-27B-Q2_g64model from the original report. - Upstream Slow prefill on small KV quants fixed聽ggml-org/llama.cpp#27140.
- Quality checks (only speed here).
If someone with a bigger card can confirm Q4_0/Q8_0 and Q8_0/Q4_0 prefill on Bonsai 2 27B with #317, I think this can close together with #267.
AI was used to help write the benchmark scripts.
馃 Tracking on behalf of the maintainers
Problem
Track investigation of the CUDA Q4 KV-cache prompt-processing slowdown reported in Bonsai-demo #145, and validation/integration of an appropriate fix in this fork.
The original measurements used the previous-generation
Ternary-Bonsai-27B-Q2_g64.ggufon upstream llama.cpp22dc605c4ead20e36f447cc67b55ef87e523bd55(b10257), RTX 4090, CUDA 12.6, full GPU offload, Flash Attention, one slot, batch 2048 / ubatch 512. They are reporter-provided results, not a fresh reproduction on currentprismor Bonsai 2.The reporter also observed approximately 49 t/s versus 3,283 t/s prefill on roughly 14K-token server prompts with Q4/Q8 versus Q8/Q8. See the original report for full settings and results.
Related work and scope
Follow-up
prism, recording exact commit, GPU, model, cache types, batch/ubatch, and prompt lengths. Check Bonsai 2 separately from the original model.No new benchmark or confirmed current-fork regression is claimed by this tracking issue.