Repository navigation
PTQ1_0 Vulkan kernel (4d583a9f3, committed as untested) — confirmed correct on gfx1152, but slow #185
Description
Activity
Adding an Intel Arc (Xe2-HPG) data point for this same kernel — also confirmed correct, also slow, so this doesn't look AMD-specific.
Hardware: Intel Arc B570 (10GB VRAM), Windows 11, Vulkan backend (bin\vulkan\llama-server.exe, build b10683-d8f26ee)
Command:
llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 --jinjaCorrectness: coherent output, sane reasoning trace, correct factual answers (spot-checked with a simple Q&A). Not garbage.
Performance:
Prompt processing: 11.9 t/s
Generation: 1.26 t/sFor reference, this is faster than the RDNA 3.5 iGPU numbers above (4.6pp/0.7tg) but nowhere near the RX 9070 XT numbers from #180 (~57pp/9tg). Confirms this is a correctness-first, not-yet-optimized kernel across vendors, not just AMD-specific.
Happy to test with different -c/batch settings or share more details if useful for narrowing down the perf gap.
Data point from a card without integer dot: AMD RX 570 (Polaris10 / gfx803, 8 GiB, RADV, Mesa 26.1.2,
RADV_PERFTEST=nogttspill), Ternary-Bonsai-2-27B-PTQ1_0,llama-server -c 16384,q8_0/q4_0KV,-fa on.Edit 2026-09-24: the prompt "before" value was wrong in the first version of this comment (3.8 tok/s, which was the 18-token prompt of the generation request). Corrected below, measured at
842b188. For generation on cards without integer dot, #252 now does the same job properly; numbers for the RX 570 are in #252.The generic
mul_mat_vec/mul_mmdecode is the bottleneck here, and it stays the path on gfx803 since the integer-dot kernel (#238) needsVK_KHR_shader_integer_dot_product. Replacing the per-element trit loop with a byte -> five trits table (256 entries, generated from the CPU codec inggml-quants.c, in shared memory viainit_iq_shmemlike the IQ grids) and reading the block as seven 32-bit words gives:stock table decode generation, 16k ctx 633 ms/token (1.58 tok/s) 143 ms/token (7.0 tok/s) generation, 48k ctx 633 ms/token 143 ms/token prompt, 847 tokens 36 tok/s 54 tok/s test-backend-ops -b Vulkan0 -p ptq1_0(MUL_MAT, MUL_MAT_ID, GET_ROWS, CPY) passes against the CPU backend, output is bit-identical to the stock shader. The generation half is superseded by #252 (7.15 tok/s on the same card). Themul_mmhalf is not covered elsewhere yet and composes with #252 (7.15 tg / 59.1 pp512 stacked).Patch, applies with
git amon1a07bfa, and withgit am -3on842b188: 0001-ptq1_0-table-decode.patchMy own words: The shader code was written with Claude Fable 5.1 and I don't have a deep understanding of it, so I won't open a PR for it. Anyone who wants the
mul_mmpart is free to take it.Thanks for testing the kernel and for the extra data points. The PTQ1_0 Vulkan mat-vec now uses an integer-dot path (#238), which is in the new release prism-b10735-842b188. On GPUs with integer dot support, including the Radeon 860M and the Arc B570 reported here, decode should be much faster. Could you rerun the same command on that build?
Cards without integer dot, like the RX 570 in the last comment, still take the generic path. #252 adds a faster fallback for those, and #231 tracks the same gap.
Reran on
prism-b10735-842b188(commit842b18804), clean Vulkan build. The integer-dot path makes a big difference on the Radeon 860M.Hardware: AMD Ryzen AI 7 350 / Radeon 860M (gfx1152, RDNA 3.5 iGPU), Mesa 26.0.8 RADV (
int dot: 1,KHR_coopmat).Same command as the original report:
./llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 4096 -p "hi" -n 64before ( 4d583a9f3)prism-b10735-842b188prompt 4.6 t/s 14.2 t/s generation 0.7 t/s 9.7 t/s llama-bench -ngl 99 -fa 1 -p 512 -n 128 -r 3:test t/s pp512 22.42 ± 0.10 tg128 9.54 ± 0.02 Output is still coherent. Decode is about 14x faster and now slightly ahead of PQ2_0 on the same iGPU (8.8 t/s tg on Vulkan). Thanks!
Builds and testing were AI-assisted.
Ran the PTQ1_0 Vulkan path added in 4d583a9 ("UNTESTED as committed") on AMD Ryzen AI 7 350 / Radeon 860M (gfx1152, RDNA 3.5 iGPU), RADV driver:
./llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 4096 -p "hi" -n 64
Output is coherent (sane greeting + reasoning trace), so not producing garbage.
Performance: 4.6 t/s prompt, 0.7 t/s generation.
For comparison, issue #180 reports ~57 pp / ~9 tg for the same file on an RX 9070 XT.