Skip to content

PTQ1_0 Vulkan kernel (4d583a9f3, committed as untested) — confirmed correct on gfx1152, but slow #185

Description

@c2p-cmd

Ran the PTQ1_0 Vulkan path added in 4d583a9 ("UNTESTED as committed") on AMD Ryzen AI 7 350 / Radeon 860M (gfx1152, RDNA 3.5 iGPU), RADV driver:

./llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 4096 -p "hi" -n 64

Output is coherent (sane greeting + reasoning trace), so not producing garbage.

Performance: 4.6 t/s prompt, 0.7 t/s generation.
For comparison, issue #180 reports ~57 pp / ~9 tg for the same file on an RX 9070 XT.

Activity

  1. DasCode-Brm commented on Sep 18, 2026

    @DasCode-Brm

    Adding an Intel Arc (Xe2-HPG) data point for this same kernel — also confirmed correct, also slow, so this doesn't look AMD-specific.

    Hardware: Intel Arc B570 (10GB VRAM), Windows 11, Vulkan backend (bin\vulkan\llama-server.exe, build b10683-d8f26ee)

    Command:
    llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 --jinja

    Correctness: coherent output, sane reasoning trace, correct factual answers (spot-checked with a simple Q&A). Not garbage.

    Performance:

    Prompt processing: 11.9 t/s
    Generation: 1.26 t/s

    For reference, this is faster than the RDNA 3.5 iGPU numbers above (4.6pp/0.7tg) but nowhere near the RX 9070 XT numbers from #180 (~57pp/9tg). Confirms this is a correctness-first, not-yet-optimized kernel across vendors, not just AMD-specific.

    Happy to test with different -c/batch settings or share more details if useful for narrowing down the perf gap.

  2. voxlo-dev commented on Sep 21, 2026

    @voxlo-dev

    Data point from a card without integer dot: AMD RX 570 (Polaris10 / gfx803, 8 GiB, RADV, Mesa 26.1.2, RADV_PERFTEST=nogttspill), Ternary-Bonsai-2-27B-PTQ1_0, llama-server -c 16384, q8_0/q4_0 KV, -fa on.

    Edit 2026-09-24: the prompt "before" value was wrong in the first version of this comment (3.8 tok/s, which was the 18-token prompt of the generation request). Corrected below, measured at 842b188. For generation on cards without integer dot, #252 now does the same job properly; numbers for the RX 570 are in #252.

    The generic mul_mat_vec / mul_mm decode is the bottleneck here, and it stays the path on gfx803 since the integer-dot kernel (#238) needs VK_KHR_shader_integer_dot_product. Replacing the per-element trit loop with a byte -> five trits table (256 entries, generated from the CPU codec in ggml-quants.c, in shared memory via init_iq_shmem like the IQ grids) and reading the block as seven 32-bit words gives:

    stock table decode
    generation, 16k ctx 633 ms/token (1.58 tok/s) 143 ms/token (7.0 tok/s)
    generation, 48k ctx 633 ms/token 143 ms/token
    prompt, 847 tokens 36 tok/s 54 tok/s

    test-backend-ops -b Vulkan0 -p ptq1_0 (MUL_MAT, MUL_MAT_ID, GET_ROWS, CPY) passes against the CPU backend, output is bit-identical to the stock shader. The generation half is superseded by #252 (7.15 tok/s on the same card). The mul_mm half is not covered elsewhere yet and composes with #252 (7.15 tg / 59.1 pp512 stacked).

    Patch, applies with git am on 1a07bfa, and with git am -3 on 842b188: 0001-ptq1_0-table-decode.patch

    My own words: The shader code was written with Claude Fable 5.1 and I don't have a deep understanding of it, so I won't open a PR for it. Anyone who wants the mul_mm part is free to take it.

  3. bri-prism commented on Sep 24, 2026

    @bri-prism
    Collaborator

    Thanks for testing the kernel and for the extra data points. The PTQ1_0 Vulkan mat-vec now uses an integer-dot path (#238), which is in the new release prism-b10735-842b188. On GPUs with integer dot support, including the Radeon 860M and the Arc B570 reported here, decode should be much faster. Could you rerun the same command on that build?

    Cards without integer dot, like the RX 570 in the last comment, still take the generic path. #252 adds a faster fallback for those, and #231 tracks the same gap.

  4. c2p-cmd commented on Sep 24, 2026

    @c2p-cmd
    Author

    Reran on prism-b10735-842b188 (commit 842b18804), clean Vulkan build. The integer-dot path makes a big difference on the Radeon 860M.

    Hardware: AMD Ryzen AI 7 350 / Radeon 860M (gfx1152, RDNA 3.5 iGPU), Mesa 26.0.8 RADV (int dot: 1, KHR_coopmat).

    Same command as the original report:

    ./llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 4096 -p "hi" -n 64
    
    before (4d583a9f3) prism-b10735-842b188
    prompt 4.6 t/s 14.2 t/s
    generation 0.7 t/s 9.7 t/s

    llama-bench -ngl 99 -fa 1 -p 512 -n 128 -r 3:

    test t/s
    pp512 22.42 ± 0.10
    tg128 9.54 ± 0.02

    Output is still coherent. Decode is about 14x faster and now slightly ahead of PQ2_0 on the same iGPU (8.8 t/s tg on Vulkan). Thanks!

    Builds and testing were AI-assisted.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions