Skip to content

cuda: use one-row scheduling for PTQ1_0 planar GEMV - #335

Closed
Max-sm-yc wants to merge 1 commit into
PrismML-Eng:prismfrom
Max-sm-yc:pr/ptq1-ampere-rows1
Closed

Max-sm-yc wants to merge 1 commit into
PrismML-Eng:prismfrom
Max-sm-yc:pr/ptq1-ampere-rows1

Conversation

@Max-sm-yc

Copy link
Copy Markdown

Overview

Use one output row per work item for the one-column PTQ1_0 planar-transposed GEMV path. This reduces register use and exposes more independent row work on the Ampere batch-1 decode path. The PTQ1_0 dot arithmetic and activation layout are unchanged; multi-column paths keep their existing 4/2-row schedule.

Results

Measured on an NVIDIA GeForce RTX 3080 (GA102, compute capability 8.6) with Ternary Bonsai 2 27B PTQ1_0:

Context ROWS=4 baseline ROWS=1 candidate Change Peak GPU memory
512 77.9966 tok/s 82.2218 tok/s +5.42% 6,805 MiB both
4096 75.6584 tok/s 79.6967 tok/s +5.34% 6,805 MiB both

This A/B used seven repetitions per process. The reported gains use medians. Context-4096 had slow-tail samples in both variants; one pair had a lower ROWS=1 mean despite its higher median, so the median improvement does not establish a tail-latency improvement.

The compiled kernel uses 76 registers per thread for ROWS=1 versus 108 for ROWS=4, with no spills.

Measurement method

The final comparison used a rebuilt source-default candidate and an isolated ROWS=4 baseline on the RTX 3080. Each llama-bench process generated 128 tokens at contexts 512 and 4096, with seven repetitions and the same 60 C / 5% idle start gate. The model was fully offloaded (99 GPU layers), with Flash Attention enabled, F16 KV cache, batch/microbatch 2048/512, and 8 CPU threads. LD_LIBRARY_PATH, ldd, and loader traces verified that each process loaded its intended CUDA library. The earlier two-pair comparisons alternated process order.

Correctness

The recorded final sm_86 ROWS=1 build passed four selected CTests, and 96/96 CUDA-versus-CPU PTQ1_0/PQ2_0 matmul cases. No new operator or quantization type is introduced. Full CI and perplexity were not rerun for this draft: this session's nvidia-smi cannot communicate with the driver. The change keeps the per-row dot and reduction order unchanged.

Scope

The performance result is from Ampere (sm_86). The one-column planar path is selected by the existing host layout policy; this change leaves dispatch and other row schedules unchanged.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - Codex assisted with isolating the one-row change and drafting this description. I have reviewed line differences in PR, which only concern changing the output row count.

@bri-prism

Copy link
Copy Markdown
Collaborator

Closing this version because the universal one-row choice causes a repeatable H200 invariant-mode decode regression. The current PR head is the version tested.

Independent whole-model validation on native H200 sm_90, CUDA 12.8.61, driver 570.172.08, using the public PTQ1_0 model referenced in the PR:

Initial KV depth Parent median tok/s PR median tok/s Throughput change
512 149.349 143.213 −4.109%
4096 145.629 139.894 −3.938%

Three alternating AB / BA / AB comparisons, seven repetitions per process. All three pairs reproduced the invariant-mode loss: −4.097/−4.110/−4.004% at depth 512 and −3.961/−4.038/−3.916% at depth 4096. Default-mode controls were effectively unchanged (+0.000% and −0.056%).

Matched llama-bench settings: generation 128, prompt 0, initial KV depths 512/4096, GPU layers 99, Flash Attention on, F16 K/V, batch 2048/microbatch 512, eight CPU threads. Each process started below 60°C and at ≤5% GPU utilization. Medians use the seven raw duration samples per process, then the three process medians per revision. Model hash and runtime-library selection were verified; no profiler was attached. These are synthetic-token decode timings, not output-quality or speculative-acceptance measurements.

Correctness passed: the fresh 648 fixed-input cases per revision/mode were bit-identical between parent and PR. Earlier validation of the same head passed 257/257 PTQ1 matrix-operation tests and 38/38 fused matvec CUDA-versus-CPU tests in each mode. Earlier H200 operation timings also showed invariant-mode throughput regressions of 9.89% and 27.71% for two representative projection shapes.

The trigger is GGML_CUDA_BATCH_INVARIANT: its presence selects the planar one-column path on Hopper too, so the scheduling change extends beyond the Ampere hardware used for the reported gain. This is a performance issue, not a demonstrated numerical-correctness failure. The RTX 3080 improvement was not independently reproduced here.

Please retain the four-row schedule on Hopper, or restrict the one-row schedule to hardware with demonstrated benefit. The shared-memory eligibility check, rows-per-CTA selection and launcher instantiation must use the same schedule. A revised PR with that hardware scope and matched validation would be welcome.

@bri-prism bri-prism closed this Oct 9, 2026
@bri-prism bri-prism mentioned this pull request Oct 10, 2026
1 task done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants