Repository navigation
Conversation
|
Closing this version because the universal one-row choice causes a repeatable H200 invariant-mode decode regression. The current PR head is the version tested. Independent whole-model validation on native H200 sm_90, CUDA 12.8.61, driver 570.172.08, using the public PTQ1_0 model referenced in the PR:
Three alternating AB / BA / AB comparisons, seven repetitions per process. All three pairs reproduced the invariant-mode loss: −4.097/−4.110/−4.004% at depth 512 and −3.961/−4.038/−3.916% at depth 4096. Default-mode controls were effectively unchanged (+0.000% and −0.056%). Matched Correctness passed: the fresh 648 fixed-input cases per revision/mode were bit-identical between parent and PR. Earlier validation of the same head passed 257/257 PTQ1 matrix-operation tests and 38/38 fused matvec CUDA-versus-CPU tests in each mode. Earlier H200 operation timings also showed invariant-mode throughput regressions of 9.89% and 27.71% for two representative projection shapes. The trigger is Please retain the four-row schedule on Hopper, or restrict the one-row schedule to hardware with demonstrated benefit. The shared-memory eligibility check, rows-per-CTA selection and launcher instantiation must use the same schedule. A revised PR with that hardware scope and matched validation would be welcome. |
Overview
Use one output row per work item for the one-column PTQ1_0 planar-transposed GEMV path. This reduces register use and exposes more independent row work on the Ampere batch-1 decode path. The PTQ1_0 dot arithmetic and activation layout are unchanged; multi-column paths keep their existing 4/2-row schedule.
Results
Measured on an NVIDIA GeForce RTX 3080 (GA102, compute capability 8.6) with Ternary Bonsai 2 27B PTQ1_0:
This A/B used seven repetitions per process. The reported gains use medians. Context-4096 had slow-tail samples in both variants; one pair had a lower ROWS=1 mean despite its higher median, so the median improvement does not establish a tail-latency improvement.
The compiled kernel uses 76 registers per thread for ROWS=1 versus 108 for ROWS=4, with no spills.
Measurement method
The final comparison used a rebuilt source-default candidate and an isolated ROWS=4 baseline on the RTX 3080. Each
llama-benchprocess generated 128 tokens at contexts 512 and 4096, with seven repetitions and the same 60 C / 5% idle start gate. The model was fully offloaded (99 GPU layers), with Flash Attention enabled, F16 KV cache, batch/microbatch 2048/512, and 8 CPU threads.LD_LIBRARY_PATH,ldd, and loader traces verified that each process loaded its intended CUDA library. The earlier two-pair comparisons alternated process order.Correctness
The recorded final sm_86 ROWS=1 build passed four selected CTests, and 96/96 CUDA-versus-CPU PTQ1_0/PQ2_0 matmul cases. No new operator or quantization type is introduced. Full CI and perplexity were not rerun for this draft: this session's
nvidia-smicannot communicate with the driver. The change keeps the per-row dot and reduction order unchanged.Scope
The performance result is from Ampere (sm_86). The one-column planar path is selected by the existing host layout policy; this change leaves dispatch and other row schedules unchanged.
Requirements