Skip to content

cuda: use the MMA flash attention kernel for GQA above 4 with quantized K/V on Ada - #307

Open
sb32445 wants to merge 3 commits into
PrismML-Eng:prismfrom
sb32445:pr/fattn-gqa-mma
Open

sb32445 wants to merge 3 commits into
PrismML-Eng:prismfrom
sb32445:pr/fattn-gqa-mma

cuda : keep the vector FA kernel when MMA cannot read quantized K/V i…

d01570f
Select commit
Loading
Failed to load commit list.
Sign in for the full log view

Annotations

1 warning
labeler
succeeded Oct 7, 2026 in 18s