Retrieved article excerpt
Open article · Retrieved 2026-09-11T10:22:56.323220+00:00
Uh oh! There was an error while loading. Please reload this page . ggml-org / llama.cpp Public Notifications You must be signed in to change notification settings Fork 23k Star 128k CUDA/HIP: Flash Attention tuning (gfx1201) - # 28102 # 28102 Merged IMbackK merged 4 commits into ggml-org:master ggml-org/llama.cpp:master from pwilkin:fattn-wmma-rdna4-256 pwilkin/llama.cpp:fattn-wmma-rdna4-256 Copy head branch name to clipboard Sep 11, 2026 Merged CUDA/HIP: Flash Attention tuning (gfx1201) # 28102 IMbackK merged 4 commits into ggml-org:master ggml-org/llama.cpp:master from pwilkin:fattn-wmma-rdna4-256 pwilkin/llama.cpp:fattn-wmma-rdna4-256 Copy head branch name to clipboard Conversation pwilkin commented Aug 31, 2026 Copy link Copy Markdown Member Overview So, got my new R9700 PRO. It's nice and shiny and has 32GB VRAM, so I decided to try out Qwen3.8 27B. That was a mistake. The prefill performance at longer contexts was abysmal, so I decided to do something about it. Also managed to fix a HS=256 bug in the general CUDA FA code which was preventing the selection of 256 kernels before. Additional information Before: model size params backend threads dev test t/s qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 1101.41 ± 221.33 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 29.81 ± 0.19 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d40000 426.36 ± 37.60 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d40000 26.75 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d150000 164.42 ± 5.79 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d150000 19.70 ± 0.20 After: model size params backend threads dev test t/s qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 1103.66 ± 254.83 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 29.87 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d40000 639.46 ± 185.36 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d40000 26.94 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d150000 399.01 ± 30.87 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d150000 19.67 ± 0.39 Requirements I have read and agree with the contributing guidelines AI usage disclosure: Yes pwilkin requested review from a team and ggerganov as code owners August 31, 2026 16:54 github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Aug 31, 2026 Geramy commented Aug 31, 2026 Copy link Copy Markdown Contributor This looks really good! Such a simple change providing extra performance, good job on finding it! IMbackK self-assigned this Aug 31, 2026 pwilkin commented Sep 1, 2026 Copy link Copy Markdown Member Author I'm not done yet! model size params backend threads dev test t/s qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 1138.55 ± 224.34 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 29.48 ± 0.83 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d40000 704.92 ± 239.12 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d40000 27.00 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d150000 456.51 ± 38.04 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d150000 19.67 ± 0.42 pwilkin force-pushed the fattn-wmma-rdna4-256 branch
from 6b24ee9 to d68f876 Compare September 1, 2026 07:54 IMbackK approved these changes Sep 3, 2026 View reviewed changes IMbackK left a comment • edited Loading Uh oh! There was an error while loading. Please reload this page . Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . nice, looks good Details GPU Model Microbatch size Test t/s master t/s fattn-wmma-rdna4-256 Speedup AI PRO R9700 gemma4 26B.A4B Q6_K 8 pp2048@d32768 327.19 308.26 0.94 AI PRO R9700 gemma4 26B.A4B Q6_K 64 pp2048@d32768 662.22 668.99 1.01 AI PRO R9700 gemma4 26B.A4B Q6_K 512 pp2048@d32768 1263.12 1256.57 0.99 AI PRO R9700 gemma4 26B.A4B Q6_K 1024 pp2048@d32768 1366.90 1495.99 1.09 AI PRO R9700 gpt-oss 20B MXFP4 MoE 8 pp2048@d32768 523.46 525.46 1.00 AI PRO R9700 gpt-oss 20B MXFP4 MoE 64 pp2048@d32768 1408.82 1364.30 0.97 AI PRO R9700 gpt-oss 20B MXFP4 MoE 512 pp2048@d32768 2588.42 2838.27 1.10 AI PRO R9700 gpt-oss 20B MXFP4 MoE 1024 pp2048@d32768 2869.95 3281.25 1.14 AI PRO R9700 lfm2moe 8B.A1B Q8_0 8 pp2048@d32768 817.76 817.97 1.00 AI PRO R9700 lfm2moe 8B.A1B Q8_0 64 pp2048@d32768 2947.91 2947.64 1.00 AI PRO R9700 lfm2moe 8B.A1B Q8_0 512 pp2048@d32768 7976.07 7986.46 1.00 AI PRO R9700 lfm2moe 8B.A1B Q8_0 1024 pp2048@d32768 9433.99 8928.84 0.95 AI PRO R9700 llama 8B Q8_0 8 pp2048@d32768 312.91 312.85 1.00 AI PRO R9700 llama 8B Q8_0 64 pp2048@d32768 1209.11 1181.70 0.98 AI PRO R9700 llama 8B Q8_0 512 pp2048@d32768 1145.38 1705.73 1.49 AI PRO R9700 llama 8B Q8_0 1024 pp2048@d32768 1064.61 1587.85 1.49 AI PRO R9700 qwen35 27B Q5_K_M 8 pp2048@d32768 75.64 72.25 0.96 AI PRO R9700 qwen35 27B Q5_K_M 64 pp2048@d32768 372.91 479.10 1.28 AI PRO R9700 qwen35 27B Q5_K_M 512 pp2048@d32768 468.45 707.39 1.51 AI PRO R9700 qwen35 27B Q5_K_M 1024 pp2048@d32768 475.40 708.34 1.49 <\details> Comment thread ggml/src/ggml-cuda/fattn-common.cuh Outdated bool use_stream_k = cc >= GGML_CUDA_CC_ADA_LOVELACE || amd_wmma_available(cc) || tiles_efficiency_percent < 75; if (amd_wmma_available(cc) && ntiles_dst >= 2*max_blocks && tiles_efficiency_percent >= 75) { use_stream_k = false; } IMbackK Sep 3, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . we might want to refactor this into a helper function with table at this point. sbstnh commented Sep 3, 2026 Copy link Copy Markdown can confirm the performance improvements on my R9700. Hope this PR will be merged soon. pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author @IMbackK added the helper. JohannesGaessler requested changes Sep 5, 2026 View reviewed changes JohannesGaessler left a comment Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Please avoid piling on many unrelated changes into a single PR like this. JohannesGaessler self-assigned this Sep 5, 2026 pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author Yeah, I just noticed the GDA changes got swept in as well, I'll move them to a separate branch. pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author @JohannesGaessler aight, cleaned it up to just the FATTN changes. JohannesGaessler commented Sep 5, 2026 Copy link Copy Markdown Contributor For this PR, please reduce it to just the configuration changes. Make a follow-up with only the changes for the stream-k logic. The other changes I'm not sure are logically correct so I would like to treat them separately. It is a lot easier for me to check the performance with separate PRs. I'm available to work on this this weekend so I think we can get it done if you are as well. pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author Aight, will try that. JohannesGaessler commented Sep 6, 2026 Copy link Copy Markdown Contributor Actually, looking at the order the commits were done in it's not clear to me that the config change and the change to stream-k can be separated. But Claude is definitely wrong about the supposed race condition it fixed. IMbackK commented Sep 6, 2026 Copy link Copy Markdown Contributor imo the state of this pr was fine at 2d55b5f IMbackK commented Sep 6, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page . Copy link Copy Markdown Contributor nvm, your right d6feff3 is not necessary its just reading its own rows in that case, got bamboozled by that one. up to 2d55b5f minus d6feff3 is good. pwilkin commented Sep 6, 2026 Copy link Copy Markdown Member Author Yeah, shouldn't have trusted Claude on that one. Reverting that. IMbackK commented Sep 6, 2026 Copy link Copy Markdown Contributor mind doing a git rebase -i instead? gets kinda confusing to review like this. JohannesGaessler reviewed Sep 6, 2026 View reviewed changes Comment thread ggml/src/ggml-cuda/fattn-mma-f16.cuh Outdated Comment on lines +582 to +587 // swizzle the tile stride for K and V based on the batch size. constexpr int stride_tile_K = ggml_cuda_fattn_smem_swizzle::tile_stride(nbatch_K2); #if defined(AMD_WMMA_AVAILABLE) constexpr int stride_tile_V = V_is_K_view ? stride_tile_K : nbatch_V2 + 6; #else constexpr int stride_tile_V = V_is_K_view ? stride_tile_K : ggml_cuda_fattn_smem_swizzle::tile_stride(nbatch_V2); #endif // defined(AMD_WMMA_AVAILABLE) JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Again, please submit changes to the actual device code beyond changes to the config and the host-side orchestration separately . I would be surprised if this is actually the correct padding to minimize LDS bank conflicts. pwilkin force-pushed the fattn-wmma-rdna4-256 branch
from f6264d8 to a3b28f9 Compare September 6, 2026 13:08 pwilkin commented Sep 6, 2026 Copy link Copy Markdown Member Author Rebased, squashed, hopefully it's good now. JohannesGaessler reviewed Sep 6, 2026 View reviewed changes Comment thread ggml/src/ggml-cuda/fattn.cu Outdated Comment on lines +224 to +240 if (GGML_CUDA_CC_IS_RDNA4(cc)) { if (use_gqa_opt && gqa_ratio % 8 == 0) { ggml_cuda_flash_attn_ext_mma_f16_switch_ncols1<DKQ, DV, 8>(ctx, dst); return; } if (use_gqa_opt && gqa_ratio % 4 == 0) { ggml_cuda_flash_attn_ext_mma_f16_switch_ncols1<DKQ, DV, 4>(ctx, dst); return; } if (use_gqa_opt && gqa_ratio % 2 == 0) { ggml_cuda_flash_attn_ext_mma_f16_switch_ncols1<DKQ, DV, 2>(ctx, dst); return; } } JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . I don't think these extra code branches are needed. JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Looking at the exact parameters for Qwen 3.5, what these seem to be doing is launch different kernels for GQA ratios that are not a power of 2, Qwen 3.5 has 6. pwilkin Sep 6, 2026 Copy link Copy Markdown Member Author There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Yeah, is that wrong? JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . I don't know yet. DanoPTT mentioned this pull request Sep 6, 2026 Eval bug: with -np 2 and --kv-unified, prompt processing drops 42-54% from the second long request onward (single GPU, sequential requests, no spill, no speculation) #28495 Open JohannesGaess