2026-10-11 16:37 UTC

llama.cpp contributor pwilkin claims the merged Flash Attention tuning in PR #28102 materially accelerates long-context prefill on AMD RDNA4 hardware, potentially improving local inference responsiveness without a comparable decode-speed gain.

state: corroboratedheat: lowuncertainty: mediumnovelscott: lowllama-cpp local-inference amd-inference flash-attentionpwilkinIMbackKJohannesGaesslerggml-org

What is this?

llama.cpp is software for running language models locally, with GPU backends including CUDA and AMD hipBLAS and an option to enable Flash Attention, according to the supplied Qwen documentation. The case identifies ggml-org/llama.cpp PR #28102 as CUDA/HIP Flash Attention tuning for gfx1201 by contributor pwilkin, reportedly addressing poor long-context prefill on an R9700 PRO. The web results corroborate pwilkin's participation in llama.cpp but do not include the PR itself, so they do not establish its merge status, benchmark gains, or the claimed difference between prefill and decode improvements; results about other hardware and optimizations cannot verify those claims.

Why it matters to Scott

This is only an example of Scott’s hardware-aware local inference practice: gamepc documents a WSL2/CUDA substrate, not an RDNA4 deployment, so the supplied hits establish no immediate build change or challenge to his claims. The radar’s “amd-llama-cpp-prefill-speedup” tracks a different branch targeting integrated GPUs, not this PR; this development is distinct, but its merge status and prefill-versus-decode gains remain unverified.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.llama-cppradar:concept.amd-inferenceradar:picchio-llama-cpp-bottleneck-diagnosticsradar:amd-llama-cpp-prefill-speedup
queries asked of Scott's wikis
  • llama.cpp local inference backend dependencies
  • AMD GPU HIP local inference hardware projects
  • long-context prefill time-to-first-token versus decode benchmarks
  • agent harness latency large prompts local models
  • RAG agent memory context budgets inference costs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 1010h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-30 14:00⭐ origin echo-reconstructedReports poor long-context prefill on an R9700 PRO and proposes Flash Attention tuning, with benchmarks showing substantial prefill gains and
pwilkin on github (echo) · attributed from reddit.post.1wdbal8
—
09-11 09:27first on r/LocalLLaMA · published · +283.5hCUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp
pmttyji
—
09-11 09:27amplified on r/LocalLLaMA 👑reddit.post.1wdbal8
pmttyji
peak 68 · 16 comments · 100% of case engagement
09-11 10:20our radar first saw it · +284.3hdiscovery anchor: reddit.post.1wdbal8—
pace: p66 vs 519 stories at the 720h mark (now 1010h old) — ahead of antfly-v02-zig-rewrite (1.0x), behind solveathome-public-agent-research (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditCUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-11T10:22:56.323220+00:00

Uh oh! There was an error while loading. Please reload this page . ggml-org / llama.cpp Public Notifications You must be signed in to change notification settings Fork 23k Star 128k CUDA/HIP: Flash Attention tuning (gfx1201) - # 28102 # 28102 Merged IMbackK merged 4 commits into ggml-org:master ggml-org/llama.cpp:master from pwilkin:fattn-wmma-rdna4-256 pwilkin/llama.cpp:fattn-wmma-rdna4-256 Copy head branch name to clipboard Sep 11, 2026 Merged CUDA/HIP: Flash Attention tuning (gfx1201) # 28102 IMbackK merged 4 commits into ggml-org:master ggml-org/llama.cpp:master from pwilkin:fattn-wmma-rdna4-256 pwilkin/llama.cpp:fattn-wmma-rdna4-256 Copy head branch name to clipboard Conversation pwilkin commented Aug 31, 2026 Copy link Copy Markdown Member Overview So, got my new R9700 PRO. It's nice and shiny and has 32GB VRAM, so I decided to try out Qwen3.8 27B. That was a mistake. The prefill performance at longer contexts was abysmal, so I decided to do something about it. Also managed to fix a HS=256 bug in the general CUDA FA code which was preventing the selection of 256 kernels before. Additional information Before: model size params backend threads dev test t/s qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 1101.41 ± 221.33 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 29.81 ± 0.19 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d40000 426.36 ± 37.60 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d40000 26.75 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d150000 164.42 ± 5.79 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d150000 19.70 ± 0.20 After: model size params backend threads dev test t/s qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 1103.66 ± 254.83 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 29.87 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d40000 639.46 ± 185.36 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d40000 26.94 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d150000 399.01 ± 30.87 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d150000 19.67 ± 0.39 Requirements I have read and agree with the contributing guidelines AI usage disclosure: Yes pwilkin requested review from a team and ggerganov as code owners August 31, 2026 16:54 github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Aug 31, 2026 Geramy commented Aug 31, 2026 Copy link Copy Markdown Contributor This looks really good! Such a simple change providing extra performance, good job on finding it! IMbackK self-assigned this Aug 31, 2026 pwilkin commented Sep 1, 2026 Copy link Copy Markdown Member Author I'm not done yet! model size params backend threads dev test t/s qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 1138.55 ± 224.34 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 29.48 ± 0.83 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d40000 704.92 ± 239.12 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d40000 27.00 ± 0.46 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 pp512 @ d150000 456.51 ± 38.04 qwen35 27B IQ4_XS - 4.25 bpw 13.26 GiB 27.32 B BLAS,CUDA,ROCm,Vulkan 8 ROCm0 tg128 @ d150000 19.67 ± 0.42 pwilkin force-pushed the fattn-wmma-rdna4-256 branch
    from 6b24ee9 to d68f876 Compare September 1, 2026 07:54 IMbackK approved these changes Sep 3, 2026 View reviewed changes IMbackK left a comment • edited Loading Uh oh! There was an error while loading. Please reload this page . Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . nice, looks good Details GPU Model Microbatch size Test t/s master t/s fattn-wmma-rdna4-256 Speedup AI PRO R9700 gemma4 26B.A4B Q6_K 8 pp2048@d32768 327.19 308.26 0.94 AI PRO R9700 gemma4 26B.A4B Q6_K 64 pp2048@d32768 662.22 668.99 1.01 AI PRO R9700 gemma4 26B.A4B Q6_K 512 pp2048@d32768 1263.12 1256.57 0.99 AI PRO R9700 gemma4 26B.A4B Q6_K 1024 pp2048@d32768 1366.90 1495.99 1.09 AI PRO R9700 gpt-oss 20B MXFP4 MoE 8 pp2048@d32768 523.46 525.46 1.00 AI PRO R9700 gpt-oss 20B MXFP4 MoE 64 pp2048@d32768 1408.82 1364.30 0.97 AI PRO R9700 gpt-oss 20B MXFP4 MoE 512 pp2048@d32768 2588.42 2838.27 1.10 AI PRO R9700 gpt-oss 20B MXFP4 MoE 1024 pp2048@d32768 2869.95 3281.25 1.14 AI PRO R9700 lfm2moe 8B.A1B Q8_0 8 pp2048@d32768 817.76 817.97 1.00 AI PRO R9700 lfm2moe 8B.A1B Q8_0 64 pp2048@d32768 2947.91 2947.64 1.00 AI PRO R9700 lfm2moe 8B.A1B Q8_0 512 pp2048@d32768 7976.07 7986.46 1.00 AI PRO R9700 lfm2moe 8B.A1B Q8_0 1024 pp2048@d32768 9433.99 8928.84 0.95 AI PRO R9700 llama 8B Q8_0 8 pp2048@d32768 312.91 312.85 1.00 AI PRO R9700 llama 8B Q8_0 64 pp2048@d32768 1209.11 1181.70 0.98 AI PRO R9700 llama 8B Q8_0 512 pp2048@d32768 1145.38 1705.73 1.49 AI PRO R9700 llama 8B Q8_0 1024 pp2048@d32768 1064.61 1587.85 1.49 AI PRO R9700 qwen35 27B Q5_K_M 8 pp2048@d32768 75.64 72.25 0.96 AI PRO R9700 qwen35 27B Q5_K_M 64 pp2048@d32768 372.91 479.10 1.28 AI PRO R9700 qwen35 27B Q5_K_M 512 pp2048@d32768 468.45 707.39 1.51 AI PRO R9700 qwen35 27B Q5_K_M 1024 pp2048@d32768 475.40 708.34 1.49 <\details> Comment thread ggml/src/ggml-cuda/fattn-common.cuh Outdated bool use_stream_k = cc >= GGML_CUDA_CC_ADA_LOVELACE || amd_wmma_available(cc) || tiles_efficiency_percent < 75; if (amd_wmma_available(cc) && ntiles_dst >= 2*max_blocks && tiles_efficiency_percent >= 75) { use_stream_k = false; } IMbackK Sep 3, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . we might want to refactor this into a helper function with table at this point. sbstnh commented Sep 3, 2026 Copy link Copy Markdown can confirm the performance improvements on my R9700. Hope this PR will be merged soon. pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author @IMbackK added the helper. JohannesGaessler requested changes Sep 5, 2026 View reviewed changes JohannesGaessler left a comment Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Please avoid piling on many unrelated changes into a single PR like this. JohannesGaessler self-assigned this Sep 5, 2026 pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author Yeah, I just noticed the GDA changes got swept in as well, I'll move them to a separate branch. pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author @JohannesGaessler aight, cleaned it up to just the FATTN changes. JohannesGaessler commented Sep 5, 2026 Copy link Copy Markdown Contributor For this PR, please reduce it to just the configuration changes. Make a follow-up with only the changes for the stream-k logic. The other changes I'm not sure are logically correct so I would like to treat them separately. It is a lot easier for me to check the performance with separate PRs. I'm available to work on this this weekend so I think we can get it done if you are as well. pwilkin commented Sep 5, 2026 Copy link Copy Markdown Member Author Aight, will try that. JohannesGaessler commented Sep 6, 2026 Copy link Copy Markdown Contributor Actually, looking at the order the commits were done in it's not clear to me that the config change and the change to stream-k can be separated. But Claude is definitely wrong about the supposed race condition it fixed. IMbackK commented Sep 6, 2026 Copy link Copy Markdown Contributor imo the state of this pr was fine at 2d55b5f IMbackK commented Sep 6, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page . Copy link Copy Markdown Contributor nvm, your right d6feff3 is not necessary its just reading its own rows in that case, got bamboozled by that one. up to 2d55b5f minus d6feff3 is good. pwilkin commented Sep 6, 2026 Copy link Copy Markdown Member Author Yeah, shouldn't have trusted Claude on that one. Reverting that. IMbackK commented Sep 6, 2026 Copy link Copy Markdown Contributor mind doing a git rebase -i instead? gets kinda confusing to review like this. JohannesGaessler reviewed Sep 6, 2026 View reviewed changes Comment thread ggml/src/ggml-cuda/fattn-mma-f16.cuh Outdated Comment on lines +582 to +587 // swizzle the tile stride for K and V based on the batch size. constexpr int stride_tile_K = ggml_cuda_fattn_smem_swizzle::tile_stride(nbatch_K2); #if defined(AMD_WMMA_AVAILABLE) constexpr int stride_tile_V = V_is_K_view ? stride_tile_K : nbatch_V2 + 6; #else constexpr int stride_tile_V = V_is_K_view ? stride_tile_K : ggml_cuda_fattn_smem_swizzle::tile_stride(nbatch_V2); #endif // defined(AMD_WMMA_AVAILABLE) JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Again, please submit changes to the actual device code beyond changes to the config and the host-side orchestration separately . I would be surprised if this is actually the correct padding to minimize LDS bank conflicts. pwilkin force-pushed the fattn-wmma-rdna4-256 branch
    from f6264d8 to a3b28f9 Compare September 6, 2026 13:08 pwilkin commented Sep 6, 2026 Copy link Copy Markdown Member Author Rebased, squashed, hopefully it's good now. JohannesGaessler reviewed Sep 6, 2026 View reviewed changes Comment thread ggml/src/ggml-cuda/fattn.cu Outdated Comment on lines +224 to +240 if (GGML_CUDA_CC_IS_RDNA4(cc)) { if (use_gqa_opt && gqa_ratio % 8 == 0) { ggml_cuda_flash_attn_ext_mma_f16_switch_ncols1<DKQ, DV, 8>(ctx, dst); return; } if (use_gqa_opt && gqa_ratio % 4 == 0) { ggml_cuda_flash_attn_ext_mma_f16_switch_ncols1<DKQ, DV, 4>(ctx, dst); return; } if (use_gqa_opt && gqa_ratio % 2 == 0) { ggml_cuda_flash_attn_ext_mma_f16_switch_ncols1<DKQ, DV, 2>(ctx, dst); return; } } JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . I don't think these extra code branches are needed. JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Looking at the exact parameters for Qwen 3.5, what these seem to be doing is launch different kernels for GQA ratios that are not a power of 2, Qwen 3.5 has 6. pwilkin Sep 6, 2026 Copy link Copy Markdown Member Author There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . Yeah, is that wrong? JohannesGaessler Sep 6, 2026 Copy link Copy Markdown Contributor There was a problem hiding this comment. Choose a reason for hiding this comment The reason will be displayed to describe this comment to others. Learn more . I don't know yet. DanoPTT mentioned this pull request Sep 6, 2026 Eval bug: with -np 2 and --kv-unified, prompt processing drops 42-54% from the second long request onward (single GPU, sequential requests, no spill, no speculation) #28495 Open JohannesGaess
pmttyji6716
🟧 echo.github ⭐Reports poor long-context prefill on an R9700 PRO and proposes Flash Attention tuning, with benchmarks showing substantial prefill gains andpwilkin——

Interpretation history

Decision trace