2026-10-11 16:38 UTC

llama.cpp contributor thelittlefireman claims the merged GCN-specific MMQ configuration improves prompt processing by about 5% in a published gfx906 benchmark, potentially accelerating local inference on older AMD MI50/MI60-class hardware.

state: watchingheat: mediumuncertainty: mediumnovelscott: lowllama-cpp amd-gpu local-inferencethelittlefiremanIMbackKggml-org

What is this?

The case concerns ggml-org/llama.cpp PR #27841, attributed to contributor thelittlefireman, reportedly adding an AMD GCN-specific configuration for quantized matrix multiplication rather than falling back to RDNA2 settings. The supplied search snippets establish community llama.cpp optimization work for gfx906 MI50/MI60 GPUs, including specialized kernels and benchmarks, but do not include this PR or verify its claimed merge and roughly 5% prompt-processing gain. The evidence title's description of MI50/MI60 as RDNA2 conflicts with the GCN/gfx906 framing in the case and supporting repository snippets; the separate forks' benchmark results should not be treated as validation of this patch.

Why it matters to Scott

Scott’s recorded local-serving setup is WSL2/CUDA with Ollama; the hits establish neither MI50/MI60 use nor a claim this AMD-specific optimization would materially extend or challenge, so the connection is topical rather than actionable. The radar’s gfx906-llama-cpp-throughput-gains page tracks related fork benchmarks, not this upstream PR, whose merge and roughly 5% gain remain unverified in the supplied grounding.
radar:gfx906-llama-cpp-throughput-gainsradar:concept.llama-cppradar:concept.amd-inference
queries asked of Scott's wikis
  • llama.cpp local inference deployment stack
  • AMD ROCm MI50 MI60 GPU hardware projects
  • local inference economics older GPUs hardware reuse
  • quantized inference prompt processing bottlenecks benchmarks
  • upstream inference support versus hardware-specific forks

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 1082h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-27 14:00⭐ origin echo-reconstructedAdds an AMD GCN wave64 MMQ configuration instead of falling back to RDNA2 settings; reports quantized matrix-multiplication gains and a gfx9
thelittlefireman on github (echo) · attributed from reddit.post.1we8elm
—
09-12 09:53first on r/LocalLLaMA · published · +379.9hggml-cuda: hip: add missing AMD GCN MMQ config by thelittlefireman · Pull Request #27841 · ggml-org/llama.cpp - PP improvements for RDNA2(MI50, MI60)
pmttyji
—
09-12 09:53amplified on r/LocalLLaMA 👑reddit.post.1we8elm
pmttyji
peak 23 · 7 comments · 100% of case engagement
09-12 10:20our radar first saw it · +380.3hdiscovery anchor: reddit.post.1we8elm—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditggml-cuda: hip: add missing AMD GCN MMQ config by thelittlefireman · Pull Request #27841 · ggml-org/llama.cpp - PP improvements for RDNA2(MI50, MI60)
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-12T10:21:53.153607+00:00

### Uh oh!

There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27841).

[ggml-org](https://github.com/ggml-org) 
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public

- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
  23.1k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
   128k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)

# ggml-cuda: hip: add missing AMD GCN MMQ config - #27841

#27841

Merged

[IMbackK](https://github.com/IMbackK) merged 1 commit into

[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom 

[thelittlefireman:feature\_GCN\_MMQ\_CONFIG](https://github.com/thelittlefireman/llama.cpp/tree/feature_GCN_MMQ_CONFIG)thelittlefireman/llama.cpp:feature\_GCN\_MMQ\_CONFIGCopy head branch name to clipboard

Sep 12, 2026

Merged

## [ggml-cuda: hip: add missing AMD GCN MMQ config](https://github.com/ggml-org/llama.cpp/pull/27841#top)#27841 [IMbackK](https://github.com/IMbackK) merged 1 commit into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [thelittlefireman:feature\_GCN\_MMQ\_CONFIG](https://github.com/thelittlefireman/llama.cpp/tree/feature_GCN_MMQ_CONFIG)thelittlefireman/llama.cpp:feature\_GCN\_MMQ\_CONFIGCopy head branch name to clipboard

## Conversation

[@thelittlefireman](https://github.com/thelittlefireman)

### @thelittlefireman **[thelittlefireman](https://github.com/thelittlefireman)** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#issue-5273787780) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27841).

Copy link
 

Copy Markdown

Contributor

## Overview

Per-arch MMQ config for AMD GCN arch to properly handle wave64(nthreads 512 (8 warps)) to avoid the fallback to RDNA2 MMQ config (wave32/nthreads 256))

## Additional information

**Additional changes:**

- avoid unnecessary MMQ fallback for I < 128
- tile widths up to J=128 (rdna2 caps at 64) for performance with a safeguard J<=64 for MoE and -sm tensor.

They could be moved to another PR, for clarity.

Everything have been tested multiples times. My main goal was to be as good/fast as master branch for VGPR spill and kernel speeds.

**test-backend-ops:**

```
$.bin/test-backend-ops perf -b ROCm0 -o MUL_MAT -p "type_a=(q1_0|q2_0|q4_0|q4_1|q5_0|q5_1|q8_0|q2_K|q3_K|q4_K|q5_K|q6_K|iq1_s|iq2_xxs|iq2_xs|iq2_s|iq3_xxs|iq3_s|iq4_nl|iq4_xs|mxfp4|nvfp4),type_b=f32,m=(4096|4160|4097),n=(1|2|3|4|5|6|7|8|16|24|32|40|48|64|129|145|161|177|193|209|225|241|512),k=14336"
```

Curent master MAT\_MUL: <https://pastebin.com/raw/3nEP6FJJ>  
PR MAT\_MUL: <https://pastebin.com/pL16aN4S>

| Quant | Average MMQ gain | `n=512` gain | Fallback `m=4097` | Adaptive `m=4160` |
| --- | --- | --- | --- | --- |
| Q1\_0 | +26.1% | +10.6% | +13.8% | +17.3% |
| Q2\_0 | +46.2% | +55.8% | +31.4% | +50.8% |
| Q4\_0 | +37.2% | +17.9% | +24.8% | +37.0% |
| Q4\_1 | +44.8% | +21.8% | +31.9% | +45.3% |
| Q5\_0 | +25.9% | +36.2% | +14.8% | +21.5% |
| Q5\_1 | +17.1% | +36.7% | +14.8% | +25.4% |
| Q8\_0 | +46.4% | +63.9% | +45.4% | +50.3% |
| Q2\_K | +395.4% | +336.0% | +382.6% | +460.7% |
| Q3\_K | +9.6% | +10.6% | +2.6% | −0.3% |
| Q4\_K | +5.7% | −0.5% | +11.2% | +15.8% |
| Q5\_K | +21.7% | +0.0% | +53.4% | +36.6% |
| Q6\_K | +34.5% | +51.5% | +38.8% | +29.7% |
| IQ1\_S | +20.9% | +43.5% | +20.0% | +31.1% |
| IQ2\_XXS | +16.0% | +26.9% | +13.4% | +20.4% |
| IQ2\_XS | +15.6% | +6.1% | +10.3% | +16.4% |
| IQ2\_S | +10.8% | +7.1% | +2.0% | +4.0% |
| IQ3\_XXS | +15.3% | +27.0% | +13.1% | +19.0% |
| IQ3\_S | +7.0% | +27.6% | +2.4% | +3.0% |
| IQ4\_NL | +31.8% | +33.2% | +16.7% | +23.8% |
| IQ4\_XS | +18.5% | +33.1% | +16.7% | +24.7% |
| MXFP4 | +26.6% | +32.1% | +15.9% | +23.4% |
| NVFP4 | +9.1% | +9.7% | +5.7% | +9.3% |

The Q3\_K `m=4160` (−0.3%) and Q4\_K `n=512` (−0.5%) differences are within measurement variance. No MMQ test shows a regression greater than 1%.

**Small benchmark:**

```
$ ./build-feature_GCN_MMQ_CONFIG/bin/llama-bench -m "/home/user/.cache/huggingface/hub/models--unsloth--Qwen3.6-27B-MTP-GGUF/snapshots/5cb35eb3dcbf52dbce5f87dbc64df6aaffadcace/Qwen3.6-27B-Q4_1.gguf" -ngl 999 -p 2048 -n 0 -r 1 --no-warmup
ggml_cuda_init: found 3 ROCm devices (Total VRAM: 98256 MiB):
  Device 0: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| qwen35 27B Q4_1                |  16.33 GiB |    27.32 B | ROCm       | 999 |          pp2048 |        297.89 ± 0.00 |

build: 17c0b05ee (10659)
$ ./build-master/bin/llama-bench -m "/home/user/.cache/huggingface/hub/models--unsloth--Qwen3.6-27B-MTP-GGUF/snapshots/5cb35eb3dcbf52dbce5f87dbc64df6aaffadcace/Qwen3.6-27B-Q4_1.gguf" -ngl 999 -p 2048 -n 0 -r 1 --no-warmup
ggml_cuda_init: found 3 ROCm devices (Total VRAM: 98256 MiB):
  Device 0: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| qwen35 27B Q4_1                |  16.33 GiB |    27.32 B | ROCm       | 999 |          pp2048 |        283.47 ± 0.00 |

build: d74a75570 (10643)
```

[#23881](https://github.com/ggml-org/llama.cpp/discussions/23881)

## Requirements

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure :   
  YES : For testing, small code generation and help to understand. Every changes have been manually edited and tested on gfx906.

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
[CUDA](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3ACUDA)
Related to the CUDA backend
labels
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#event-30139930857)

[@thelittlefireman](https://github.com/thelittlefireman)

[thelittlefireman](https://github.com/thelittlefireman)
marked this pull request as ready for review
[August 28, 2026 01:38](https://github.com/ggml-org/llama.cpp/pull/27841#event-30139932427)

[@thelittlefireman](https://github.com/thelittlefireman)

[thelittlefireman](https://github.com/thelittlefireman)
requested review from
a team and
[ggerganov](https://github.com/ggerganov)
as [code owners](https://github.com/ggml-org/llama.cpp/blob/ca3d5a3e10d53f7ea672cb9b6178faca3e2807bc/CODEOWNERS#L61)
[August 28, 2026 01:38](https://github.com/ggml-org/llama.cpp/pull/27841#event-30139932736)

[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)

### **[ggml-gh-bot](https://github.com/apps/ggml-gh-bot) Bot** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#issuecomment-5447344871)

Copy link
 

Copy Markdown

|  |
| --- |
| Hi [@thelittlefireman](https://github.com/thelittlefireman), thanks for your contribution!  Per our [contribution guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md), the automated PR checker found the following issue(s) that need your attention:   - **PR Template not respected**: Please respect the [template](https://github.com/ggml-org/llama.cpp/blob/master/.github/pull_request_template.md?plain=1) when creating a new pull request. Make sure to fill out all required sections.  ---   *Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.* |

[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
[ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
Bot
added
the [draft](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Adraft)
PR will be changed to draft by github-actions bot
label
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#event-30140088292)

[@github-actions](https://github.com/apps/github-actions)

[github-actions](https://github.com/apps/github-actions)
Bot
marked this pull request as draft
[August 28, 2026 01:43](https://github.com/ggml-org/llama.cpp/pull/27841#event-30140092967)

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
removed
the [draft](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Adraft)
PR will be changed to draft by github-actions bot
label
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#event-30140094131)

[@thelittlefireman](https://github.com/thelittlefireman)
[thelittlefireman](https://github.com/thelittlefireman)
mentioned this pull request
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#ref-pullrequest-5043984353)

[ggml-cuda: HIP replace \_\_shfl\_xor\_sync with dpp instructions
#26466](https://github.com/ggml-org/llama.cpp/pull/26466)

Closed

[@thelittlefireman](https://github.com/thelittlefireman)

[thelittlefireman](https://github.com/thelittlefireman)
marked this pull request as ready for review
[August 28, 2026 02:00](https://github.com/ggml-org/llama.cpp/pull/27841#event-30140738129)

[@IMbackK](https://github.com/IMbackK)
[IMbackK](https://github.com/IMbackK)
self-assigned this
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#event-30148430582)

[@IMbackK](https://github.com/IMbackK)

### **[IMbackK](https://github.com/IMbackK)** commented [Sep 2, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#issuecomment-5512378460)

Copy link
 

Copy Markdown

Contributor

|  |
| --- |
| There seams to be code left over from perf testing in this pr.  Please remove those changes. |

[@thelittlefireman](https://github.com/thelittlefireman)

### **[thelittlefireman](https://github.com/thelittlefireman)** commented [Sep 2, 2026](https://github.com/ggml-org/llama.cpp/pull/27841#issuecomment-5516608218)

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| There seams to be code left over from perf testing in this pr.  Please remove those changes.  I left them cause I think they are missing for proper testing.  But i will remove them |

[@thelittlefireman](https://github.com/thelittlefireman)

[thelittlefireman](https://github.com/thelittlefireman)
[force-pushed](https://github.com/ggml-org/llama.cpp/compare/15ac7dfd750f4bb04d084672c00a745f75d25fb4..b79594dc6eff78d65ee7fad7e379c35a626d85ab)
the

feature\_GCN\_MMQ\_CONFIG
branch
from
[`15ac7df`](https://github.com/ggml-org/llama.cpp/commit/15ac7dfd750f4bb04d084672c00a745f75d25fb4) to
[`b79594d`](https://github.com/ggml-org/llama.cpp/commit/b79594dc6eff78d65ee7fad7e379c35a626d85ab)  [Compare](https://github.com/ggml-org/llama.cpp/compare/15ac7dfd750f4bb04d084672c00a745f75d25fb4..b79594dc6eff78d65ee7fad7e379c35a626d85ab)
[September 2, 2026 21:24](https://github.com/ggml-org/llama.cpp/pull/27841#event-30446628067)

[@thelitt
pmttyji237
🟧 echo.github ⭐Adds an AMD GCN wave64 MMQ configuration instead of falling back to RDNA2 settings; reports quantized matrix-multiplication gains and a gfx9thelittlefireman——

Interpretation history

Decision trace