2026-10-11 17:15 UTC

llama.cpp contributor pratiknarola-t's merged PR #29869 claims few-row MMA Metal mat-mul and batched-copy kernels turn DFlash2 speculative decoding of Qwen3.8-27B from slower than serial decoding (~30 tok/s) into ~110 tok/s on an M3 Ultra โ€” with the kernels, tests, and benchmarks disclosed as Claude Code-written โ€” and the gains replicating across Apple GPUs, models, and specdec drafters would establish few-row matmul optimization as the standard enabler of speculative decoding on non-tensor-API Apple Silicon.

state: seedheat: lowuncertainty: mediumconvergesscott: highllama-cpp speculative-decoding local-inference apple-siliconpratiknarola-tggerganov

What is this?

llama.cpp โ€” the ggml-org open-source GGUF inference engine โ€” merged contributor pratiknarola-t's PR #29869 on Oct 5, 2026, adding few-row MMA mat-mul and batched-copy kernels to the Metal backend, targeting the small-batch shapes that speculative decoding hits hardest. The reported effect (echoed second-hand on X) is that DFlash2 โ€” a block-diffusion drafter that predicts whole token blocks for Qwen3.8-27B in one pass and is lossless against the target model โ€” goes from ~30 tok/s (slower than plain decoding) to ~110 tok/s on an M3 Ultra. The surrounding evidence is mixed: DFlash2 shows 3.4x server-side gains on H200s and community PR-build benchmarks report 2.26x on real coding prompts, but pre-kernel Apple Silicon hands-ons measured only 11โ€“35 tok/s with Ollama/MLX beating the llama.cpp build, and gains are known to compress sharply under concurrency. The snippets do not include the PR text itself, so the disclosed Claude Code authorship of the kernels/tests and any cross-device/cross-drafter replication rest on the contributor's report and are unverified here.

Why it matters to Scott

A first-party llama.cpp merge of disclosed-Claude-Code Metal kernels shipping with tests and benchmarks is a dated receipt for his verification-ebook thesis โ€” trust the tests, not the developer โ€” now operating at the GPU-kernel layer in a major inference engine, and the reported ~30โ†’110 tok/s DFlash2 result, if replicated, flips the MLX-vs-llama.cpp serving tradeoff his local stack (MLX Mac mini, Ollama endpoint) rides on while answering his own 'when does specdec pay off locally' question (pre-kernel: it didn't). Caveats: kernel authorship and cross-device/cross-drafter replication rest on the contributor's report, and the sibling ynankani CUDA MoE-fusion case remains a distinct open episode โ€” lineage, not corroboration.
ip:source.custom-software-verification-ebookip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookip:concept.model-barbelldev:concept.hardware-aware-local-inferencedev:technology.mlxdev:technology.ollamaradar:concept.llama-cppradar:concept.speculative-decodingradar:concept.local-inferenceradar:concept.apple-siliconradar:concept.qwen38radar:llama-cpp-specdec-moe-fusionradar:dflash-2-parallel-drafting-validationradar:qwen38-dflash2-long-context-speedupradar:mlx-dspark-muse-glimmer-speedupradar:llama-cpp-metal-iq3-moe-speedupradar:magnitude-self-optimizing-inference-engineradar:inco-splash-apple-silicon-inferenceradar:mlxfast-agent-engine-rewriteradar:coding-agent-pr-merge-rates
queries asked of Scott's wikis
  • llama.cpp Metal backend performance Apple Silicon local inference notes
  • speculative decoding drafters โ€” when specdec pays off locally vs plain decode
  • coding agents writing GPU kernels / systems code โ€” Claude Code-authored PRs disclosure
  • MLX vs GGUF/llama.cpp local serving tradeoffs for coding agents
  • Qwen 27B-class models as local coding-agent workhorse
  • open-source maintainer review of AI-generated contributions

Measured heat

now 0 pts/hpeak 5 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 242h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-01 14:00โญ origin echo-reconstructedPR opened Oct 2 and merged into llama.cpp master Oct 5, 2026: 'metal : few-row MMA mat-mul and batched copies for speculative decoding' โ€” ad
pratiknarola-t (merged into master by ggerganov) on github (echo) ยท attributed from reddit.post.1wyurmi
โ€”
10-06 05:44first on r/LocalLLaMA ยท published ยท +111.7hmetal : few-row MMA mat-mul and batched copies for speculative decoding by pratiknarola-t ยท Pull Request #29869 ยท ggml-org/llama.cpp
pmttyji
โ€”
10-06 05:44amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyurmi
pmttyji
peak 3 ยท 3 comments ยท 100% of case engagement
10-06 06:20our radar first saw it ยท +112.3hdiscovery anchor: reddit.post.1wyurmiโ€”
pace: p38 vs 1188 stories at the 168h mark (now 242h old) โ€” ahead of agentsec-static-config-auditing (1.2x), behind agent-trace-tampering (0.8x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditmetal : few-row MMA mat-mul and batched copies for speculative decoding by pratiknarola-t ยท Pull Request #29869 ยท ggml-org/llama.cpp
LocalLLaMA
Retrieved article excerpt

Open article ยท Retrieved 2026-10-06T06:24:40.428478+00:00

[ggml-org](https://github.com/ggml-org) 
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public

- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
  24.1k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
   130k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)

# metal : few-row MMA mat-mul and batched copies for speculative decoding - #29869

#29869

Merged

[ggerganov](https://github.com/ggerganov) merged 10 commits into

[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom 

[pratiknarola-t:metal-few-row-mma](https://github.com/pratiknarola-t/llama.cpp/tree/metal-few-row-mma)pratiknarola-t/llama.cpp:metal-few-row-mmaCopy head branch name to clipboard

Oct 5, 2026

Merged

## [metal : few-row MMA mat-mul and batched copies for speculative decoding](https://github.com/ggml-org/llama.cpp/pull/29869#top)#29869 [ggerganov](https://github.com/ggerganov) merged 10 commits into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [pratiknarola-t:metal-few-row-mma](https://github.com/pratiknarola-t/llama.cpp/tree/metal-few-row-mma)pratiknarola-t/llama.cpp:metal-few-row-mmaCopy head branch name to clipboard

## Conversation

[@pratiknarola-t](https://github.com/pratiknarola-t)

### @pratiknarola-t **[pratiknarola-t](https://github.com/pratiknarola-t)** commented [Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#issue-5681163061) โ€ข edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/29869).

Copy link
 

Copy Markdown

Contributor

## Overview

Speculative decoding verifies a few draft tokens per step, and batched decoding runs one row per sequence, so the model runs mat-muls with 2..16 src1 rows. On Apple GPUs without the tensor API (M1 to M4), Metal runs these with the mat-vec kernels, whose time grows with each row. On an M3 Ultra, DFlash2 decoding of Qwen3.8-27B is therefore slower than serial decoding on master.

This PR adds:

- Mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices. Each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4\_0, Q8\_0 and Q5\_K have their own kernels. F32, F16, Q4\_1, Q5\_0, Q5\_1, Q4\_K and Q6\_K use a generic path over the existing 16-weight dequantizers. Q4\_0 at 2 rows uses a 2-row variant of the mat-vec kernel.
- The kernels fill simdgroup matrices per lane, so they run only on MTLGPUFamilyApple7+ without the tensor API. They start at the row count where they beat the mat-vec kernels on an M3 Ultra: 6 for F32, 3 for F16, Q4\_K, Q5\_0 and Q5\_1, and 2 for the other types.
- A fusion table entry: MUL\_MAT + ADD writes the residual sum from the MMA store.
- CONCAT splits long rows across threadgroups when there are few rows.

## Additional information

Changes to the fusion and reorder code:

- The fusion checks, `ggml_metal_fusion_next`, `ggml_metal_fusion_max` and `ggml_graph_optimize` take the device props, so the reorder packs MUL\_MAT + ADD only on devices that fuse it. The pack does not depend on the src1 row count, so graphs of different batch sizes get the same node order (no reallocation with `GGML_SCHED_NO_REALLOC`).
- The alloc-deps pass also matches the new MUL\_MAT + ADD pattern, so on every Metal device the mat-mul inputs stay alive until the ADD.

Performance on an Apple M3 Ultra (60-core GPU), macOS 15.7.9, Qwen3.8-27B Q4\_0, master [836d571](https://github.com/ggml-org/llama.cpp/commit/836d57176dc699a726c55418e4f96b8ca628e1bf).

`test-backend-ops perf -o MUL_MAT`, m=4096, k=14336, time per run of this PR divided by master (two interleaved runs each). This table is from master [2923cf2](https://github.com/ggml-org/llama.cpp/commit/2923cf2862ad0afa159444cf07fec7600d755fe1). The kernels of this PR and the mat-vec kernels of master did not change after that.

| type | n=2 | n=3 | n=4 | n=5 | n=8 |
| --- | --- | --- | --- | --- | --- |
| F32 | 1.00 | 0.97 | 1.00 | 0.99 | 0.57 |
| F16 | 0.99 | 0.97 | 0.71 | 0.59 | 0.34 |
| Q4\_0 | 0.58 | 0.58 | 0.44 | 0.36 | 0.22 |
| Q4\_1 | 0.93 | 0.65 | 0.50 | 0.41 | 0.25 |
| Q5\_0 | 1.01 | 0.95 | 0.75 | 0.63 | 0.38 |
| Q5\_1 | 0.99 | 0.94 | 0.75 | 0.63 | 0.37 |
| Q8\_0 | 0.86 | 0.59 | 0.46 | 0.38 | 0.24 |
| Q4\_K | 1.00 | 0.97 | 0.56 | 0.45 | 0.29 |
| Q5\_K | 0.81 | 0.56 | 0.49 | 0.45 | 0.25 |
| Q6\_K | 0.92 | 0.64 | 0.48 | 0.39 | 0.21 |

n=1 and n=512 are unchanged (0.99 to 1.01). A sweep of 9..16 rows gives 0.25 to 0.40 for all ten types.

`llama-batched-bench -npp 512 -ntg 32 -npl 1,2,3,4,8 -c 32768 -pps -kvu`, TG t/s (mean of two interleaved runs each). PP is 317 t/s on both builds.

| B | master | this PR |
| --- | --- | --- |
| 1 | 32.9 | 32.9 |
| 2 | 32.7 | 49.7 |
| 3 | 36.2 | 62.2 |
| 4 | 36.3 | 76.1 |
| 8 | 38.6 | 119.7 |

llama-server with a Qwen3.8-27B DFlash2 Q8\_0 drafter, `-ngl 99 -fa on -c 8192 -np 1 --jinja`, DFlash2 with `--spec-type draft-dflash --spec-draft-n-max 7`. 64 generated tokens, median of 5 requests after one warm-up, mean of two server runs. Decode tok/s:

| mode | prompt | T | master | this PR |
| --- | --- | --- | --- | --- |
| serial | code | 0 | 32.1 | 32.0 |
| serial | code | 1 | 32.1 | 32.0 |
| serial | prose | 0 | 32.1 | 32.0 |
| serial | prose | 1 | 32.1 | 32.0 |
| DFlash2 | code | 0 | 30.2 | 110.0 |
| DFlash2 | code | 1 | 24.3 | 80.9 |
| DFlash2 | prose | 0 | 16.8 | 62.6 |
| DFlash2 | prose | 1 | 13.9 | 48.8 |

At T=0 the DFlash2 replies are the same as the serial replies (first 160 characters). At T=1 the two builds accept the same draft tokens.

`llama-bench -fa 1` gives pp512 317.2 t/s on both builds, and tg128 32.45 t/s on master and 32.36 t/s with this PR (two interleaved runs each). `llama-perplexity` (10 chunks of 512 tokens) gives 5.1775 on both builds with `-ub 512`, and 5.1773 on master and 5.1774 with this PR with `-ub 8`.

Tests:

- `test-backend-ops -b MTL0`: 16806/16806 on the M3 Ultra, and on an Apple M5 with the tensor API on and off (`GGML_METAL_TENSOR_DISABLE=1`).
- `test-metal-graph-optimize` passes on both devices.
- With `-DGGML_SCHED_NO_REALLOC=ON`, as in the Metal CI job: `test-llama-archs -s 1` with 1 to 4 devices gives the same results as master, with no failures, and `test-fusion --check tests/fusion/MTL.csv` passes. The test-fusion graphs have 1 or 32 src1 rows, so the new fusion does not fire there.

The row thresholds come from the M3 Ultra. With the tensor API turned off, an M5 reaches parity one or two rows later for some types. I did not tune for that, because the M5 uses the tensor API by default.

This ports the Metal part of [tetherto#321](https://github.com/tetherto/qvac-fabric-llm.cpp/pull/321) to master.

## Requirements

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: YES - the kernels, the port to master, the tests, the benchmarks and this description were written with Claude Code (Anthropic).

[@pratiknarola-t](https://github.com/pratiknarola-t)

`metal : few-row MMA mat-mul and batched copies for speculative decoding`

 โ€ฆ

`cc470b6`

```
Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding.

- add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel
- use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2)
- fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch
- the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count
- views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group
- the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read
- CONCAT splits long rows across threadgroups when there are few rows
- tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias
```

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
[Apple Metal](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3A%22Apple%20Metal%22)
https://en.wikipedia.org/wiki/Metal\_(API)
labels
[Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#event-32353932404)

[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)

### **[ggml-gh-bot](https://github.com/apps/ggml-gh-bot) Bot** commented [Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#issuecomment-5958107365)

Copy link
 

Copy Markdown

|  |
| --- |
| Hi [@pratiknarola-t](https://github.com/pratiknarola-t), thanks for your contribution!  Per our [contribution guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md), the automated PR checker found the following issue(s) that need your attention:   - **Large PR**: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.  ---   *Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.* |

[@ggerganov](https://github.com/ggerganov)
[ggerganov](https://github.com/ggerganov)
self-assigned this
[Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#event-32354789927)

[ggerganov](https://github.com/ggerganov)

**[ggerganov](https://github.com/ggerganov)**
reviewed
[Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#pullrequestreview-5395590974)

[View reviewed changes](https://github.com/ggml-org/llama.cpp/pull/29869/files/cc470b64e205d7ac09684fe9ae85e4a48d9d2ce6)

Comment thread

[ggml/src/ggml-metal/ggml-metal-fusion.cpp](https://github.com/ggml-org/llama.cpp/pull/29869/files/cc470b64e205d7ac09684fe9ae85e4a48d9d2ce6#diff-8372ceb3f0618f985e825802fe46a025e433bbb594987e7370583d6a40b2a427)

Outdated

Comment on lines
+853
 to 
+854

|  |  |  |
| --- | --- | --- |
|  |  | // longest batch first, so ggml\_metal\_fusion\_next checks no shorter batch once one matches |
|  |  | { GGML\_METAL\_FUSION\_CPY\_BATCH, std::vector<ggml\_op>(16, GGML\_OP\_CPY), {}, true, ggml\_metal\_fusion\_check\_cpy\_batch }, |

### @ggerganov **[ggerganov](https://github.com/ggerganov)** [Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#discussion_r4168831964)

Copy link
 

Copy Markdown

Member

There was a problem hiding this comment.

### Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. [Learn more](https://docs.github.com/articles/managing-disruptive-comments/#hiding-a-comment).

Do you have an estimate of how big is the impact of the 
pmttyji23
๐ŸŸง echo.github โญPR opened Oct 2 and merged into llama.cpp master Oct 5, 2026: 'metal : few-row MMA mat-mul and batched copies for speculative decoding' โ€” adpratiknarola-t (merged into master by ggerganov)โ€”โ€”

Interpretation history

Decision trace