Retrieved article excerpt
Open article ยท Retrieved 2026-10-06T06:24:40.428478+00:00
[ggml-org](https://github.com/ggml-org)
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public
- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
24.1k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
130k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
# metal : few-row MMA mat-mul and batched copies for speculative decoding - #29869
#29869
Merged
[ggerganov](https://github.com/ggerganov) merged 10 commits into
[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom
[pratiknarola-t:metal-few-row-mma](https://github.com/pratiknarola-t/llama.cpp/tree/metal-few-row-mma)pratiknarola-t/llama.cpp:metal-few-row-mmaCopy head branch name to clipboard
Oct 5, 2026
Merged
## [metal : few-row MMA mat-mul and batched copies for speculative decoding](https://github.com/ggml-org/llama.cpp/pull/29869#top)#29869 [ggerganov](https://github.com/ggerganov) merged 10 commits into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [pratiknarola-t:metal-few-row-mma](https://github.com/pratiknarola-t/llama.cpp/tree/metal-few-row-mma)pratiknarola-t/llama.cpp:metal-few-row-mmaCopy head branch name to clipboard
## Conversation
[@pratiknarola-t](https://github.com/pratiknarola-t)
### @pratiknarola-t **[pratiknarola-t](https://github.com/pratiknarola-t)** commented [Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#issue-5681163061) โข edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/29869).
Copy link
Copy Markdown
Contributor
## Overview
Speculative decoding verifies a few draft tokens per step, and batched decoding runs one row per sequence, so the model runs mat-muls with 2..16 src1 rows. On Apple GPUs without the tensor API (M1 to M4), Metal runs these with the mat-vec kernels, whose time grows with each row. On an M3 Ultra, DFlash2 decoding of Qwen3.8-27B is therefore slower than serial decoding on master.
This PR adds:
- Mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices. Each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4\_0, Q8\_0 and Q5\_K have their own kernels. F32, F16, Q4\_1, Q5\_0, Q5\_1, Q4\_K and Q6\_K use a generic path over the existing 16-weight dequantizers. Q4\_0 at 2 rows uses a 2-row variant of the mat-vec kernel.
- The kernels fill simdgroup matrices per lane, so they run only on MTLGPUFamilyApple7+ without the tensor API. They start at the row count where they beat the mat-vec kernels on an M3 Ultra: 6 for F32, 3 for F16, Q4\_K, Q5\_0 and Q5\_1, and 2 for the other types.
- A fusion table entry: MUL\_MAT + ADD writes the residual sum from the MMA store.
- CONCAT splits long rows across threadgroups when there are few rows.
## Additional information
Changes to the fusion and reorder code:
- The fusion checks, `ggml_metal_fusion_next`, `ggml_metal_fusion_max` and `ggml_graph_optimize` take the device props, so the reorder packs MUL\_MAT + ADD only on devices that fuse it. The pack does not depend on the src1 row count, so graphs of different batch sizes get the same node order (no reallocation with `GGML_SCHED_NO_REALLOC`).
- The alloc-deps pass also matches the new MUL\_MAT + ADD pattern, so on every Metal device the mat-mul inputs stay alive until the ADD.
Performance on an Apple M3 Ultra (60-core GPU), macOS 15.7.9, Qwen3.8-27B Q4\_0, master [836d571](https://github.com/ggml-org/llama.cpp/commit/836d57176dc699a726c55418e4f96b8ca628e1bf).
`test-backend-ops perf -o MUL_MAT`, m=4096, k=14336, time per run of this PR divided by master (two interleaved runs each). This table is from master [2923cf2](https://github.com/ggml-org/llama.cpp/commit/2923cf2862ad0afa159444cf07fec7600d755fe1). The kernels of this PR and the mat-vec kernels of master did not change after that.
| type | n=2 | n=3 | n=4 | n=5 | n=8 |
| --- | --- | --- | --- | --- | --- |
| F32 | 1.00 | 0.97 | 1.00 | 0.99 | 0.57 |
| F16 | 0.99 | 0.97 | 0.71 | 0.59 | 0.34 |
| Q4\_0 | 0.58 | 0.58 | 0.44 | 0.36 | 0.22 |
| Q4\_1 | 0.93 | 0.65 | 0.50 | 0.41 | 0.25 |
| Q5\_0 | 1.01 | 0.95 | 0.75 | 0.63 | 0.38 |
| Q5\_1 | 0.99 | 0.94 | 0.75 | 0.63 | 0.37 |
| Q8\_0 | 0.86 | 0.59 | 0.46 | 0.38 | 0.24 |
| Q4\_K | 1.00 | 0.97 | 0.56 | 0.45 | 0.29 |
| Q5\_K | 0.81 | 0.56 | 0.49 | 0.45 | 0.25 |
| Q6\_K | 0.92 | 0.64 | 0.48 | 0.39 | 0.21 |
n=1 and n=512 are unchanged (0.99 to 1.01). A sweep of 9..16 rows gives 0.25 to 0.40 for all ten types.
`llama-batched-bench -npp 512 -ntg 32 -npl 1,2,3,4,8 -c 32768 -pps -kvu`, TG t/s (mean of two interleaved runs each). PP is 317 t/s on both builds.
| B | master | this PR |
| --- | --- | --- |
| 1 | 32.9 | 32.9 |
| 2 | 32.7 | 49.7 |
| 3 | 36.2 | 62.2 |
| 4 | 36.3 | 76.1 |
| 8 | 38.6 | 119.7 |
llama-server with a Qwen3.8-27B DFlash2 Q8\_0 drafter, `-ngl 99 -fa on -c 8192 -np 1 --jinja`, DFlash2 with `--spec-type draft-dflash --spec-draft-n-max 7`. 64 generated tokens, median of 5 requests after one warm-up, mean of two server runs. Decode tok/s:
| mode | prompt | T | master | this PR |
| --- | --- | --- | --- | --- |
| serial | code | 0 | 32.1 | 32.0 |
| serial | code | 1 | 32.1 | 32.0 |
| serial | prose | 0 | 32.1 | 32.0 |
| serial | prose | 1 | 32.1 | 32.0 |
| DFlash2 | code | 0 | 30.2 | 110.0 |
| DFlash2 | code | 1 | 24.3 | 80.9 |
| DFlash2 | prose | 0 | 16.8 | 62.6 |
| DFlash2 | prose | 1 | 13.9 | 48.8 |
At T=0 the DFlash2 replies are the same as the serial replies (first 160 characters). At T=1 the two builds accept the same draft tokens.
`llama-bench -fa 1` gives pp512 317.2 t/s on both builds, and tg128 32.45 t/s on master and 32.36 t/s with this PR (two interleaved runs each). `llama-perplexity` (10 chunks of 512 tokens) gives 5.1775 on both builds with `-ub 512`, and 5.1773 on master and 5.1774 with this PR with `-ub 8`.
Tests:
- `test-backend-ops -b MTL0`: 16806/16806 on the M3 Ultra, and on an Apple M5 with the tensor API on and off (`GGML_METAL_TENSOR_DISABLE=1`).
- `test-metal-graph-optimize` passes on both devices.
- With `-DGGML_SCHED_NO_REALLOC=ON`, as in the Metal CI job: `test-llama-archs -s 1` with 1 to 4 devices gives the same results as master, with no failures, and `test-fusion --check tests/fusion/MTL.csv` passes. The test-fusion graphs have 1 or 32 src1 rows, so the new fusion does not fire there.
The row thresholds come from the M3 Ultra. With the tensor API turned off, an M5 reaches parity one or two rows later for some types. I did not tune for that, because the M5 uses the tensor API by default.
This ports the Metal part of [tetherto#321](https://github.com/tetherto/qvac-fabric-llm.cpp/pull/321) to master.
## Requirements
- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: YES - the kernels, the port to master, the tests, the benchmarks and this description were written with Claude Code (Anthropic).
[@pratiknarola-t](https://github.com/pratiknarola-t)
`metal : few-row MMA mat-mul and batched copies for speculative decoding`
โฆ
`cc470b6`
```
Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding.
- add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel
- use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2)
- fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch
- the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count
- views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group
- the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read
- CONCAT splits long rows across threadgroups when there are few rows
- tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias
```
[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
[Apple Metal](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3A%22Apple%20Metal%22)
https://en.wikipedia.org/wiki/Metal\_(API)
labels
[Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#event-32353932404)
[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
### **[ggml-gh-bot](https://github.com/apps/ggml-gh-bot) Bot** commented [Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#issuecomment-5958107365)
Copy link
Copy Markdown
| |
| --- |
| Hi [@pratiknarola-t](https://github.com/pratiknarola-t), thanks for your contribution! Per our [contribution guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md), the automated PR checker found the following issue(s) that need your attention: - **Large PR**: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs. --- *Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.* |
[@ggerganov](https://github.com/ggerganov)
[ggerganov](https://github.com/ggerganov)
self-assigned this
[Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#event-32354789927)
[ggerganov](https://github.com/ggerganov)
**[ggerganov](https://github.com/ggerganov)**
reviewed
[Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#pullrequestreview-5395590974)
[View reviewed changes](https://github.com/ggml-org/llama.cpp/pull/29869/files/cc470b64e205d7ac09684fe9ae85e4a48d9d2ce6)
Comment thread
[ggml/src/ggml-metal/ggml-metal-fusion.cpp](https://github.com/ggml-org/llama.cpp/pull/29869/files/cc470b64e205d7ac09684fe9ae85e4a48d9d2ce6#diff-8372ceb3f0618f985e825802fe46a025e433bbb594987e7370583d6a40b2a427)
Outdated
Comment on lines
+853
to
+854
| | | |
| --- | --- | --- |
| | | // longest batch first, so ggml\_metal\_fusion\_next checks no shorter batch once one matches |
| | | { GGML\_METAL\_FUSION\_CPY\_BATCH, std::vector<ggml\_op>(16, GGML\_OP\_CPY), {}, true, ggml\_metal\_fusion\_check\_cpy\_batch }, |
### @ggerganov **[ggerganov](https://github.com/ggerganov)** [Oct 2, 2026](https://github.com/ggml-org/llama.cpp/pull/29869#discussion_r4168831964)
Copy link
Copy Markdown
Member
There was a problem hiding this comment.
### Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. [Learn more](https://docs.github.com/articles/managing-disruptive-comments/#hiding-a-comment).
Do you have an estimate of how big is the impact of the