Retrieved article excerpt
Open article ยท Retrieved 2026-09-26T06:23:03.629681+00:00
[ggml-org](https://github.com/ggml-org)
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public
- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
23.7k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
130k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
# ggml-cpu: tiled mul\_mat for k-quants - #27851
#27851
Merged
[ggerganov](https://github.com/ggerganov) merged 40 commits into
[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom
[jbooth:master](https://github.com/jbooth/llama.cpp/tree/master)jbooth/llama.cpp:masterCopy head branch name to clipboard
Sep 26, 2026
Merged
## [ggml-cpu: tiled mul\_mat for k-quants](https://github.com/ggml-org/llama.cpp/pull/27851#top)#27851 [ggerganov](https://github.com/ggerganov) merged 40 commits into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [jbooth:master](https://github.com/jbooth/llama.cpp/tree/master)jbooth/llama.cpp:masterCopy head branch name to clipboard
## Conversation
[@jbooth](https://github.com/jbooth)
### @jbooth **[jbooth](https://github.com/jbooth)** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27851#issue-5274566732) โข edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27851).
Copy link
Copy Markdown
Contributor
## Overview
TL;DR: 3-7x faster CPU mul\_mat using VNNI with IMO minimal complexity
I noticed that the vec\_dot approach currently taken for mul\_mat duplicates the work of unpacking quants a bunch of times. I put together a generic tiled mul\_mat that works with 256x256 windows of int8. We slide a 256-long window across the source tensors, unpacking quants 256x256 at a time, (this amortizes the unpacking work 256x compared to baseline), and then we sweep an optimized 16x16 microkernel across them to build up a 256x256 float32 output tile. If this results in too little work for our number of CPUs, I step down the chunk size to 128, 64 etc so that we have more chunks to saturate available cores.
For tall src1 (prefills), this results in 3-7x performance gain (2x compared to repack) on my x86 machine using the optimized microkernel. The gain goes away as we shrink the rows in src1, reaching break-even at 32 rows and falling behind to 80% of stock performance at pure GEMV. Error rates are on the order of 5e-04 for max single scalar error, 6e-05 for RMSE across the whole tensor compared to standard vec\_dot based mul\_mat (microscopically better than repack).
Integration with the main mul\_mat path is similar to llamafile, we return false if we don't support or if we think it will be unprofitable (< 64 rows).
The majority of the code is portable C++ with only the microkernel and 3 inlined bit-unpacking routines as pieces to fill in for a new architecture. This means we could expand the approach to ARM and other architectures with minimal code addition.
So far, this only supports the k-quant types. I'm reasonably sure we could expand it to the iq-quants as well, but I haven't gone through all of them to ensure their codebooks stay in bounds of the range my unpacked tiles support.
## Additional information
The VNNI path has some additional complexity. VNNI takes one chunk of 4 scalars from src0, broadcasted 16x to a 512 bit vector from one row, and another vector of 4 scalars each from 16 columns of src1 all packed together to execute 16 partial dot products. This requires a somewhat ugly additional repack of the q8\_k quants to enable contiguous reads in the VNNI kernel. I gated it away from everything else so it only exists for VNNI (and other future archs which may require bespoke packing). The normal path for AVX2 and other future archs can ignore it.
We assume that K % 256 == 0, this seems safe given that all of the quant types we support use QK\_K = 256.
Benchmarks all run using 8 threads on an AMD 9950x3D. Used 8 instead of 16 to reduce impact of the rest of the system on the benchmarks (it's my desktop).
```
BENCH 8192x8192 * 8192x8192, min of 5 timings, 8 threads
type std TF repack TF tiled TF repack/std tiled/std max_err(repack) rmse(repack) max_err(tiled) rmse(tiled)
q2_K 1.198 1.856 4.511 1.55 3.77 9.15527e-04 8.55308e-05 5.87463e-04 6.45327e-05
q3_K 0.733 n/a 5.164 n/a 7.04 n/a n/a 3.05176e-04 2.94283e-05
q4_K 1.064 2.474 4.973 2.32 4.67 1.34277e-03 1.08380e-04 5.03540e-04 5.79592e-05
q5_K 0.711 n/a 4.933 n/a 6.94 n/a n/a 5.18799e-04 5.71286e-05
q6_K 1.009 n/a 5.095 n/a 5.05 n/a n/a 4.27246e-04 5.21808e-05
BENCH 4096x4096 * 4096x4096, min of 5 timings, 8 threads
type std TF repack TF tiled TF repack/std tiled/std max_err(repack) rmse(repack) max_err(tiled) rmse(tiled)
q2_K 1.177 1.899 4.174 1.61 3.55 4.57764e-04 4.36373e-05 2.44141e-04 3.39860e-05
q3_K 0.728 n/a 4.635 n/a 6.37 n/a n/a 1.83105e-04 1.58651e-05
q4_K 1.059 2.437 4.507 2.30 4.26 6.71387e-04 5.56072e-05 2.59399e-04 3.19465e-05
q5_K 0.706 n/a 4.490 n/a 6.36 n/a n/a 2.74658e-04 3.11155e-05
q6_K 0.998 n/a 4.573 n/a 4.58 n/a n/a 2.13623e-04 2.77119e-05
BENCH 4096x4096 * 4096x64, min of 5 timings, 8 threads
type std TF repack TF tiled TF repack/std tiled/std max_err(repack) rmse(repack) max_err(tiled) rmse(tiled)
q2_K 0.425 0.529 0.446 1.25 1.05 3.05176e-04 4.35196e-05 2.13623e-04 3.39202e-05
q3_K 0.347 n/a 0.480 n/a 1.38 n/a n/a 1.22070e-04 1.58227e-05
q4_K 0.418 0.460 0.468 1.10 1.12 4.88281e-04 5.54726e-05 1.83105e-04 3.20078e-05
q5_K 0.348 n/a 0.472 n/a 1.35 n/a n/a 2.13623e-04 3.10485e-05
q6_K 0.405 n/a 0.474 n/a 1.17 n/a n/a 1.60217e-04 2.76515e-05
BENCH 4096x4096 * 4096x32, min of 5 timings, 8 threads
type std TF repack TF tiled TF repack/std tiled/std max_err(repack) rmse(repack) max_err(tiled) rmse(tiled)
q2_K 0.254 0.266 0.240 1.05 0.95 3.05176e-04 4.34596e-05 2.13623e-04 3.38719e-05
q3_K 0.234 n/a 0.230 n/a 0.98 n/a n/a 1.22070e-04 1.57968e-05
q4_K 0.263 0.233 0.238 0.89 0.90 4.27246e-04 5.55013e-05 1.83105e-04 3.20242e-05
q5_K 0.242 n/a 0.246 n/a 1.02 n/a n/a 2.13623e-04 3.10980e-05
q6_K 0.246 n/a 0.247 n/a 1.00 n/a n/a 1.60217e-04 2.76375e-05
BENCH 4096x4096 * 4096x16, min of 5 timings, 8 threads
type std TF repack TF tiled TF repack/std tiled/std max_err(repack) rmse(repack) max_err(tiled) rmse(tiled)
q2_K 0.152 0.133 0.124 0.88 0.82 3.05176e-04 4.36595e-05 2.13623e-04 3.39212e-05
q3_K 0.136 n/a 0.125 n/a 0.93 n/a n/a 9.15527e-05 1.57678e-05
q4_K 0.149 0.114 0.127 0.77 0.85 4.27246e-04 5.55512e-05 1.75476e-04 3.20304e-05
q5_K 0.134 n/a 0.127 n/a 0.94 n/a n/a 2.13623e-04 3.10842e-05
q6_K 0.143 n/a 0.126 n/a 0.88 n/a n/a 1.52588e-04 2.74798e-05
BENCH 4096x4096 * 4096x1, min of 5 timings, 8 threads
type std TF repack TF tiled TF repack/std tiled/std max_err(repack) rmse(repack) max_err(tiled) rmse(tiled)
q2_K 0.010 n/a 0.008 n/a 0.83 n/a n/a 1.37329e-04 3.37436e-05
q3_K 0.010 n/a 0.008 n/a 0.78 n/a n/a 9.15527e-05 1.62198e-05
q4_K 0.010 n/a 0.008 n/a 0.78 n/a n/a 1.52946e-04 3.14343e-05
q5_K 0.010 n/a 0.008 n/a 0.84 n/a n/a 2.13623e-04 3.14237e-05
q6_K 0.010 n/a 0.008 n/a 0.79 n/a n/a 1.22070e-04 2.72714e-05
```
## Requirements
- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
AI usage disclosure: After a manual first draft getting my outer loops right, I used AI to write unpackers for the various quant types and the microkernel, both of which have very tight contracts. The original idea and first draft proof of concept were mine. I did spend time understanding the microkernel afterwards and edited the comments into my own words.
[@jbooth](https://github.com/jbooth)
[jbooth](https://github.com/jbooth)
requested a review
from [ggerganov](https://github.com/ggerganov)
as a [code owner](https://github.com/ggml-org/llama.cpp/blob/ca3d5a3e10d53f7ea672cb9b6178faca3e2807bc/CODEOWNERS#L59)
[August 28, 2026 04:01](https://github.com/ggml-org/llama.cpp/pull/27851#event-30144287036)
[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
labels
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27851#event-30144310561)
[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)
### **[ggml-gh-bot](https://github.com/apps/ggml-gh-bot) Bot** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27851#issuecomment-5448260996)
Copy link
Copy Markdown
| |
| --- |
| Hi [@jbooth](https://github.com/jbooth), thanks for your contribution! Per our [contribution guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md), the automated PR checker found the following issue(s) that need your attention: - **Large PR**: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs. --- *Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.* |
[@taronaeo](https://github.com/taronaeo)
### **[taronaeo](https://github.com/taronaeo)** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/2785