2026-10-11 16:38 UTC

jbooth's merged llama.cpp PR #27851 claims a tiled VNNI mul_mat path accelerates CPU k-quant prompt processing 3-7x on x86 (about 2x over repack) with microscopic error, and confirmation of the gains on broader hardware plus shipping in releases would make tiled CPU prefill a standard optimization for CPU-served local inference.

state: watchingheat: lowuncertainty: mediumconvergesscott: highllama-cpp cpu-inference quantization local-inferencejboothggerganov

What is this?

jbooth's merged PR #27851 in ggml-org/llama.cpp adds a tiled 256ร—256 int8 mul_mat kernel to the CPU backend that uses x86 VNNI instructions and amortizes k-quant unpacking across the tile, claiming 3โ€“7x faster k-quant prompt processing (~2x over the existing repack route) at negligible numerical error. The web snippets contain no independent coverage of this specific PR โ€” the PR page and its TL;DR are the primary evidence โ€” but they confirm the terrain: mul_mat is llama.cpp's dominant CPU hotspot, and k-quant prefill acceleration is landing across other backends in parallel (a WebGPU k-quant prefill change shows up to 3.78x on M2 Pro; a ROCm prompt-processing PR claims ~15%). Tiled hand-written CPU matmul kernels also have direct precedent for outsized prefill gains โ€” llamafile's 84 tiled kernels delivered 30โ€“500% faster prompt eval on F16/Q8_0 with the same tiling/cache-locality ideas. Whether jbooth's specific gains hold on broader x86 hardware and ship in a tagged release is not settled by the supplied material.

Why it matters to Scott

First-party llama.cpp has merged โ€” in the engine under Scott's Ollama/gamepc serving stack โ€” the exact priority his hardware-aware-local-inference notes already hold: k-quant prompt-processing throughput as the local serving bottleneck, now attacked with a claimed 3-7x tiled VNNI kernel, upgrading the radar's open VNNI/AVX2 CPU-prefill thread (radar:llama-cpp-x86-vnni-q2-speedup and the AVX2 IQ pair) from contributor claim to merged artifact. It also hands him a live instance of deterministic-verification-before-assertion: one fixed prefill benchmark on gamepc before/after the next tagged release cheaply confirms or kills the 3-7x figure on hardware broader than the author's, and if it holds it shifts his CPU-only serving and k-quant selection calculus โ€” a dated-receipts opportunity to publish that his notes named the bottleneck the ecosystem's first-party arm is now shipping against.
dev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcip:concept.deterministic-verification-before-assertionradar:concept.llama-cppradar:llama-cpp-x86-vnni-q2-speedupradar:llama-cpp-avx2-iq-prompt-speedupradar:llama-cpp-avx2-iq-batch-speedupradar:concept.cpu-inferenceradar:concept.inference-kernelsradar:concept.quantization
queries asked of Scott's wikis
  • CPU-only local inference: serving when GPU offload isn't an option
  • prefill / prompt processing throughput as the local serving bottleneck
  • k-quant GGUF selection: quality vs speed tradeoffs
  • llama.cpp usage in local model tooling or RAG pipeline projects
  • long-context document ingestion throughput on local hardware
  • hand-tuned SIMD kernel optimization claims and how to verify them

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1082h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

08-27 14:00โญ origin echo-reconstructed"TL;DR: 3-7x faster CPU mul_mat using VNNI with IMO minimal complexity" โ€” tiled 256x256 int8 mul_mat amortizing quant unpacking, benchmarks
jbooth (PR opened Aug 28, 2026; merged by ggerganov Sep 26, 2026) on github (echo) ยท attributed from reddit.post.1wqj4zg
โ€”
09-26 06:18first on r/LocalLLaMA ยท published ยท +712.3hggml-cpu: tiled mul_mat for k-quants by jbooth ยท Pull Request #27851 ยท ggml-org/llama.cpp
jacek2023
โ€”
10-02 21:38first on hacker news ยท published ยท +871.6hLLaMA Now Goes Faster on CPUs
luu
โ€”
09-26 06:18amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wqj4zg
jacek2023
peak 49 ยท 20 comments ยท 85% of case engagement
10-02 21:38amplified on hacker newshn.story.49938916
luu
peak 4 ยท 3 comments ยท 15% of case engagement
09-26 06:20our radar first saw it ยท +712.3hdiscovery anchor: reddit.post.1wqj4zgโ€”

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditggml-cpu: tiled mul_mat for k-quants by jbooth ยท Pull Request #27851 ยท ggml-org/llama.cpp
LocalLLaMA
Retrieved article excerpt

Open article ยท Retrieved 2026-09-26T06:23:03.629681+00:00

[ggml-org](https://github.com/ggml-org) 
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public

- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
  23.7k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
   130k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)

# ggml-cpu: tiled mul\_mat for k-quants - #27851

#27851

Merged

[ggerganov](https://github.com/ggerganov) merged 40 commits into

[ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom 

[jbooth:master](https://github.com/jbooth/llama.cpp/tree/master)jbooth/llama.cpp:masterCopy head branch name to clipboard

Sep 26, 2026

Merged

## [ggml-cpu: tiled mul\_mat for k-quants](https://github.com/ggml-org/llama.cpp/pull/27851#top)#27851 [ggerganov](https://github.com/ggerganov) merged 40 commits into [ggml-org:master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [jbooth:master](https://github.com/jbooth/llama.cpp/tree/master)jbooth/llama.cpp:masterCopy head branch name to clipboard

## Conversation

[@jbooth](https://github.com/jbooth)

### @jbooth **[jbooth](https://github.com/jbooth)** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27851#issue-5274566732) โ€ข edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/27851).

Copy link
 

Copy Markdown

Contributor

## Overview

TL;DR: 3-7x faster CPU mul\_mat using VNNI with IMO minimal complexity

I noticed that the vec\_dot approach currently taken for mul\_mat duplicates the work of unpacking quants a bunch of times. I put together a generic tiled mul\_mat that works with 256x256 windows of int8. We slide a 256-long window across the source tensors, unpacking quants 256x256 at a time, (this amortizes the unpacking work 256x compared to baseline), and then we sweep an optimized 16x16 microkernel across them to build up a 256x256 float32 output tile. If this results in too little work for our number of CPUs, I step down the chunk size to 128, 64 etc so that we have more chunks to saturate available cores.

For tall src1 (prefills), this results in 3-7x performance gain (2x compared to repack) on my x86 machine using the optimized microkernel. The gain goes away as we shrink the rows in src1, reaching break-even at 32 rows and falling behind to 80% of stock performance at pure GEMV. Error rates are on the order of 5e-04 for max single scalar error, 6e-05 for RMSE across the whole tensor compared to standard vec\_dot based mul\_mat (microscopically better than repack).

Integration with the main mul\_mat path is similar to llamafile, we return false if we don't support or if we think it will be unprofitable (< 64 rows).

The majority of the code is portable C++ with only the microkernel and 3 inlined bit-unpacking routines as pieces to fill in for a new architecture. This means we could expand the approach to ARM and other architectures with minimal code addition.

So far, this only supports the k-quant types. I'm reasonably sure we could expand it to the iq-quants as well, but I haven't gone through all of them to ensure their codebooks stay in bounds of the range my unpacked tiles support.

## Additional information

The VNNI path has some additional complexity. VNNI takes one chunk of 4 scalars from src0, broadcasted 16x to a 512 bit vector from one row, and another vector of 4 scalars each from 16 columns of src1 all packed together to execute 16 partial dot products. This requires a somewhat ugly additional repack of the q8\_k quants to enable contiguous reads in the VNNI kernel. I gated it away from everything else so it only exists for VNNI (and other future archs which may require bespoke packing). The normal path for AVX2 and other future archs can ignore it.

We assume that K % 256 == 0, this seems safe given that all of the quant types we support use QK\_K = 256.

Benchmarks all run using 8 threads on an AMD 9950x3D. Used 8 instead of 16 to reduce impact of the rest of the system on the benchmarks (it's my desktop).

```
BENCH 8192x8192 * 8192x8192, min of 5 timings, 8 threads
type         std TF    repack TF     tiled TF  repack/std   tiled/std   max_err(repack)      rmse(repack)    max_err(tiled)       rmse(tiled)
q2_K          1.198        1.856        4.511        1.55        3.77       9.15527e-04       8.55308e-05       5.87463e-04       6.45327e-05
q3_K          0.733          n/a        5.164         n/a        7.04               n/a               n/a       3.05176e-04       2.94283e-05
q4_K          1.064        2.474        4.973        2.32        4.67       1.34277e-03       1.08380e-04       5.03540e-04       5.79592e-05
q5_K          0.711          n/a        4.933         n/a        6.94               n/a               n/a       5.18799e-04       5.71286e-05
q6_K          1.009          n/a        5.095         n/a        5.05               n/a               n/a       4.27246e-04       5.21808e-05

BENCH 4096x4096 * 4096x4096, min of 5 timings, 8 threads
type         std TF    repack TF     tiled TF  repack/std   tiled/std   max_err(repack)      rmse(repack)    max_err(tiled)       rmse(tiled)
q2_K          1.177        1.899        4.174        1.61        3.55       4.57764e-04       4.36373e-05       2.44141e-04       3.39860e-05
q3_K          0.728          n/a        4.635         n/a        6.37               n/a               n/a       1.83105e-04       1.58651e-05
q4_K          1.059        2.437        4.507        2.30        4.26       6.71387e-04       5.56072e-05       2.59399e-04       3.19465e-05
q5_K          0.706          n/a        4.490         n/a        6.36               n/a               n/a       2.74658e-04       3.11155e-05
q6_K          0.998          n/a        4.573         n/a        4.58               n/a               n/a       2.13623e-04       2.77119e-05

BENCH 4096x4096 * 4096x64, min of 5 timings, 8 threads
type         std TF    repack TF     tiled TF  repack/std   tiled/std   max_err(repack)      rmse(repack)    max_err(tiled)       rmse(tiled)
q2_K          0.425        0.529        0.446        1.25        1.05       3.05176e-04       4.35196e-05       2.13623e-04       3.39202e-05
q3_K          0.347          n/a        0.480         n/a        1.38               n/a               n/a       1.22070e-04       1.58227e-05
q4_K          0.418        0.460        0.468        1.10        1.12       4.88281e-04       5.54726e-05       1.83105e-04       3.20078e-05
q5_K          0.348          n/a        0.472         n/a        1.35               n/a               n/a       2.13623e-04       3.10485e-05
q6_K          0.405          n/a        0.474         n/a        1.17               n/a               n/a       1.60217e-04       2.76515e-05

BENCH 4096x4096 * 4096x32, min of 5 timings, 8 threads
type         std TF    repack TF     tiled TF  repack/std   tiled/std   max_err(repack)      rmse(repack)    max_err(tiled)       rmse(tiled)
q2_K          0.254        0.266        0.240        1.05        0.95       3.05176e-04       4.34596e-05       2.13623e-04       3.38719e-05
q3_K          0.234          n/a        0.230         n/a        0.98               n/a               n/a       1.22070e-04       1.57968e-05
q4_K          0.263        0.233        0.238        0.89        0.90       4.27246e-04       5.55013e-05       1.83105e-04       3.20242e-05
q5_K          0.242          n/a        0.246         n/a        1.02               n/a               n/a       2.13623e-04       3.10980e-05
q6_K          0.246          n/a        0.247         n/a        1.00               n/a               n/a       1.60217e-04       2.76375e-05

BENCH 4096x4096 * 4096x16, min of 5 timings, 8 threads
type         std TF    repack TF     tiled TF  repack/std   tiled/std   max_err(repack)      rmse(repack)    max_err(tiled)       rmse(tiled)
q2_K          0.152        0.133        0.124        0.88        0.82       3.05176e-04       4.36595e-05       2.13623e-04       3.39212e-05
q3_K          0.136          n/a        0.125         n/a        0.93               n/a               n/a       9.15527e-05       1.57678e-05
q4_K          0.149        0.114        0.127        0.77        0.85       4.27246e-04       5.55512e-05       1.75476e-04       3.20304e-05
q5_K          0.134          n/a        0.127         n/a        0.94               n/a               n/a       2.13623e-04       3.10842e-05
q6_K          0.143          n/a        0.126         n/a        0.88               n/a               n/a       1.52588e-04       2.74798e-05

BENCH 4096x4096 * 4096x1, min of 5 timings, 8 threads
type         std TF    repack TF     tiled TF  repack/std   tiled/std   max_err(repack)      rmse(repack)    max_err(tiled)       rmse(tiled)
q2_K          0.010          n/a        0.008         n/a        0.83               n/a               n/a       1.37329e-04       3.37436e-05
q3_K          0.010          n/a        0.008         n/a        0.78               n/a               n/a       9.15527e-05       1.62198e-05
q4_K          0.010          n/a        0.008         n/a        0.78               n/a               n/a       1.52946e-04       3.14343e-05
q5_K          0.010          n/a        0.008         n/a        0.84               n/a               n/a       2.13623e-04       3.14237e-05
q6_K          0.010          n/a        0.008         n/a        0.79               n/a               n/a       1.22070e-04       2.72714e-05
```

## Requirements

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)

AI usage disclosure: After a manual first draft getting my outer loops right, I used AI to write unpackers for the various quant types and the microkernel, both of which have very tight contracts. The original idea and first draft proof of concept were mine. I did spend time understanding the microkernel afterwards and edited the comments into my own words.

[@jbooth](https://github.com/jbooth)

[jbooth](https://github.com/jbooth)
requested a review
from [ggerganov](https://github.com/ggerganov)
as a [code owner](https://github.com/ggml-org/llama.cpp/blob/ca3d5a3e10d53f7ea672cb9b6178faca3e2807bc/CODEOWNERS#L59)
[August 28, 2026 04:01](https://github.com/ggml-org/llama.cpp/pull/27851#event-30144287036)

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
labels
[Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27851#event-30144310561)

[@ggml-gh-bot](https://github.com/apps/ggml-gh-bot)

### **[ggml-gh-bot](https://github.com/apps/ggml-gh-bot) Bot** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/27851#issuecomment-5448260996)

Copy link
 

Copy Markdown

|  |
| --- |
| Hi [@jbooth](https://github.com/jbooth), thanks for your contribution!  Per our [contribution guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md), the automated PR checker found the following issue(s) that need your attention:   - **Large PR**: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.  ---   *Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.* |

[@taronaeo](https://github.com/taronaeo)

### **[taronaeo](https://github.com/taronaeo)** commented [Aug 28, 2026](https://github.com/ggml-org/llama.cpp/pull/2785
jacek20234920
๐ŸŸง echo.github โญ"TL;DR: 3-7x faster CPU mul_mat using VNNI with IMO minimal complexity" โ€” tiled 256x256 int8 mul_mat amortizing quant unpacking, benchmarks jbooth (PR opened Aug 28, 2026; merged by ggerganov Sep 26, 2026)โ€”โ€”
๐ŸŸง hnLLaMA Now Goes Faster on CPUsluu43

Interpretation history

Decision trace