2026-10-11 17:15 UTC

Hugging Face claims its new Transformers GGUF integration runs packed Qwen3.5 weights on Apple Silicon near llama.cpp throughput using ggml kernels, enabling quantized local inference and evaluation through standard PyTorch and Transformers interfaces.

state: watchingheat: highuncertainty: mediumconvergesscott: mediumlocal-inference model-interoperability inference-economicsHugging FaceMarc SunArthur ZuckerLysandre
Surfaced 2026-09-25T14:32:10Z — GGUFs in transformers natively! — The announcement's own run has completed: the three velocity-spike triggers were the Reddit post's original ramp (peak ~44 pts/h), already priced at medium — current rate is now 0/h, the HN resubmits sit at single-digit points with zero comments, and Reddit has decayed to long-tail (+5 pts, +3 comments over ~a day). The case's meaning shifts from fresh first-party claim to established first-party capability awaiting independent validation: seed→watching, heat cools to low. The magnitude-valve spread reading is one announcement mirrored across platforms, not expanding periphery — no independent benchmarks, implementations, or new communities have appeared. Only substantive new signal is a comment-claimed, unverified PoC of LoRA training over packed GGUF via Unsloth/Axolotl.

What is this?

Hugging Face has shipped native GGUF inference in the Transformers library (on main / v5.17 docs), authored by members of the Transformers team (Marc Sun, Arthur Zucker, Lysandre are named in the case). Instead of dequantizing GGUF checkpoints at load time as before, Transformers now fetches llama.cpp ggml kernels from the Hub and runs matmuls directly on packed quantized weight blocks on Metal/MPS (Apple Silicon), including a ggml flash-attention kernel for decode and prefill β€” with throughput reported near llama.cpp. The fast compressed path is currently limited to Qwen3.5 on MPS; other architectures (Llama, Mistral, Qwen2, Phi3, etc.) or devices fall back to a legacy loader that dequantizes, and GGUF models are also servable via `transformers serve` using a `<repo>:<file>.gguf` naming scheme.

Why it matters to Scott

Hugging Face independently built what Scott's own concepts treat as the right architecture: hardware-aware quantized inference where packed weights and ggml kernels become explicit runtime policy under a standard interface, and quantized models that can be evaluated directly through the PyTorch/Transformers harness rather than dequantized first β€” answering the exact 'can quants be evaluated with standard interfaces' question his model-plus-harness benchmark unit poses. It bears directly on his active local-inference stack choices (Ollama/llama.cpp vs Transformers vs MLX on his Mac mini and gamepc box) and lands Qwen-first, his most-used open model family β€” though it's a first-party throughput claim needing independent validation, and the fast path is currently limited to Qwen3.5 on MPS while his GPU box is CUDA.
ip:concept.model-plus-harness-benchmark-unitdev:concept.hardware-aware-local-inferencedev:technology.ollamaradar:concept.ggufradar:concept.quantizationradar:concept.local-inferenceradar:concept.llama-cppradar:concept.apple-siliconradar:person.qwen
queries asked of Scott's wikis
  • local inference stack choices: llama.cpp vs transformers vs MLX on Apple Silicon
  • quantized model evaluation harness β€” can quants be evaluated with standard PyTorch interfaces?
  • agent harness memory/runtime requirements for on-device models
  • open-weights local model strategy and inference economics positions
  • Qwen model family usage in his projects and tools
  • tooling abstractions over inference backends (interoperability layers)

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 482h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-23 05:57⭐ origin directly observedGGUFs in transformers natively!
Disastrous-Work-1632 on r/LocalLLaMA
β€”
09-21 14:00first on blog (echo) Β· published Β· +-40.0hAnnounces efficient GGUF inference in Transformers main using ggml kernels, initially targeting Qwen3.5 on Apple Silicon, and reports throug
Hugging Face
β€”
09-22 10:41first on hacker news Β· published Β· +-19.3hTransformers now runs llama.cpp quants
vertigoruntime
β€”
09-22 10:41amplified on hacker newshn.story.49799125
vertigoruntime
peak 2 Β· 0 comments Β· 1% of case engagement
09-23 05:57amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wnxm0r
Disastrous-Work-1632
peak 289 Β· 39 comments Β· 97% of case engagement
09-23 06:18amplified on hacker newshn.story.49812339
mmis1000
peak 4 Β· 0 comments Β· 2% of case engagement
09-22 11:20our radar first saw it Β· +-18.6hdiscovery anchor: hn.story.49799125β€”
09-25 14:32reached heat=high Β· +56.6h Β· via queue+ledgerβ€”β€”
pace: p80 vs 1032 stories at the 336h mark (now 482h old) β€” ahead of microsoft-copilot-consumer-retreat (1.0x), behind qubes-copy-vm-backchannel-rce (1.0x)

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTransformers now runs llama.cpp quants
Retrieved article excerpt

Open article Β· Retrieved 2026-09-23T12:21:57.520418+00:00

[Back to Articles](https://huggingface.co/blog)

# Transformers now runs llama.cpp quants

Published
September 22, 2026

[Update on GitHub](https://github.com/huggingface/blog/blob/main/transformers-llama-cpp-quants.md)

[Upvote

43](https://huggingface.co/login?next=%2Fblog%2Ftransformers-llama-cpp-quants)  







 - +37

[Marc Sun's avatar](https://huggingface.co/marcsun13) 

[Marc Sun

marcsun13 

Follow](https://huggingface.co/marcsun13)

[Arthur Zucker's avatar](https://huggingface.co/ArthurZ) 

[Arthur Zucker

ArthurZ 

Follow](https://huggingface.co/ArthurZ)

[Lysandre's avatar](https://huggingface.co/lysandre) 

[Lysandre

lysandre 

Follow](https://huggingface.co/lysandre)

 

**We're adding support for running GGUF models efficiently in transformers**, so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with `from_pretrained`, and start generating on your own machine.

[

](https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/transformers-llama-cpp-quants/transformers-gguf.mp4)

Running AI models on your laptop has become much easier, and [llama.cpp](https://github.com/ggml-org/llama.cpp) has been a big part of that. Its inference engine powers local AI tools such as Ollama, LM Studio, and Jan. Alongside projects like [MLX](https://github.com/ml-explore/mlx), it has helped make local inference a practical option for everyday use.

*A recent example of what local AI can feel like:*

> This is where we are right now. And i’m not gonna lie it feels pretty magical πŸ§™β€β™€οΈ  
>   
> Qwen3.6 27B running inside of Pi coding agent via Llama.cpp on the MacBook Pro  
>   
> For non-trivial tasks on the [@huggingface](https://x.com/huggingface?ref_src=twsrc%5Etfw) codebases, this feels very, very close to hitting the latest Opus in Claude… [pic.twitter.com/lsIxLoUneU](https://t.co/lsIxLoUneU)
>
> β€” Julien Chaumond (@julien\_c) [April 24, 2026](https://x.com/julien_c/status/2047647522173104145?ref_src=twsrc%5Etfw)

**GGUF**, developed by the llama.cpp team, is a widely used format for local inference. The team also shares quantized checkpoints under [ggml-org on the Hub](https://huggingface.co/ggml-org). Publishers such as [Unsloth](https://huggingface.co/unsloth), [LM Studio Community](https://huggingface.co/lmstudio-community), and [bartowski](https://huggingface.co/bartowski) also provide ready-to-use GGUF checkpoints in a range of quantizations, so users can pick the version that fits their machine. GGUF models have been downloaded millions of times.

We want to make it easier to run these models locally with transformers, too. Compatibility is only useful if the model is pleasant to run. To bring performance close to llama.cpp, we're reusing its underlying ggml kernels through the [`kernels`](https://huggingface.co/docs/kernels/index) library, and reducing overhead in `generate`. Our initial focus is local inference on Apple Silicon, starting with the Qwen3.5 architecture.

## What is the GGUF file format?

[GGUF](https://github.com/ggml-org/ggml/blob/master/docs/gguf.md) packages model weights and metadata, including tokenizer information and an optional chat template, in one file. It supports different quantization levels, letting you trade some precision for a smaller memory footprint. Variants such as `Q4_K_M` mix tensor precisions, using mostly 4-bit weights while keeping sensitive tensors at higher precision.

Here's how quantization changes the file size of [Unsloth's Qwen3.5-4B](https://huggingface.co/unsloth/Qwen3.5-4B-GGUF/tree/main):

| GGUF variant | File size | Tradeoff |
| --- | --- | --- |
| `BF16` | 8.42 GB | Unquantized reference |
| `Q6_K` | 3.53 GB | More precision than the smaller variants |
| `Q5_K_M` | 3.14 GB | A middle ground between size and precision |
| `Q4_K_M` | 2.74 GB | A practical starting point for local inference |

We suggest starting with `Q4_K_M`, then trying `Q5_K_M` or `Q6_K` if you have more memory available. More aggressive quantization can help larger models fit, but the quality tradeoff depends on the model and the task. Evaluate it on the work you actually want the model to do. The [Hub's GGUF documentation](https://huggingface.co/docs/hub/gguf#quantization-types) describes the available quantization types.

## Load GGUF with transformers

To get started, you need:

- **An Apple Silicon Mac**.
- **A PyTorch version supported by the published [ggml-quantization kernel builds](https://huggingface.co/kernels/ggml-org/ggml-quantization)**, usually the two latest PyTorch releases.
- **The latest version of transformers (main for now, until the next release) and a compatible version of `kernels`**.

```
pip install -U "git+https://github.com/huggingface/transformers.git" kernels
```

To load a GGUF model, pass its Hub `model_id` and filename as `gguf_file` to `from_pretrained`.

No extra configuration is needed: when the weights stay packed on Metal, transformers automatically loads the compatible ggml/Metal layer kernels and uses `ggml-org/ggml-attn` as the attention implementation. If that kernel cannot be fetched, the model falls back to `"sdpa"` with a warning, and you can always force `"sdpa"` by passing `attn_implementation="sdpa"` explicitly. See the [GGUF documentation](https://huggingface.co/docs/transformers/main/en/quantization/gguf) for more loading options.

```
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"

tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    gguf_file=filename
)
```

That is the only GGUF-specific step. Everything after it is the standard transformers API:

```
messages = [{"role": "user", "content": "Explain why the sky is blue in a few sentences."}]
inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=256)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

> Without a compatible quantization kernel, the loader falls back to dequantizing the model and uses more memory.

## Serve GGUF with your preferred interface

You can also use the same checkpoint with [`transformers serve`](https://huggingface.co/docs/transformers/main/en/serve-cli/serving), which exposes an OpenAI-compatible API:

```
pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels

transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"
```

The model argument uses `<model_id>:<filename>.gguf`: before the colon is the Hub repository (`unsloth/Qwen3.5-4B-GGUF`), and after it is the file to load (`Qwen3.5-4B-Q4_K_M.gguf`). This selects a specific quantization from a repository that may contain several.

For models whose chat template supports thinking, add `--reasoning off` to skip it or `--reasoning on` to enable it. The default, `--reasoning auto`, follows the chat template’s default. See the [reasoning options](https://huggingface.co/docs/transformers/main/en/serve-cli/serving#enable-reasoning-on-the-server) for details.

You can connect a client such as [Jan](https://www.jan.ai/docs/desktop/remote-models/custom-endpoint) or [Pi](https://pi.dev) by adding a custom OpenAI-compatible provider with these settings:

| Setting | Value |
| --- | --- |
| Base URL | `http://localhost:8000/v1` |
| Model ID | `unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf` |

transformers runs the model on your Mac, while the client provides the conversation interface. The same endpoint can be used by other clients that support this API.

## Benchmarking against llama.cpp

Our reference for local inference performance is llama.cpp. The comparison below focuses on three GGUF checkpoints: a small dense model, a larger dense model, and a mixture-of-experts model.

The llama.cpp column comes from the [`llama-bench`](https://github.com/ggml-org/llama.cpp/tree/master/tools/llama-bench) tool (build `5f55650a7`, release b10200, Metal backend from ggml 0.18.0), run as `llama-bench -m <file> -p 0 -n 128 -r 3`, which reports `tg128`: the token-generation rate over 128 decoded tokens, averaged across three repetitions, with prompt processing excluded. The transformers column is `generate` producing the same 128 tokens from a 12-token prompt, best of three warmed runs, and it includes prefill.

Measured on a MacBook Pro M2 Max, 32 GB unified memory, macOS 26.6, PyTorch 2.12.1, kernels 0.17.0,
plugged in.

The benchmark script

```
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id, filename = "unsloth/Qwen3.5-4B-GGUF", "Qwen3.5-4B-Q4_K_M.gguf"

model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
inputs = tokenizer("The capital of France is Paris. The capital of Germany is", return_tensors="pt")
inputs = inputs.to(model.device)


with torch.inference_mode():
    model.generate(**inputs, max_new_tokens=8, min_new_tokens=8, do_sample=False)  # warm up
    torch.mps.synchronize()
    for _ in range(3):
        time.sleep(90)  # let the machine cool: back-to-back runs decay by 10% or more
        start = time.perf_counter()
        model.generate(**inputs, max_new_tokens=128, min_new_tokens=128, do_sample=False)
        torch.mps.synchronize()
        print(f"{128 / (time.perf_counter() - start):.1f} tok/s")
```

For the other column:

```
llama-bench -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M -p 0 -n 128 -r 3
```

[GGUF generation throughput compared with llama.cpp](https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/transformers-llama-cpp-quants/benchmark-comparison.svg)

Transformers is close to llama.cpp across all three checkpoints. The chart uses the same measurements described above; it does not imply identical benchmark conditions, since the Transformers measurement includes prefill while `llama-bench` reports decode-only throughput.

## transformers and llama.cpp

When [GGML and llama.cpp joined Hugging Face](https://huggingface.co/blog/ggml-joins-hf), we described their complementary roles: llama.cpp provides a foundation for local inference, while transformers provides a foundation for model definition. GGUF support brings those two closer together.

**llama.cpp remains our recommended engine when your priority is efficient local inference.** Its dedicated runtime, memory management, and broad hardware support are built around that goal. This integration gives developers a convenient way to work with the same GGUF checkpoints inside transformers:

- **Experiment with GGUF in Python and PyTorch.** Inspect intermediate activations with hooks, modify a model's forward pass, or prototype custom layers using familiar PyTorch tools.
- **Evaluate GGUF models.** Use your existing transformers evaluation workflows to measure the quality of quantized checkpoints.
- **Validate GGUF conversions.** For us as developers, loading the original checkpoint and its GGUF conversion in transformers makes it easier to check that the weights were converted correctly, accounting for quantization error.
- **Try new decoding ideas.** Use custom logits processors and stopping criteria with `generate`, or write your own generation loop in Python.
- **Fine-tune from a GGUF checkpoint.** Dequantize the weights and continue with a standard transformers training workflow.

For that last case, use `GgufConfig(dequantize=True)`:

```
import torch
from transformers import AutoModelForCausalLM, GgufConfig

model = AutoModelForCausalLM.from_pretrained(
    "unsloth/Qwen3.5-4B-GGUF",
    gguf_file="Qwen3.5-4B-Q4_K_M.gguf",
    quantization_config=GgufConfig(dequantize=True),
    dtype=
vertigoruntime20
🟧 echo.blogAnnounces efficient GGUF inference in Transformers main using ggml kernels, initially targeting Qwen3.5 on Apple Silicon, and reports througHugging Faceβ€”β€”
🟧 hnTransformers now runs llama.cpp quantsmmis100040
🟠 reddit ⭐GGUFs in transformers natively!
LocalLLaMA
Disastrous-Work-163228939

Interpretation history

Decision trace