2026-10-11 16:38 UTC

Luigi reports that Vulkan runs a quantized Qwen3.6-35B MoE at 32.75 generation tokens per second on a Panther Lake laptop, outperforming CPU and the tested SYCL configuration while OpenVINO fails, making backend choice consequential for this local-inference setup.

state: watchingheat: lowuncertainty: mediumknownscott: lowlocal-inference llama-cpp intelLuigi

What is this?

Luigi, writing on grigio.org, benchmarked all four llama.cpp compute backends (Vulkan, SYCL, OpenVINO, CPU) on an Intel Panther Lake laptop — per the case, a Core Ultra X7 358H with Arc iGPU and ~32 GiB shared LPDDR5 — running the 22 GiB Q4_K_M quant of Alibaba's Qwen3.6-35B-A3B MoE (35B total / 3B active params). Vulkan won generation at 32.75 tok/s versus SYCL's 30.62 and CPU's 19.44 (with a far larger ~6x prompt-processing edge), while OpenVINO failed to run at all on this model; the author attributes Vulkan's token-generation edge to more efficient memory access in the Mesa ANV driver, noting generation is memory-bandwidth-bound on unified memory. The surrounding snippet set shows Qwen3.6-35B-A3B is currently a hot local-inference target across very different hardware (24GB cards, 8GB and even 6GB GPUs via MoE CPU-offload, Macs via LM Studio), and independent reporting (The Agent Times) confirms that backend winners and techniques like MTP shift per architecture and context length — MTP helps Qwen3.6-27B on Strix Halo but collapses throughput for the 35B on Apple Metal. The snippet of the grigio.org post is truncated before its full Vulkan-vs-SYCL reasoning, and nothing in the snippets supplies third-party replication of Luigi's specific numbers or covers the case's noted SYCL Level Zero / F16-KV caveats.

Why it matters to Scott

known — Scott's dev:concept.hardware-aware-local-inference page already holds the exact position this case re-instances (backend choice, KV precision, and memory pressure as explicit runtime policy), on Intel iGPU hardware his stack doesn't run (gamepc is CUDA, the Mac mini is MLX), and the Vulkan-over-SYCL/OpenVINO pattern is already established radar lineage; with the follow-up corrected to same-author, the numbers remain single-source and add nothing new on either side. The one forward-looking datum — OpenVINO failing on a MoE+Mamba2+MTP architecture while its 2026.4 release expands coverage — bears on the radar's OpenVINO-coverage cases, not on anything Scott builds or argues.
dev:concept.hardware-aware-local-inferenceradar:strix-point-qwen36-local-inferenceradar:llama-cpp-sycl-battlemage-kv-decoderadar:lemonade-vulkan-rocm-dropradar:openvino-2026-4-model-coverageradar:xeboostlm-intel-native-inferenceradar:concept.memory-bandwidth
queries asked of Scott's wikis
  • hardware-aware local inference backend selection runtime policy
  • Intel Arc iGPU SYCL OpenVINO local model serving
  • MoE 35B model local inference laptop UMA feasibility
  • llama.cpp backend benchmark notes Vulkan SYCL comparison
  • UMA shared memory bandwidth token generation bottleneck
  • open-weights local stack hardware guidance consumer laptops

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 626h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedReports Vulkan at 624.53 prompt-processing and 32.75 generation tokens per second, versus SYCL at 609.36 and 30.62 and CPU at 101.57 and 19.
Luigi on blog (echo) · attributed from hn.story.49726446
—
09-16 13:08first on hacker news · published · +23.1hBenchmark Llamacpp on Intel Panther Lake: Vulkan vs. SYCL vs. OpenVINO vs. CPU
grigio
—
09-16 13:08amplified on hacker newshn.story.49726446
grigio
peak 1 · 0 comments · 21% of case engagement
09-28 12:16amplified on hacker news 👑hn.story.49876817
grigio
peak 4 · 0 comments · 79% of case engagement
09-16 13:21our radar first saw it · +23.4hdiscovery anchor: hn.story.49726446—
pace: p39 vs 1032 stories at the 336h mark (now 626h old) — ahead of agentsec-static-config-auditing (1.2x), behind anthropic-meta-lawsuit (0.8x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBenchmark Llamacpp on Intel Panther Lake: Vulkan vs. SYCL vs. OpenVINO vs. CPU
Retrieved article excerpt

Open article · Retrieved 2026-09-16T13:22:48.865692+00:00

[llama](https://grigio.org/tag/llama/)

# Benchmarking llama.cpp Backends on Intel Panther Lake: Vulkan vs SYCL vs OpenVINO vs CPU

- [Luigi](https://grigio.org/author/luigi/)

#### [Luigi](https://grigio.org/author/luigi/)

16 Sep 2026
• 5 min read

[Share](https://grigio.org/benchmarking-llama-cpp-backends-on-intel-panther-lake-vulkan-vs-sycl-vs-openvino-vs-cpu/#/share)

I benchmarked the Qwen3.6-35B-A3B-MTP-UD-Q4\_K\_M.gguf model (a 35B Mixture-of-Experts model with 3B active parameters, 22 GiB quantized) across all available llama.cpp compute backends on an Intel Panther Lake laptop. The goal: find out which backend delivers the best inference performance for large MoE models on integrated Intel GPU hardware.

## Hardware

| Component | Spec |
| --- | --- |
| **CPU** | Intel Core Ultra X7 358H — 16c/16t, Panther Lake, max 4.8 GHz, AVX2 + AVX-VNNI (no AVX-512) |
| **GPU** | Intel Arc B390 (integrated, Panther Lake, 12 Xe3 EUs) |
| **RAM** | ~32 GiB LPDDR5 (UMA — unified memory with GPU) + zram swap (zstd) |
| **Storage** | 953 GiB NVMe SSD (YMTC PC411), LUKS encrypted, btrfs with zstd:1 compression |
| **OS** | CachyOS (Arch-based, rolling), kernel `linux-cachyos-rc` 7.13 |
| **Display driver** | Mesa Vulkan (ANV) 26.2.2, Vulkan API 1.4.357 |

The key architectural detail here is **UMA** (Unified Memory Architecture). The Intel Arc B390 shares system RAM with the CPU — there's no dedicated VRAM. This means the 22 GiB model lives in the same 30 GiB pool as the KV cache and OS, making context size and cache quantization critical to avoid OOM.

## Software Setup

- **llama.cpp** v0.4.1-dev (commit 83078fe), cloned fresh from master
- **Model**: Qwen3.6-35B-A3B-MTP-UD-Q4\_K\_M.gguf — 35B MoE (A3B active), 256 experts x 8, 41 blocks, hybrid Attention (full attn every 4 layers) + Mamba2 SSM, MTP nextn=1
- **Builds**: 4 separate CMake builds in isolated directories:
  - `build-vulkan/` — `-DGGML_VULKAN=ON -DGGML_NATIVE=ON`
  - `build-sycl/` — `-DGGML_SYCL=ON -DGGML_SYCL_F16=ON`, compiled with Intel `icpx`
  - `build-openvino/` — `-DGGML_OPENVINO=ON`
  - `build-cpu/` — `-DGGML_NATIVE=ON` (pure CPU, no GPU offload)
- **Benchmark**: `llama-bench` with pp512 (prompt processing 512 tokens), tg128 (token generation 128 tokens), KV cache q4\_0, flash-attn on, 5 repetitions, 16 threads

## Issues Encountered

### 1. OpenVINO: Model Incompatibility

The OpenVINO backend failed with two distinct errors:

1. **`quantized V cache requires flash_attn to be enabled`** — When using q4\_0 KV cache without flash-attn, OpenVINO can't create the context because it doesn't support flash attention for this model architecture.
2. **`test_prompt: failed to decode prompt batch, res = -3`** — When flash-attn is enabled (which would allow quantized KV), OpenVINO still fails because it doesn't support the hybrid Attention + Mamba2 SSM + MTP architecture of Qwen3.6-35B.

The root cause: OpenVINO's llama.cpp backend is primarily validated on dense models under 8B parameters. MoE routing, recurrent state models (Mamba2/SSM), and multi-token prediction heads are outside its current scope. With f16 KV cache, the 22 GiB model + ~10 GiB KV cache exceeds the 30 GiB RAM limit, causing OOM.

**Verdict**: OpenVINO is not viable for this model class. It would require significant upstream work to support MoE offloading and SSM state management.

### 2. SYCL: Level Zero API Missing

During SYCL build configuration, CMake warned:

```
CMake Warning: Level Zero loader or development headers not found,
Level Zero API support disabled.
```

Level Zero is Intel's low-level GPU API that provides additional performance features. The SYCL backend compiled and works without it, but operates in a degraded mode:

- Device-to-device memory copies use SYCL API instead of Level Zero
- Memory allocation uses SYCL API instead of VMM
- Some optimization paths are disabled

This means the SYCL results represent a "good but not optimal" configuration. Installing the Level Zero SDK could improve SYCL performance further.

### 3. OneAPI Installation

Intel's oneAPI toolkit is 2.1 GiB (offline installer) and requires specific silent install syntax:

```
sudo sh ./intel-oneapi-toolkit-2026.1.1.33_offline.sh \
  -a -s --eula accept --install-dir /opt/intel/oneapi
```

The `--eula` flag must come after `-a` (argument passthrough), not as a top-level flag. The installer also warns about missing VTune GUI dependencies (NSS, XCB, DRM, GTK3, etc.) which are harmless on a headless/server setup but noisy in logs.

### 4. OpenCL Headers

OpenVINO build initially failed with `CL/cl2.hpp: No such file or directory`. The fix was installing both `opencl-headers` and `opencl-clhpp` packages.

## Results

| Backend | pp512 (tok/s) | tg128 (tok/s) | Speedup vs CPU |
| --- | --- | --- | --- |
| **Vulkan** | **624.53 +/- 7.06** | **32.75 +/- 0.09** | **6.1x pp, 1.7x tg** |
| **SYCL** | 609.36 +/- 4.65 | 30.62 +/- 0.10 | 6.0x pp, 1.6x tg |
| **CPU** | 101.57 +/- 1.87 | 19.44 +/- 0.06 | baseline |
| OpenVINO | -- | -- | N/A |

### Prompt Processing (pp512)

Vulkan and SYCL are neck-and-neck, with Vulkan holding a slight edge:

- **Vulkan**: 624.53 tok/s (best case)
- **SYCL**: 609.36 tok/s (-2.4%)
- **CPU**: 101.57 tok/s (-83.7%)

The 6x speedup over CPU confirms that GPU offload is essential for this model. The gap between Vulkan and SYCL is small (~2.5%), suggesting both backends are well-optimized for the compute-bound prompt processing phase.

### Token Generation (tg128)

The gap widens slightly in token generation:

- **Vulkan**: 32.75 tok/s
- **SYCL**: 30.62 tok/s (-6.5%)
- **CPU**: 19.44 tok/s

Token generation is memory-bandwidth bound (each step reads the full model weights). On UMA, GPU and CPU share the same memory bus, so the 1.7x speedup over CPU is smaller than the 6x for prompt processing. The Vulkan backend's slightly better tg performance likely comes from more efficient memory access patterns in the Mesa ANV driver.

### Why Vulkan Beats SYCL

Several factors contribute:

1. **Mesa ANV maturity**: The open-source Mesa Vulkan driver for Intel GPUs has years of optimization. SYCL's Level Zero backend, while functional, is newer.
2. **Level Zero disabled**: Without Level Zero, SYCL falls back to less efficient memory management. Enabling it could close the gap.
3. **F16 vs native precision**: SYCL was built with `GGML_SYCL_F16=ON` (half-precision compute), which trades some accuracy for throughput. Vulkan uses the driver's native precision selection.
4. **KHR\_coopmat**: Both backends use cooperative matrix extensions, but Vulkan's implementation may have better-tuned tile sizes for this specific hardware.

## Key Takeaways

1. **For Panther Lake with MoE models, use Vulkan.** Zero setup, best performance, no oneAPI required.
2. **SYCL is competitive but needs tuning.** The Level Zero SDK would help, and future oneAPI releases may close the gap.
3. **OpenVINO is not ready for MoE+SSM models.** It's a capable framework for smaller dense models but lacks the operator support for modern architectures.
4. **GPU offload is non-negotiable on UMA.** A 6x speedup on prompt processing and 1.7x on token generation makes the difference between usable and sluggish for a 35B model.
5. **Context size is the hidden bottleneck.** With 30 GiB shared between model (22 GiB) and KV cache, even q4\_0 quantization limits context to ~35k tokens. Users who need long contexts must reduce offloaded layers or accept smaller context windows.

## Reproduction

All builds, scripts, and results are in `/home/g/Code/llama-bench/`:

```
# Run the full benchmark suite
./bench.sh

# Or run individual backends
build-vulkan/bin/llama-bench -m models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf -p 512 -n 128 -ngl 99 -t 16 -r 5 -fa on -ctk q4_0 -ctv q4_0 -o md

source /opt/intel/oneapi/setvars.sh && build-sycl/bin/llama-bench -m models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf -p 512 -n 128 -ngl 99 -t 16 -r 5 -fa on -ctk q4_0 -ctv q4_0 -o md

build-cpu/bin/llama-bench -m models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf -p 512 -n 128 -ngl 0 -t 16 -r 5 -fa on -ctk q4_0 -ctv q4_0 -o md
```
grigio10
🟧 echo.blog ⭐Reports Vulkan at 624.53 prompt-processing and 32.75 generation tokens per second, versus SYCL at 609.36 and 30.62 and CPU at 101.57 and 19.Luigi——
🟧 hnI got 2.2x more tokens per second from llama.cpp on Intel Arcgrigio40

Interpretation history

Decision trace