2026-10-11 17:12 UTC

Independent reproduction will determine whether AMD Strix Point integrated graphics can sustain roughly 20 tokens per second on Qwen3.6-35B-A3B using shared system memory, enabling practical local coding workloads.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference amd-igpu qwen inference-economicsjstormesAMDQwen

What is this?

Qwen3.6-35B-A3B is an open model with 35B total parameters and roughly 3B active per token, positioned for coding and agentic workloads. A community measurement attributed to jstormes’s ThinkPad P14 repository claims about 449 tokens/s prefill and 23 tokens/s generation on AMD Strix Point integrated graphics, with `amd_iommu=off` reportedly producing a 23–26% improvement. The supplied web snippets establish the model and its coding use cases but do not independently reproduce the AMD iGPU result; AMD’s own snippet concerns discrete Instinct datacenter GPUs, while the search summary’s pessimistic conclusion is unsupported by a cited test.

Why it matters to Scott

The claimed laptop-iGPU throughput and material effect of `amd_iommu=off` converge with Scott’s hardware-aware local-inference practice: runtime and hardware policy can determine whether local coding inference reaches a usable operating point. If independently reproduced, it could affect his local-serving hardware and unit-economics choices, but the result remains unverified and his recorded substrate is currently CUDA-based.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsip:concept.operating-pointradar:llama-cpp-rocm-714-validationradar:concept.amd-inferenceradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.coding-models
queries asked of Scott's wikis
  • shared-memory iGPU local inference economics
  • local coding model minimum usable tokens per second
  • MoE active parameters versus memory bandwidth
  • AMD ROCm and consumer iGPU inference
  • local AI hardware sovereignty and commodity laptops
  • reproducible inference benchmarks and tuning flags

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen3.6-35B on a ThinkPad AMD iGPU: 449 t/s prefill, 23 t/s gen — plus the one boot flag that was worth 26%
LocalLLaMA
jstormes013
🟧 echo.github ⭐The primary measurement record is jstormes's p14 repository. The cited commit says: “Confirmed: the 23% was amd_iommu=off” and reports matchjstormes——
🟠 redditWhy does t/s go down as offload more to egpu?
LocalLLaMA
Pyrolistical00

Interpretation history

Decision trace