Qwen3.6-35B-A3B is an open model with 35B total parameters and roughly 3B active per token, positioned for coding and agentic workloads. A community measurement attributed to jstormes’s ThinkPad P14 repository claims about 449 tokens/s prefill and 23 tokens/s generation on AMD Strix Point integrated graphics, with `amd_iommu=off` reportedly producing a 23–26% improvement. The supplied web snippets establish the model and its coding use cases but do not independently reproduce the AMD iGPU result; AMD’s own snippet concerns discrete Instinct datacenter GPUs, while the search summary’s pessimistic conclusion is unsupported by a cited test.
The claimed laptop-iGPU throughput and material effect of `amd_iommu=off` converge with Scott’s hardware-aware local-inference practice: runtime and hardware policy can determine whether local coding inference reaches a usable operating point. If independently reproduced, it could affect his local-serving hardware and unit-economics choices, but the result remains unverified and his recorded substrate is currently CUDA-based.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsip:concept.operating-pointradar:llama-cpp-rocm-714-validationradar:concept.amd-inferenceradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.coding-models
queries asked of Scott's wikis
- shared-memory iGPU local inference economics
- local coding model minimum usable tokens per second
- MoE active parameters versus memory bandwidth
- AMD ROCm and consumer iGPU inference
- local AI hardware sovereignty and commodity laptops
- reproducible inference benchmarks and tuning flags
2026-08-28T05:24:34Z
Repeated review windows produced no matched independent reproduction or additional technical record, leaving the result an isolated operator benchmark. The episode has faded without establishing a transferable AMD-iGPU operating point.
2026-08-26T04:32:39Z
No independent benchmark or new technical evidence arrived within the review window, so the case remains a credible but isolated measurement rather than evidence of a transferable AMD-iGPU operating point.
2026-08-24T03:30:15Z
A second user’s claim of roughly 20 t/s on a different AMD iGPU weakly supports a memory-bandwidth-limited operating point, but lacks configuration details and a measurement record. It is not a matched independent reproduction of the Strix Point result.
2026-08-23T03:23:03Z
The Strix Halo plus PCIe-attached eGPU report adds a clue that split-device placement and transfer overhead can reduce throughput, but it is neither comparable enough nor complete enough to reproduce the Strix Point iGPU claim. The case remains a credible single-operator benchmark awaiting an independent matched test.
2026-08-23T03:22:08Z
evidence attached: reddit.post.1vvuypl — The firsthand benchmark investigates how Strix Point and discrete-GPU memory placement affect Qwen3.6 local inference performance.
2026-08-22T21:34:10Z
No independent reproduction or new technical evidence has arrived; the lone additional comment does not change the single-operator status of the benchmark. Cool the case while retaining the public measurement record as a credible seed.
2026-08-22T21:30:22Z
grounded: converges/medium — The claimed laptop-iGPU throughput and material effect of `amd_iommu=off` converge with Scott’s hardware-aware local-inference practice: runtime and hardware po
2026-08-22T21:26:25Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vvo7s9 -> echo.github.696a235211 by jstormes
2026-08-22T21:24:52Z
case created — The post reports a reproducible hardware configuration, public repository, and sustained measurements for a materially useful local-inference claim.