2026-10-11 17:11 UTC

Independent benchmarks will determine whether the reported open kernels can sustain roughly 78,500 output tokens per second for Qwen3.6-35B-A3B on eight AMD MI350X GPUs under practically comparable serving conditions.

state: expiredheat: lowuncertainty: highknownscott: lowamd-gpu inference-kernels llm-servingAMD

What is this?

Qwen3.6-35B-A3B is an Alibaba Cloud open-weight multimodal mixture-of-experts model with 35 billion total parameters and roughly 3 billion active per token. AMD documents Qwen3.6 support through ROCm with vLLM/SGLang on Instinct GPUs, but the supplied snippets do not establish the claimed open-kernel result of 78,498 output tokens per second on eight MI350X GPUs; AMD’s snippet instead names MI300X and MI355X. Independent testing with matched batching, concurrency, precision, context, output length, and serving configuration is therefore needed to determine whether the headline throughput is practically reproducible.

Why it matters to Scott

This is another unvalidated AMD open-kernel throughput claim in a territory already tracked by `radar:netra-amdgcn-inference-kernels` and the AMD-inference/LLM-serving concept pages. It touches Scott’s hardware-aware inference work, but the supplied evidence neither validates the result nor establishes a practical change to his current local-serving stack, so it adds little beyond the existing benchmark-validation pattern.
dev:concept.hardware-aware-local-inferenceradar:netra-amdgcn-inference-kernelsradar:concept.amd-inferenceradar:concept.llm-serving
queries asked of Scott's wikis
  • open inference kernels and hardware sovereignty
  • ROCm versus CUDA for local LLM serving
  • LLM throughput benchmark comparability
  • batching concurrency and aggregate token throughput
  • vLLM SGLang serving benchmarks
  • open-model inference economics on AMD GPUs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs
LocalLLaMA
SmilingGen5920

Interpretation history

Decision trace