Independent benchmarks will determine whether the reported open kernels can sustain roughly 78,500 output tokens per second for Qwen3.6-35B-A3B on eight AMD MI350X GPUs under practically comparable serving conditions.
state: expiredheat: lowuncertainty: highknownscott: lowamd-gpu inference-kernels llm-servingAMD
What is this?
Qwen3.6-35B-A3B is an Alibaba Cloud open-weight multimodal mixture-of-experts model with 35 billion total parameters and roughly 3 billion active per token. AMD documents Qwen3.6 support through ROCm with vLLM/SGLang on Instinct GPUs, but the supplied snippets do not establish the claimed open-kernel result of 78,498 output tokens per second on eight MI350X GPUs; AMD’s snippet instead names MI300X and MI355X. Independent testing with matched batching, concurrency, precision, context, output length, and serving configuration is therefore needed to determine whether the headline throughput is practically reproducible.
Why it matters to Scott
This is another unvalidated AMD open-kernel throughput claim in a territory already tracked by `radar:netra-amdgcn-inference-kernels` and the AMD-inference/LLM-serving concept pages. It touches Scott’s hardware-aware inference work, but the supplied evidence neither validates the result nor establishes a practical change to his current local-serving stack, so it adds little beyond the existing benchmark-validation pattern.
dev:concept.hardware-aware-local-inferenceradar:netra-amdgcn-inference-kernelsradar:concept.amd-inferenceradar:concept.llm-serving
queries asked of Scott's wikis
- open inference kernels and hardware sovereignty
- ROCm versus CUDA for local LLM serving
- LLM throughput benchmark comparability
- batching concurrency and aggregate token throughput
- vLLM SGLang serving benchmarks
- open-model inference economics on AMD GPUs
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-28T15:39:33Z
After repeated checks and 48 hours without methodology, code, configuration details, or independent reproduction, the claim has faded as an actionable episode. It remains an unverified self-report and can be reopened if a reproducible benchmark appears.
2026-08-26T15:34:57Z
The refreshed comments remain adjacent technical discussion rather than evidence about the claimed MI350X result. With no methodology, artifact, clarification of aggregate versus single-stream throughput, or independent reproduction, the case’s meaning is unchanged and increasingly repetitive.
2026-08-26T11:29:48Z
The refreshed comments add adjacent AMD optimization testimony but no artifact, methodology, or reproduction of the MI350X result. Discussion remains repetitive around the unresolved aggregate-throughput ambiguity, so the case has not advanced beyond a single-source claim.
2026-08-26T08:29:11Z
The refreshed discussion identifies the decisive ambiguity—whether 78,498 tok/s is aggregate batched throughput or single-stream performance—but supplies no answer, configuration details, artifact, or independent validation. The case therefore remains an unverified self-report rather than evidence of a practical AMD serving breakthrough.
2026-08-26T04:34:26Z
The modest engagement increase adds no independent benchmark, reproducible artifact, or serving-configuration detail. The throughput claim remains a single-source self-report and does not yet advance the broader AMD inference-kernel story.
2026-08-26T04:26:43Z
grounded: known/low — This is another unvalidated AMD open-kernel throughput claim in a territory already tracked by `radar:netra-amdgcn-inference-kernels` and the AMD-inference/LLM-
2026-08-26T04:24:31Z
case created — The specific throughput claim is technically consequential and testable, but currently rests on a single lightly discussed self-report.
Decision trace
- 08-29 01:39expireAfter repeated checks and 48 hours without methodology, code, configuration details, or independent reproduction, the claim has faded as an actionable episode. It remains an unverified self-report and
- 08-29 01:39alert_silentThe staleness trigger adds no consequential evidence; absent an artifact or independent benchmark, this does not merit Scott's attention or another scheduled review.
- 08-29 01:39alert_routeThe staleness trigger adds no consequential evidence; absent an artifact or independent benchmark, this does not merit Scott's attention or another scheduled review.
- 08-27 17:21sensor_dirtyengagement_update
- 08-27 05:21sensor_dirtyengagement_update
- 08-27 03:21sensor_dirtyengagement_update
- 08-27 01:34repriceThe refreshed comments remain adjacent technical discussion rather than evidence about the claimed MI350X result. With no methodology, artifact, clarification of aggregate versus single-stream through
- 08-27 01:34alert_silentNo consequential new fact has emerged; another discussion refresh does not justify attention before benchmark details, code, or an independent reproduction appears.
- 08-27 01:34alert_routeNo consequential new fact has emerged; another discussion refresh does not justify attention before benchmark details, code, or an independent reproduction appears.
- 08-27 00:21sensor_dirtycomment_update
- 08-26 21:29repriceThe refreshed comments add adjacent AMD optimization testimony but no artifact, methodology, or reproduction of the MI350X result. Discussion remains repetitive around the unresolved aggregate-through
- 08-26 21:29alert_silentNo new evidence establishes what the 78,498 tok/s figure measures or whether it is reproducible under comparable serving conditions; discussion alone can wait for code, benchmark details, or independe
- 08-26 21:29alert_routeNo new evidence establishes what the 78,498 tok/s figure measures or whether it is reproducible under comparable serving conditions; discussion alone can wait for code, benchmark details, or independe
- 08-26 21:21sensor_dirtycomment_update
- 08-26 19:21sensor_dirtyengagement_update
- 08-26 18:29repriceThe refreshed discussion identifies the decisive ambiguity—whether 78,498 tok/s is aggregate batched throughput or single-stream performance—but supplies no answer, configuration details, artifact, or
- 08-26 18:29alert_silentThe new comments only ask the benchmark-comparability question already central to the case; they do not establish a new event or result. This can wait for an inspectable implementation, complete bench
- 08-26 18:29alert_routeThe new comments only ask the benchmark-comparability question already central to the case; they do not establish a new event or result. This can wait for an inspectable implementation, complete bench
- 08-26 18:21sensor_dirtycomment_update
- 08-26 16:21sensor_dirtyengagement_update
- 08-26 15:21sensor_dirtyengagement_update
- 08-26 14:34repriceThe modest engagement increase adds no independent benchmark, reproducible artifact, or serving-configuration detail. The throughput claim remains a single-source self-report and does not yet advance
- 08-26 14:34alert_silentOnly Reddit engagement changed; no new evidence establishes reproducibility or practical comparability, so this can wait for an independent benchmark or inspectable implementation.
- 08-26 14:34alert_routeOnly Reddit engagement changed; no new evidence establishes reproducibility or practical comparability, so this can wait for an independent benchmark or inspectable implementation.
- 08-26 14:30alert_silentA low-engagement Reddit post reports an open-source MI350X kernel and unusually high aggregate throughput, but the supplied evidence does not verify the repository contents, benchmark configuration, p
- 08-26 14:30alert_routeA low-engagement Reddit post reports an open-source MI350X kernel and unusually high aggregate throughput, but the supplied evidence does not verify the repository contents, benchmark configuration, p
- 08-26 14:26groundThis is another unvalidated AMD open-kernel throughput claim in a territory already tracked by `radar:netra-amdgcn-inference-kernels` and the AMD-inference/LLM-serving concept pages. It touches Scott’
- 08-26 14:24createThe specific throughput claim is technically consequential and testable, but currently rests on a single lightly discussed self-report.