2026-10-11 18:04 UTC

Independent benchmarks will reproduce NInfer's reported roughly 542-token-per-second long decode for Qwen3.6-35B-A3B on one RTX 5090 and establish whether its checkpoint-specific design offers practical gains over general-purpose runtimes.

state: expiredheat: lowuncertainty: highconvergesscott: mediuminference-optimization qwen local-inferenceNInfer

What is this?

Qwen3.6-35B-A3B is described as an open-weight, multimodal mixture-of-experts model from Qwen with 35 billion total parameters and roughly 3 billion activated during inference. NInfer’s benchmark document reportedly measured 542.8 ± 12.5 decode tokens per second for a 65,536-token completion on one RTX 5090 using MTP3, but the supplied results do not independently reproduce that figure or explain who is behind NInfer. The only separate user benchmark shown reports about 205 tokens per second at roughly 125K context with a GPTQ-Int4 checkpoint, so NInfer’s claimed advantage over general-purpose runtimes remains unestablished here.

Why it matters to Scott

NInfer’s checkpoint-specific optimization claim converges with Scott’s hardware-aware local-inference approach and could identify a new operating point where a specialized runtime materially outperforms general-purpose serving on consumer GPUs. It matters to his self-hosted GPU model substrate, but the gain remains unreplicated, the test uses an RTX 5090 rather than his documented hardware, and the supplied evidence does not establish NInfer’s credibility.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.operating-pointradar:concept.local-inference
queries asked of Scott's wikis
  • checkpoint-specific inference optimization versus general-purpose runtimes
  • speculative decoding and multi-token prediction tradeoffs
  • local MoE inference economics on consumer GPUs
  • long-context decode benchmark methodology
  • local coding-agent throughput versus model quality
  • RTX 5090 inference runtime optimization

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decode
LocalLLaMA
FormOne2615212112
🟧 echo.github ⭐The benchmark document reports Qwen3.6-35B-A3B on one RTX 5090 using MTP3: a 65,536-token completion averaged 542.8 ± 12.5 decode tok/s, witNeroued——
🟧 hnNInfer: 271 tok/s (543 with speculative) for Qwen3.6-35B-A3B on one RTX 5090opwizardx10

Interpretation history

Decision trace