Qwen3.6-35B-A3B is described as an open-weight, multimodal mixture-of-experts model from Qwen with 35 billion total parameters and roughly 3 billion activated during inference. NInfer’s benchmark document reportedly measured 542.8 ± 12.5 decode tokens per second for a 65,536-token completion on one RTX 5090 using MTP3, but the supplied results do not independently reproduce that figure or explain who is behind NInfer. The only separate user benchmark shown reports about 205 tokens per second at roughly 125K context with a GPTQ-Int4 checkpoint, so NInfer’s claimed advantage over general-purpose runtimes remains unestablished here.
NInfer’s checkpoint-specific optimization claim converges with Scott’s hardware-aware local-inference approach and could identify a new operating point where a specialized runtime materially outperforms general-purpose serving on consumer GPUs. It matters to his self-hosted GPU model substrate, but the gain remains unreplicated, the test uses an RTX 5090 rather than his documented hardware, and the supplied evidence does not establish NInfer’s credibility.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.operating-pointradar:concept.local-inference
queries asked of Scott's wikis
- checkpoint-specific inference optimization versus general-purpose runtimes
- speculative decoding and multi-token prediction tradeoffs
- local MoE inference economics on consumer GPUs
- long-context decode benchmark methodology
- local coding-agent throughput versus model quality
- RTX 5090 inference runtime optimization
2026-07-26T15:23:56Z
The active window has produced only repeated amplification of NInfer’s first-party result, with no independent reproduction or controlled runtime comparison. The case has faded without advancing and can be reopened if a substantive third-party benchmark appears.
2026-07-26T14:23:51Z
Minor comment and score growth remains repetitive amplification, with no measured third-party reproduction or controlled comparison against general-purpose runtimes. The case is dormant but still testable; revisit only when an independent benchmark appears.
2026-07-23T13:37:23Z
The latest attachment still adds no independent measurement or controlled comparison; the claim remains a precise first-party benchmark surrounded by repetitive amplification. Keep the case dormant until a third party publishes reproducible throughput and quality results.
2026-07-23T12:28:47Z
The HN attachment merely repeats NInfer’s first-party figure and has attracted no discussion, while refreshed Reddit comments add enthusiasm and a claimed Windows port but no measured reproduction or controlled runtime comparison. The case remains dormant and should only revive on a substantive third-party benchmark.
2026-07-23T12:22:20Z
evidence attached: hn.story.49020083 — Directly supports the open case's claimed Qwen3.6-35B-A3B throughput on one RTX 5090, though it is not independent validation.
2026-07-21T18:30:17Z
The purported new evidence is an empty reobservation, leaving the claim entirely dependent on NInfer’s own benchmark without independent reproduction or a controlled runtime comparison. The case is dormant rather than disproved; only a substantive third-party benchmark should revive it.
2026-07-21T17:40:22Z
The new trigger is another empty reobservation, leaving NInfer’s precise first-party result without independent reproduction or a controlled comparison against general-purpose runtimes. The discussion cycle is exhausted; retain the case but check only for substantive third-party benchmark results.
2026-07-21T15:33:26Z
Another empty reobservation adds no independent benchmark or controlled runtime comparison; repeated amplification has fully exhausted its informational value. Keep the claim dormant until a third-party reproduction or implementation result appears.
2026-07-21T13:26:00Z
The latest attachment is another empty reobservation, so the case remains wholly dependent on NInfer’s first-party benchmark without an independent reproduction or controlled runtime comparison. The attention cycle is exhausted; revisit only after a substantive third-party result.
2026-07-21T11:28:11Z
The latest trigger is another empty reobservation, leaving the claim dependent entirely on NInfer’s own benchmark with no independent reproduction or controlled runtime comparison. The active discussion cycle is exhausted; retain the case but wait for a substantive third-party result.
2026-07-21T09:24:38Z
The latest trigger adds no evidence beyond NInfer’s first-party benchmark, so the reproduction and runtime-comparison hypothesis remains entirely untested. Repetitive engagement has no further informational value; revisit only when an independent benchmark or implementation result appears.
2026-07-21T07:24:08Z
No substantive new evidence has appeared: the signal still consists of NInfer’s first-party benchmark and repetitive amplification, without an independent reproduction or controlled runtime comparison. The claim remains worth retaining, but only a third-party benchmark should trigger near-term repricing.
2026-07-21T06:29:05Z
The latest trigger is another empty reobservation, not an independent benchmark or runtime comparison; the case remains a precise first-party claim whose repeated amplification no longer warrants frequent checks.
2026-07-21T04:22:30Z
The latest attachment is another reobservation of NInfer’s first-party result, not an independent reproduction or controlled runtime comparison. Repetitive amplification has exhausted the near-term signal; wait for a substantive third-party benchmark before repricing again.
2026-07-21T03:26:45Z
The new attachment is only another reobservation and provides no independent reproduction or controlled runtime comparison. The case remains an unvalidated first-party benchmark, so monitoring should stay slow until substantive third-party results appear.
2026-07-21T02:23:14Z
No independent benchmark or controlled runtime comparison has appeared; the new attachment is another reobservation of NInfer’s own claim. Attention remains repetitive amplification, so further checks should wait for a substantive third-party reproduction.
2026-07-21T00:21:35Z
The evidence still reduces to NInfer’s own benchmark and Reddit amplification, with no independent reproduction or controlled comparison against a general-purpose runtime. Repeated reobservation is not changing the case’s meaning, so monitoring can slow pending an actual third-party benchmark.
2026-07-20T22:23:04Z
The new attachment still adds no independent reproduction or controlled comparison with a general-purpose runtime. Continued attention is repetitive amplification of NInfer’s own benchmark, so the claim remains precise but uncorroborated.
2026-07-20T21:23:43Z
The attached material adds no independent reproduction or controlled comparison; it remains NInfer’s own precise but unvalidated benchmark. Repeated engagement is amplification rather than progress on the hypothesis.
2026-07-20T20:27:27Z
The attached material still provides no evidence independent of NInfer’s own benchmark and no controlled comparison with a general-purpose runtime. Repeated amplification is not advancing the reproduction hypothesis, so the case remains uncorroborated and can be checked less frequently.
2026-07-20T19:23:56Z
The newly attached evidence still collapses to NInfer’s own benchmark and adds no independent reproduction or controlled runtime comparison. Discussion remains plausibility and amplification rather than validation, so the claim’s meaning is unchanged.
2026-07-20T18:27:13Z
The attached evidence still resolves to NInfer’s own benchmark, with no independent reproduction or controlled comparison against general-purpose runtimes. Additional attention is repetitive amplification rather than validation, so the case remains an interesting but uncorroborated claim.
2026-07-20T17:31:17Z
The newly attached material still traces to NInfer’s own benchmark, while discussion offers only plausibility judgments and requests for comparisons—not an independent reproduction. The claim remains technically interesting but unvalidated against general-purpose runtimes.
2026-07-20T16:24:30Z
The primary benchmark remains precise but unreplicated; the slight engagement increase adds no independent validation or runtime comparison, so the case’s meaning is unchanged.
2026-07-20T15:30:47Z
grounded: converges/medium — NInfer’s checkpoint-specific optimization claim converges with Scott’s hardware-aware local-inference approach and could identify a new operating point where a
2026-07-20T15:28:57Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1v1no8e -> echo.github.c2067d2dc7 by Neroued
2026-07-20T15:26:42Z
case created — The open-source engine makes an unusually large, precise, and reproducible throughput claim, though its specialization to two checkpoints limits the immediate scope.