A reported v1.0 implementation called “v100-skinny” claims that four Tesla V100 GPUs can serve a Qwen3.6-27B NVFP4 model at up to 366 tokens per second, apparently using speculative decoding. The supplied sources establish that the checkpoint includes an MTP module for self-speculation and that NVFP4 is a compact 4-bit format, but they primarily discuss modern NVIDIA hardware and do not independently verify the V100 benchmark. The project’s authorship, test methodology, quality impact, workload conditions, and practical cost comparison are not established here, so the central throughput claim remains awaiting reproducible, apples-to-apples testing.
This is another unverified throughput-claim benchmark case in a pattern the radar already tracks extensively — radar:qwen36-quant-specdecode-scaling covers the same Qwen3.6-27B quant/speculative-decoding scaling story, and radar:concept.speculative-decoding, radar:concept.quantization, radar:concept.local-inference, radar:concept.inference-efficiency already aggregate dozens of similar 'independent testing will determine whether X GPU config delivers claimed inference throughput' episodes. It illustrates Scott's local-inference/hardware-aware-inference interests (dev:concept.hardware-aware-local-inference, dev:project.gpt4all) but adds no new claim, technology, or challenge beyond what the radar's existing cluster already holds.
dev:concept.hardware-aware-local-inferenceradar:qwen36-quant-specdecode-scalingradar:concept.speculative-decodingradar:concept.quantizationradar:concept.local-inferenceradar:concept.inference-efficiency
queries asked of Scott's wikis
- commodity GPU inference economics
- extending legacy GPU life for local inference
- quantization quality versus throughput tradeoffs
- speculative decoding and self-drafting models
- open-model hardware sovereignty
- reproducible inference benchmark methodology
2026-08-17T03:29:07Z
The promised independent testing has not appeared after repeated review, and the latest change is engagement-only. The episode has faded without advancing beyond its original first-party benchmark claim and can be reopened if substantive reproduction arrives.
2026-08-15T03:22:26Z
The latest activity is still comment churn rather than independent reproduction, prefill data, or methodological clarification. The testable first-party claim remains open but cold and should not be revisited on engagement alone.
2026-08-13T02:33:03Z
The refreshed comments add no independent reproduction, prefill result, or methodological clarification. This remains repetitive amplification of a first-party benchmark and should stay cold until substantive testing arrives.
2026-08-13T00:24:19Z
The refreshed discussion still provides no independent reproduction, prefill benchmark, or methodological clarification. Comment churn is repetitive amplification, so the case remains a cold first-party claim pending substantive external testing.
2026-08-12T21:36:50Z
The refreshed discussion adds no independent reproduction, prefill benchmark, or methodological clarification. Repeated comment churn is no longer changing the case; it remains a cold first-party throughput claim awaiting substantive external testing.
2026-08-12T19:27:30Z
The latest comment refresh again adds no independent reproduction, prefill measurement, or methodological clarification. The case remains a cold first-party benchmark claim and should await substantive external testing rather than further discussion updates.
2026-08-12T15:41:54Z
The refreshed comments still provide no independent reproduction, prefill measurement, or methodological clarification. This remains a cold first-party benchmark claim; further discussion-only updates should not trigger review.
2026-08-12T11:41:35Z
The refreshed discussion adds no independent reproduction, prefill measurement, or methodological clarification. Repeated amplification is no longer changing the case’s meaning; it remains a first-party benchmark claim awaiting substantive external testing.
2026-08-12T09:26:50Z
The latest comment refresh still adds no independent reproduction, prefill measurement, or methodological clarification. Repetitive discussion no longer changes the case’s meaning; it remains a first-party benchmark claim awaiting external testing.
2026-08-12T08:39:29Z
Two additional comments add no independent reproduction, prefill measurement, or methodological clarification. The discussion remains repetitive amplification of the first-party benchmark, so the case should stay cold pending substantive test results.
2026-08-12T07:31:44Z
The comment refresh adds no independent reproduction, prefill benchmark, or methodological clarification. Repeated amplification has stopped changing the case’s meaning, so it should remain cold until substantive test results arrive.
2026-08-12T06:34:40Z
The refreshed comments add no independent reproduction, prefill measurement, or methodological clarification; the case remains a first-party benchmark claim despite sustained interest. Repeated discussion updates no longer merit frequent review absent substantive test results.
2026-08-12T05:34:24Z
The refreshed discussion remains enthusiasm and repeated questions about prefill, without the promised independent test or new methodological evidence. The case still represents a reproducible first-party throughput claim awaiting external validation.
2026-08-12T04:30:58Z
The refreshed discussion still supplies no independent benchmark, prefill result, or methodological clarification. It remains a reproducible but first-party throughput claim, with repeated community interest adding no validation.
2026-08-12T03:25:25Z
The velocity spike is amplification of the original benchmark claim, not validation; no independent result, prefill measurement, or methodological clarification changes the case’s meaning.
2026-08-12T02:31:46Z
The refreshed comments remain repetitive enthusiasm and prefill questions; no promised independent benchmark or methodological clarification has appeared. The case still means only a reproducible first-party throughput claim awaiting external validation.
2026-08-11T23:25:27Z
Refreshed discussion remains hardware enthusiasm and questions about omitted prefill performance; the previously mentioned independent test has not produced results. The case is still a reproducible first-party claim awaiting external benchmarks.
2026-08-11T22:24:31Z
Discussion is now probing missing prefill data and includes one commenter’s intent to test, but no independent results have arrived. The case remains an unverified first-party throughput claim rather than emerging corroboration.
2026-08-11T21:52:09Z
Engagement on the original Reddit post increased (19 upvotes, 20 comments), but no new evidence or independent verification; the claim remains unconfirmed.
2026-08-11T21:30:43Z
grounded: known/low — This is another unverified throughput-claim benchmark case in a pattern the radar already tracks extensively — radar:qwen36-quant-specdecode-scaling covers the
2026-08-11T21:27:36Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vlt0lj -> echo.github.051df353d3 by dnv2003
2026-08-11T21:25:28Z
case created — Single first-party reddit post with concrete single-stream benchmark numbers for a novel low-cost inference path; no corroboration yet.