2026-10-11 17:12 UTC

Independent testing will determine whether the v100-skinny kernels make NVFP4 weights and speculative decoding a practically high-throughput inference path for Qwen-class models on inexpensive V100 GPUs.

state: expiredheat: lowuncertainty: highknownscott: lownvfp4 speculative-decoding inference-economics

What is this?

A reported v1.0 implementation called “v100-skinny” claims that four Tesla V100 GPUs can serve a Qwen3.6-27B NVFP4 model at up to 366 tokens per second, apparently using speculative decoding. The supplied sources establish that the checkpoint includes an MTP module for self-speculation and that NVFP4 is a compact 4-bit format, but they primarily discuss modern NVIDIA hardware and do not independently verify the V100 benchmark. The project’s authorship, test methodology, quality impact, workload conditions, and practical cost comparison are not established here, so the central throughput claim remains awaiting reproducible, apples-to-apples testing.

Why it matters to Scott

This is another unverified throughput-claim benchmark case in a pattern the radar already tracks extensively — radar:qwen36-quant-specdecode-scaling covers the same Qwen3.6-27B quant/speculative-decoding scaling story, and radar:concept.speculative-decoding, radar:concept.quantization, radar:concept.local-inference, radar:concept.inference-efficiency already aggregate dozens of similar 'independent testing will determine whether X GPU config delivers claimed inference throughput' episodes. It illustrates Scott's local-inference/hardware-aware-inference interests (dev:concept.hardware-aware-local-inference, dev:project.gpt4all) but adds no new claim, technology, or challenge beyond what the radar's existing cluster already holds.
dev:concept.hardware-aware-local-inferenceradar:qwen36-quant-specdecode-scalingradar:concept.speculative-decodingradar:concept.quantizationradar:concept.local-inferenceradar:concept.inference-efficiency
queries asked of Scott's wikis
  • commodity GPU inference economics
  • extending legacy GPU life for local inference
  • quantization quality versus throughput tradeoffs
  • speculative decoding and self-drafting models
  • open-model hardware sovereignty
  • reproducible inference benchmark methodology

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit366 t/s Qwen3.6 27B NVFP4 on v100s
LocalLLaMA
Simple_Library_2700124114
🟧 echo.github ⭐The v1.0 commit publishes the benchmark implementation and results. Its README says: “Four Tesla V100 cards serve Qwen3.6-27B at up to 366 tdnv2003——

Interpretation history

Decision trace