2026-10-11 17:11 UTC

Independent benchmarks will determine whether v100-skinny can run unchanged NVFP4 models on Tesla V100 GPUs with practically competitive decode performance and economics.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference inference-economics gpu-optimizationv100-skinnyNInferNVIDIAQwen

What is this?

v100-skinny is presented as a GPU-optimization implementation that claims to run Qwen3.8-27B mixed NVFP4/FP8 weights on four 2017-era Tesla V100 GPUs, despite NVFP4 being designed for newer Blackwell hardware, and to approach the decode performance of a roughly $6,000 RTX 5090 system. The supplied results establish that V100s can remain economical for moderate inference when models fit in memory, but they do not independently verify v100-skinny’s unchanged-model compatibility, benchmark methodology, performance parity, power costs, or total economics. The project’s creator and any relationship to NInfer are not established by the snippets, so the v1.1 commit and reproducible third-party benchmarks remain central evidence.

Why it matters to Scott

The radar already tracks this exact development in `radar:v100-skinny-nvfp4-speculative-decoding`, including the need for independent performance and economic verification. The result could bear on Scott’s hardware-aware local-inference practice and AI unit-economics decisions by showing whether legacy accelerators can economically run newer quantized weights, but this case adds no new evidence beyond the existing open question.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsip:concept.evidence-class-ladderradar:v100-skinny-nvfp4-speculative-decodingradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.nvfp4
queries asked of Scott's wikis
  • legacy GPU local inference economics
  • quantization formats across GPU generations
  • hand-written kernels for unsupported inference formats
  • decode throughput versus total cost of ownership
  • local model hardware sovereignty and reused accelerators
  • reproducible benchmarking for inference optimization claims

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.
LocalLLaMA
Simple_Library_2700178138
🟧 echo.github ⭐The v1.1 commit is the primary artifact for the claim. It states that Qwen3.8-27B’s mixed NVFP4/FP8 weights run on four V100s using hand-wridnv2003——

Interpretation history

Decision trace