v100-skinny is presented as a GPU-optimization implementation that claims to run Qwen3.8-27B mixed NVFP4/FP8 weights on four 2017-era Tesla V100 GPUs, despite NVFP4 being designed for newer Blackwell hardware, and to approach the decode performance of a roughly $6,000 RTX 5090 system. The supplied results establish that V100s can remain economical for moderate inference when models fit in memory, but they do not independently verify v100-skinny’s unchanged-model compatibility, benchmark methodology, performance parity, power costs, or total economics. The project’s creator and any relationship to NInfer are not established by the snippets, so the v1.1 commit and reproducible third-party benchmarks remain central evidence.
2026-09-05T00:28:51Z
The case has gone dormant after nearly two weeks without evidence beyond the kernel-level replication. End-to-end unchanged-model execution, decode parity, power use, and economics remain unresolved, but further monitoring should be evidence-triggered rather than scheduled.
2026-09-03T00:22:45Z
Another staleness pass adds no technical evidence beyond the existing kernel-level replication; independent end-to-end compatibility, decode, power, and economic validation remain absent, so the case should move to a longer evidence-triggered cadence.
2026-08-31T23:35:05Z
This staleness pass adds no evidence beyond the prior kernel-level replication, so the broader unchanged-model execution, end-to-end decode, power, and economics claims remain open. Further engagement-only checks should wait behind substantive independent testing.
2026-08-29T22:37:26Z
The refreshed discussion adds no technical evidence beyond the previously counted kernel-level replication. The central end-to-end compatibility, decode parity, power, and economic claims remain open, and discussion-only updates no longer justify frequent review.
2026-08-27T21:40:33Z
This staleness check adds only negligible engagement and no technical evidence beyond the prior kernel-level replication. The broader unchanged-model, end-to-end decode, power, and economic claims remain open, so the case should wait for substantive independent testing.
2026-08-25T21:34:26Z
The refreshed discussion adds no evidence beyond the existing kernel-level replication. Unchanged-model execution, end-to-end decode parity, power consumption, and practical economics remain independently unverified, so comment-only updates no longer merit close monitoring.
2026-08-24T03:30:06Z
The refreshed comments add no evidence beyond the already-counted third-party kernel-level replication. End-to-end unchanged-model execution, decode performance, power use, and economics remain unverified.
2026-08-22T19:38:22Z
A third party now reports reproducing the V100 kernel throughput in GCP to within roughly 1%, providing the first independent technical support for the implementation. This is only a kernel-level replication, so unchanged-model execution, end-to-end decode parity, power use, and economics remain unverified.
2026-08-20T18:36:57Z
The refreshed discussion remains prospective and repetitive, with no independent execution, compatibility result, benchmark, power measurement, or cost comparison. The case still rests entirely on the author’s implementation claims, so further discussion-only updates should not increase monitoring cadence.
2026-08-20T10:36:46Z
The refreshed discussion remains repetitive prospective interest, with no independent execution result, compatibility finding, benchmark, power measurement, or economic comparison. The case still depends entirely on the author’s reported implementation, so comment and engagement updates no longer warrant frequent review.
2026-08-20T05:29:08Z
The refreshed discussion remains repetitive prospective interest, with no independent run, failure report, benchmark, power measurement, or cost comparison. The case still rests entirely on the author’s implementation and reported results, so further comment-only updates do not change its meaning.
2026-08-20T03:28:34Z
The refreshed discussion still adds no independent run, compatibility result, benchmark, power measurement, or cost analysis. Repeated prospective interest does not change the case from an author-reported implementation claim.
2026-08-20T02:30:48Z
The refreshed comments add no completed replication, failure report, benchmark, power data, or economic comparison. This remains an author-reported implementation claim, and further discussion-only refreshes do not merit tighter monitoring.
2026-08-20T01:24:46Z
Refreshed comments still offer only prospective testing and general enthusiasm, not an independent run, benchmark, compatibility result, power measurement, or cost comparison. Repeated discussion updates no longer justify frequent review absent technical evidence.
2026-08-19T22:37:07Z
The refreshed discussion still contains only prospective testing and repetition of the author’s claim, with no independent run, failure report, power data, or economic comparison. The case remains an uncorroborated but testable implementation claim.
2026-08-19T20:42:13Z
The latest comment refresh adds no completed replication, failure report, benchmark, power measurement, or cost analysis; it is repetitive prospective interest rather than new evidence. The unchanged-model compatibility and competitive economics claims remain dependent on the author's results.
2026-08-19T19:37:30Z
The refreshed discussion remains repetitive amplification and prospective testing, with no completed independent run or new performance, power, compatibility, or cost evidence. The claim stays technically plausible but wholly dependent on author-reported results.
2026-08-19T18:35:26Z
Refreshed discussion remains prospective replication interest rather than independent technical evidence. The implementation claim is still uncorroborated, though multiple owners of suitable hardware make a near-term test result plausible.
2026-08-19T16:51:08Z
Discussion now includes several prospective replicators, but no one has reported an independent run, benchmark, power measurement, or cost comparison. The case remains a testable author claim rather than corroborated evidence.
2026-08-19T16:31:22Z
grounded: known/medium — The radar already tracks this exact development in `radar:v100-skinny-nvfp4-speculative-decoding`, including the need for independent performance and economic v
2026-08-19T16:28:14Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vsq3zg -> echo.github.0e92fb940d by dnv2003
2026-08-19T16:27:17Z
case created — The released implementation makes a consequential, independently testable claim about repurposing inexpensive legacy GPUs for modern low-bit inference.