LocalLLaMA user Biomass23 claims zero-padding model-weight dimensions to divisible sizes makes vLLM tensor parallelism work on six non-power-of-two GPUs (reported Qwen 3.8 27B BF16 at ~50 tok/s with 256k context on six 7900 XTXs), and independent replication or upstream vLLM support would establish odd-GPU-count padding as a standard local-inference technique.
state: seedheat: lowuncertainty: highnovelscott: mediumlocal-inference vllm amd-gpus tensor-parallelism
What is this?
LocalLLaMA user Biomass23 posted 'tp=6 can work on vLLM, with padding,' claiming that zero-padding model-weight dimensions to GPU-divisible sizes lets vLLM tensor parallelism run on six non-power-of-two AMD 7900 XTXs โ reportedly ~50 tok/s on Qwen 3.8 27B BF16 with 256k context. The supplied material carries only the post title: no post body, code, or independent replication was available, so the technique claim itself is unverified in what was reviewed. The surrounding premises do check out in the snippets: vLLM users treat tensor-parallel size as needing power-of-two / divisible-by-two values (a 3x A6000 build hangs on tp=3 with commenters citing that rule), Qwen 3.8 27B is a heavily benchmarked local model on vLLM across many hardware configs, AMD ROCm vLLM serving of a 27B BF16 checkpoint at 256k context is active territory, and the standard workaround is the inverse move โ scale up to a power-of-two GPU count rather than reshape the model (Qwen's reference config uses TP=8 for the full 262k context).
Why it matters to Scott
Novel technique claim, not a held position: nothing in the supplied wiki hits shows Scott's canon arguing odd-count TP or padding-for-TP, and no radar page asserts the power-of-two rule this would challenge โ so no canon position is contradicted, converged with, or already held. But it bears directly on his active stack: if replication or upstream vLLM support lands, it extends the gamepc vLLM tensor-parallel multi-GPU setup notes and his weight-surgery/padding repertoire with tp=3/6 rigs instead of scaling to power-of-two GPU counts, and it feeds the radar's Qwen3.8-27B-on-AMD serving thread. Evidence is title-only and unverified โ replication (or a vLLM merge) is the hinge on which any upgrade to high turns.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.vllmradar:concept.local-inferenceradar:concept.amd-gpuradar:vllm-amd-speculative-decodingradar:r9700-nvfp4-mxfp4-fast-path
queries asked of Scott's wikis
- vLLM tensor parallelism multi-GPU setup notes
- local inference rig consumer GPU serving experiments
- AMD ROCm inference tooling and pain points
- weight surgery checkpoint modification padding sharding
- open-weight model in local coding agent harness
- vLLM vs llama.cpp engine choice tradeoffs
Measured heat
now 0 pts/hpeak 2 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 266h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p56 vs 1188 stories at the 168h mark (now 266h old) โ ahead of chatgpt-word-integration (1.0x), behind opencontext-project-local-agent-memory (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-02T13:32:42Z
First community-response window closed without replication: the only new substantive comment is a detail-free tp=3-on-3060s anecdote that never mentions padding (and, if taken at face value without padding, could even mean the divisibility constraint is looser than claimed), while the 'logits match 1:1?' correctness question sits unanswered and one commenter reads the technique as 'adding fake layers'. The claim remains a single-author, LLM-scripted weight-surgery writeup on a thread that has cooled to rest (0 pts/h, 10th percentile at 47h) โ stays a cold seed on an expiry watch unless a logits-verified follow-up, independent replication, or upstream vLLM movement appears.
2026-10-01T03:22:19Z
origin walked (opencode/cheap-glm, conf 0.9): anchor reddit.post.1wullxh -> echo.reddit.462e5c2924 by Biomass23
2026-10-01T02:57:57Z
grounded: novel/medium โ Novel technique claim, not a held position: nothing in the supplied wiki hits shows Scott's canon arguing odd-count TP or padding-for-TP, and no radar page asse
2026-10-01T02:48:24Z
case created โ Concrete, replication-resolvable technique claim squarely in Scott's local-inference wheelhouse with no overlapping open case.
Decision trace
- 10-02 23:32repriceFirst community-response window closed without replication: the only new substantive comment is a detail-free tp=3-on-3060s anecdote that never mentions padding (and, if taken at face value without pa
- 10-02 23:31jev_reprice_gatechanges_anything noul=0.16 would_skip=False
- 10-02 23:31review_screenA new comment reports TP working on three RTX 3060s, but it lacks any indication the zero-padding technique was used versus vanilla vLLM TP, offers no model, throughput, or correctness details, and th
- 10-02 23:30review_screenjev screen borderline (noul=0.40) โ luna review
- 10-01 18:20sensor_dirtycomment_update
- 10-01 13:22promote_anchororigin walk conf 0.9
- 10-01 12:57groundNovel technique claim, not a held position: nothing in the supplied wiki hits shows Scott's canon arguing odd-count TP or padding-for-TP, and no radar page asserts the power-of-two rule this woul
- 10-01 12:48createConcrete, replication-resolvable technique claim squarely in Scott's local-inference wheelhouse with no overlapping open case.