2026-10-11 16:36 UTC

tensor-parallelism

band: warmmomentum: stable score: 0.418
temperature history

Episodes (2)

am17an's merged llama.cpp PR #26610 adds RPC '-sm tensor' โ€” tensor parallelism across networked machines (author-demonstrated on 2x DGX Sparks over RDMA, independently confirmed by ryan5rdx on 2x M3 Ultra, running a 284B-param MoE at 619 pp2048 / ~20 tg128) โ€” and becomes a practical multi-machine local-inference pattern if outside users adopt it across their own clusters with replicated throughput; quiet disuse after the merge closes it.
watchingconvergesscott: high
LocalLLaMA user Biomass23 claims zero-padding model-weight dimensions to divisible sizes makes vLLM tensor parallelism work on six non-power-of-two GPUs (reported Qwen 3.8 27B BF16 at ~50 tok/s with 256k context on six 7900 XTXs), and independent replication or upstream vLLM support would establish odd-GPU-count padding as a standard local-inference technique.
seednovelscott: medium