2026-10-11 17:11 UTC

mattescala claims a llama.cpp NUMA weight-mirroring implementation improves dual-socket CPU decode throughput by 64–137% by replicating weights per NUMA node, trading doubled weight memory for materially better local-inference performance.

state: expiredheat: lowuncertainty: highconvergesscott: mediumllama-cpp local-inference inference-economicsmattescalallama.cpp

What is this?

mattescala proposed a llama.cpp `--numa mirror` implementation that replicates model weights on each NUMA node, reporting roughly 64–137% higher decode throughput on a dual-EPYC system at the cost of doubling weight memory. The supplied snippets support the underlying rationale—dual-socket inference is constrained by NUMA locality and memory bandwidth, and per-node data replication can reduce remote-memory access—but they do not independently verify mattescala’s benchmark figures, broader hardware results, or whether the implementation was merged upstream.

Why it matters to Scott

The proposed NUMA mirroring concretely converges with Scott’s “Hardware-aware local inference” position by making memory placement and pressure explicit runtime policy, and it could materially change dual-socket CPU inference economics. It is not yet high relevance because the reported gains are unverified, upstream status is unknown, and the hits do not show Scott currently runs local models on comparable dual-socket hardware.
dev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:concept.cpu-inferenceradar:concept.memory-bandwidthradar:concept.llama-cppradar:concept.inference-economicsradar:picchio-llama-cpp-bottleneck-diagnostics
queries asked of Scott's wikis
  • NUMA locality and memory-bandwidth limits in local LLM inference
  • CPU inference economics versus GPU inference
  • trading model-memory duplication for decode throughput
  • dual-socket hardware strategy for local models
  • llama.cpp performance tuning and benchmark methodology

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit--numa mirror for llama.cpp: replicate weights per NUMA node, +64% to +137% decode on my dual EPYC. Need people with 2-socket boxes to test it.
LocalLLaMA
mattescala53
🟧 echo.github ⭐Earliest primary artifact found. The author proposed: “Replicate models on each NUMA,” reporting CPU inference gains from ~6.6 to ~10.7 tok/wkgcass——

Interpretation history

Decision trace