mattescala claims a llama.cpp NUMA weight-mirroring implementation improves dual-socket CPU decode throughput by 64–137% by replicating weights per NUMA node, trading doubled weight memory for materially better local-inference performance.
state: expiredheat: lowuncertainty: highconvergesscott: mediumllama-cpp local-inference inference-economicsmattescalallama.cpp
What is this?
mattescala proposed a llama.cpp `--numa mirror` implementation that replicates model weights on each NUMA node, reporting roughly 64–137% higher decode throughput on a dual-EPYC system at the cost of doubling weight memory. The supplied snippets support the underlying rationale—dual-socket inference is constrained by NUMA locality and memory bandwidth, and per-node data replication can reduce remote-memory access—but they do not independently verify mattescala’s benchmark figures, broader hardware results, or whether the implementation was merged upstream.
Why it matters to Scott
The proposed NUMA mirroring concretely converges with Scott’s “Hardware-aware local inference” position by making memory placement and pressure explicit runtime policy, and it could materially change dual-socket CPU inference economics. It is not yet high relevance because the reported gains are unverified, upstream status is unknown, and the hits do not show Scott currently runs local models on comparable dual-socket hardware.
dev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:concept.cpu-inferenceradar:concept.memory-bandwidthradar:concept.llama-cppradar:concept.inference-economicsradar:picchio-llama-cpp-bottleneck-diagnostics
queries asked of Scott's wikis
- NUMA locality and memory-bandwidth limits in local LLM inference
- CPU inference economics versus GPU inference
- trading model-memory duplication for decode throughput
- dual-socket hardware strategy for local models
- llama.cpp performance tuning and benchmark methodology
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-03T20:30:10Z
Repeated checks have produced no independent benchmark, corrected baseline, or upstream action, leaving the contested single-system result without momentum. Close the episode as faded; a replication or maintainer decision would warrant reopening it.
2026-09-01T19:53:34Z
No replication, corrected benchmark, or upstream maintainer action has appeared after 48 hours; the claimed optimization remains an interesting but methodologically contested single-system result.
2026-08-30T19:43:17Z
A new technical objection suggests the reported gains may rely on a poorly configured `--numa distribute` baseline and insufficient cache/mmap controls. This weakens confidence in the benchmark comparison but neither disproves the implementation nor independently validates mirror performance.
2026-08-30T14:31:29Z
No independent benchmark, additional implementation, or upstream decision has appeared; the case remains a promising but hardware-specific author report rather than a validated llama.cpp optimization. The unchanged discussion warrants cooling while awaiting replication or maintainer action.
2026-08-30T14:29:19Z
grounded: converges/medium — The proposed NUMA mirroring concretely converges with Scott’s “Hardware-aware local inference” position by making memory placement and pressure explicit runtime
2026-08-30T14:27:12Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1w2hlm8 -> echo.github.ebd58e5e30 by wkgcass
2026-08-30T14:25:35Z
case created — This is a concrete implementation with quantified results and a clear memory-throughput tradeoff relevant to multi-socket local inference.
Decision trace
- 09-04 06:30expireRepeated checks have produced no independent benchmark, corrected baseline, or upstream action, leaving the contested single-system result without momentum. Close the episode as faded; a replication o
- 09-04 06:30alert_silentThe staleness check contains no new consequential evidence, so Scott gains nothing from an alert; any future independent benchmark or upstream action should be treated as a fresh delta.
- 09-04 06:30alert_routeThe staleness check contains no new consequential evidence, so Scott gains nothing from an alert; any future independent benchmark or upstream action should be treated as a fresh delta.
- 09-02 05:53repriceNo replication, corrected benchmark, or upstream maintainer action has appeared after 48 hours; the claimed optimization remains an interesting but methodologically contested single-system result.
- 09-02 05:53alert_silentThe staleness check adds no consequential evidence, and Scott can wait for controlled testing or an upstream decision.
- 09-02 05:53alert_routeThe staleness check adds no consequential evidence, and Scott can wait for controlled testing or an upstream decision.
- 08-31 05:43repriceA new technical objection suggests the reported gains may rely on a poorly configured `--numa distribute` baseline and insufficient cache/mmap controls. This weakens confidence in the benchmark compar
- 08-31 05:43alert_silentThe methodological critique is useful counterevidence, but it is an unverified comment rather than a replication, corrected benchmark, or maintainer decision. It can wait for the next briefing while t
- 08-31 05:43alert_routeThe methodological critique is useful counterevidence, but it is an unverified comment rather than a replication, corrected benchmark, or maintainer decision. It can wait for the next briefing while t
- 08-31 05:21sensor_dirtycomment_update
- 08-31 04:21sensor_dirtyengagement_update
- 08-31 00:31repriceNo independent benchmark, additional implementation, or upstream decision has appeared; the case remains a promising but hardware-specific author report rather than a validated llama.cpp optimization.
- 08-31 00:31alert_silentThis look adds no consequential delta beyond the already assessed implementation claim, so there is nothing new that Scott needs before the next briefing.
- 08-31 00:31alert_routeThis look adds no consequential delta beyond the already assessed implementation claim, so there is nothing new that Scott needs before the next briefing.
- 08-31 00:31alert_silentA concrete llama.cpp implementation reportedly mirrors model weights across NUMA nodes and shows large decode gains on one dual-socket EPYC system, establishing a substantive hardware-aware optimizati
- 08-31 00:31surface_candidateA concrete llama.cpp implementation reportedly mirrors model weights across NUMA nodes and shows large decode gains on one dual-socket EPYC system, establishing a substantive hardware-aware optimizati
- 08-31 00:31alert_routeA concrete llama.cpp implementation reportedly mirrors model weights across NUMA nodes and shows large decode gains on one dual-socket EPYC system, establishing a substantive hardware-aware optimizati
- 08-31 00:29groundThe proposed NUMA mirroring concretely converges with Scott’s “Hardware-aware local inference” position by making memory placement and pressure explicit runtime policy, and it could materially change
- 08-31 00:27promote_anchororigin walk conf 0.98
- 08-31 00:25createThis is a concrete implementation with quantified results and a clear memory-throughput tradeoff relevant to multi-socket local inference.