2026-10-11 16:37 UTC

am17an's merged llama.cpp PR #26610 adds RPC '-sm tensor' — tensor parallelism across networked machines (author-demonstrated on 2x DGX Sparks over RDMA, independently confirmed by ryan5rdx on 2x M3 Ultra, running a 284B-param MoE at 619 pp2048 / ~20 tg128) — and becomes a practical multi-machine local-inference pattern if outside users adopt it across their own clusters with replicated throughput; quiet disuse after the merge closes it.

state: watchingheat: mediumuncertainty: mediumconvergesscott: highllama-cpp distributed-inference tensor-parallelism local-inferenceam17anrgerganovggerganovryan5rdx

What is this?

llama.cpp PR #26610 (merged) adds `-sm tensor` — a tensor-parallelism mode for the existing RPC subsystem that splits individual model layers across networked machines. The author (am17an) demonstrated it on 2× NVIDIA DGX Spark (GB10) nodes connected via RDMA; ryan5rdx independently reproduced it on 2× Apple M3 Ultra machines running a 284B-parameter MoE at ~619 tokens/s prefill (pp2048) and ~20 tokens/s decode (tg128). This upgrades llama.cpp from single-machine multi-GPU to true multi-machine tensor-parallel inference, making very large models runnable on commodity clusters without layer-splitting latency penalties. The web snippets confirm the merge and the two hardware-class reproductions; they do not yet show wider community adoption or throughput replication on user-owned clusters.

Why it matters to Scott

llama.cpp upstream has independently merged cross-machine tensor parallelism via RPC (-sm tensor) — a concrete capability that directly embodies Scott's hardware-aware local inference framework (explicit runtime policy over accelerator placement, multi-machine patterns, hardware heterogeneity across RDMA/NVIDIA/Apple Silicon). Independent reproduction on two hardware classes (DGX Spark + M3 Ultra) validates the heterogeneity thesis. This isn't merely an example of his pattern; it's a new primitive in the dominant local runtime that could change what he deploys on gamepc and strengthens the sovereign-software-assurance and local-inference-economics positions he already holds.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:framework.sovereign-software-assuranceip:concept.ai-unit-economicsradar:amd-llama-cpp-prefill-speedupradar:aa-agentperf-local-benchmarkradar:adaptive-speculative-decoding-300-gpu
queries asked of Scott's wikis
  • distributed inference tensor parallelism local clusters
  • llama.cpp ggml RPC multi-machine inference patterns
  • open weights model sovereignty local inference infrastructure
  • agent memory distributed compute tensor parallel
  • hardware heterogeneity RDMA Apple Silicon NVIDIA inference
  • local inference economics consumer clusters vs cloud

Measured heat

now 0 pts/hpeak 9 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 1634h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-04 14:00⭐ origin echo-reconstructedRPC: add '-sm tensor' — tensor parallelism over llama.cpp RPC: 'This is on 2x Sparks connected via RDMA', async graph_compute, custom all_re
am17an on github (echo) · attributed from reddit.post.1wz5z7n
—
10-06 15:45first on r/LocalLLaMA · published · +1513.8hRPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp
jacek2023
—
10-06 15:45amplified on r/LocalLLaMA 👑reddit.post.1wz5z7n
jacek2023
peak 17 · 10 comments · 100% of case engagement
10-06 16:23our radar first saw it · +1514.4hdiscovery anchor: reddit.post.1wz5z7n—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditRPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-10-06T16:42:22.152631+00:00

[ggml-org](https://github.com/ggml-org) 
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public

- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
  24.1k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
   130k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)

# RPC: add `-sm tensor` - #26610

#26610

Merged

1/2

[am17an](https://github.com/am17an) merged 8 commits into

[master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom 

[rpc\_tensor](https://github.com/ggml-org/llama.cpp/tree/rpc_tensor)ggml-org/llama.cpp:rpc\_tensorCopy head branch name to clipboard

Oct 6, 2026

Merged

## [RPC: add `-sm tensor`](https://github.com/ggml-org/llama.cpp/pull/26610#top)#26610 [am17an](https://github.com/am17an) merged 8 commits into [master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [rpc\_tensor](https://github.com/ggml-org/llama.cpp/tree/rpc_tensor)ggml-org/llama.cpp:rpc\_tensorCopy head branch name to clipboard

## Conversation

[@am17an](https://github.com/am17an)

### @am17an **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issue-5067469637) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).

Copy link
 

Copy Markdown

Contributor

## Overview

Add RPC `-sm tensor`. This is on 2x Sparks connected via RDMA.

For RPC following changes are required:

1. async graph\_compute
2. custom `all_reduce`
3. graph uid cache like CUDA
4. `set_tensor_2d`, `get_tensor_2d`

Looking for feedback [@ggerganov](https://github.com/ggerganov) [@rgerganov](https://github.com/rgerganov)

| model | size | params | backend | ngl | n\_ubatch | sm | lm | test | t/s |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| deepseek4 ?B MXFP4 MoE | 145.63 GiB | 284.33 B | RPC | -1 | 2048 | tensor | dio | pp2048 | 619.36 ± 15.42 |
| deepseek4 ?B MXFP4 MoE | 145.63 GiB | 284.33 B | RPC | -1 | 2048 | tensor | dio | tg128 | 19.75 ± 0.39 |

## Additional information

```
  sequenceDiagram
      participant C as Client (RPC backend)
      box Server A (rank 0)
          participant A as A: rpc port
          participant Ac as A: comm port
      end
      box Server B (rank 1)
          participant Bc as B: comm port
          participant B as B: rpc port
      end

      Note over C,B: initialization
      C->>A: COMM_INIT (rank 0)
      C->>B: COMM_INIT (rank 1)
      Note over Ac: listen on comm port
      Bc-->>Ac: connect + caps negotiation (transport upgrade, e.g. RDMA)
      A->>C: response (ok)
      B->>C: response (ok)

      Note over C,B: for each subgraph (split at reduction boundaries)
      C->>A: GRAPH_COMPUTE (uid) [GRAPH_RECOMPUTE on reuse]
      C->>B: GRAPH_COMPUTE (uid)
      Note over A: compute subgraph async
      Note over B: compute subgraph async
  
      C->>A: COMM_ALLREDUCE (partial tensor) [fire and forget, no response]
      C->>B: COMM_ALLREDUCE (partial tensor)

      Note over A: sync pending graph<br/>cast F32→BF16 if ne ≥ 32768<br/>copy partial to send buffer
      Note over B: sync pending graph<br/>cast F32→BF16 if ne ≥ 32768<br/>copy partial to send buffer

      Ac-->>Bc: partial A (rank 0 sends first)
      Bc-->>Ac: partial B (rank 1 receives first)

      Note over A: upload partial B<br/>dst = dst + partial B (async ADD)
      Note over B: upload partial A<br/>dst = dst + partial A (async ADD)

      Note over C,B: read back the output
      C->>A: GET_TENSOR (output)
      Note over A: sync all backends
      A->>C: data
```

 Loading

## Requirements

- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: Yes, paired with codex and claude

---

Stack created with [GitHub Stacks CLI](https://github.com/github/gh-stack) • [Give Feedback 💬](https://gh.io/stacks-feedback)

[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
labels
[Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#event-28989160580)

[@am17an](https://github.com/am17an)
[am17an](https://github.com/am17an)
mentioned this pull request
[Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#ref-pullrequest-5039609470)

[rpc: support apple RDMA as an RPC transport
#26421](https://github.com/ggml-org/llama.cpp/pull/26421)

Merged

[@ryan5rdx](https://github.com/ryan5rdx)

### **[ryan5rdx](https://github.com/ryan5rdx)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188516731) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).

Copy link
 

Copy Markdown

Contributor

|  |
| --- |
| confirmed working on metal(RDMA 2 x M3 Ultra with [#26421](https://github.com/ggml-org/llama.cpp/pull/26421)), testing same model (ds4 MXFP4):  tg2048: 9.61 t/s  pp2048: 166.05 t/s  it does however break with dspark applied because some ops it depends on appear to not be supported with TP (add across the split), so this is without any mtp. |

[@am17an](https://github.com/am17an)

### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188600939)

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| [@ryan5rdx](https://github.com/ryan5rdx) yeah I know it doesn't work with dspark, but what do you get when for just `-sm layer` as tg2048? It seems kinda low based on the numbers you posted in the other PR |

[@ggerganov](https://github.com/ggerganov)

### **[ggerganov](https://github.com/ggerganov)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188695004)

Copy link
 

Copy Markdown

Member

|  |
| --- |
| This was supposed on top of the [#25860](https://github.com/ggml-org/llama.cpp/pull/25860) but that didn't happen.  Shouldn't you stack it on top of [#26490](https://github.com/ggml-org/llama.cpp/pull/26490), instead of [#25860](https://github.com/ggml-org/llama.cpp/pull/25860)? |

[@am17an](https://github.com/am17an)

### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188720492) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| Sorry it is on top of the that  [image](https://private-user-images.githubusercontent.com/2929750/631596765-ea104131-6201-4891-8a05-c91f93b07538.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3OTEzMDUyMzksIm5iZiI6MTc5MTMwNDkzOSwicGF0aCI6Ii8yOTI5NzUwLzYzMTU5Njc2NS1lYTEwNDEzMS02MjAxLTQ4OTEtOGEwNS1jOTFmOTNiMDc1MzgucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MTAwNiUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjEwMDZUMTY0MjE5WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NjgwMzQ2ZmJjN2RkYTQ0OThmODFiNDI5NWZlNzQ4MTAyNmEzNDAyMjljYTk2ZTgwNjhjZTVkMWFmZmNhZDY3ZSZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QmcmVzcG9uc2UtY29udGVudC10eXBlPWltYWdlJTJGcG5nIn0.30JyCi75P9QVU9F4bzCqE8D1G6NPDQ_p97ywDXS9QVc) |

[@ryan5rdx](https://github.com/ryan5rdx)

### **[ryan5rdx](https://github.com/ryan5rdx)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188724149)

Copy link
 

Copy Markdown

Contributor

|  |
| --- |
| [@ryan5rdx](https://github.com/ryan5rdx) yeah I know it doesn't work with dspark, but what do you get when for just `-sm layer` as tg2048? It seems kinda low based on the numbers you posted in the other PR  command for reference, let me know happy to test an alternate config:  ``` ./bin/llama-server -m ~/Downloads/DeepSeek-V4-Flash-0731-MXFP4.gguf -c 1048576 --reasoning on --rpc 192.168.0.13:50052   -np 1 --reasoning-preserve -ub 2048 -b 4096  --no-mmap -ngl 999 -fa on  --fit off -ts 1,1 --host 0.0.0.0 -kvu -sm tensor ```  and yup just confirmed - with `-sm layer` numbers align with what I have in [#26421](https://github.com/ggml-org/llama.cpp/pull/26421):  tg2048: 23.05 t/s  pp2048: 270.3 t/s |

[@am17an](https://github.com/am17an)

### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188777367)

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| [@ryan5rdx](https://github.com/ryan5rdx) try using two rpc servers, one on each machine and connect via  `llama-server --rpc <ip_1, ip_2> --device RPC0, RPC1` |

[@Kononnable](https://github.com/Kononnable)

### **[Kononnable](https://github.com/Kononnable)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188788627)

Copy link
 

Copy Markdown

Contributor

|  |
| --- |
| It might be worth to add `-sm layer` results to bench results table - just to have performance baseline in a single place.  I don't know if this would be observable with RDNA, but sometimes manually moving tensors instead of standard `-sm layer` can increase performance when tcp/ip is used (just a sidenote for 'base' performance).  `-sm layer -ts 0,1 -ot 'blk\.[0-1][0-9]?\.ffn_(up|down|gate|gate_up)_(ch|)exps=RPC0[127.0.0.1:50052]'` |

[@ggerganov](https://github.com/ggerganov)

### **[ggerganov](https://github.com/ggerganov)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188798792)

Copy link
 

Copy Markdown

Member

|  |
| --- |
| [@am17an](https://github.com/am17an) The shared commits in the 2 branches differ:   - `dsv4-sm-tensor`:   - [813cfe7](https://github.com/ggml-org/llama.cpp/commit/813cfe7f2cdf5177f17afd6eec9cbdc047d7b89f)   - [59c2539](https://github.com/ggml-org/llama.cpp/commit/59c2539009622a933b4a5cfee58ab0eb5e70d62b) - `rpc_tensor`:   - [38b3cdc](https://github.com/ggml-org/llama.cpp/commit/38b3cdc47447e5cb9bd992fd2526b38f807e22eb)   - [80b2f51](https://github.com/ggml-org/llama.cpp/commit/80b2f518cb4630100f97cb6f219d8210a97b845e)   Likely you've made changes to the `dsv4-sm-tensor` branch after you created the `rpc_tensor` branch. That's why the PR stack does not work. |

[@am17an](https://github.com/am17an)

### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188813222)

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| Yeah I messed up, I think [#26490](https://github.com/ggml-org/llama.cpp/pull/26490) should be okay to merge though |

[@ggerganov](https://github.com/ggerganov)

### **[ggerganov](https://github.com/ggerganov)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188841248)

Copy link
 

Copy Markdown

Member

|  |
| --- |
| Yeah I messed up, I think [#26490](https://github.com/ggml-org/llama.cpp/pull/26490) should be okay to merge though  Don't we want to fix the DSpark support first? |

[@am17an](https://github.com/am17an)

### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188866941) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).

Copy link
 

Copy Markdown

Contributor

Author

|  |
| --- |
| The support is broken over RPC I think(i.e. this PR), 
jacek20231710
🟧 echo.github ⭐RPC: add '-sm tensor' — tensor parallelism over llama.cpp RPC: 'This is on 2x Sparks connected via RDMA', async graph_compute, custom all_ream17an——

Interpretation history

Decision trace