Retrieved article excerpt
Open article · Retrieved 2026-10-06T16:42:22.152631+00:00
[ggml-org](https://github.com/ggml-org)
/
**[llama.cpp](https://github.com/ggml-org/llama.cpp)**
Public
- [Notifications](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp) You must be signed in to change notification settings
- [Fork
24.1k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
- [Star
130k](https://github.com/login?return_to=%2Fggml-org%2Fllama.cpp)
# RPC: add `-sm tensor` - #26610
#26610
Merged
1/2
[am17an](https://github.com/am17an) merged 8 commits into
[master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom
[rpc\_tensor](https://github.com/ggml-org/llama.cpp/tree/rpc_tensor)ggml-org/llama.cpp:rpc\_tensorCopy head branch name to clipboard
Oct 6, 2026
Merged
## [RPC: add `-sm tensor`](https://github.com/ggml-org/llama.cpp/pull/26610#top)#26610 [am17an](https://github.com/am17an) merged 8 commits into [master](https://github.com/ggml-org/llama.cpp/tree/master)ggml-org/llama.cpp:masterfrom [rpc\_tensor](https://github.com/ggml-org/llama.cpp/tree/rpc_tensor)ggml-org/llama.cpp:rpc\_tensorCopy head branch name to clipboard
## Conversation
[@am17an](https://github.com/am17an)
### @am17an **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issue-5067469637) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).
Copy link
Copy Markdown
Contributor
## Overview
Add RPC `-sm tensor`. This is on 2x Sparks connected via RDMA.
For RPC following changes are required:
1. async graph\_compute
2. custom `all_reduce`
3. graph uid cache like CUDA
4. `set_tensor_2d`, `get_tensor_2d`
Looking for feedback [@ggerganov](https://github.com/ggerganov) [@rgerganov](https://github.com/rgerganov)
| model | size | params | backend | ngl | n\_ubatch | sm | lm | test | t/s |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| deepseek4 ?B MXFP4 MoE | 145.63 GiB | 284.33 B | RPC | -1 | 2048 | tensor | dio | pp2048 | 619.36 ± 15.42 |
| deepseek4 ?B MXFP4 MoE | 145.63 GiB | 284.33 B | RPC | -1 | 2048 | tensor | dio | tg128 | 19.75 ± 0.39 |
## Additional information
```
sequenceDiagram
participant C as Client (RPC backend)
box Server A (rank 0)
participant A as A: rpc port
participant Ac as A: comm port
end
box Server B (rank 1)
participant Bc as B: comm port
participant B as B: rpc port
end
Note over C,B: initialization
C->>A: COMM_INIT (rank 0)
C->>B: COMM_INIT (rank 1)
Note over Ac: listen on comm port
Bc-->>Ac: connect + caps negotiation (transport upgrade, e.g. RDMA)
A->>C: response (ok)
B->>C: response (ok)
Note over C,B: for each subgraph (split at reduction boundaries)
C->>A: GRAPH_COMPUTE (uid) [GRAPH_RECOMPUTE on reuse]
C->>B: GRAPH_COMPUTE (uid)
Note over A: compute subgraph async
Note over B: compute subgraph async
C->>A: COMM_ALLREDUCE (partial tensor) [fire and forget, no response]
C->>B: COMM_ALLREDUCE (partial tensor)
Note over A: sync pending graph<br/>cast F32→BF16 if ne ≥ 32768<br/>copy partial to send buffer
Note over B: sync pending graph<br/>cast F32→BF16 if ne ≥ 32768<br/>copy partial to send buffer
Ac-->>Bc: partial A (rank 0 sends first)
Bc-->>Ac: partial B (rank 1 receives first)
Note over A: upload partial B<br/>dst = dst + partial B (async ADD)
Note over B: upload partial A<br/>dst = dst + partial A (async ADD)
Note over C,B: read back the output
C->>A: GET_TENSOR (output)
Note over A: sync all backends
A->>C: data
```
Loading
## Requirements
- I have read and agree with the [contributing guidelines](https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md)
- AI usage disclosure: Yes, paired with codex and claude
---
Stack created with [GitHub Stacks CLI](https://github.com/github/gh-stack) • [Give Feedback 💬](https://gh.io/stacks-feedback)
[@github-actions](https://github.com/apps/github-actions)
[github-actions](https://github.com/apps/github-actions)
Bot
added
[testing](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Atesting)
Everything test related
[ggml](https://github.com/ggml-org/llama.cpp/issues?q=state%3Aopen%20label%3Aggml)
changes relating to the ggml tensor library for machine learning
labels
[Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#event-28989160580)
[@am17an](https://github.com/am17an)
[am17an](https://github.com/am17an)
mentioned this pull request
[Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#ref-pullrequest-5039609470)
[rpc: support apple RDMA as an RPC transport
#26421](https://github.com/ggml-org/llama.cpp/pull/26421)
Merged
[@ryan5rdx](https://github.com/ryan5rdx)
### **[ryan5rdx](https://github.com/ryan5rdx)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188516731) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).
Copy link
Copy Markdown
Contributor
| |
| --- |
| confirmed working on metal(RDMA 2 x M3 Ultra with [#26421](https://github.com/ggml-org/llama.cpp/pull/26421)), testing same model (ds4 MXFP4): tg2048: 9.61 t/s pp2048: 166.05 t/s it does however break with dspark applied because some ops it depends on appear to not be supported with TP (add across the split), so this is without any mtp. |
[@am17an](https://github.com/am17an)
### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188600939)
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| [@ryan5rdx](https://github.com/ryan5rdx) yeah I know it doesn't work with dspark, but what do you get when for just `-sm layer` as tg2048? It seems kinda low based on the numbers you posted in the other PR |
[@ggerganov](https://github.com/ggerganov)
### **[ggerganov](https://github.com/ggerganov)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188695004)
Copy link
Copy Markdown
Member
| |
| --- |
| This was supposed on top of the [#25860](https://github.com/ggml-org/llama.cpp/pull/25860) but that didn't happen. Shouldn't you stack it on top of [#26490](https://github.com/ggml-org/llama.cpp/pull/26490), instead of [#25860](https://github.com/ggml-org/llama.cpp/pull/25860)? |
[@am17an](https://github.com/am17an)
### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188720492) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| Sorry it is on top of the that [image](https://private-user-images.githubusercontent.com/2929750/631596765-ea104131-6201-4891-8a05-c91f93b07538.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3OTEzMDUyMzksIm5iZiI6MTc5MTMwNDkzOSwicGF0aCI6Ii8yOTI5NzUwLzYzMTU5Njc2NS1lYTEwNDEzMS02MjAxLTQ4OTEtOGEwNS1jOTFmOTNiMDc1MzgucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MTAwNiUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjEwMDZUMTY0MjE5WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NjgwMzQ2ZmJjN2RkYTQ0OThmODFiNDI5NWZlNzQ4MTAyNmEzNDAyMjljYTk2ZTgwNjhjZTVkMWFmZmNhZDY3ZSZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QmcmVzcG9uc2UtY29udGVudC10eXBlPWltYWdlJTJGcG5nIn0.30JyCi75P9QVU9F4bzCqE8D1G6NPDQ_p97ywDXS9QVc) |
[@ryan5rdx](https://github.com/ryan5rdx)
### **[ryan5rdx](https://github.com/ryan5rdx)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188724149)
Copy link
Copy Markdown
Contributor
| |
| --- |
| [@ryan5rdx](https://github.com/ryan5rdx) yeah I know it doesn't work with dspark, but what do you get when for just `-sm layer` as tg2048? It seems kinda low based on the numbers you posted in the other PR command for reference, let me know happy to test an alternate config: ``` ./bin/llama-server -m ~/Downloads/DeepSeek-V4-Flash-0731-MXFP4.gguf -c 1048576 --reasoning on --rpc 192.168.0.13:50052 -np 1 --reasoning-preserve -ub 2048 -b 4096 --no-mmap -ngl 999 -fa on --fit off -ts 1,1 --host 0.0.0.0 -kvu -sm tensor ``` and yup just confirmed - with `-sm layer` numbers align with what I have in [#26421](https://github.com/ggml-org/llama.cpp/pull/26421): tg2048: 23.05 t/s pp2048: 270.3 t/s |
[@am17an](https://github.com/am17an)
### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188777367)
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| [@ryan5rdx](https://github.com/ryan5rdx) try using two rpc servers, one on each machine and connect via `llama-server --rpc <ip_1, ip_2> --device RPC0, RPC1` |
[@Kononnable](https://github.com/Kononnable)
### **[Kononnable](https://github.com/Kononnable)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188788627)
Copy link
Copy Markdown
Contributor
| |
| --- |
| It might be worth to add `-sm layer` results to bench results table - just to have performance baseline in a single place. I don't know if this would be observable with RDNA, but sometimes manually moving tensors instead of standard `-sm layer` can increase performance when tcp/ip is used (just a sidenote for 'base' performance). `-sm layer -ts 0,1 -ot 'blk\.[0-1][0-9]?\.ffn_(up|down|gate|gate_up)_(ch|)exps=RPC0[127.0.0.1:50052]'` |
[@ggerganov](https://github.com/ggerganov)
### **[ggerganov](https://github.com/ggerganov)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188798792)
Copy link
Copy Markdown
Member
| |
| --- |
| [@am17an](https://github.com/am17an) The shared commits in the 2 branches differ: - `dsv4-sm-tensor`: - [813cfe7](https://github.com/ggml-org/llama.cpp/commit/813cfe7f2cdf5177f17afd6eec9cbdc047d7b89f) - [59c2539](https://github.com/ggml-org/llama.cpp/commit/59c2539009622a933b4a5cfee58ab0eb5e70d62b) - `rpc_tensor`: - [38b3cdc](https://github.com/ggml-org/llama.cpp/commit/38b3cdc47447e5cb9bd992fd2526b38f807e22eb) - [80b2f51](https://github.com/ggml-org/llama.cpp/commit/80b2f518cb4630100f97cb6f219d8210a97b845e) Likely you've made changes to the `dsv4-sm-tensor` branch after you created the `rpc_tensor` branch. That's why the PR stack does not work. |
[@am17an](https://github.com/am17an)
### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188813222)
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| Yeah I messed up, I think [#26490](https://github.com/ggml-org/llama.cpp/pull/26490) should be okay to merge though |
[@ggerganov](https://github.com/ggerganov)
### **[ggerganov](https://github.com/ggerganov)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188841248)
Copy link
Copy Markdown
Member
| |
| --- |
| Yeah I messed up, I think [#26490](https://github.com/ggml-org/llama.cpp/pull/26490) should be okay to merge though Don't we want to fix the DSpark support first? |
[@am17an](https://github.com/am17an)
### **[am17an](https://github.com/am17an)** commented [Aug 5, 2026](https://github.com/ggml-org/llama.cpp/pull/26610#issuecomment-5188866941) • edited Loading Uh oh! There was an error while loading. [Please reload this page](https://github.com/ggml-org/llama.cpp/pull/26610).
Copy link
Copy Markdown
Contributor
Author
| |
| --- |
| The support is broken over RPC I think(i.e. this PR),