2026-10-11 17:18 UTC

pd-bridge's maintainer claims its released DeepSeek-V4-Flash bridge combines NVIDIA prefill with Apple Silicon decode over 10GbE to reduce measured cold long-prompt latency by 1.5โ€“3.7 times versus Mac-only serving while preserving decode throughput, potentially accelerating mixed-hardware local inference.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumdisaggregated-inference local-inference inference-economicschadhurley25075-pngpd-bridge

What is this?

pd-bridge is an Apache-2.0 reference implementation by GitHub user chadhurley25075-png that splits serving of DeepSeek-V4-Flash (a 284B-total/13B-active MoE released 2026-04-24 under MIT) across vendors: prefill runs on two NVIDIA DGX Sparks under vLLM FP8, decode on a Mac Studio M3 Ultra under oMLX MXFP4. Rather than transferring a KV cache between incompatible formats, it recomputes the decoder's finished cache on the prefill machine using the decoder's own weights and writes it directly into the decoder's prefix-cache store, with a verdict header intended to block silent native fallbacks. The maintainer's README reports cold long-prompt latency falling 1.5x at ~25K tokens to 3.7x at ~241K tokens versus Mac-only serving with decode unchanged at 23-25 tok/s, and later updates describe 700K-token prompts bridging end-to-end in ~11.5 minutes plus a RoCE/RDMA fabric replacing the original plain 10GbE link. Prefill/decode disaggregation itself is established datacenter practice (now default in vLLM/SGLang/Dynamo), but pd-bridge's cross-vendor prosumer-hardware variant sits beside several independent heterogeneous-PD projects, and every performance figure remains maintainer-reported with no independent validation in the available sources.

Why it matters to Scott

Two independent implementations now embody mixed-hardware prefill acceleration โ€” pd-bridge's NVIDIA-prefill/Mac-oMLX-decode split plus the consumer-scale Mac+iPhone layer demo โ€” converging with Scott's hardware-aware-local-inference position: accelerator placement as explicit runtime policy, here extended across machine boundaries, which is a dated-receipts shape for that page. The warm-turns-bypass finding also restates his prefix-caching-economics boundary at the hardware layer (cold prompts are exactly where the bridge pays), and the decode side runs oMLX inside his own MLX-serving territory. Claims remain maintainer-measured only, so the action this warrants is evaluation on his own gear (gamepc CUDA box beside a Mac MLX runtime), not adoption.
dev:concept.hardware-aware-local-inferencedev:technology.mlxdev:project.gamepcip:concept.prefix-caching-economicsradar:concept.disaggregated-inferenceradar:inferpd-shared-prefix-disaggregationradar:concept.kv-cacheradar:concept.distributed-inferenceradar:exo-heterogeneous-device-inferenceradar:swarmllm-browser-distributed-inference
queries asked of Scott's wikis
  • MLX oMLX Mac Studio local model serving
  • prefill decode disaggregation KV cache transfer
  • long-context prefill cost RAG retrieval tradeoff
  • local vs cloud inference economics cost per token
  • DGX Spark GB10 desktop inference hardware
  • heterogeneous mixed-device model partitioning

Measured heat

now 0 pts/hpeak 134 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 610h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-16 06:25 (minted)โญ origin echo-reconstructedThe reference implementation reconstructs the decoder's cache on NVIDIA hardware for oMLX consumption over 10GbE, reporting cold-prompt spee
chadhurley25075-png on github (echo) ยท attributed from hn.story.49722597 ยท published time unknown
โ€”
09-16 05:57first on hacker news ยท published ยท lag ?Prefill a 284B model on Nvidia. Decode it on Apple Silicon. Over plain 10GbE
jnaina
โ€”
10-02 16:59first on r/LocalLLaMA ยท published ยท lag ?I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29โ€“44% faster & my holds part of the CTX window.
StayLameBro
โ€”
09-16 05:57amplified on hacker newshn.story.49722597
jnaina
peak 3 ยท 1 comments ยท 0% of case engagement
10-02 16:59amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wvz1ex
StayLameBro
peak 2337 ยท 338 comments ยท 100% of case engagement
09-16 06:21our radar first saw it ยท lag ?discovery anchor: hn.story.49722597โ€”
pace: p32 vs 1032 stories at the 336h mark (now 610h old) โ€” ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnPrefill a 284B model on Nvidia. Decode it on Apple Silicon. Over plain 10GbE
Retrieved article excerpt

Open article ยท Retrieved 2026-09-16T06:22:06.923616+00:00

# pd-bridge โ€” heterogeneous prefill/decode for DeepSeek-V4-Flash

[License: Apache-2.0](https://github.com/chadhurley25075-png/pd-bridge/blob/main/LICENSE)
[status: reference implementation](https://github.com/chadhurley25075-png/pd-bridge#status-honestly)
[model: DeepSeek-V4-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash)
[prefill: vLLM / CUDA](https://github.com/chadhurley25075-png/pd-bridge#running-it)
[decode: oMLX / Metal](https://github.com/chadhurley25075-png/pd-bridge#running-it)

**Prefill a 284B model on NVIDIA. Decode it on Apple Silicon. Over plain 10GbE.**

[How the bridge works](https://github.com/chadhurley25075-png/pd-bridge/blob/main/docs/architecture.svg)

Two production inference engines that share no cache format, no framework, no vendor and no
quantization, serving one request together. DeepSeek-V4-Flash is 284B total / 13B active,
256 routed experts, MLA + sparse attention, 149 GB resident on the prefill side and 156 GB on
the decode side:

- **Prefill:** 2ร— NVIDIA DGX Spark (GB10), vLLM TP2, official `deepseek-ai/DeepSeek-V4-Flash` **FP8**
- **Decode:** 1ร— Mac Studio M3 Ultra, oMLX, `DV4-Flash-MXFP4-MLX` (**MXFP4**)
- **Link:** ordinary 10 gigabit Ethernet. No RDMA, no Thunderbolt.

```
cold prompt        Mac Studio alone    Sparks prefill -> Mac decode
   ~25K tokens          42.6 s               28.2 s     1.5x
   ~82K tokens         205.8 s               72.9 s     2.8x
  ~105K tokens         245.6 s               75.5 s     3.3x
  ~241K tokens         732.3 s              200.3 s     3.7x
decode rate unchanged (23-25 tok/s both ways); warm turns bypass the bridge (4.9 s / 19.3 s at 241K)
```

Every row verdict-checked, 2026-09-06. The bridged leg scores **5/5 on the judged quality eval**,
same as native.

**Since then the served window went 262,144 -> 2,097,152 and the ceiling moved with it.**
A **700,630-token cold prompt now bridges end to end in 11 minutes 26 seconds** โ€” `verdict complete`,
342 blocks, 1,117 tok/s of prefill. That is **2.9x past the largest prompt in the table above**, and
the pair has accepted over a million.

```
cold prompt     engine   tok/s    end-to-end   blocks   verdict
   192,099      119.2 s   1,611      143.1 s      93    complete
   385,838      288.7 s   1,337      466.6 s     188    complete
   677,069      596.8 s   1,134      655.9 s     330    complete
   700,630      627.0 s   1,117      685.5 s     342    complete
   709,055      635.4 s   1,116      694.6 s     345    partial 706,560/709,055 (99.6%)
   988,487    1,037.8 s     952    1,142.5 s     376    partial 770,048/988,487 (78%)
 1,006,172    1,070.5 s     940    1,171.7 s     346    partial 708,608/1,006,172 (70%)
```

All 2026-09-08, one prefill pair, same 10GbE. `partial` is not a failure: past roughly 772K tokens the
capture crosses the prefill box's free-memory floor and **seals a valid contiguous prefix** instead of
dying; the decoder mounts what arrived and natively prefills the remainder. See *Known limits*.

**We have not run a Mac-alone control at these sizes**, so those rows carry no ratio โ€” they are what
the bridge does, not a claimed speedup. The measured ratios stop at 241K and are shown above.

Full numbers and methodology: [RESULTS.md](https://github.com/chadhurley25075-png/pd-bridge/blob/main/RESULTS.md) ยท [bench/BENCHMARK-PROTOCOL.md](https://github.com/chadhurley25075-png/pd-bridge/blob/main/bench/BENCHMARK-PROTOCOL.md)

---

## The idea

Prefill/decode disaggregation is well established, and so is the hardware argument for it: prefill is
compute-bound, decode is memory-bandwidth-bound, so run each phase where it is cheapest. Existing
systems do this by **transferring the KV cache** from the prefill worker to the decode worker.

That requires both ends to agree on a cache format. Ours never can. One side is CUDA/vLLM with an
FP8 paged cache; the other is Metal/MLX with its own block layout. Worse, DeepSeek-V4-Flash does not
have "a KV cache" โ€” each layer carries a rotating 128-token window, a compressor pool (ratio 4 with
overlap carry, and ratio 128), and an indexer pool, all with layer-dependent RoPE.

**So we don't transfer a cache. We compute the decoder's finished cache on the prefill machine,
using the decoder's own weights, and write it straight into the decoder's prefix-cache store.**

The prefill engine already computes the exact tensor those pools are a pure function of โ€” the
attention input. A hook takes it there, applies the *Mac's* projection and pooling math on the GPU,
and emits the finished pools. The decode side assembles them into MLX cache objects and hands them
to oMLX's own block writer. oMLX then sees a normal prefix-cache hit and only decodes. Neither
engine is modified in its hot path; the decoder does not know a bridge exists.

Payload: **~10 KB per token** โ€” 0.80 GB for an 81K-token prompt, pulled in 1.08 s. The network
stopped being the bottleneck; the prefill engine is now 65% of wall time, which is where you want it.

### Why the reconstruction is trustworthy

The whole design rests on the pooled tensors being *the same tensors* the decoder would have
computed. That is tested, not assumed:

| check | result |
| --- | --- |
| Cache arrays rebuilt on the Mac vs. a full native forward | **313/313 bit-exact** |
| Blocks written by the bridge vs. blocks oMLX writes itself | **11/11 identical** (only the `created_at` stamp differs) |
| Torch pooling port vs. MLX ground truth (T=23,217) | projections, window, carries **bit-exact**; pooled tensors 99.95โ€“99.96% identical, worst delta **one bf16 ulp** |
| In-container hook selftest (chunked == one-shot) | **52/52** |
| Needle retrieval through a fully reconstructed 81K cache | correct on every benchmark run |

`studio/verify_blocks.py`, `studio/pd_diff_state.py`, `spark/pd_pool_validate.py` and
`spark/pd_pool_selftest.py` reproduce these. Compare tensors, never file hashes โ€” `created_at` means
a bridge-written block can never be byte-identical as a *file*.

---

## Known limits โ€” read this before you run it

### The memory floor, and why big prompts come back `partial` (2026-09-08)

The capture lives in the prefill box's memory for the whole request and costs about **11.8 KB of
unified memory per token** on rank 0 only (rank 1 stays flat โ€” it does its half of the attention and
holds no capture). A GB10 has ONE 121 GB pool shared by weights, the vLLM KV arena and everything
else, so steady-state free memory with the model up is ~22 GB.

Measured: the capture crosses a 5.0 GB free-memory floor at roughly **772,000 tokens**. Past that the
hook **seals a valid contiguous prefix `[0,T)` and reports it** rather than dying โ€” the decoder mounts
what arrived and natively prefills the tail. That is the `partial` verdict in the table above. A
700,630-token wake cleared the floor with **100 MB to spare**; a 1,006,172-token wake sealed at 70%.

This replaced a real failure. Before the fix, `_finish()` copied every layer to the host **without
freeing the layers it had already written**, so the full device capture and the growing host copies
were alive at once. At 1,021,199 tokens it died after 11 of 43 layers and drove the box into swap
thrash โ€” no sshd even over a 200G fabric. It needed a physical power button. The hook now releases
each layer as its file lands and checks `MemAvailable` every 64 layer-chunks.

**If you are memory-tight, this is your limit, not the window.** The window is a config number; the
floor is physics on your box. Measure `MemAvailable` during a long prefill before trusting either.

**The ~82K-token ceiling was a real bug. It is fixed.** (History kept because the failure mode is
instructive and the arithmetic still matters on smaller machines.)

`omlx_block_writer` used to hold one *materialised cumulative* cache snapshot per 2048-token
boundary until `finalize()`: peak memory grew **quadratically** with prompt length (~`N(N+1)/2`
blocks' worth of arrays โ€” ~16 GB at 39 boundaries, ~23 GB at 47, which exhausted a 256 GB M3 Ultra
holding a 156 GB model and wrote zero blocks at 97,848 tokens).

The writer now **streams**: each boundary is stored through oMLX's own pipeline and released the
moment it is snapshotted (`begin_stream`/`store_boundary`; `finalize` drains and verifies). Peak
memory is ONE boundary snapshot (~20 MB ร— boundary index / N โ€” tens of MB, not tens of GB). It is
validated against a real oMLX reference block (synthetic layout match) and live at 52 boundaries /
109,085 tokens โ€” see RESULTS.md. `PD_STREAM_BOUNDARIES=0` restores the batched path for comparison.

`PD_MAX_BRIDGE_TOKENS` remains as a configurable envelope guard, not a bug workaround: raise it to
your machine's measured headroom.

Two related behaviours worth knowing:

- **The fallback works, and it can no longer lie.** When a bridge fails the reply still comes back
  correct โ€” the decoder serves natively. Since the bench4 autopsy (docs/FINDING-bench4-cold-fallback.md),
  every response carries an `X-PD-Bridge` verdict (`complete` / `partial B/T` / declined with reason),
  and `bench_cold.py` records it โ€” a silent native fallback can never again enter a results table
  as a bridged number.
- **The "cold-start variance" at 20K was not variance.** The 25.5 s vs 55.4 s spread was the capture
  hook flushing *mid-request* during chunked prefill (see the FINDING): the 55 s runs were native
  fallbacks wearing a bridge label. The hook now guards its idle flush with a CUDA-event query and a
  chunk-alignment check, and the front validates every capture manifest before trusting it.

## Status, honestly

This is a **reference implementation, not a library.** It is pinned hard and it is young.

- **One model.** DeepSeek-V4-Flash. The pooling math is specific to its sparse attention.
- **Pinned stacks.** oMLX 0.6.4; vLLM 0.21.1rc1 with the DeepSeek-V4 plugin (sparkrun image).
- **It monkey-patches private internals of both engines** โ€” a `sitecustomize` hook onto
  `DeepseekV4MultiHeadLatentAttentionWrapper.attention_impl` on the vLLM side, and a
  filesystem-fallback patch to oMLX's `PagedSSDCacheIndex` on the MLX side (oMLX indexes SSD blocks
  at model load only, so externally written blocks are otherwise invisible). **Expect this to break
  when either project moves.**
- **The judged quality eval is five questions on one document.** The bridged leg scores 5/5 on it,
  twice (once from a fresh cold v3 bridge), same as native. Prefill runs FP8 weights and decode runs
  MXFP4, so bridged output is *not* token-identical to native; it is factually faithful on what we
  checked, which is a smaller claim than "equivalent".
- **The flush signal was unreliable until 2026-09-06 17:52 โ€” fixed.** The hook's watcher ran in three
  processes and two of them deleted the signal before the capturing worker saw it (~1 in 3 hit rate). Hook v5
  fixes it; captures now close 0.6โ€“0.9 s after the engine returns (4.8 s at 236K, which is the block write).
  Autopsy: `docs/FINDING-flush-signal-three-watchers.md`. Unit test: `spark/test_flush_decision.py`.
- **The front door is threaded, with one caveat.** HTTP handlers run in threads (health, model list and
  oMLX passthrough answer immediately, and concurrent decodes overlap because the decoder batches them),
  but every `bridge()` call โ€” the MLX cache assembly โ€” is marshalled to the main thread and runs one at a
  time, because MLX streams are thread-local and the model lives there. Measured 2026-09-06: a short
  request completed in 47 s while an 86K-token cold bridge was in flight, instead of waiting it out.
  Two clients do slow each other down; they no longer block each other.
- Only cold, long prompts benefit. Warm turns bypass the bridge by design and are served natively.

**The transferable idea is bigger than this code:** when two engines cannot share a cache format,
compute the *consumer's* finished cache on the *producer*, using the consumer's weights. That
generalizes past this model and this hardware, and it is the part worth stealin
jnaina31
๐ŸŸง echo.github โญThe reference implementation reconstructs the decoder's cache on NVIDIA hardware for oMLX consumption over 10GbE, reporting cold-prompt speechadhurley25075-pngโ€”โ€”
๐ŸŸ  redditI made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29โ€“44% faster & my holds part of the CTX window.
LocalLLaMA
StayLameBro2337338

Interpretation history

Decision trace