Retrieved article excerpt
Open article ยท Retrieved 2026-09-16T06:22:06.923616+00:00
# pd-bridge โ heterogeneous prefill/decode for DeepSeek-V4-Flash
[License: Apache-2.0](https://github.com/chadhurley25075-png/pd-bridge/blob/main/LICENSE)
[status: reference implementation](https://github.com/chadhurley25075-png/pd-bridge#status-honestly)
[model: DeepSeek-V4-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash)
[prefill: vLLM / CUDA](https://github.com/chadhurley25075-png/pd-bridge#running-it)
[decode: oMLX / Metal](https://github.com/chadhurley25075-png/pd-bridge#running-it)
**Prefill a 284B model on NVIDIA. Decode it on Apple Silicon. Over plain 10GbE.**
[How the bridge works](https://github.com/chadhurley25075-png/pd-bridge/blob/main/docs/architecture.svg)
Two production inference engines that share no cache format, no framework, no vendor and no
quantization, serving one request together. DeepSeek-V4-Flash is 284B total / 13B active,
256 routed experts, MLA + sparse attention, 149 GB resident on the prefill side and 156 GB on
the decode side:
- **Prefill:** 2ร NVIDIA DGX Spark (GB10), vLLM TP2, official `deepseek-ai/DeepSeek-V4-Flash` **FP8**
- **Decode:** 1ร Mac Studio M3 Ultra, oMLX, `DV4-Flash-MXFP4-MLX` (**MXFP4**)
- **Link:** ordinary 10 gigabit Ethernet. No RDMA, no Thunderbolt.
```
cold prompt Mac Studio alone Sparks prefill -> Mac decode
~25K tokens 42.6 s 28.2 s 1.5x
~82K tokens 205.8 s 72.9 s 2.8x
~105K tokens 245.6 s 75.5 s 3.3x
~241K tokens 732.3 s 200.3 s 3.7x
decode rate unchanged (23-25 tok/s both ways); warm turns bypass the bridge (4.9 s / 19.3 s at 241K)
```
Every row verdict-checked, 2026-09-06. The bridged leg scores **5/5 on the judged quality eval**,
same as native.
**Since then the served window went 262,144 -> 2,097,152 and the ceiling moved with it.**
A **700,630-token cold prompt now bridges end to end in 11 minutes 26 seconds** โ `verdict complete`,
342 blocks, 1,117 tok/s of prefill. That is **2.9x past the largest prompt in the table above**, and
the pair has accepted over a million.
```
cold prompt engine tok/s end-to-end blocks verdict
192,099 119.2 s 1,611 143.1 s 93 complete
385,838 288.7 s 1,337 466.6 s 188 complete
677,069 596.8 s 1,134 655.9 s 330 complete
700,630 627.0 s 1,117 685.5 s 342 complete
709,055 635.4 s 1,116 694.6 s 345 partial 706,560/709,055 (99.6%)
988,487 1,037.8 s 952 1,142.5 s 376 partial 770,048/988,487 (78%)
1,006,172 1,070.5 s 940 1,171.7 s 346 partial 708,608/1,006,172 (70%)
```
All 2026-09-08, one prefill pair, same 10GbE. `partial` is not a failure: past roughly 772K tokens the
capture crosses the prefill box's free-memory floor and **seals a valid contiguous prefix** instead of
dying; the decoder mounts what arrived and natively prefills the remainder. See *Known limits*.
**We have not run a Mac-alone control at these sizes**, so those rows carry no ratio โ they are what
the bridge does, not a claimed speedup. The measured ratios stop at 241K and are shown above.
Full numbers and methodology: [RESULTS.md](https://github.com/chadhurley25075-png/pd-bridge/blob/main/RESULTS.md) ยท [bench/BENCHMARK-PROTOCOL.md](https://github.com/chadhurley25075-png/pd-bridge/blob/main/bench/BENCHMARK-PROTOCOL.md)
---
## The idea
Prefill/decode disaggregation is well established, and so is the hardware argument for it: prefill is
compute-bound, decode is memory-bandwidth-bound, so run each phase where it is cheapest. Existing
systems do this by **transferring the KV cache** from the prefill worker to the decode worker.
That requires both ends to agree on a cache format. Ours never can. One side is CUDA/vLLM with an
FP8 paged cache; the other is Metal/MLX with its own block layout. Worse, DeepSeek-V4-Flash does not
have "a KV cache" โ each layer carries a rotating 128-token window, a compressor pool (ratio 4 with
overlap carry, and ratio 128), and an indexer pool, all with layer-dependent RoPE.
**So we don't transfer a cache. We compute the decoder's finished cache on the prefill machine,
using the decoder's own weights, and write it straight into the decoder's prefix-cache store.**
The prefill engine already computes the exact tensor those pools are a pure function of โ the
attention input. A hook takes it there, applies the *Mac's* projection and pooling math on the GPU,
and emits the finished pools. The decode side assembles them into MLX cache objects and hands them
to oMLX's own block writer. oMLX then sees a normal prefix-cache hit and only decodes. Neither
engine is modified in its hot path; the decoder does not know a bridge exists.
Payload: **~10 KB per token** โ 0.80 GB for an 81K-token prompt, pulled in 1.08 s. The network
stopped being the bottleneck; the prefill engine is now 65% of wall time, which is where you want it.
### Why the reconstruction is trustworthy
The whole design rests on the pooled tensors being *the same tensors* the decoder would have
computed. That is tested, not assumed:
| check | result |
| --- | --- |
| Cache arrays rebuilt on the Mac vs. a full native forward | **313/313 bit-exact** |
| Blocks written by the bridge vs. blocks oMLX writes itself | **11/11 identical** (only the `created_at` stamp differs) |
| Torch pooling port vs. MLX ground truth (T=23,217) | projections, window, carries **bit-exact**; pooled tensors 99.95โ99.96% identical, worst delta **one bf16 ulp** |
| In-container hook selftest (chunked == one-shot) | **52/52** |
| Needle retrieval through a fully reconstructed 81K cache | correct on every benchmark run |
`studio/verify_blocks.py`, `studio/pd_diff_state.py`, `spark/pd_pool_validate.py` and
`spark/pd_pool_selftest.py` reproduce these. Compare tensors, never file hashes โ `created_at` means
a bridge-written block can never be byte-identical as a *file*.
---
## Known limits โ read this before you run it
### The memory floor, and why big prompts come back `partial` (2026-09-08)
The capture lives in the prefill box's memory for the whole request and costs about **11.8 KB of
unified memory per token** on rank 0 only (rank 1 stays flat โ it does its half of the attention and
holds no capture). A GB10 has ONE 121 GB pool shared by weights, the vLLM KV arena and everything
else, so steady-state free memory with the model up is ~22 GB.
Measured: the capture crosses a 5.0 GB free-memory floor at roughly **772,000 tokens**. Past that the
hook **seals a valid contiguous prefix `[0,T)` and reports it** rather than dying โ the decoder mounts
what arrived and natively prefills the tail. That is the `partial` verdict in the table above. A
700,630-token wake cleared the floor with **100 MB to spare**; a 1,006,172-token wake sealed at 70%.
This replaced a real failure. Before the fix, `_finish()` copied every layer to the host **without
freeing the layers it had already written**, so the full device capture and the growing host copies
were alive at once. At 1,021,199 tokens it died after 11 of 43 layers and drove the box into swap
thrash โ no sshd even over a 200G fabric. It needed a physical power button. The hook now releases
each layer as its file lands and checks `MemAvailable` every 64 layer-chunks.
**If you are memory-tight, this is your limit, not the window.** The window is a config number; the
floor is physics on your box. Measure `MemAvailable` during a long prefill before trusting either.
**The ~82K-token ceiling was a real bug. It is fixed.** (History kept because the failure mode is
instructive and the arithmetic still matters on smaller machines.)
`omlx_block_writer` used to hold one *materialised cumulative* cache snapshot per 2048-token
boundary until `finalize()`: peak memory grew **quadratically** with prompt length (~`N(N+1)/2`
blocks' worth of arrays โ ~16 GB at 39 boundaries, ~23 GB at 47, which exhausted a 256 GB M3 Ultra
holding a 156 GB model and wrote zero blocks at 97,848 tokens).
The writer now **streams**: each boundary is stored through oMLX's own pipeline and released the
moment it is snapshotted (`begin_stream`/`store_boundary`; `finalize` drains and verifies). Peak
memory is ONE boundary snapshot (~20 MB ร boundary index / N โ tens of MB, not tens of GB). It is
validated against a real oMLX reference block (synthetic layout match) and live at 52 boundaries /
109,085 tokens โ see RESULTS.md. `PD_STREAM_BOUNDARIES=0` restores the batched path for comparison.
`PD_MAX_BRIDGE_TOKENS` remains as a configurable envelope guard, not a bug workaround: raise it to
your machine's measured headroom.
Two related behaviours worth knowing:
- **The fallback works, and it can no longer lie.** When a bridge fails the reply still comes back
correct โ the decoder serves natively. Since the bench4 autopsy (docs/FINDING-bench4-cold-fallback.md),
every response carries an `X-PD-Bridge` verdict (`complete` / `partial B/T` / declined with reason),
and `bench_cold.py` records it โ a silent native fallback can never again enter a results table
as a bridged number.
- **The "cold-start variance" at 20K was not variance.** The 25.5 s vs 55.4 s spread was the capture
hook flushing *mid-request* during chunked prefill (see the FINDING): the 55 s runs were native
fallbacks wearing a bridge label. The hook now guards its idle flush with a CUDA-event query and a
chunk-alignment check, and the front validates every capture manifest before trusting it.
## Status, honestly
This is a **reference implementation, not a library.** It is pinned hard and it is young.
- **One model.** DeepSeek-V4-Flash. The pooling math is specific to its sparse attention.
- **Pinned stacks.** oMLX 0.6.4; vLLM 0.21.1rc1 with the DeepSeek-V4 plugin (sparkrun image).
- **It monkey-patches private internals of both engines** โ a `sitecustomize` hook onto
`DeepseekV4MultiHeadLatentAttentionWrapper.attention_impl` on the vLLM side, and a
filesystem-fallback patch to oMLX's `PagedSSDCacheIndex` on the MLX side (oMLX indexes SSD blocks
at model load only, so externally written blocks are otherwise invisible). **Expect this to break
when either project moves.**
- **The judged quality eval is five questions on one document.** The bridged leg scores 5/5 on it,
twice (once from a fresh cold v3 bridge), same as native. Prefill runs FP8 weights and decode runs
MXFP4, so bridged output is *not* token-identical to native; it is factually faithful on what we
checked, which is a smaller claim than "equivalent".
- **The flush signal was unreliable until 2026-09-06 17:52 โ fixed.** The hook's watcher ran in three
processes and two of them deleted the signal before the capturing worker saw it (~1 in 3 hit rate). Hook v5
fixes it; captures now close 0.6โ0.9 s after the engine returns (4.8 s at 236K, which is the block write).
Autopsy: `docs/FINDING-flush-signal-three-watchers.md`. Unit test: `spark/test_flush_decision.py`.
- **The front door is threaded, with one caveat.** HTTP handlers run in threads (health, model list and
oMLX passthrough answer immediately, and concurrent decodes overlap because the decoder batches them),
but every `bridge()` call โ the MLX cache assembly โ is marshalled to the main thread and runs one at a
time, because MLX streams are thread-local and the model lives there. Measured 2026-09-06: a short
request completed in 47 s while an 86K-token cold bridge was in flight, instead of waiting it out.
Two clients do slow each other down; they no longer block each other.
- Only cold, long prompts benefit. Warm turns bypass the bridge by design and are served natively.
**The transferable idea is bigger than this code:** when two engines cannot share a cache format,
compute the *consumer's* finished cache on the *producer*, using the consumer's weights. That
generalizes past this model and this hardware, and it is the part worth stealin