2026-10-11 16:37 UTC

Shubh Mehta reports that inferpd's shared host-memory KV store enables prefix reuse across disaggregated DeepSeek workers, delivering 0.52-second median TTFT on V2-Lite at 8.5 requests per second but substantially lower capacity on V3.2, providing a reproducible serving architecture rather than demonstrated production-scale economics.

state: seedheat: mediumuncertainty: mediumconvergesscott: highllm-serving disaggregated-inference kv-cacheShubh Mehta

What is this?

Shubh Mehta's inferpd project demonstrates a disaggregated DeepSeek serving architecture using a shared LMCache multiprocess setup as an L1 KV store in host RAM, enabling prefix reuse across workers. The LinkedIn snippet confirms the basic setup: DeepSeek V2-Lite with 1 prefiller + 2 workers using shared host-memory KV caching. The specific performance claims (0.52s median TTFT at 8.5 req/s on V2-Lite, substantially lower capacity on V3.2) and the 'eight deployment experiments and runbooks' are referenced in the case's evidence titles but do not appear in the retrieved web snippets — the snippets are thin on quantitative results and the V3.2 comparison.

Why it matters to Scott

Shubh Mehta's inferpd artifacts — reproducible disaggregated DeepSeek runs with shared host-memory KV caching, explicit TTFT/capacity numbers, and runbooks — independently arrive at positions Scott's canon already holds: prefix-caching economics (near-linear agent-loop cost), inference unit economics (net-value-per-transaction lens), verification loops (generate→check→repair with receipts), spec-as-asset (deployment artifacts as durable spec), and hardware-aware local inference (accelerator placement, memory pressure as runtime policy). The concrete V2-Lite vs V3.2 capacity gap (0.52s median TTFT at 8.5 req/s vs substantially lower) directly bears on his capacity-planning and platform-economics frameworks, and the LMCache multiprocess L1 KV store maps to his singleton GPU job queue and addressable cognitive heap concepts.
ip:concept.prefix-caching-economicsip:concept.ai-unit-economicsip:concept.inference-fieldip:framework.context-engineeringip:concept.verification-loopsip:concept.spec-as-assetip:concept.platform-economicsdev:concept.hardware-aware-local-inferencedev:concept.singleton-gpu-job-queuedev:project.beamdev:project.gamepcdev:technology.litellmdev:technology.deepseekradar:concept.kv-cacheradar:concept.disaggregated-inferenceradar:concept.inference-servingradar:concept.inference-economicsradar:concept.reproducibilityradar:concept.inference-runtimesradar:concept.deepseekradar:concept.inference-capacityradar:concept.inference-costsradar:narwhal-prefill-decode-disaggregationradar:pd-bridge-heterogeneous-prefill-decoderadar:lmcache-hybrid-restore-corruption
queries asked of Scott's wikis
  • disaggregated inference prefix caching KV cache sharing
  • inference serving architecture LMCache host memory KV store
  • DeepSeek V2 V3 serving throughput TTFT capacity planning
  • reproducible deployment artifacts runbooks inference economics
  • prefill decode disaggregation shared prefix optimization

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 502h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-20 17:47⭐ origin directly observedFrom zero to disaggregated DeepSeek deployment
dockerd on hacker news
—
09-20 22:23first on blog (echo) · first seen by us · +4.6hPublishes eight deployment experiments and runbooks, reporting successful shared-prefix caching but throughput below the production target a
Shubh Mehta
—
09-20 17:47amplified on hacker news 👑hn.story.49778115
dockerd
peak 2 · 0 comments · 98% of case engagement
09-20 18:20our radar first saw it · +0.6hdiscovery anchor: hn.story.49778115—
pace: p23 vs 1032 stories at the 336h mark (now 502h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐From zero to disaggregated DeepSeek deployment
Retrieved article excerpt

Open article · Retrieved 2026-09-20T18:22:56.205722+00:00

Problem

## What had to be solved

Self-hosted DeepSeek-class inference for real chat traffic: long prompts, heavy prefix reuse,
streaming API, TTFT matters more than end-to-end latency.

~80%
Prefix cache hits

~7.9k
Input tokens (P50)

~60
Peak QPS target

~2s
Prod TTFT P50 ballpark

What I tried

## Eight setups, in the order I actually ran them

Chat traffic reuses the same system prompt about 80% of the time. The question was how to stop
recomputing that prefix on every request. I started on the production-class model, burned money,
then dropped to a smaller model and compared three ways of moving the KV cache between GPUs.

Shorthand: **1P2D** means 1 GPU does prefill (read the prompt) and 2 GPUs do decode
(stream tokens). **LMCache MP** is a shared KV store in host RAM that every GPU talks to,
instead of handing cache privately from one GPU to another.

| What I tried | Model and GPUs | How KV moved | What happened |
| --- | --- | --- | --- |
| BF16 baseline | DeepSeek V3.2 on 8× H200 | All 8 GPUs used for tensor parallelism. No split between prefill and decode on one node. | The model ran, but the cluster was too expensive (~$60/hr class) to iterate on daily. |
| Classic LMCache P/D | DeepSeek V2-Lite on A100s | Prefiller pushed KV to the decoder per request, using LMCache’s built-in P/D connector. | Worked in smoke tests. Decoder ran out of memory under real concurrency. |
| NIXL P/D | V2-Lite on 2× A100 1 prefiller + 1 decoder | GPU-to-GPU KV handoff over NVLink/UCX. Each request copies cache privately to its decoder. | First stable split. Fast for a single request, but later chats could not reuse that prefix. |
| NIXL scale-out | V2-Lite on A100s 2 prefillers + 2 decoders | Same GPU-to-GPU handoff, more workers. | Throughput went up. Still no shared prefix cache across requests. |
| Shared L1 | V2-Lite on A100s 1 prefiller + 1 decoder | One central LMCache multiprocess server. Prefiller writes prefix KV once; decoder reads it back. | Prefix reuse started working. This is the architecture that matched the 80% hit rate. |
| Best V2-Lite run | V2-Lite on 3× A100 1 prefiller + 2 decoders | Same shared LMCache store, extra decode GPUs so token streaming can scale independently. | Asked 10 req/s, delivered 8.5. TTFT P50 0.52s vs ~2s in production. The headline result. |
| Production-class model | V3.2 NVFP4 on 8× B200 1 prefiller + 1 decoder | Same shared LMCache store, now on the real model and Blackwell GPUs. | The path works end to end. Sub-second TTFT only held to ~2 req/s, not the V2-Lite curve. |
| Cross-node scale-out | V3.2 NVFP4 on two 8× B200 rentals 1 prefiller + 3 decoders | Same design, split across two machines so decode can grow past one node. | Runbook is done. The provider had no private network between rentals, so the sweep never ran. |

What worked

## Shared L1 via LMCache MP + ZMQ

One **central** LMCache Multiprocess server (ZMQ port 6000) as a shared L1 KV layer in host RAM.
Every prefiller and decoder uses `LMCacheMPConnector` against the *same* store.
Prefiller writes prefix KV once; later requests with the same prefix skip full prefill.

Disagg proxy :9000

Prefiller :8100

Decoders :8200+

LMCache MP · shared L1 KV · ZMQ :6000

Results

## Asked vs delivered

Best V2-Lite stack: 1 prefiller + 2 decoders on 3× A100, shared LMCache store.
Production wanted ~2s time-to-first-token and ~60 req/s. Here is what that stack actually did.

0.52s
TTFT P50 when asked for 10 req/s
vs ~2s production ballpark

8.5 / 10
Req/s delivered vs asked
Sustained at the headline point

~11 / 60
Ceiling vs peak target
Saturates well below 60 QPS

Measured on **DeepSeek V2-Lite** (3× A100), not a production-class model.
On **V3.2 NVFP4** (8× B200) the same architecture held sub-second TTFT only to ~2 req/s.
This proves a *shared L1 KV layer is the right shape* for prefix-heavy chat, not production capacity.

### Could it keep up?

Gold: load we requested. Blue: load the stack actually served. After 10 they split:
served stalls near 11.\* That is this box’s ceiling, not 60.

### How long until the first token?

Wait time in seconds, not speed. Lower is better. Green: typical request (median).
Red: slow tail. Dashed lines: production (~2s / ~3.5s). At 10 requested, median is 0.52s.

\* One 1P2D box (3 GPUs). ~11 req/s is that hardware ceiling, not a failed design.
More QPS means more replicas at about 1:2 to 1:3 prefill:decode, e.g. 3 prefiller + 6 decoder GPUs.

Cost

## Total GPU spend

~$1,205
RunPod + Vast.ai (spot rentals)

Insights

## What I would do again

Small first
V2-Lite on cheap A100 before V3.2 on B200

Move fast
Experiments per week beat perfect plans

Shrink it
One request, one role, one path, then scale up

Technical depth

## Everything else is in the repo

Full write-ups, sweep tables, per-phase runbooks, and the resource bibliography.

- [Deep-dive index ↗](https://github.com/shubhmehta3121/inferpd/blob/master/docs/deep-dives/README.md)
- [Engineering phases & architecture ↗](https://github.com/shubhmehta3121/inferpd/blob/master/docs/deep-dives/BUILD.md)
- [Full benchmark tables ↗](https://github.com/shubhmehta3121/inferpd/blob/master/docs/deep-dives/RESULTS.md)
- [Cost breakdown ↗](https://github.com/shubhmehta3121/inferpd/blob/master/docs/deep-dives/COSTS.md)
- [Extended lessons ↗](https://github.com/shubhmehta3121/inferpd/blob/master/docs/deep-dives/INSIGHTS.md)
- [Learning resources ↗](https://github.com/shubhmehta3121/inferpd/blob/master/docs/deep-dives/RESOURCES.md)
- [Repo (runbooks under `models/`) ↗](https://github.com/shubhmehta3121/inferpd)

[Open repo ↗](https://github.com/shubhmehta3121/inferpd)
[LinkedIn ↗](https://www.linkedin.com/in/shubh-mehta-2bb173221/)
dockerd20
🟧 echo.blogPublishes eight deployment experiments and runbooks, reporting successful shared-prefix caching but throughput below the production target aShubh Mehta——

Interpretation history

Decision trace