2026-10-11 16:37 UTC

disaggregated-inference

band: warmmomentum: stable score: 0.56
temperature history

Episodes (2)

Shubh Mehta reports that inferpd's shared host-memory KV store enables prefix reuse across disaggregated DeepSeek workers, delivering 0.52-second median TTFT on V2-Lite at 8.5 requests per second but substantially lower capacity on V3.2, providing a reproducible serving architecture rather than demonstrated production-scale economics.
seedconvergesscott: high
pd-bridge's maintainer claims its released DeepSeek-V4-Flash bridge combines NVIDIA prefill with Apple Silicon decode over 10GbE to reduce measured cold long-prompt latency by 1.5โ€“3.7 times versus Mac-only serving while preserving decode throughput, potentially accelerating mixed-hardware local inference.
corroboratedconvergesscott: medium