Retrieved article excerpt
Open article · Retrieved 2026-09-14T12:22:03.329405+00:00
- [Server](https://www.servethehome.com/category/server-parts/)
- [Accelerators](https://www.servethehome.com/category/server-parts/accelerators/)
# d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
By
[Patrick Kennedy](https://www.servethehome.com/author/patrick/)
-
August 23, 2026
[0](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/#respond)
[Facebook](https://www.facebook.com/sharer.php?u=https%3A%2F%2Fwww.servethehome.com%2Fd-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026%2F "Facebook")[X](https://x.com/intent/post?text=d-Matrix+Raptor+3D-DRAM+Accelerator+for+Generative+Inference+at+Hot+Chips+2026&url=https%3A%2F%2Fwww.servethehome.com%2Fd-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026%2F&via=ServeTheHome "X")[Pinterest](https://pinterest.com/pin/create/button/?url=https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/&media=https://www.servethehome.com/wp-content/uploads/2026/08/meta-3d-dram-based-accelerator-for-genai-slide-17.jpg&description=At Hot Chips 2026, d-Matrix showed off its Raptor 3D-DRAM accelerator for AI breaking free of using HBM for memory by stacking DRAM and logic "Pinterest")[Linkedin](https://www.linkedin.com/shareArticle?mini=true&url=https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/&title=d-Matrix+Raptor+3D-DRAM+Accelerator+for+Generative+Inference+at+Hot+Chips+2026 "Linkedin")[ReddIt](https://reddit.com/submit?url=https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/&title=d-Matrix+Raptor+3D-DRAM+Accelerator+for+Generative+Inference+at+Hot+Chips+2026 "ReddIt")Email[Print](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/ "Print")[Copy URL](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/ "Copy URL")
[d-Matrix Raptor 3D-DRAM](https://www.servethehome.com/wp-content/uploads/2026/08/meta-3d-dram-based-accelerator-for-genai-slide-17.jpg)
d-Matrix Raptor 3D-DRAM
Next up, d-Matrix is presenting its Raptor 3D-DRAM accelerator for generative inference at Hot Chips 2026. The company has made waves, and we have covered it before, including the [d-Matrix Corsair In-Memory Computing for AI Inference at Hot Chips 2025](https://www.servethehome.com/d-matrix-corsair-in-memory-computing-for-ai-inference-at-hot-chips-2025/). We also found they were doing networking in [The New d-Matrix JetStream 400G Ethernet Card for Data Center Scale AI Inference](https://www.servethehome.com/the-new-d-matrix-jetstream-400g-ethernet-card-for-data-center-scale-ai-inference/). Let us see what they have going on this year.
This is being done live, so please excuse typos.
## d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
Model weights keep growing, and the KV cache scales with context length multiplied by batch size. So 64 users at 1M context can mean roughly 935 GB of KV cache. Weights and cache together create a problem that is both a capacity problem and a bandwidth problem, and both sides keep growing.
[d-Matrix The Growing Data Problem](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-02/)
d-Matrix The Growing Data Problem
SRAM meets the bandwidth target, but only on a tiny scale. A Corsair SRAM accelerator card pair reaches roughly 300 TB/s at about 1 ns latency, yet holds only about 4 GB. A 6T SRAM cell is around 10 times larger than a DRAM cell, and leakage runs to tens of watts at GB scale. This makes SRAM suitable for a draft model in speculative decoding, not for holding frontier model weights. That seems to be what NVIDIA is using Groq for as an example.
[d-Matrix SRAM: Bandwidth Advantage](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-04/)
d-Matrix SRAM: Bandwidth Advantage
HBM solves the capacity half but struggles on bandwidth. Pin speed and I/O width per base die improve slowly, and the number of stacks is limited by available package beachfront, roughly 8-16 stacks per package. d-Matrix cites a practical bandwidth ceiling around 20 TB/s for HBM4 packages such as the NVIDIA Vera Rubin and AMD Instinct MI455.
[d-Matrix HBM: The Bandwidth Issue](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-05/)
d-Matrix HBM: The Bandwidth Issue
Bandwidth that high carries a power price. At 2.4 pJ/bit, pushing 100 TB/s through HBM eats about 1.92 kW before any fabric traffic is counted. Packages today lack both the beachfront and the power budget to reach SRAM-class bandwidth with HBM.
[d-Matrix HBM: The Power Problem](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-06/)
d-Matrix HBM: The Power Problem
d-Matrix’s answer is to stack compute directly on top of DRAM dies. Stacking creates a thermal challenge because hundreds of watts must escape through TSVs in a temperature-sensitive DRAM stack, plus a power-delivery challenge from IR drop. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm2 can be liquid cooled and keep DRAM under 100 C.
[d-Matrix 3D-DRAM: An idea whose time has come](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-07/)
d-Matrix 3D-DRAM: An idea whose time has come
3D DRAM lands between the two extremes on an energy ladder. On-die SRAM costs roughly 50 fJ, while 2.5D HBM4 systems run in the 2.5 to 5 pJ range when chip-level energy is included. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, about 10 times lower than HBM, because it is a PHY-less millimeter-scale path rather than a centimeter-scale interposer route. Fewer stacked layers than HBM also means a larger die and better yield.
[d-Matrix Why 3D-DRAM?](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-08/)
d-Matrix Why 3D-DRAM?
d-Matrix is now mapping that view of technologies onto how LLM inference workloads behave. Prefill processes many prompt tokens in parallel and is compute-throughput-bound, whereas decode produces one token at a time and is typically memory-bandwidth-bound. Attention can flip to compute-bound with high GQA and speculative decoding, and MoE stays memory-bound even at modest batch sizes. Decode is the phase that wants huge bandwidth. If you saw our [NVIDIA GB10](https://www.servethehome.com/nvidia-dgx-spark-review-the-gb10-machine-is-so-freaking-cool/) or [AMD Strix Halo](https://www.servethehome.com/amd-ryzen-ai-halo-developer-system-review-amd-goes-for-local-ai/) coverage, memory bandwidth is the big challenge with those types of systems.
[d-Matrix LLM Inference](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-10/)
d-Matrix LLM Inference
Since decode dominates wall-clock runtime, the memory-bound portion matters most. d-Matrix highlights that most inference time is spent in the decode phase, so improving decode bandwidth improves overall inference performance.
[d-Matrix Majority of wall-clock inference time is spent in decode](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-11/)
d-Matrix Majority of wall-clock inference time is spent in decode
At 32GB per card, with 4-bit weights and an 8-bit KV cache, d-Matrix sizes to fit in one rack. A 72-card scale-up can host a frontier model such as Kimi K3 at 1M context. Disaggregation and multi-rack extend beyond a single Raptor rack.
[d-Matrix 1Hi 32GB 3D-DRAM: Frontier LLMs fit in one Raptor Rack](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-12/)
d-Matrix 1Hi 32GB 3D-DRAM: Frontier LLMs fit in one Raptor Rack
Building the system around this memory is a co-design exercise across the memory subsystem, data movement fabric, and workload mapping.
[d-Matrix Building a 3D-DRAM Based Inference System](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-14/)
d-Matrix Building a 3D-DRAM Based Inference System
Now d-Matrix is showing its topology using the full mesh package and discussing its communication protocol.
[d-Matrix Low Latency Fabric (Intra-Card and Inter-Card)](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-15/)
d-Matrix Low Latency Fabric Intra-Card and Inter-Card
d-Matrix’s specific implementation is called Raptor. A TSMC N4 logic die sits on top of a 3D DRAM die using 36 um face-to-face stacking, a process d-Matrix describes as proven, low-cost, high-volume, and high-yield.
[d-Matrix Raptor 3D-DRAM](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-17/)
d-Matrix Raptor 3D-DRAM
Turning the dies into a working system exposes a broad set of integration challenges. d-Matrix is highlighting four here.
[d-Matrix The 3D-DRAM Integration Landscape](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-18/)
d-Matrix The 3D-DRAM Integration Landscape
Those four problems are not independent. d-Matrix walks through three entangled challenges in bank mapping, I/O power, and thermal reliability, noting that a solution to any one constrains the design space of the other two.
[d-Matrix These Challenges Are Entangled](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-19/)
d-Matrix Challenges Are Entangled
Each tensor engine needs a 128B flit per access, and with 32B delivered per column access from 32B banks, that works out to needing 4 banks per channel. d-Matrix’s die has 840 banks, 768 after 72 spares, spread across 256 channels for just 3 banks per channel. This flit does not divide evenly across what is available.
[d-Matrix Challenge 1: The Bank-to-Channel Mapping Problem](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-20/)
d-Matrix Challenge 1: The Bank-to-Channel Mapping Problem
With 3 banks per channel, a single access returns 96B, so delivering a 128B flit takes two accesses and fetches 192B, wasting about 33 percent of bandwidth near 33 TB/s. Column staggering could pack flits but needs a 192B shifting buffer and complicates timing and verification.
[d-Matrix Challenge 1: The Overfetch Dilemma](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-21/)
d-Matrix Challenge 1: The Overfetch Dilemma
Stream blocking reclaims that waste. d-Matrix shares one partial 32B access across three flits, so 4 accesses at 96B feed 3 flits at 128B, with 384B in, matching 384B out. Overfetch drops to zero, every column access is used, and no shifting network is required.
[d-Matrix Solution: Stream Blocking](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-22/)
d-Matrix Solution: Stream Blocking
Moving 100 TB/s at 0.37 pJ/bit works out to 296 W just for I/O, and conventional DBI could save 20 percent. HBM gets away with DBI because its multi-cycle bursts let the PHY see the full burst, but d-Matrix’s single-cycle 256-bit 3D-DRAM link has no burst and no sideband pin to signal the inversion choice.
[d-Matrix Challenge 2: The I/O Power Wall](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-23/)
d-Matrix Challenge 2: The I/O Power Wall
Stream flipping delivers that 20 percent without the pin. Each flit is compared to the previous one and inverted when needed, cutting toggles to near zero with a single metadata bit per flit carried alongside ECC. d-Matrix puts the overhead at 0.8 percent with no PHY change.
[d-Matrix Solution: Stream Flipping (Pinless DBI)](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-24/)
d-Matrix Solution: Stream Flipping Pinless DBI
Heat poses the third challenge at a 105C junction temperature. Yield matters because 840 banks mean even a 1 percent fault rate threatens whole channels, and disca