2026-10-11 16:38 UTC

d-Matrix claims its Raptor logic-on-DRAM architecture reduces memory-transfer energy to roughly one-tenth of HBM while targeting 100 TB/s bandwidth, potentially easing the bandwidth and power constraints of LLM decoding.

state: watchingheat: lowuncertainty: highnovelscott: lowinference-accelerators 3d-dram ai-infrastructure inference-economicsd-Matrix

What is this?

d-Matrix’s Raptor is a generative-inference accelerator architecture that stacks compute logic directly atop custom DRAM, shortening memory connections to address bandwidth-limited LLM decoding. The supplied Hot Chips 2026 coverage reports early silicon with roughly 100 TB/s bandwidth and company-reported measured transfer energy of 0.37 pJ/bit. The claimed roughly tenfold energy advantage depends on the comparison: snippets give about sixfold versus HBM transfer energy alone and larger gains with additional data movement included, not independently demonstrated whole-system inference savings. Commercial readiness is unclear: one snippet says it is heading into production, while another describes an unfinished product without a specified shipping timeline.

Why it matters to Scott

Raptor loosely touches Scott’s Cost of Cognition lens, but component-level transfer-energy claims do not establish cheaper useful cognitive output or a change to his CUDA/Ollama deployments; shipping availability and whole-system savings remain unestablished. The supplied radar hits track memory bandwidth and inference economics, not this Raptor development, and no substantive convergence with or challenge to Scott’s positions is established.
ip:concept.cost-of-cognitionradar:concept.memory-bandwidthradar:concept.inference-economics
queries asked of Scott's wikis
  • LLM decode memory bandwidth bottlenecks prefill separation
  • inference economics energy per token data movement costs
  • local inference memory capacity bandwidth hardware selection
  • agent workloads inference latency accelerator deployment
  • hardware software co-design inference efficiency

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1202h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-22 14:00⭐ origin echo-reconstructedThe official Hot Chips 2026 program identifies this presentation as “3D DRAM based Accelerator for Generative Inference,” presented by Sudee
d-Matrix (Sudeep Bhoja; co-presenter Aayush Ankit, Meta) on other (echo) · attributed from hn.story.49695065
—
09-14 11:27first on hacker news · published · +549.5hD-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
rbanffy
—
09-14 19:27first on r/singularity · published · +557.5hD-Matrix is basically plugging its inference chips straight into Nvidia's own server racks now
ocean_protocol
—
09-14 11:27amplified on hacker news 👑hn.story.49695065
rbanffy
peak 17 · 6 comments · 80% of case engagement
09-14 19:27amplified on r/singularityreddit.post.1wgdevt
ocean_protocol
peak 7 · 3 comments · 20% of case engagement
09-14 12:20our radar first saw it · +550.3hdiscovery anchor: hn.story.49695065—

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnD-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
Retrieved article excerpt

Open article · Retrieved 2026-09-14T12:22:03.329405+00:00

- [Server](https://www.servethehome.com/category/server-parts/)
- [Accelerators](https://www.servethehome.com/category/server-parts/accelerators/)

# d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

By

[Patrick Kennedy](https://www.servethehome.com/author/patrick/)

-

August 23, 2026 

[0](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/#respond)

[Facebook](https://www.facebook.com/sharer.php?u=https%3A%2F%2Fwww.servethehome.com%2Fd-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026%2F "Facebook")[X](https://x.com/intent/post?text=d-Matrix+Raptor+3D-DRAM+Accelerator+for+Generative+Inference+at+Hot+Chips+2026&url=https%3A%2F%2Fwww.servethehome.com%2Fd-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026%2F&via=ServeTheHome "X")[Pinterest](https://pinterest.com/pin/create/button/?url=https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/&media=https://www.servethehome.com/wp-content/uploads/2026/08/meta-3d-dram-based-accelerator-for-genai-slide-17.jpg&description=At Hot Chips 2026, d-Matrix showed off its Raptor 3D-DRAM accelerator for AI breaking free of using HBM for memory by stacking DRAM and logic "Pinterest")[Linkedin](https://www.linkedin.com/shareArticle?mini=true&url=https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/&title=d-Matrix+Raptor+3D-DRAM+Accelerator+for+Generative+Inference+at+Hot+Chips+2026 "Linkedin")[ReddIt](https://reddit.com/submit?url=https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/&title=d-Matrix+Raptor+3D-DRAM+Accelerator+for+Generative+Inference+at+Hot+Chips+2026 "ReddIt")Email[Print](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/ "Print")[Copy URL](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/ "Copy URL")

[d-Matrix Raptor 3D-DRAM](https://www.servethehome.com/wp-content/uploads/2026/08/meta-3d-dram-based-accelerator-for-genai-slide-17.jpg)

d-Matrix Raptor 3D-DRAM

Next up, d-Matrix is presenting its Raptor 3D-DRAM accelerator for generative inference at Hot Chips 2026. The company has made waves, and we have covered it before, including the [d-Matrix Corsair In-Memory Computing for AI Inference at Hot Chips 2025](https://www.servethehome.com/d-matrix-corsair-in-memory-computing-for-ai-inference-at-hot-chips-2025/). We also found they were doing networking in [The New d-Matrix JetStream 400G Ethernet Card for Data Center Scale AI Inference](https://www.servethehome.com/the-new-d-matrix-jetstream-400g-ethernet-card-for-data-center-scale-ai-inference/). Let us see what they have going on this year.

This is being done live, so please excuse typos.

## d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

Model weights keep growing, and the KV cache scales with context length multiplied by batch size. So 64 users at 1M context can mean roughly 935 GB of KV cache. Weights and cache together create a problem that is both a capacity problem and a bandwidth problem, and both sides keep growing.

[d-Matrix The Growing Data Problem](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-02/)

d-Matrix The Growing Data Problem

SRAM meets the bandwidth target, but only on a tiny scale. A Corsair SRAM accelerator card pair reaches roughly 300 TB/s at about 1 ns latency, yet holds only about 4 GB. A 6T SRAM cell is around 10 times larger than a DRAM cell, and leakage runs to tens of watts at GB scale. This makes SRAM suitable for a draft model in speculative decoding, not for holding frontier model weights. That seems to be what NVIDIA is using Groq for as an example.

[d-Matrix SRAM: Bandwidth Advantage](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-04/)

d-Matrix SRAM: Bandwidth Advantage

HBM solves the capacity half but struggles on bandwidth. Pin speed and I/O width per base die improve slowly, and the number of stacks is limited by available package beachfront, roughly 8-16 stacks per package. d-Matrix cites a practical bandwidth ceiling around 20 TB/s for HBM4 packages such as the NVIDIA Vera Rubin and AMD Instinct MI455.

[d-Matrix HBM: The Bandwidth Issue](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-05/)

d-Matrix HBM: The Bandwidth Issue

Bandwidth that high carries a power price. At 2.4 pJ/bit, pushing 100 TB/s through HBM eats about 1.92 kW before any fabric traffic is counted. Packages today lack both the beachfront and the power budget to reach SRAM-class bandwidth with HBM.

[d-Matrix HBM: The Power Problem](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-06/)

d-Matrix HBM: The Power Problem

d-Matrix’s answer is to stack compute directly on top of DRAM dies. Stacking creates a thermal challenge because hundreds of watts must escape through TSVs in a temperature-sensitive DRAM stack, plus a power-delivery challenge from IR drop. d-Matrix says a 1-Hi logic-on-top stack at no more than 0.5 W/mm2 can be liquid cooled and keep DRAM under 100 C.

[d-Matrix 3D-DRAM: An idea whose time has come](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-07/)

d-Matrix 3D-DRAM: An idea whose time has come

3D DRAM lands between the two extremes on an energy ladder. On-die SRAM costs roughly 50 fJ, while 2.5D HBM4 systems run in the 2.5 to 5 pJ range when chip-level energy is included. Vertical 3D IO comes in at around 0.3 to 0.4 pJ, about 10 times lower than HBM, because it is a PHY-less millimeter-scale path rather than a centimeter-scale interposer route. Fewer stacked layers than HBM also means a larger die and better yield.

[d-Matrix Why 3D-DRAM?](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-08/)

d-Matrix Why 3D-DRAM?

d-Matrix is now mapping that view of technologies onto how LLM inference workloads behave. Prefill processes many prompt tokens in parallel and is compute-throughput-bound, whereas decode produces one token at a time and is typically memory-bandwidth-bound. Attention can flip to compute-bound with high GQA and speculative decoding, and MoE stays memory-bound even at modest batch sizes. Decode is the phase that wants huge bandwidth. If you saw our [NVIDIA GB10](https://www.servethehome.com/nvidia-dgx-spark-review-the-gb10-machine-is-so-freaking-cool/) or [AMD Strix Halo](https://www.servethehome.com/amd-ryzen-ai-halo-developer-system-review-amd-goes-for-local-ai/) coverage, memory bandwidth is the big challenge with those types of systems.

[d-Matrix LLM Inference](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-10/)

d-Matrix LLM Inference

Since decode dominates wall-clock runtime, the memory-bound portion matters most. d-Matrix highlights that most inference time is spent in the decode phase, so improving decode bandwidth improves overall inference performance.

[d-Matrix Majority of wall-clock inference time is spent in decode](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-11/)

d-Matrix Majority of wall-clock inference time is spent in decode

At 32GB per card, with 4-bit weights and an 8-bit KV cache, d-Matrix sizes to fit in one rack. A 72-card scale-up can host a frontier model such as Kimi K3 at 1M context. Disaggregation and multi-rack extend beyond a single Raptor rack.

[d-Matrix 1Hi 32GB 3D-DRAM: Frontier LLMs fit in one Raptor Rack](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-12/)

d-Matrix 1Hi 32GB 3D-DRAM: Frontier LLMs fit in one Raptor Rack

Building the system around this memory is a co-design exercise across the memory subsystem, data movement fabric, and workload mapping.

[d-Matrix Building a 3D-DRAM Based Inference System](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-14/)

d-Matrix Building a 3D-DRAM Based Inference System

Now d-Matrix is showing its topology using the full mesh package and discussing its communication protocol.

[d-Matrix Low Latency Fabric (Intra-Card and Inter-Card)](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-15/)

d-Matrix Low Latency Fabric Intra-Card and Inter-Card

d-Matrix’s specific implementation is called Raptor. A TSMC N4 logic die sits on top of a 3D DRAM die using 36 um face-to-face stacking, a process d-Matrix describes as proven, low-cost, high-volume, and high-yield.

[d-Matrix Raptor 3D-DRAM](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-17/)

d-Matrix Raptor 3D-DRAM

Turning the dies into a working system exposes a broad set of integration challenges. d-Matrix is highlighting four here.

[d-Matrix The 3D-DRAM Integration Landscape](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-18/)

d-Matrix The 3D-DRAM Integration Landscape

Those four problems are not independent. d-Matrix walks through three entangled challenges in bank mapping, I/O power, and thermal reliability, noting that a solution to any one constrains the design space of the other two.

[d-Matrix These Challenges Are Entangled](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-19/)

d-Matrix Challenges Are Entangled

Each tensor engine needs a 128B flit per access, and with 32B delivered per column access from 32B banks, that works out to needing 4 banks per channel. d-Matrix’s die has 840 banks, 768 after 72 spares, spread across 256 channels for just 3 banks per channel. This flit does not divide evenly across what is available.

[d-Matrix Challenge 1: The Bank-to-Channel Mapping Problem](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-20/)

d-Matrix Challenge 1: The Bank-to-Channel Mapping Problem

With 3 banks per channel, a single access returns 96B, so delivering a 128B flit takes two accesses and fetches 192B, wasting about 33 percent of bandwidth near 33 TB/s. Column staggering could pack flits but needs a 192B shifting buffer and complicates timing and verification.

[d-Matrix Challenge 1: The Overfetch Dilemma](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-21/)

d-Matrix Challenge 1: The Overfetch Dilemma

Stream blocking reclaims that waste. d-Matrix shares one partial 32B access across three flits, so 4 accesses at 96B feed 3 flits at 128B, with 384B in, matching 384B out. Overfetch drops to zero, every column access is used, and no shifting network is required.

[d-Matrix Solution: Stream Blocking](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-22/)

d-Matrix Solution: Stream Blocking

Moving 100 TB/s at 0.37 pJ/bit works out to 296 W just for I/O, and conventional DBI could save 20 percent. HBM gets away with DBI because its multi-cycle bursts let the PHY see the full burst, but d-Matrix’s single-cycle 256-bit 3D-DRAM link has no burst and no sideband pin to signal the inversion choice.

[d-Matrix Challenge 2: The I/O Power Wall](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-23/)

d-Matrix Challenge 2: The I/O Power Wall

Stream flipping delivers that 20 percent without the pin. Each flit is compared to the previous one and inverted when needed, cutting toggles to near zero with a single metadata bit per flit carried alongside ECC. d-Matrix puts the overhead at 0.8 percent with no PHY change.

[d-Matrix Solution: Stream Flipping (Pinless DBI)](https://www.servethehome.com/meta-3d-dram-based-accelerator-for-genai-slide-24/)

d-Matrix Solution: Stream Flipping Pinless DBI

Heat poses the third challenge at a 105C junction temperature. Yield matters because 840 banks mean even a 1 percent fault rate threatens whole channels, and disca
rbanffy176
🟧 echo.other ⭐The official Hot Chips 2026 program identifies this presentation as “3D DRAM based Accelerator for Generative Inference,” presented by Sudeed-Matrix (Sudeep Bhoja; co-presenter Aayush Ankit, Meta)——
🟠 redditD-Matrix is basically plugging its inference chips straight into Nvidia's own server racks now
singularity
ocean_protocol73

Interpretation history

Decision trace