2026-10-11 16:38 UTC

Leroux and coauthors claim a gain-cell analog in-memory architecture computes attention in place, cutting attention latency by up to two and energy by up to four orders of magnitude versus GPU KV-cache transfer at GPT-2-comparable quality; whether it scales beyond small models to practical LLMs determines if analog in-memory computing becomes a credible alternative inference substrate.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumanalog-in-memory-computing llm-inference-hardware inference-economicsNathan LerouxEmre Neftci

What is this?

A 2025 Nature Computational Science paper from Forschungszentrum Jülich (Leroux, Manea, Sudarshan, Finkbeiner, Siegel, Strachan, Neftci; with RWTH Aachen) presents an analog in-memory computing (AIMC) architecture that stores KV-cache key/value projections in gain-cell crossbar arrays built on oxide-semiconductor (IGZO/ITO) transistors and computes attention dot-products directly in memory, eliminating the off-chip KV-cache transfers that dominate GPU decode latency. Reported headline numbers: ~65 ns and 6.1 nJ per token for a GPT-2 attention head — up to two orders of magnitude lower latency and up to four orders lower energy for attention versus GPUs (the arXiv/preprint versions claimed five orders of energy reduction; the published figure reads four), with adapted models reaching GPT-2-comparable quality on ARC-Easy, WinoGrande, and WikiText-2 without training from scratch. The paper has ~55 citations and follow-on accelerator work (LOKI, DirectGeMM, algorithm-hardware co-design studies) already cites it; every quality claim in the supplied material is at GPT-2 scale, so nothing shown establishes scaling to practical modern LLMs.

Why it matters to Scott

A first-party Nature artifact independently converges on the KV-cache memory-bound decode bottleneck Scott's canon treats as first-order — hardware-aware local inference makes memory pressure explicit runtime policy, and prefix-caching economics and context engineering are premised on exactly that constraint — then attacks it at the substrate level rather than the software level, making it both independent validation of his diagnosis and the analog sibling of the radar's d-Matrix digital near-memory case. It stays at medium rather than high because the artifact is lab-stage, quality-proven only at GPT-2 scale, a year old, and engages nobody: it extends his inference-hardware-economics territory and hands the radar a clean scaling-to-real-LLMs monitor, but changes nothing he builds or argues today beyond adding a cost-of-cognition floor scenario if the four-orders energy claim ever survives scale-up.
dev:concept.hardware-aware-local-inferenceip:concept.prefix-caching-economicsip:framework.context-engineeringip:concept.cost-of-cognitionradar:dmatrix-raptor-3d-dramradar:concept.kv-cacheradar:concept.ai-hardwareradar:samsung-pim-ai-memory-bandwidthradar:xcena-mx1-cxl-computational-memoryradar:concept.inference-economics
queries asked of Scott's wikis
  • KV cache memory-bound decode inference bottleneck
  • inference economics cost per token trends
  • analog in-memory computing neuromorphic exotic hardware
  • local inference hardware edge accelerators sovereignty
  • near-memory compute digital inference d-Matrix
  • attention quadratic cost scaling limits

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 9578h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-07 14:00⭐ origin echo-reconstructedNature Computational Science paper: a gain-cell analog in-memory architecture stores token projections and computes attention dot-products i
Nathan Leroux, Paul-Philipp Manea, Chirag Sudarshan, et al. on paper (echo) · attributed from hn.story.49830214
—
09-24 13:20first on hacker news · published · +9167.3hAnalog in-memory computing attention mechanism for fast and energy-efficient LLM
bilsbie
—
09-24 13:20amplified on hacker news 👑hn.story.49830214
bilsbie
peak 2 · 0 comments · 98% of case engagement
09-24 13:20our radar first saw it · +9167.4hdiscovery anchor: hn.story.49830214—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAnalog in-memory computing attention mechanism for fast and energy-efficient LLM
Retrieved article excerpt

Open article · Retrieved 2026-09-24T13:28:18.205211+00:00

[Download PDF](https://www.nature.com/articles/s43588-025-00854-1.pdf)

- Article
- [Open access](https://www.springernature.com/gp/open-science/about/the-fundamentals-of-open-access-and-open-research)
- Published: 08 September 2025

# Analog in-memory computing attention mechanism for fast and energy-efficient large language models

- [Nathan Leroux](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Nathan-Leroux-Aff1) 
  [ORCID: orcid.org/0000-0003-3672-0870](https://orcid.org/0000-0003-3672-0870)[1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff1)[na1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#na1),
- [Paul-Philipp Manea](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Paul_Philipp-Manea-Aff2-Aff3) 
  [ORCID: orcid.org/0000-0001-6998-3066](https://orcid.org/0000-0001-6998-3066)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3)[na1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#na1),
- [Chirag Sudarshan](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Chirag-Sudarshan-Aff2)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),
- [Jan Finkbeiner](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Jan-Finkbeiner-Aff1-Aff3) 
  [ORCID: orcid.org/0000-0003-4556-3758](https://orcid.org/0000-0003-4556-3758)[1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff1),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3),
- [Sebastian Siegel](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Sebastian-Siegel-Aff2)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),
- [John Paul Strachan](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-John_Paul-Strachan-Aff2-Aff3)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3) &
- …
- [Emre Neftci](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Emre-Neftci-Aff1-Aff3) 
  [ORCID: orcid.org/0000-0002-0332-3273](https://orcid.org/0000-0002-0332-3273)[1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff1),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3)

Show authors

[*Nature Computational Science*](https://www.nature.com/natcomputsci)
**volume 5**, pages 813–824 (2025) [Cite this article](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#citeas)

[Save article](https://www.nature.com/articles/s43588-025-00854-1/save-research?_csrf=kMDYzC9CDfw2-NbiuKvgJHQGqSOcW0C1)

[View saved research](https://www.nature.com/saved-research)

- 59k Accesses
- 28 Citations
- 83 Altmetric
- [Metrics details](https://www.nature.com/articles/s43588-025-00854-1/metrics)

A [preprint version](https://arxiv.org/abs/2409.19315) of the article is available at arXiv.

## Abstract

Transformer networks, driven by self-attention, are central to large language models. In generative transformers, self-attention uses cache memory to store token projections, avoiding recomputation at each time step. However, graphics processing unit (GPU)-stored projections must be loaded into static random-access memory for each new generation step, causing latency and energy bottlenecks. Here we present a custom self-attention in-memory computing architecture based on emerging charge-based memories called gain cells, which can be efficiently written to store new tokens during sequence generation and enable parallel analog dot-product computation required for self-attention. However, the analog gain-cell circuits introduce non-idealities and constraints preventing the direct mapping of pre-trained models. To circumvent this problem, we design an initialization algorithm achieving text-processing performance comparable to GPT-2 without training from scratch. Our architecture reduces attention latency and energy consumption by up to two and four orders of magnitude, respectively, compared with GPUs, marking a substantial step toward ultrafast, low-power generative transformers.

### Similar content being viewed by others

### [Back to recurrent processing at the crossroad of transformers and state-space models](https://www.nature.com/articles/s42256-025-01034-6?fromPaywallRec=false)

Article
15 May 2025

### [Arsenic-free Ge-Te-based ovonic threshold switching material with reduced leakage current](https://www.nature.com/articles/s41598-025-01323-5?fromPaywallRec=false)

Article
Open access
01 July 2025

### [Efficient nonlinear function approximation in analog resistive crossbars for recurrent neural networks](https://www.nature.com/articles/s41467-025-56254-6?fromPaywallRec=false)

Article
Open access
29 January 2025

### Explore related subjects

Discover the latest articles and news in related subjects.

- [Computational science](https://www.nature.com/subjects/computational-science)
- [Electrical and electronic engineering](https://www.nature.com/subjects/electrical-and-electronic-engineering)
- [Algorithm-Hardware Co-Design for Neural Network Acceleration](https://www.nature.com/subjects/algorithm-hardware-co-design-for-neural-network-acceleration)

## Main

Transformers[1](https://www.nature.com/articles/s43588-025-00854-1#ref-CR1 "Vaswani, A. et al. Attention is all you need. In Proc. 31st International Conference on Neural Information Processing Systems, NIPS’17 6000–6010 (Curran Associates, 2017).") are central to modern artificial intelligence (AI), powering advances in language models, image processing and beyond. However, their high computational demands lead to substantial energy consumption. Enhancing their efficiency is essential to reduce environmental impact and to keep pace with the exponentially growing size of AI models. The success of transformers as state of the art in sequence processing and generation is enabled by their attention mechanism[2](https://www.nature.com/articles/s43588-025-00854-1#ref-CR2 "Bahdanau, D., Cho, K. & Bengio, Y. Neural machine translation by jointly learning to align and translate. Preprint at 
                  http://arxiv.org/abs/1409.0473
                  
                 (2016)."). To capture dependencies across sequences, the attention mechanism performs dot products between different projections of multiple sequence elements, known as tokens. For generative tasks, the best performance is achieved by autoregressive, decoder-only transformers[3](https://www.nature.com/articles/s43588-025-00854-1#ref-CR3 "Lin, T., Wang, Y., Liu, X. & Qiu, X. A survey of transformers. AI Open 3, 111–132 (2022)."). At each inference step, the decoder generates a token, which is then appended to the input sequence, forming the input for the subsequent step. To avoid recomputing the keys and values (KV cache) projections of the previously generated tokens, the so-called KV-caching method stores the projections from previous tokens in memory and updates the KV cache with the new projections[4](https://www.nature.com/articles/s43588-025-00854-1#ref-CR4 "Pope, R. et al. Efficiently scaling transformer inference. Proc. Mach. Learn. Syst. 5, 606–624 (2023).").

In a graphics processing unit (GPU), for each token, the entire KV cache must be transferred from main high-bandwidth memory to cache memory (static random-access memory (SRAM)). In addition, the KV cache is often much larger than the available SRAM memory owing to the dimensions of the stored projections and the sequence length[5](https://www.nature.com/articles/s43588-025-00854-1#ref-CR5 "Liu, Z. et al. KIVI: a tuning-free asymmetric 2bit quantization for KV cache. In Proc. 41st International Conference on Machine Learning, ICML’24 Vol. 235, 32332–32344 (JMLR.org, 2024)."). For instance, the entire KV cache of the model Mistral 7B[6](https://www.nature.com/articles/s43588-025-00854-1#ref-CR6 "Jiang, A.Q. et al. Mistral 7B. Preprint at 
                  http://arxiv.org/abs/2310.06825
                  
                 (2023).") requires 8 Gb for a batch size of 1, as necessary for inference workloads. In recent technologies, the energy for data access exceeds the energy required for computations[7](https://www.nature.com/articles/s43588-025-00854-1#ref-CR7 "Jouppi, N. P. et al. Ten lessons from three generations shaped Google’s TPUv4i: industrial product. In Proc. 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) 1–14 (IEEE, 2021); 
                  https://doi.org/10.1109/ISCA52012.2021.00010
                  
                "). Loading the KV cache for the attention mechanism is thus a major bottleneck, causing increased energy consumption and latency in large language models (LLMs)[8](https://www.nature.com/articles/s43588-025-00854-1#ref-CR8 "Fu, Y. Challenges in deploying long-context transformers: a theoretical peak performance analysis. Preprint at 
                  https://arxiv.org/abs/2405.08944
                  
                 (2024)."). To mitigate this bottleneck, a wide body of literature explores resource-efficient algorithms[9](https://www.nature.com/articles/s43588-025-00854-1#ref-CR9 "Xu, M. et al. Resource-efficient algorithms and systems of foundation models: a survey. ACM Comput. Surv. 57, 110–111039 (2025)."). Alternative architectures to transformers with linear time complexity are developed to improve long-sequence processing efficiency[10](https://www.nature.com/articles/s43588-025-00854-1#ref-CR10 "Katharopoulos, A., Vyas, A., Pappas, N. & Fleuret, F. Transformers are RNNs: fast autoregressive transformers with linear attention. In Proc. 37th International Conference on Machine Learning, ICML’20 Vol. 119, 5156–5165 (JMLR.org, 2020); 
                  https://doi.org/10.5555/3524938.3525416
                  
                "),[11](https://www.nature.com/articles/s43588-025-00854-1#ref-CR11 "Gu, A. & Dao, T. Mamba: linear-time sequence modeling with selective state spaces. In Proc. Conference on Language Modeling (2024); 
                  https://openreview.net/forum?id=tEYskw1VY2
                  
                "). However, transformers continue to exhibit more stable training at scale than alternatives such as Mamba[11](https://www.nature.com/articles/s43588-025-00854-1#ref-CR11 "Gu, A. & Dao, T. Mamba: linear-time sequence modeling with selective state spaces. In Proc. Conference on Language Modeling (2024); 
                  https://openreview.net/forum?id=tEYskw1VY2
                  
                "), which contributes to their ongoing dominance despite the efficiency of state-space models. Alternatively, different methods have been d
bilsbie20
🟧 echo.paper ⭐Nature Computational Science paper: a gain-cell analog in-memory architecture stores token projections and computes attention dot-products iNathan Leroux, Paul-Philipp Manea, Chirag Sudarshan, et al.——

Interpretation history

Decision trace