Retrieved article excerpt
Open article · Retrieved 2026-09-24T13:28:18.205211+00:00
[Download PDF](https://www.nature.com/articles/s43588-025-00854-1.pdf)
- Article
- [Open access](https://www.springernature.com/gp/open-science/about/the-fundamentals-of-open-access-and-open-research)
- Published: 08 September 2025
# Analog in-memory computing attention mechanism for fast and energy-efficient large language models
- [Nathan Leroux](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Nathan-Leroux-Aff1)
[ORCID: orcid.org/0000-0003-3672-0870](https://orcid.org/0000-0003-3672-0870)[1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff1)[na1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#na1),
- [Paul-Philipp Manea](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Paul_Philipp-Manea-Aff2-Aff3)
[ORCID: orcid.org/0000-0001-6998-3066](https://orcid.org/0000-0001-6998-3066)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3)[na1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#na1),
- [Chirag Sudarshan](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Chirag-Sudarshan-Aff2)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),
- [Jan Finkbeiner](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Jan-Finkbeiner-Aff1-Aff3)
[ORCID: orcid.org/0000-0003-4556-3758](https://orcid.org/0000-0003-4556-3758)[1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff1),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3),
- [Sebastian Siegel](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Sebastian-Siegel-Aff2)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),
- [John Paul Strachan](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-John_Paul-Strachan-Aff2-Aff3)[2](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff2),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3) &
- …
- [Emre Neftci](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#auth-Emre-Neftci-Aff1-Aff3)
[ORCID: orcid.org/0000-0002-0332-3273](https://orcid.org/0000-0002-0332-3273)[1](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff1),[3](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#Aff3)
Show authors
[*Nature Computational Science*](https://www.nature.com/natcomputsci)
**volume 5**, pages 813–824 (2025) [Cite this article](https://www.nature.com/articles/s43588-025-00854-1?error=cookies_not_supported&code=aa2e2f7e-9ade-434c-85bc-afd5856249c7#citeas)
[Save article](https://www.nature.com/articles/s43588-025-00854-1/save-research?_csrf=kMDYzC9CDfw2-NbiuKvgJHQGqSOcW0C1)
[View saved research](https://www.nature.com/saved-research)
- 59k Accesses
- 28 Citations
- 83 Altmetric
- [Metrics details](https://www.nature.com/articles/s43588-025-00854-1/metrics)
A [preprint version](https://arxiv.org/abs/2409.19315) of the article is available at arXiv.
## Abstract
Transformer networks, driven by self-attention, are central to large language models. In generative transformers, self-attention uses cache memory to store token projections, avoiding recomputation at each time step. However, graphics processing unit (GPU)-stored projections must be loaded into static random-access memory for each new generation step, causing latency and energy bottlenecks. Here we present a custom self-attention in-memory computing architecture based on emerging charge-based memories called gain cells, which can be efficiently written to store new tokens during sequence generation and enable parallel analog dot-product computation required for self-attention. However, the analog gain-cell circuits introduce non-idealities and constraints preventing the direct mapping of pre-trained models. To circumvent this problem, we design an initialization algorithm achieving text-processing performance comparable to GPT-2 without training from scratch. Our architecture reduces attention latency and energy consumption by up to two and four orders of magnitude, respectively, compared with GPUs, marking a substantial step toward ultrafast, low-power generative transformers.
### Similar content being viewed by others
### [Back to recurrent processing at the crossroad of transformers and state-space models](https://www.nature.com/articles/s42256-025-01034-6?fromPaywallRec=false)
Article
15 May 2025
### [Arsenic-free Ge-Te-based ovonic threshold switching material with reduced leakage current](https://www.nature.com/articles/s41598-025-01323-5?fromPaywallRec=false)
Article
Open access
01 July 2025
### [Efficient nonlinear function approximation in analog resistive crossbars for recurrent neural networks](https://www.nature.com/articles/s41467-025-56254-6?fromPaywallRec=false)
Article
Open access
29 January 2025
### Explore related subjects
Discover the latest articles and news in related subjects.
- [Computational science](https://www.nature.com/subjects/computational-science)
- [Electrical and electronic engineering](https://www.nature.com/subjects/electrical-and-electronic-engineering)
- [Algorithm-Hardware Co-Design for Neural Network Acceleration](https://www.nature.com/subjects/algorithm-hardware-co-design-for-neural-network-acceleration)
## Main
Transformers[1](https://www.nature.com/articles/s43588-025-00854-1#ref-CR1 "Vaswani, A. et al. Attention is all you need. In Proc. 31st International Conference on Neural Information Processing Systems, NIPS’17 6000–6010 (Curran Associates, 2017).") are central to modern artificial intelligence (AI), powering advances in language models, image processing and beyond. However, their high computational demands lead to substantial energy consumption. Enhancing their efficiency is essential to reduce environmental impact and to keep pace with the exponentially growing size of AI models. The success of transformers as state of the art in sequence processing and generation is enabled by their attention mechanism[2](https://www.nature.com/articles/s43588-025-00854-1#ref-CR2 "Bahdanau, D., Cho, K. & Bengio, Y. Neural machine translation by jointly learning to align and translate. Preprint at
http://arxiv.org/abs/1409.0473
(2016)."). To capture dependencies across sequences, the attention mechanism performs dot products between different projections of multiple sequence elements, known as tokens. For generative tasks, the best performance is achieved by autoregressive, decoder-only transformers[3](https://www.nature.com/articles/s43588-025-00854-1#ref-CR3 "Lin, T., Wang, Y., Liu, X. & Qiu, X. A survey of transformers. AI Open 3, 111–132 (2022)."). At each inference step, the decoder generates a token, which is then appended to the input sequence, forming the input for the subsequent step. To avoid recomputing the keys and values (KV cache) projections of the previously generated tokens, the so-called KV-caching method stores the projections from previous tokens in memory and updates the KV cache with the new projections[4](https://www.nature.com/articles/s43588-025-00854-1#ref-CR4 "Pope, R. et al. Efficiently scaling transformer inference. Proc. Mach. Learn. Syst. 5, 606–624 (2023).").
In a graphics processing unit (GPU), for each token, the entire KV cache must be transferred from main high-bandwidth memory to cache memory (static random-access memory (SRAM)). In addition, the KV cache is often much larger than the available SRAM memory owing to the dimensions of the stored projections and the sequence length[5](https://www.nature.com/articles/s43588-025-00854-1#ref-CR5 "Liu, Z. et al. KIVI: a tuning-free asymmetric 2bit quantization for KV cache. In Proc. 41st International Conference on Machine Learning, ICML’24 Vol. 235, 32332–32344 (JMLR.org, 2024)."). For instance, the entire KV cache of the model Mistral 7B[6](https://www.nature.com/articles/s43588-025-00854-1#ref-CR6 "Jiang, A.Q. et al. Mistral 7B. Preprint at
http://arxiv.org/abs/2310.06825
(2023).") requires 8 Gb for a batch size of 1, as necessary for inference workloads. In recent technologies, the energy for data access exceeds the energy required for computations[7](https://www.nature.com/articles/s43588-025-00854-1#ref-CR7 "Jouppi, N. P. et al. Ten lessons from three generations shaped Google’s TPUv4i: industrial product. In Proc. 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) 1–14 (IEEE, 2021);
https://doi.org/10.1109/ISCA52012.2021.00010
"). Loading the KV cache for the attention mechanism is thus a major bottleneck, causing increased energy consumption and latency in large language models (LLMs)[8](https://www.nature.com/articles/s43588-025-00854-1#ref-CR8 "Fu, Y. Challenges in deploying long-context transformers: a theoretical peak performance analysis. Preprint at
https://arxiv.org/abs/2405.08944
(2024)."). To mitigate this bottleneck, a wide body of literature explores resource-efficient algorithms[9](https://www.nature.com/articles/s43588-025-00854-1#ref-CR9 "Xu, M. et al. Resource-efficient algorithms and systems of foundation models: a survey. ACM Comput. Surv. 57, 110–111039 (2025)."). Alternative architectures to transformers with linear time complexity are developed to improve long-sequence processing efficiency[10](https://www.nature.com/articles/s43588-025-00854-1#ref-CR10 "Katharopoulos, A., Vyas, A., Pappas, N. & Fleuret, F. Transformers are RNNs: fast autoregressive transformers with linear attention. In Proc. 37th International Conference on Machine Learning, ICML’20 Vol. 119, 5156–5165 (JMLR.org, 2020);
https://doi.org/10.5555/3524938.3525416
"),[11](https://www.nature.com/articles/s43588-025-00854-1#ref-CR11 "Gu, A. & Dao, T. Mamba: linear-time sequence modeling with selective state spaces. In Proc. Conference on Language Modeling (2024);
https://openreview.net/forum?id=tEYskw1VY2
"). However, transformers continue to exhibit more stable training at scale than alternatives such as Mamba[11](https://www.nature.com/articles/s43588-025-00854-1#ref-CR11 "Gu, A. & Dao, T. Mamba: linear-time sequence modeling with selective state spaces. In Proc. Conference on Language Modeling (2024);
https://openreview.net/forum?id=tEYskw1VY2
"), which contributes to their ongoing dominance despite the efficiency of state-space models. Alternatively, different methods have been d