2026-10-11 16:38 UTC

Fangzhou Liang and coauthors claim SSD-LLaMA runs a trillion-parameter MoE above one token per second on one RTX 5090 with at most 32GB RAM while executing every selected expert, potentially making full-expert large-model inference feasible on consumer PCs.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediuminference-economics local-inference ai-infrastructureFangzhou LiangZili Meng

What is this?

The case describes SSD-LLaMA as an SSD-native mixture-of-experts inference system attributed to Fangzhou Liang and coauthors, with Zili Meng named as a key person; it claims over one token per second for a trillion-parameter model on one RTX 5090 and at most 32GB RAM while executing every selected expert. Its evidence titles describe SSD expert delivery, a three-tier storage hierarchy, and CPU–GPU execution, but none of the supplied web results directly identifies the project, verifies its authorship, or corroborates the benchmark. The snippets provide only surrounding context on consumer-GPU memory constraints and MoE inference, with conflicting capacity and throughput estimates; the case's benchmark conditions, implementation availability, and quality implications remain unestablished.

Why it matters to Scott

SSD-LLaMA’s claimed SSD/CPU/GPU execution approach converges with Scott’s hardware-aware local inference practice and could expand the models worth testing on gamepc, where Ollama already serves cheap bulk classification and generation. The radar tracks related approaches in hotpin-lossless-moe-streaming and kimi-k3-nvme-expert-streaming, but not this specific development; the uncorroborated benchmark, unknown implementation availability, and unestablished workload quality make this a conditional evaluation lead rather than evidence to change his serving stack.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:hotpin-lossless-moe-streamingradar:kimi-k3-nvme-expert-streamingradar:concept.expert-streamingradar:concept.moe-offloading
queries asked of Scott's wikis
  • local inference economics cloud tradeoffs latency
  • SSD expert offloading tiered memory CPU GPU inference
  • mixture of experts routing fidelity quantization quality
  • consumer hardware large model memory constraints
  • self-hosted coding agents throughput requirements

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 626h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedSSD-LLaMA combines SSD expert delivery, a three-tier storage hierarchy, and CPU–GPU execution, reporting above one token per second for a tr
Fangzhou Liang and coauthors on paper (echo) · attributed from hn.story.49758274
—
09-18 18:28first on hacker news · published · +76.5hSSD-Llama: SSD-Native Inference for Trillion-Parameter Moe on a Consumer PC
AlmostCosmo79
—
10-07 17:33first on r/LocalLLaMA · published · +531.6hA 341 GB DeepSeek on a 128 GB Mac: 2x decode and first token in 0.2 s instead of 2.5 s, streaming from two SSDs, same tokens as stock ds4. The trick wasn't a kernel: I made the codebase small enough for an agent
Chida82
—
09-18 18:28amplified on hacker newshn.story.49758274
AlmostCosmo79
peak 3 · 0 comments · 33% of case engagement
10-07 17:33amplified on r/LocalLLaMA 👑reddit.post.1x02t8t
Chida82
peak 0 · 11 comments · 66% of case engagement
09-18 19:20our radar first saw it · +77.3hdiscovery anchor: hn.story.49758274—
pace: p23 vs 1032 stories at the 336h mark (now 626h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnSSD-Llama: SSD-Native Inference for Trillion-Parameter Moe on a Consumer PC
Retrieved article excerpt

Open article · Retrieved 2026-09-18T19:22:33.396576+00:00

# Computer Science > Distributed, Parallel, and Cluster Computing

**arXiv:2609.18110** (cs)

[Submitted on 16 Sep 2026]

# Title:SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC

Authors:[Fangzhou Liang](https://arxiv.org/search/cs?searchtype=author&query=Liang,+F), [Yibin Shen](https://arxiv.org/search/cs?searchtype=author&query=Shen,+Y), [Jianmin Hu](https://arxiv.org/search/cs?searchtype=author&query=Hu,+J), [Jiayang Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+J), [Hanchi Gao](https://arxiv.org/search/cs?searchtype=author&query=Gao,+H), [Minxian Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+M), [Zili Meng](https://arxiv.org/search/cs?searchtype=author&query=Meng,+Z)

View a PDF of the paper titled SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC, by Fangzhou Liang and 5 other authors

[View PDF](https://arxiv.org/pdf/2609.18110)
[HTML (experimental)](https://arxiv.org/html/2609.18110v1)
> Abstract:Frontier open-weight language models increasingly use Mixture-of-Experts (MoE) architectures to expand model capacity while activating only a small subset of experts per token. Local inference must nevertheless keep the complete expert pool available, which remains far beyond consumer-grade RAM and VRAM capacity even after quantization. SSDs provide practical capacity at this scale, but turning that capacity into executable model memory requires efficient expert delivery, coordinated management of SSD, RAM, and VRAM, and CPU--GPU hybrid execution under bounded bandwidth. We present \textit{SSD-LLaMA}, an SSD-native local MoE inference system that addresses these challenges with an SSD I/O pipeline optimized for expert delivery, a native three-tier storage hierarchy that delivers and retains experts dynamically, and balanced CPU--GPU hybrid execution. \textit{SSD-LLaMA} executes every selected expert without pruning or substitution. Across three frontier MoE model families, \textit{SSD-LLaMA} improves prefill token rate by 1.52$\times$--4.19$\times$ and decode token rate by 2.10$\times$--15.58$\times$ over the evaluated baselines. We also achieve higher than 1 token/s for running trillion-parameter model with a single RTX 5090 and no more than 32GB RAM.

|  |  |
| --- | --- |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC) |
| Cite as: | [arXiv:2609.18110](https://arxiv.org/abs/2609.18110) [cs.DC] |
|  | (or  [arXiv:2609.18110v1](https://arxiv.org/abs/2609.18110v1) [cs.DC] for this version) |
|  | <https://doi.org/10.48550/arXiv.2609.18110> Focus to learn more  arXiv-issued DOI via DataCite (pending registration) |

## Submission history

From: Zili Meng [[view email](https://arxiv.org/show-email/57effe33/2609.18110)]   
 **[v1]**
Wed, 16 Sep 2026 04:28:47 UTC (578 KB)

Full-text links:

## Access Paper:

View a PDF of the paper titled SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC, by Fangzhou Liang and 5 other authors

- [View PDF](https://arxiv.org/pdf/2609.18110)
- [HTML (experimental)](https://arxiv.org/html/2609.18110v1)
- [TeX Source](https://arxiv.org/src/2609.18110)

[license icon](http://creativecommons.org/licenses/by-nc-nd/4.0/ "Rights to this article")

### Current browse context:

cs.DC

[< prev](https://arxiv.org/prevnext?id=2609.18110&function=prev&context=cs.DC "previous in cs.DC (accesskey p)")
  |   
[next >](https://arxiv.org/prevnext?id=2609.18110&function=next&context=cs.DC "next in cs.DC (accesskey n)")

[new](https://arxiv.org/list/cs.DC/new)
 | 
[recent](https://arxiv.org/list/cs.DC/recent)
 | [2026-09](https://arxiv.org/list/cs.DC/2026-09)

Change to browse by:

[cs](https://arxiv.org/abs/2609.18110?context=cs)

### References & Citations

- [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2609.18110)
- [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2609.18110)
- [Semantic Scholar](https://api.semanticscholar.org/arXiv:2609.18110)

export BibTeX citation
Loading...

## BibTeX formatted citation

×

loading...

Data provided by:

### Bookmark

[BibSonomy](http://www.bibsonomy.org/BibtexHandler?requTask=upload&url=https://arxiv.org/abs/2609.18110&description=SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC "Bookmark on BibSonomy")
[Reddit](https://reddit.com/submit?url=https://arxiv.org/abs/2609.18110&title=SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC "Bookmark on Reddit")



Bibliographic Tools

# Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer *([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))*

Connected Papers Toggle

Connected Papers *([What is Connected Papers?](https://www.connectedpapers.com/about))*

Litmaps Toggle

Litmaps *([What is Litmaps?](https://www.litmaps.co/))*

scite.ai Toggle

scite Smart Citations *([What are Smart Citations?](https://www.scite.ai/))*

Code, Data, Media

# Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv *([What is alphaXiv?](https://alphaxiv.org/))*

Links to Code Toggle

CatalyzeX Code Finder for Papers *([What is CatalyzeX?](https://www.catalyzex.com))*

DagsHub Toggle

DagsHub *([What is DagsHub?](https://dagshub.com/))*

GotitPub Toggle

Gotit.pub *([What is GotitPub?](http://gotit.pub/faq))*

Huggingface Toggle

Hugging Face *([What is Huggingface?](https://huggingface.co/huggingface))*

ScienceCast Toggle

ScienceCast *([What is ScienceCast?](https://sciencecast.org/welcome))*

Demos

# Demos

Replicate Toggle

Replicate *([What is Replicate?](https://replicate.com/docs/arxiv/about))*

Spaces Toggle

Hugging Face Spaces *([What is Spaces?](https://huggingface.co/docs/hub/spaces))*

Spaces Toggle

TXYZ.AI *([What is TXYZ.AI?](https://txyz.ai))*

Related Papers

# Recommenders and Search Tools

Link to Influence Flower

Influence Flower *([What are Influence Flowers?](https://influencemap.cmlab.dev/))*

Core recommender toggle

CORE Recommender *([What is CORE?](https://core.ac.uk/services/recommender))*

- Author
- Venue
- Institution
- Topic


About arXivLabs

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.18110) |
Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
AlmostCosmo7930
🟧 echo.paper ⭐SSD-LLaMA combines SSD expert delivery, a three-tier storage hierarchy, and CPU–GPU execution, reporting above one token per second for a trFangzhou Liang and coauthors——
🟠 redditA 341 GB DeepSeek on a 128 GB Mac: 2x decode and first token in 0.2 s instead of 2.5 s, streaming from two SSDs, same tokens as stock ds4. The trick wasn't a kernel: I made the codebase small enough for an agent
LocalLLaMA
Chida82011

Interpretation history

Decision trace