2026-10-11 17:12 UTC

Independent benchmarks will determine whether AirLLM’s layer and expert streaming can run very large dense and sparse-MoE models on 4–12 GB GPUs with correct outputs and practically useful throughput.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference memory-efficiency sparse-moeAirLLMlyogavin

What is this?

AirLLM is an open-source inference library from the lyogavin GitHub account that claims to run models larger than GPU memory by streaming dense-model layers or individual MoE experts, using disk as an extension of memory. Its repository claims examples ranging from 70B dense models on 4 GB GPUs to much larger sparse-MoE models on 4–12 GB, without quantization, distillation, or pruning. The supplied results mostly repeat project and tutorial claims; they do not provide independent AirLLM benchmarks establishing output correctness, end-to-end latency, token throughput, storage requirements, or performance across hardware configurations.

Why it matters to Scott

The core position is already explicit in Scott’s Hardware-aware local inference and Evaluation-Driven Development pages, while the radar already has near-identical open validation cases for HotPin and Slipstream expert streaming. AirLLM is still relevant as a potentially testable runtime for Scott’s gamepc model-serving stack, but without independent correctness and throughput results it is another implementation candidate rather than a new conclusion.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentip:concept.operating-pointradar:hotpin-lossless-moe-streamingradar:slipstream-ssd-moe-streamingradar:concept.expert-streamingradar:concept.local-inference
queries asked of Scott's wikis
  • local inference memory hierarchy and model offloading
  • disk streaming versus quantization economics
  • sparse MoE expert streaming
  • minimum useful throughput for local LLMs
  • inference correctness under memory optimization
  • consumer hardware and model sovereignty

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAirLLM - Recent Updates - with Qwen3.8-27B, Kimi-K3 too
LocalLLaMA
pmttyji298
🟧 echo.github ⭐AirLLM claims to reduce inference memory through layer offloading and per-expert streaming, including running 70B dense models on 4 GB GPUs lyogavin——

Interpretation history

Decision trace