Independent benchmarks will determine whether AirLLM’s layer and expert streaming can run very large dense and sparse-MoE models on 4–12 GB GPUs with correct outputs and practically useful throughput.
state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference memory-efficiency sparse-moeAirLLMlyogavin
What is this?
AirLLM is an open-source inference library from the lyogavin GitHub account that claims to run models larger than GPU memory by streaming dense-model layers or individual MoE experts, using disk as an extension of memory. Its repository claims examples ranging from 70B dense models on 4 GB GPUs to much larger sparse-MoE models on 4–12 GB, without quantization, distillation, or pruning. The supplied results mostly repeat project and tutorial claims; they do not provide independent AirLLM benchmarks establishing output correctness, end-to-end latency, token throughput, storage requirements, or performance across hardware configurations.
Why it matters to Scott
The core position is already explicit in Scott’s Hardware-aware local inference and Evaluation-Driven Development pages, while the radar already has near-identical open validation cases for HotPin and Slipstream expert streaming. AirLLM is still relevant as a potentially testable runtime for Scott’s gamepc model-serving stack, but without independent correctness and throughput results it is another implementation candidate rather than a new conclusion.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentip:concept.operating-pointradar:hotpin-lossless-moe-streamingradar:slipstream-ssd-moe-streamingradar:concept.expert-streamingradar:concept.local-inference
queries asked of Scott's wikis
- local inference memory hierarchy and model offloading
- disk streaming versus quantization economics
- sparse MoE expert streaming
- minimum useful throughput for local LLMs
- inference correctness under memory optimization
- consumer hardware and model sovereignty
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-22T13:37:21Z
Repeated checks produced no independent correctness or throughput evidence, and the discussion has faded without a concrete benchmark expected. The current attention episode can expire; a reproducible benchmark or material release would justify reopening it.
2026-08-20T13:27:46Z
Refreshed discussion continues to ask for tokens-per-second rather than supplying measurements, reinforcing that practical throughput remains the decisive missing evidence. This is repetitive amplification of the existing validation gap, not independent corroboration.
2026-08-20T11:30:58Z
The reobservation is only minor engagement growth around the same first-party claims; no independent correctness or throughput benchmark has appeared. AirLLM remains an unvalidated implementation candidate rather than evidence that extreme low-VRAM streaming is practically useful.
2026-08-20T11:27:43Z
grounded: known/medium — The core position is already explicit in Scott’s Hardware-aware local inference and Evaluation-Driven Development pages, while the radar already has near-identi
2026-08-20T11:24:39Z
case created — The maintained implementation makes unusually strong, independently testable memory-efficiency claims for current large-model local inference.
Decision trace
- 08-22 23:37expireRepeated checks produced no independent correctness or throughput evidence, and the discussion has faded without a concrete benchmark expected. The current attention episode can expire; a reproducible
- 08-22 23:37alert_silentThe only trigger is staleness, with no new evidence or consequential change since the prior review; Scott can wait unless an independent benchmark or documented release appears.
- 08-22 23:37alert_routeThe only trigger is staleness, with no new evidence or consequential change since the prior review; Scott can wait unless an independent benchmark or documented release appears.
- 08-21 13:21sensor_dirtyengagement_update
- 08-21 06:21sensor_dirtyengagement_update
- 08-21 02:21sensor_dirtyengagement_update
- 08-20 23:27repriceRefreshed discussion continues to ask for tokens-per-second rather than supplying measurements, reinforcing that practical throughput remains the decisive missing evidence. This is repetitive amplific
- 08-20 23:27alert_silentNo consequential delta occurred: engagement and benchmark requests increased, but no independent correctness result, throughput measurement, reproducible implementation, or release change appeared.
- 08-20 23:27alert_routeNo consequential delta occurred: engagement and benchmark requests increased, but no independent correctness result, throughput measurement, reproducible implementation, or release change appeared.
- 08-20 23:21sensor_dirtyengagement_update
- 08-20 22:21sensor_dirtycomment_update
- 08-20 21:30repriceThe reobservation is only minor engagement growth around the same first-party claims; no independent correctness or throughput benchmark has appeared. AirLLM remains an unvalidated implementation cand
- 08-20 21:30alert_silentNothing consequential changed beyond engagement, so this can wait for an independent benchmark, reproducible implementation result, or documented release change.
- 08-20 21:30alert_routeNothing consequential changed beyond engagement, so this can wait for an independent benchmark, reproducible implementation result, or documented release change.
- 08-20 21:28alert_silentThe only new evidence is a low-engagement Reddit post and a mirrored project claim, with no tagged release, verifiable publication date, independent correctness result, or usable-throughput benchmark.
- 08-20 21:28alert_routeThe only new evidence is a low-engagement Reddit post and a mirrored project claim, with no tagged release, verifiable publication date, independent correctness result, or usable-throughput benchmark.
- 08-20 21:27groundThe core position is already explicit in Scott’s Hardware-aware local inference and Evaluation-Driven Development pages, while the radar already has near-identical open validation cases for HotPin and
- 08-20 21:24createThe maintained implementation makes unusually strong, independently testable memory-efficiency claims for current large-model local inference.