2026-10-11 17:12 UTC

Independent benchmarks will determine whether Swiftlet can reproducibly run an 80B Qwen model in roughly 4.3GB of Mac memory and a 35B model on an iPhone at practically useful speed and output quality.

state: expiredheat: lowuncertainty: highconvergesscott: mediummemory-efficient-inference mobile-llm local-inferenceLeon NicksonSwiftlet

What is this?

Swiftlet, attributed in the case to Leon Nickson, is an implementation claiming to run sparse Qwen models on Apple devices by keeping memory use unusually low: Qwen3-Next-80B-A3B in roughly 4.3GB on a Mac and Qwen3.6-35B-A3B on an iPhone. The supplied web material does not independently verify those Swiftlet-specific claims; it only shows that an aggressively quantized related 35B-A3B model can run on a 16GB M4 Mac at about 5–6 tokens per second, with meaningful quality loss. The snippets provide no independent iPhone benchmark and suggest the 80B memory claim remains uncertain, so reproducibility, speed, and output quality are unresolved.

Why it matters to Scott

Swiftlet’s claimed memory, speed, and quality tradeoffs directly extend Scott’s hardware-aware local-inference practice and his preference for cheap, deployable capability over nominal model power. If independently reproduced, it could materially broaden his deployment options beyond GPU workstations; for now it remains an unverified implementation claim, and the radar tracks closely related MoE streaming and compression cases but not this Swiftlet development itself.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.usable-mass-over-unusable-powerip:concept.evaluation-driven-developmentradar:concept.local-inferenceradar:slipstream-ssd-moe-streamingradar:compressed-llm-fidelity-safety-gapradar:concept.moe-inferenceradar:concept.model-evaluation
queries asked of Scott's wikis
  • memory-efficient local inference architecture
  • extreme quantization quality tradeoffs
  • Apple Silicon and iPhone LLM inference
  • sparse MoE inference memory economics
  • local model benchmark reproducibility
  • on-device AI model sovereignty

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhoneleonickson310141
🟧 echo.github ⭐The initial Swiftlet commit presents the implementation itself: “Runs Qwen3.6-35B-A3B and Qwen3-Next-80B-A3B on Apple devices” by keeping deLeonickson——

Interpretation history

Decision trace