2026-10-11 17:12 UTC

WARP’s creator claims the engine can run GLM-5.3-Flash using as little as 5.14GB of memory and reach about 3.3 tokens per second on a 64GB Apple Silicon Mac, making very large sparse models locally runnable with modest memory.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference open-models inference-economicsWARPMarco BambiniZ.ai

What is this?

WARP is presented as a local inference engine whose creator, Marco Bambini, claims it can run Z.ai’s sparse 313B-parameter GLM-5.3-Flash with a reported 5.14GB memory footprint and roughly 3.3 tokens per second on a 64GB Apple Silicon Mac. The cited README performance table reportedly lists 3.32 tok/s over 64 tokens and 3.86 tok/s over 200 tokens. However, the supplied web results discuss general Apple Silicon inference and the smaller GLM-4.7-Flash rather than independently establishing WARP, its architecture, or the GLM-5.3 benchmark, so the exceptional memory claim remains unverified here.

Why it matters to Scott

The claimed engine converges with Scott’s hardware-aware local-inference work and Usable Mass thesis by suggesting that sparse frontier-scale models can become deployable on commodity unified-memory hardware. If independently reproduced, the 5.14GB/3.3 tok/s result would materially change local-inference economics and warrant testing against his model-serving substrate, but the supplied evidence is currently creator testimony rather than a reproducible benchmark.
ip:concept.usable-mass-over-unusable-powerip:concept.ai-unit-economicsip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentdev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.local-inferenceradar:concept.sparse-moeradar:concept.apple-silicon-inferenceradar:concept.inference-economicsradar:deepseek-v4-flash-expert-streamingradar:freetoken-290b-moe-local-inferenceradar:swiftlet-ultralow-memory-inference
queries asked of Scott's wikis
  • sparse MoE local inference memory economics
  • Apple Silicon inference engines and unified memory
  • local model sovereignty through commodity hardware
  • extreme quantization versus model quality
  • open-weight frontier models for coding agents
  • local inference benchmarks and reproducibility

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGLM-5.3-Flash at 3.3 tok/smarcobambini10
🟧 echo.github ⭐The README’s performance table states: “GLM-5.3-Flash 313B … 3.32 tok/s” over 64 tokens and “3.86 tok/s” over 200. The commit message says tMarco Bambini——
🟧 hnGLM-5.3-Flash on Apple Siliconmarcobambini20
🟠 redditIs anyone successfully running GLM 5.3 Flash locally yet?
LocalLLaMA
CentrifugalMalaise250

Interpretation history

Decision trace