2026-10-11 17:11 UTC

Independent benchmarks will determine whether llama.cpp’s GLM-4.5-Air MTP support materially improves local inference speed across memory-rich, compute-limited hardware without reducing output quality or stability.

state: expiredheat: lowuncertainty: highknownscott: mediumllama-cpp speculative-decoding local-inferencellama.cppZ.ai

What is this?

llama.cpp has added Multi-Token Prediction (MTP) support for Z.ai’s GLM-4.5-Air, an inference optimization intended to accelerate generation through speculative token acceptance. The supplied snippets show that speculative decoding can substantially raise throughput and that GLM-4.5-Air can run on single-GPU or hybrid CPU/GPU systems, but they do not provide independent, controlled benchmarks of this specific implementation. Reported MTP gains elsewhere range from roughly 2× to 3× under short-test conditions, while warnings about longer contexts and a separate llama.cpp performance regression leave speed, quality, and stability across memory-rich but compute-limited hardware unresolved.

Why it matters to Scott

The radar already tracks llama.cpp MTP as an open performance-and-regression story in `radar:llama-cpp-adaptive-mtp` and `radar:llama-cpp-mtp-default-memory-regression`; GLM-4.5-Air support is a model-specific extension rather than a new position. It matters to Scott’s hardware-aware local-inference work and self-hosted GPU substrate because controlled throughput, quality, context-length, and memory tests could affect serving choices, though the supplied material does not show that he currently runs GLM-4.5-Air.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentradar:llama-cpp-adaptive-mtpradar:llama-cpp-mtp-default-memory-regressionradar:concept.llama-cppradar:concept.speculative-decoding
queries asked of Scott's wikis
  • speculative decoding acceptance rate and quality tradeoffs
  • local inference on memory-rich compute-limited hardware
  • llama.cpp performance benchmarking methodology
  • hybrid CPU GPU inference optimization
  • local model throughput versus long-context stability
  • multi-token prediction in open-model serving

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddityou can now use MTP in GLM-Air
LocalLLaMA
jacek202310717
🟧 echo.github ⭐The primary source is PR #26534, titled “model : support MTP in GLM-4.5-Air.” It implements glm4moe graph_mtp support, adds converter and lojacekpoplawski——

Interpretation history

Decision trace