2026-10-11 17:11 UTC

Perplexity claims its open-sourced Lily server provides a model-specific inference path that makes Qwen deployment faster and more practical on Apple Silicon Macs.

state: expiredheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference apple-silicon inference-runtimePerplexity

What is this?

Perplexity introduced Lily, a small inference server specialized for running Qwen models on Apple Silicon, and reportedly open-sourced it through its repository. Its architecture uses a Rust runtime, an OpenAI-compatible streaming API, and custom Metal kernels, with neither PyTorch nor MLX in the execution path. Perplexity reports that Lily outperformed the compared path at every recorded prompt and context length for a quantized Qwen3.6-35B-A3B model on a 128 GB M5 Max, though the supplied snippets do not establish broader performance across other Macs, models, or workloads.

Why it matters to Scott

Perplexity’s model-specific Metal runtime converges with Scott’s hardware-aware local-inference practice, while its OpenAI-compatible API preserves the swappable gateway boundary used in his systems. It is actionable as a serving-path benchmark, but the narrow Qwen/M5 Max evidence and model-specific implementation also raise Scott’s model-perishability concern rather than establishing a broadly superior Mac runtime.
dev:concept.hardware-aware-local-inferencedev:technology.litellmip:concept.model-perishabilitydev:technology.ollamaradar:concept.inference-enginesradar:concept.apple-silicon-inferenceradar:m5-w8a8-prefill-kernelsradar:llama-cpp-metal-iq3-moe-speedup
queries asked of Scott's wikis
  • model-specific inference runtimes vs general frameworks
  • Apple Silicon local inference strategy
  • custom Metal kernels for LLM serving
  • OpenAI-compatible local model servers
  • local inference performance and hardware economics
  • specialized runtimes vs MLX and PyTorch

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditPerplexity open-sourced their Mac inference server for Qwen 3.6
LocalLLaMA
Specter_Origin10220
🟧 echo.github ⭐The PerplexityAI repository commit “Add Lily Metal inference engine (#19)” introduced Lily, described in its README as “a small Metal infereYibo Wu——

Interpretation history

Decision trace