Perplexity claims its open-sourced Lily server provides a model-specific inference path that makes Qwen deployment faster and more practical on Apple Silicon Macs.
state: expiredheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference apple-silicon inference-runtimePerplexity
What is this?
Perplexity introduced Lily, a small inference server specialized for running Qwen models on Apple Silicon, and reportedly open-sourced it through its repository. Its architecture uses a Rust runtime, an OpenAI-compatible streaming API, and custom Metal kernels, with neither PyTorch nor MLX in the execution path. Perplexity reports that Lily outperformed the compared path at every recorded prompt and context length for a quantized Qwen3.6-35B-A3B model on a 128 GB M5 Max, though the supplied snippets do not establish broader performance across other Macs, models, or workloads.
Why it matters to Scott
Perplexity’s model-specific Metal runtime converges with Scott’s hardware-aware local-inference practice, while its OpenAI-compatible API preserves the swappable gateway boundary used in his systems. It is actionable as a serving-path benchmark, but the narrow Qwen/M5 Max evidence and model-specific implementation also raise Scott’s model-perishability concern rather than establishing a broadly superior Mac runtime.
dev:concept.hardware-aware-local-inferencedev:technology.litellmip:concept.model-perishabilitydev:technology.ollamaradar:concept.inference-enginesradar:concept.apple-silicon-inferenceradar:m5-w8a8-prefill-kernelsradar:llama-cpp-metal-iq3-moe-speedup
queries asked of Scott's wikis
- model-specific inference runtimes vs general frameworks
- Apple Silicon local inference strategy
- custom Metal kernels for LLM serving
- OpenAI-compatible local model servers
- local inference performance and hardware economics
- specialized runtimes vs MLX and PyTorch
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-07T04:25:07Z
Lily remains a reportedly released, single-checkpoint serving option, but repeated checks have produced no independent benchmark or deployment evidence establishing its practical advantage. With no concrete follow-up expected, retire this episode without treating the performance claim as disproved; independent results or expanded hardware support would justify reopening.
2026-09-05T03:22:28Z
No new substantive evidence changes Lily’s interpretation: it remains a reportedly released, single-checkpoint runtime rather than a demonstrated improvement to Scott’s local-serving options. The echoed repository description and repeated discussion still provide no independent performance validation; an M5 Pro benchmark or actual deployment would warrant another look.
2026-09-03T02:37:40Z
The refreshed discussion still adds only compatibility questions and reactions to Perplexity’s published numbers, not independent benchmarks or deployments. Lily remains a concrete but narrowly constrained release whose practical advantage beyond the specified M5/Qwen path is uncorroborated.
2026-09-02T23:38:49Z
Refreshed comments remain user curiosity and repetition rather than independent benchmarking or implementation evidence. Lily is still a concrete but narrowly specialized release, with its performance advantage and broader Mac applicability uncorroborated.
2026-09-02T22:47:14Z
The refreshed discussion shows early user curiosity but provides no independent benchmarks, implementations, or broader hardware results. Lily remains a concrete narrow release whose claimed speed advantage is still unvalidated outside Perplexity’s comparison.
2026-09-02T22:29:07Z
grounded: converges/medium — Perplexity’s model-specific Metal runtime converges with Scott’s hardware-aware local-inference practice, while its OpenAI-compatible API preserves the swappabl
2026-09-02T22:25:19Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1w5ozl4 -> echo.github.3fd9d19cba by Yibo Wu
2026-09-02T22:24:09Z
case created — The linked Perplexity repository is a concrete open-source runtime release, but performance evidence and community uptake are not yet visible.
Decision trace
- 09-07 14:25expireLily remains a reportedly released, single-checkpoint serving option, but repeated checks have produced no independent benchmark or deployment evidence establishing its practical advantage. With no co
- 09-07 14:25alert_silentThere is no new release, access, compatibility, or performance delta to surface. The existing narrow release can remain in Scott’s benchmark backlog without an interruption.
- 09-07 14:25alert_routeThere is no new release, access, compatibility, or performance delta to surface. The existing narrow release can remain in Scott’s benchmark backlog without an interruption.
- 09-05 13:22repriceNo new substantive evidence changes Lily’s interpretation: it remains a reportedly released, single-checkpoint runtime rather than a demonstrated improvement to Scott’s local-serving options. The echo
- 09-05 13:22alert_silentThis look adds no release, access, capability, or deployment delta. The narrow runtime remains suitable for a briefing or benchmark backlog, without a reason to interrupt Scott today.
- 09-05 13:22alert_routeThis look adds no release, access, capability, or deployment delta. The narrow runtime remains suitable for a briefing or benchmark backlog, without a reason to interrupt Scott today.
- 09-04 03:22sensor_dirtyengagement_update
- 09-03 23:21sensor_dirtyengagement_update
- 09-03 21:21sensor_dirtyengagement_update
- 09-03 20:21sensor_dirtyengagement_update
- 09-03 18:21sensor_dirtyengagement_update
- 09-03 16:21sensor_dirtyengagement_update
- 09-03 15:21sensor_dirtyengagement_update
- 09-03 14:21sensor_dirtyengagement_update
- 09-03 13:21sensor_dirtyengagement_update
- 09-03 12:37repriceThe refreshed discussion still adds only compatibility questions and reactions to Perplexity’s published numbers, not independent benchmarks or deployments. Lily remains a concrete but narrowly constr
- 09-03 12:37alert_silentThe release has already been routed, and the new comments do not change confidence in performance, compatibility, or adoption; wait for an independent benchmark or implementation report.
- 09-03 12:37alert_routeThe release has already been routed, and the new comments do not change confidence in performance, compatibility, or adoption; wait for an independent benchmark or implementation report.
- 09-03 12:21sensor_dirtycomment_update
- 09-03 10:21sensor_dirtyengagement_update
- 09-03 09:38repriceRefreshed comments remain user curiosity and repetition rather than independent benchmarking or implementation evidence. Lily is still a concrete but narrowly specialized release, with its performance
- 09-03 09:38alert_silentThe release was already routed, and the new comments add no consequential capability, compatibility, or adoption evidence; this can wait for an independent benchmark or implementation report.
- 09-03 09:38alert_routeThe release was already routed, and the new comments add no consequential capability, compatibility, or adoption evidence; this can wait for an independent benchmark or implementation report.
- 09-03 09:21sensor_dirtycomment_update
- 09-03 08:47repriceThe refreshed discussion shows early user curiosity but provides no independent benchmarks, implementations, or broader hardware results. Lily remains a concrete narrow release whose claimed speed adv
- 09-03 08:47alert_silentThe open-source release was already routed; the new delta is only light discussion and does not materially change capability confidence or practical adoption.
- 09-03 08:47alert_routeThe open-source release was already routed; the new delta is only light discussion and does not materially change capability confidence or practical adoption.
- 09-03 08:42alert_shadowPerplexity’s repository establishes an Apache-2.0, OpenAI-compatible serving path specialized for Qwen3.6-35B-A3B on Apple Silicon, making it available for comparison with Scott’s local-inference stac
- 09-03 08:42alert_routePerplexity’s repository establishes an Apache-2.0, OpenAI-compatible serving path specialized for Qwen3.6-35B-A3B on Apple Silicon, making it available for comparison with Scott’s local-inference stac
- 09-03 08:29groundPerplexity’s model-specific Metal runtime converges with Scott’s hardware-aware local-inference practice, while its OpenAI-compatible API preserves the swappable gateway boundary used in his systems.
- 09-03 08:25promote_anchororigin walk conf 0.98
- 09-03 08:24createThe linked Perplexity repository is a concrete open-source runtime release, but performance evidence and community uptake are not yet visible.