2026-10-11 16:38 UTC

Redditor ResearchCrafty1804 claims the released Inco Splash engine runs Qwen3.8-27B at 144 tokens per second on an M5 Max and delivers up to threefold Ollama decode speed, potentially making local coding-agent inference substantially more responsive on supported Macs.

state: corroboratedheat: lowuncertainty: highconvergesscott: mediumapple-silicon-inference local-model-serving coding-agents inference-economicsIncoResearchCrafty1804
Surfaced 2026-09-19T19:30:16Z — Inco AI’s launch post introduces Splash, an open-source Apple-silicon inference engine. It says Splash is built around each model, delivers — The velocity spike and loud cross-platform spread reading raise Splash’s attention priority, but the supplied evidence adds no new implementation result or benchmark validation. This is an expanding launch to watch, not yet a corroborated runtime advantage or a reason for Scott to change serving choices.

What is this?

Splash is an open-source C++/Metal inference engine from Inco AI, built specifically around supported models and Apple Silicon; its Qwen3.8-27B package combines a 4-bit target with a dedicated DFlash 2 draft model for speculative decoding. Inco claims 144 tok/s on an M5 Max, up to 3× Ollama decode speed, and—on an M5 Pro—74 tok/s for short prompts and 170 tok/s aggregate across four requests; LM Studio now offers Splash as an experimental backend on M3-or-newer Macs with macOS 26.4+ and at least 36 GB unified memory. The supplied case records independent users achieving roughly 37–129 tok/s and creating 8-bit and 4-bit extensions, corroborating that the runtime works and has an emerging implementation periphery, but not independently validating the 144 tok/s headline, matched-runtime multipliers, quality parity, or coding-agent gains.

Why it matters to Scott

The working Metal runtime and independent 8-bit/4-bit extensions converge with Scott’s hardware-aware local-inference approach by treating model architecture, precision, memory, and accelerator placement as explicit serving policy. This now bears on his active Ollama-based local-serving work enough to justify evaluation, but Splash is Mac-specific while his documented host is CUDA-based, and neither the Ollama speedup nor compatibility with Ask has been established.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaip:concept.latencyradar:dflash-2-parallel-drafting-validationradar:perplexity-lily-apple-silicon-inferenceradar:concept.speculative-decodingradar:concept.apple-siliconradar:concept.inference-engines
queries asked of Scott's wikis
  • hardware-aware model-specific inference
  • speculative decoding for local coding agents
  • Apple Silicon versus CUDA serving strategy
  • local inference latency and agent responsiveness
  • quantization quality versus throughput tradeoffs
  • parallel local agent inference economics

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 602h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 14:00⭐ origin echo-reconstructedInco AI’s launch post introduces Splash, an open-source Apple-silicon inference engine. It says Splash is built around each model, delivers
Inco AI on blog (echo) · attributed from reddit.post.1wk9hze
—
09-19 02:11first on r/LocalLLaMA · published · +60.2hQwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro
ResearchCrafty1804
—
09-19 11:54first on hacker news · published · +69.9hSplash Engine – The Fastest Local Qwen3.8 on Apple Silicon
sebiw
—
09-19 02:11amplified on r/LocalLLaMA 👑reddit.post.1wk9hze
ResearchCrafty1804
peak 221 · 87 comments · 48% of case engagement
09-19 11:54amplified on hacker newshn.story.49765787
sebiw
peak 3 · 0 comments · 1% of case engagement
09-20 13:55amplified on r/LocalLLaMAreddit.post.1wlhqbr
JLeonsarmiento
peak 0 · 2 comments · 0% of case engagement
09-21 12:25amplified on r/LocalLLaMAreddit.post.1wmbbf9
SnooPredictions515
peak 63 · 34 comments · 15% of case engagement
09-22 10:33amplified on r/LocalLLaMAreddit.post.1wn5u74
SeveralViolins
peak 11 · 36 comments · 7% of case engagement
09-25 13:13amplified on r/LocalLLaMAreddit.post.1wpw0jn
SeveralViolins
peak 13 · 11 comments · 4% of case engagement
3 more amplifiers in ainews.case_chain
09-19 02:20our radar first saw it · +60.3hdiscovery anchor: reddit.post.1wk9hze—
09-19 19:30reached heat=high · +77.5h · via ledger——
pace: p86 vs 1032 stories at the 336h mark (now 602h old) — ahead of cactus-needle3-on-device-automation (1.0x), behind world-labs-joins-amd (1.0x)

Evidence (10) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-19T02:21:41.660953+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
ResearchCrafty180422187
🟧 echo.blog ⭐Inco AI’s launch post introduces Splash, an open-source Apple-silicon inference engine. It says Splash is built around each model, delivers Inco AI——
🟧 hnSplash Engine – The Fastest Local Qwen3.8 on Apple Siliconsebiw30
🟠 redditinco.AI: is it worth to jump from non AI-polluted Sequoia to GoldenGate-AI-slop-fest just to run that engine?
LocalLLaMA
JLeonsarmiento02
🟠 reddit[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"
LocalLLaMA
SnooPredictions5156334
🟠 redditSiliconSpecies/Swift-Qwen3.8-27B-Splash
LocalLLaMA
SeveralViolins1136
🟠 redditSplash on a 40-core M5 Max: +20% decode by tuning the kernels for your own chip
LocalLLaMA
SeveralViolins1311
🟠 redditSplash 1.1.0 released, GGUF quants support, MLX import and more
LocalLLaMA
wojtek153811
🟠 redditRun Qwen3.8+Flash-Next and tiny models on Apple Silicon up to 3x faster
LocalLLaMA
VagabondTruffle3317
🟠 redditSplash fork optimised for M5 Max: ~1.5× faster (1.25× single request)
LocalLLaMA
SeveralViolins3420

Interpretation history

Decision trace