Splash is an open-source C++/Metal inference engine from Inco AI, built specifically around supported models and Apple Silicon; its Qwen3.8-27B package combines a 4-bit target with a dedicated DFlash 2 draft model for speculative decoding. Inco claims 144 tok/s on an M5 Max, up to 3× Ollama decode speed, and—on an M5 Pro—74 tok/s for short prompts and 170 tok/s aggregate across four requests; LM Studio now offers Splash as an experimental backend on M3-or-newer Macs with macOS 26.4+ and at least 36 GB unified memory. The supplied case records independent users achieving roughly 37–129 tok/s and creating 8-bit and 4-bit extensions, corroborating that the runtime works and has an emerging implementation periphery, but not independently validating the 144 tok/s headline, matched-runtime multipliers, quality parity, or coding-agent gains.
The working Metal runtime and independent 8-bit/4-bit extensions converge with Scott’s hardware-aware local-inference approach by treating model architecture, precision, memory, and accelerator placement as explicit serving policy. This now bears on his active Ollama-based local-serving work enough to justify evaluation, but Splash is Mac-specific while his documented host is CUDA-based, and neither the Ollama speedup nor compatibility with Ask has been established.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaip:concept.latencyradar:dflash-2-parallel-drafting-validationradar:perplexity-lily-apple-silicon-inferenceradar:concept.speculative-decodingradar:concept.apple-siliconradar:concept.inference-engines
queries asked of Scott's wikis
- hardware-aware model-specific inference
- speculative decoding for local coding agents
- Apple Silicon versus CUDA serving strategy
- local inference latency and agent responsiveness
- quantization quality versus throughput tradeoffs
- parallel local agent inference economics
2026-09-30T10:04:23Z
A sixth independent hand validated the tuned-kernel work cross-chip: MikeBuckets171 (20-core M5 Pro 48 GB) ran Splish v8's M5 Max kernel tables through all 24 kernel-bench checks cleanly and got Pro-side gains from v8's new 3–4-request concurrency entries — a real portability result that deepens the periphery (~5–6 independent builders spanning M1 through M5 Max) without broadening it beyond one community. Attention sits at absolute zero (0 pts/h, 0 comments/h; the magnitude-valve cross-platform reading remains the Sept 19–22 launch-wave artifact), so the case holds at corroborated and cold: a maintained, build-on-able engine whose headline multipliers and quality parity are still unverified.
2026-09-27T14:41:16Z
The re-fired velocity spike is yesterday's artifact class again: ~3.5 pts/h absolute on the 21-point Splish fork post (4.2× only off a 0.83 pt/h baseline) plus a comment-ratio dip on the ishizuki post, with no new implementation, benchmark, or substantive thread — momentum reads 'accelerating' solely because the base is near zero, and the 75th peer percentile sits inside a mostly-dormant 11-day cohort, so this is attention noise, not a re-entering wave. One soft update to the periphery picture: the Splish post references a second party's M1 port (Erp4759, 'part 2'), so hands-on uptake spans the chip line across ~4–5 independent builders — still one community at small engagement, sustaining corroborated without acceleration.
2026-09-27T07:55:45Z
The Splish fork adds concrete follow-on engineering uptake — a third party found the engine worth forking and measured 1.5× concurrency / 1.25× single-request gains on an M5 Max — but it comes from the same author behind the Swift converter and kernel-tuning posts, so the periphery has deepened rather than broadened. The case holds at corroborated: a maintained, build-on-able engine in its quiet post-launch phase (the loud cross-platform magnitude reading remains a Sept 19–22 wave artifact; current rates ~1 pt/h and cooling), with the headline multipliers and quality parity still unverified.
2026-09-27T07:22:26Z
evidence attached: reddit.post.1wrd1p1 — A third-party M5 Max fork (Splish) with measured 1.5x concurrency speedups is concrete follow-on uptake of the same Splash-on-Apple-Silicon episode.
2026-09-26T21:37:15Z
The sensor's 'velocity spike' is a 3.2× multiple on a ~1 pt/h peer baseline — ~4 pts/h absolute with ~1 comment/h and no new substantive discussion, so it buys attention, not belief. With the implementation periphery narrowed to single-author, single-digit-engagement posts and the new ishizuki evidence adding wave-context but no validation of Splash's numbers, the case settles from accelerating to corroborated: a maintained, independently corroborated Apple-Silicon engine whose headline multipliers remain unverified.
2026-09-26T21:23:10Z
evidence attached: reddit.post.1wr1qlz — A second Apple-Silicon engine (MLX-based ishizuki) claiming up-to-3x Qwen-family gains contextualizes Splash's 3x decode claim as a broader engine wave rather than a one-off.
2026-09-26T18:55:41Z
Splash 1.1.0 ships GGUF quant support and MLX import, dissolving the friction that defined the community periphery — the reverse-engineered byte-level converters are no longer the price of admission, and standard quants now run on the engine (one report of comfortable 50 tok/s agentic Qwen3.8-27B on an M5 Pro). The case consolidates from 'promising runtime with hacky extensions' toward 'maintained tool with standard-format support'; upstream substance keeps moving while launch-wave attention stays drained, so the magnitude-valve's cross-platform reading remains a Sept 19–22 artifact and heat prices low.
2026-09-26T18:25:02Z
evidence attached: reddit.post.1wqw9rn — Splash 1.1.0 with GGUF quant and MLX import support plus a user report of comfortable 50 tok/s agentic Qwen3.8-27B use on an M5 Pro is continued development and early adoption evidence for the exact engine this case tracks.
2026-09-25T13:43:37Z
The kernel-tuning measurement reframes Splash performance as chip-dependent and tunable (+20% decode on a 40-core M5 Max via make tune-kernels, upstream issue filed), lending indirect plausibility to Inco's 144 tok/s headline without reproducing it. The implementation periphery keeps its stock but its flow has narrowed — the two most recent additions are from a single author at single-digit engagement — so despite the magnitude-valve's loud cross-platform spread reading (a Sept 19–22 wave artifact: now 0.17 pts/h vs a 38 peak, zero comments/h), heat prices low: a quiet, real implementation ecosystem, not an active story.
2026-09-25T13:25:17Z
evidence attached: reddit.post.1wpw0jn — First-person measurement that Splash's default kernels are mistuned on 40-core M5 Max with a documented +20% decode from make tune-kernels, filed upstream — independent corroboration/context for an accelerating performance case.
2026-09-24T00:48:48Z
grounded: converges/medium — The working Metal runtime and independent 8-bit/4-bit extensions converge with Scott’s hardware-aware local-inference approach by treating model architecture, p
2026-09-24T00:45:48Z
A second independent extension now adds a custom Splash-compatible 4-bit artifact and reports 79–129 tok/s on an M5 Max, broadening Splash from a working release into an emerging implementation ecosystem. This materially strengthens the practical-runtime case, but current activity has cooled and neither 144 tok/s nor the Ollama/oMLX comparisons are independently established.
2026-09-22T11:23:45Z
evidence attached: reddit.post.1wn5u74 — This independently corroborates unusually fast Splash-based Qwen3.8-27B inference on Apple Silicon with concrete context-length measurements.
2026-09-21T14:22:10Z
An independent hands-on benchmark and native 8-bit extension now corroborate that Splash is a working, extensible Apple Silicon runtime delivering useful Qwen3.8-27B throughput, rather than merely a launch claim. It still does not validate 144 tok/s, the claimed Ollama/oMLX multipliers, quality parity, or end-to-end coding-agent gains, but the expanding implementation periphery keeps attention high.
2026-09-21T13:22:10Z
evidence attached: reddit.post.1wmbbf9 — This supplies additional hands-on Splash measurements for Qwen3.8 on Apple Silicon, materially informing the existing local-inference case.
2026-09-20T14:27:10Z
The new Reddit question exposes hesitation about upgrading macOS to use Splash, but supplies neither an installation failure nor a performance result; it does not establish broader adoption friction. The loud spread reading continues to justify high attention without corroborating the claimed speed advantage.
2026-09-20T14:22:09Z
evidence attached: reddit.post.1wlhqbr — This is adoption friction and user evaluation surrounding the existing Inco Splash local-inference release, though evidence is weak.
2026-09-19T21:46:42Z
Scott’s explicit up-vote raises Splash from a merely adjacent optimization example to an expressed evaluation interest, without validating its performance claims or establishing fit with his stack. The prior loud spread reading still warrants high attention; this update adds no technical evidence.
2026-09-19T19:30:16Z
The velocity spike and loud cross-platform spread reading raise Splash’s attention priority, but the supplied evidence adds no new implementation result or benchmark validation. This is an expanding launch to watch, not yet a corroborated runtime advantage or a reason for Scott to change serving choices.
2026-09-19T12:22:12Z
The HN submission extends Splash’s footprint to another community but supplies no independent testing, despite its attachment being labeled independent coverage. A commenter’s unsubstantiated q4_0 attribution sharpens the need for matched quantization and quality comparisons; it does not establish or refute the claimed runtime advantage.
2026-09-19T12:21:38Z
evidence attached: hn.story.49765787 — This is direct independent coverage of Splash Engine’s claimed Qwen3.8 speedup on Apple Silicon.
2026-09-19T09:24:08Z
A firsthand user report moves Splash beyond launch-only testimony toward evidence of usable local coding inference, but does not reproduce the headline benchmark or establish an advantage over other runtimes. Missing quantization and matched baselines keep the claimed speedup—and its value for coding agents—unsettled.
2026-09-19T02:26:55Z
grounded: converges/low — Splash’s claimed model-specific optimization converges with Scott’s Hardware-aware local inference approach, but currently adds only another example rather than
2026-09-19T02:23:48Z
origin walked (codex/luna, conf 0.86): anchor reddit.post.1wk9hze -> echo.blog.db6a2bf60f by Inco AI
2026-09-19T02:21:49Z
case created — Concrete installation commands and a named model artifact make this a distinct runtime-release episode, although the throughput comparisons remain unvalidated.