2026-10-11 17:10 UTC

Redditor a300a300's linked mlxfast project claims coding agents (mostly Opus 5.5) rewrote a 27B model's MLX inference engine on a Mac, raising decode from 66 to ~580 tok/s in three days β€” verification of the numbers on the project page or independent replication would establish agent-driven engine optimization as a demonstrated route to order-of-magnitude local-inference speedups.

state: resolvedheat: lowuncertainty: mediumconvergesscott: mediumagent-driven-inference-optimization local-inference coding-agents apple-siliconAnthropic

What is this?

A Reddit post by u/a300a300 links an 'mlxfast' project page claiming that coding agents β€” mostly Anthropic's Opus 5.5 β€” rewrote a 27B model's MLX inference engine on a Mac, lifting decode from 66 to ~580 tok/s (~9x) over three days. The supplied search did not surface the mlxfast page itself or any independent replication, so the claim currently rests on the Reddit echo of the page's self-reported numbers; model identity, quantization, Mac chip, and measurement setup are all unspecified in what we have. What the snippets do establish is the surrounding envelope for 27B-class dense models on Apple silicon: ~19–40 tok/s decode on mainstream engines (uzu, Ollama, oMLX on an M5 Pro), ~50–75 tok/s with model-specific engines or MTP/speculative decoding, and 144 tok/s as the best figure claimed anywhere in this corpus (Inco Splash on an M5 Max). A 580 tok/s result would sit roughly 4x above that best claim, so the load-bearing question is whether the page's methodology supports the number β€” which is exactly the verification or replication the hypothesis says would settle it.

Why it matters to Scott

A practitioner independently arrives at the self-equipping-agent thesis β€” an agent manufacturing its own missing capability at the point of need, here by re-engineering its MLX inference engine β€” so a verified result would be a dated receipt for Scott's position; but the self-reported 580 tok/s sits ~4x above the best figure in the corpus envelope (Inco Splash's 144) and can only be defended at the bottom of his evidence-class ladder until someone runs the repo, a cheap deterministic falsifier. It bears directly on his hardware-aware local-inference work and the open agent-kernel lineage (codex-autoresearch's 232x claim, the qwen3090 run that produced kernels but no perf win), and if it holds it would change local agentic-coding economics on Macs.
ip:source.the-self-equipping-agent-ebookip:concept.runtime-capability-synthesisip:concept.evidence-class-ladderdev:concept.hardware-aware-local-inferenceradar:concept.mlxradar:concept.inference-kernelsradar:concept.apple-siliconradar:inco-splash-apple-silicon-inferenceradar:codex-autoresearch-gpu-kernel-speedupradar:qwen3090-three-week-cuda-agentradar:anthropic-claude-ai-measured-optimization
queries asked of Scott's wikis
  • agents rewriting inference kernels or engines speedup
  • MLX Apple Silicon local inference benchmarks tok/s
  • local models powering coding agents harness
  • self-reported benchmark claims verification replication
  • coding agents doing real systems engineering
  • agents compounding improvements on AI infrastructure

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

09-27 22:26 (minted)⭐ origin echo-reconstructedPer the Reddit echo, the mlxfast page presents AI agents (mostly Opus 5.5) raising a 27B model from 66 tok/s to 580 tok/s on a Mac in 3 days
? on blog (echo) Β· attributed from reddit.post.1wrwn8h Β· published time unknown
β€”
09-27 21:45first on r/singularity Β· published Β· lag ?AI agents (mostly Opus 5.5) have been speeding up a 27B model from 66tok/s to 580tok/s on a Mac in 3 days by rewriting its inference engine
a300a300
β€”
09-27 21:45amplified on r/singularity πŸ‘‘reddit.post.1wrwn8h
a300a300
peak 368 Β· 31 comments Β· 100% of case engagement
09-27 22:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1wrwn8hβ€”

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAI agents (mostly Opus 5.5) have been speeding up a 27B model from 66tok/s to 580tok/s on a Mac in 3 days by rewriting its inference engine
singularity
Retrieved article excerpt

Open article Β· Retrieved 2026-09-27T22:24:48.622812+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
a300a30036530
🟧 echo.blog ⭐Per the Reddit echo, the mlxfast page presents AI agents (mostly Opus 5.5) raising a 27B model from 66 tok/s to 580 tok/s on a Mac in 3 daysβ€”β€”

Interpretation history

Decision trace