2026-10-11 16:37 UTC

Magnitude (YC S25) claims its open-source engine tunes kernels on-device to run open models up to 2x faster than llama.cpp (92% faster Metal decode in its benchmarks), and cross-hardware replication plus adoption by local-agent builders would establish self-optimizing serving as a practical local-inference alternative.

state: watchingheat: mediumuncertainty: mediumconvergesscott: highlocal-inference inference-engines agent-harnessesMagnitude (YC S25)
Surfaced 2026-10-02T11:01:23Z โ€” Open source inference engine for agents that optimizes itself for your exact hardware; compiles and tunes kernels on-device so open models r โ€” First independent replication arrived and went against the headline: an HN user benchmarked six Qwen models on an M5 Max 128GB and reports the engine 'adds next to nothing' โ€” one terse machine with no posted numbers, so the 2x/92% claims stay announcement-class but now carry a first negative data point. The Reddit crosspost belatedly caught a real wave (6โ†’89 pts, 24 comments) with matching skepticism (the '2x' is Metal-best-case vs +19% CUDA; 2-devs/2-months immaturity, no multi-GPU) before cooling; the case settles into a quiet wait for MLX baselines and further replications.

What is this?

Magnitude (YC S25) is an open-source local inference engine (GitHub magnitudedev/magnitude, ~5k stars, TypeScript/bun monorepo) that profiles the user's machine, recommends open models that fit it, and compiles/tunes kernels on-device for that exact hardware โ€” Apple Silicon, NVIDIA, AMD, or CPU-only โ€” with one-click connection to existing coding agents (it names Pi, OpenCode, Hermes). Its Launch HN claims up to 2x faster than llama.cpp with 92% faster Metal decode in its own benchmarks; the supplied snippets do not independently corroborate those numbers, so they rest on the launch's self-reporting pending community replication. Note a wrinkle: an Aug 2026 Medium piece described Magnitude as a fully-local coding agent with its own inference engine rather than an engine alone, so the snippets are ambiguous about whether the engine is the whole product or the serving layer of a broader agent story.

Why it matters to Scott

Converges with dev:concept.hardware-aware-local-inference: a YC-backed team has independently productized Scott's 'treat compilation as explicit runtime policy' stance into a self-optimizing engine, and its 2x/92%-Metal head-to-head is aimed directly at the llama.cpp substrate under his Ollama endpoint on gamepc โ€” so the obvious next act is replicating the claim there, with the Evidence Class Ladder dictating the 2x number stays at announcement-class until independently benchmarked. The one-click Pi/OpenCode hook also continues the radar's llama.cpp-as-agent-serving thread rather than being just another speedup claim.
dev:concept.hardware-aware-local-inferencedev:technology.ollamadev:project.gamepcip:concept.evidence-class-ladderradar:concept.local-firstradar:opengemm-b200-kernel-releaseradar:kairo-measured-cuda-graph-routingradar:rooflang-inference-architecture-searchradar:pi-native-llama-cpp-runtimeradar:aa-agentperf-local-benchmarkradar:adaptive-speculative-decoding-300-gpu
queries asked of Scott's wikis
  • local inference server llama.cpp integration in agent harness
  • prefill vs decode agent latency bottleneck measurements
  • kernel auto-tuning / hardware-specific compilation positions
  • local-first agent stack privacy and sovereignty arguments
  • benchmark verification of speedup claims against incumbents
  • OpenAI-compatible endpoint support in my harness model providers

Measured heat

now 0 pts/hpeak 64 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 290h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-29 14:00โญ origin echo-reconstructedOpen source inference engine for agents that optimizes itself for your exact hardware; compiles and tunes kernels on-device so open models r
magnitudedev (Anders and Tom) on github (echo) ยท attributed from hn.story.49911995
โ€”
09-30 17:37first on hacker news ยท published ยท +27.6hLaunch HN: Magnitude (YC S25) โ€“ Self-optimizing inference engine for agents
anerli
โ€”
09-30 22:46first on r/LocalLLaMA ยท published ยท +32.8hOpen source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU.
paranoidray
โ€”
09-30 17:37amplified on hacker news ๐Ÿ‘‘hn.story.49911995
anerli
peak 194 ยท 99 comments ยท 75% of case engagement
09-30 22:46amplified on r/LocalLLaMAreddit.post.1wuj70v
paranoidray
peak 96 ยท 24 comments ยท 17% of case engagement
10-03 01:19amplified on r/LocalLLaMAreddit.post.1wwamjw
cezarducatti
peak 6 ยท 55 comments ยท 9% of case engagement
09-30 18:21our radar first saw it ยท +28.4hdiscovery anchor: hn.story.49911995โ€”
10-02 11:01reached heat=high ยท +69.0h ยท via ledgerโ€”โ€”
pace: p83 vs 1188 stories at the 168h mark (now 290h old) โ€” ahead of gpt6-prompt-cache-controls (1.0x), behind ftc-frontier-ai-probe (1.0x)

Evidence (4) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnLaunch HN: Magnitude (YC S25) โ€“ Self-optimizing inference engine for agents
Retrieved article excerpt

Open article ยท Retrieved 2026-09-30T18:42:56.779852+00:00

Magnitude icon

# Magnitude

**Run open models as fast as your hardware allows**

[Download Magnitude](https://magnitude.dev/download)
[Documentation](https://docs.magnitude.dev)
[Discord](https://discord.gg/EHt48pPWdC)
[Follow Magnitude on Twitter](https://x.com/usemagnitude)
[GitHub Repo stars](https://github.com/magnitudedev/magnitude/stargazers)

Magnitude is an open source inference engine for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the agent you already use (Pi, OpenCode, Hermes, Codex, and more). Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU.

**[Download Magnitude for macOS, Windows, or Linux](https://magnitude.dev/download)**

โญ Help us reach more developers and grow the Magnitude community. Star this repo!

demo-9-29.mp4

[
](https://private-user-images.githubusercontent.com/32600629/661966322-983328c8-93e8-4360-bfef-11e93ff76035.mp4?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3OTA3OTQwNzYsIm5iZiI6MTc5MDc5Mzc3NiwicGF0aCI6Ii8zMjYwMDYyOS82NjE5NjYzMjItOTgzMzI4YzgtOTNlOC00MzYwLWJmZWYtMTFlOTNmZjc2MDM1Lm1wND9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNjA5MzAlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjYwOTMwVDE4NDI1NlomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPWY1MTAyM2YzZTFkYjZmMWIyYzEzOWI3YWM0MjdlYWI5ZmY1MzM0OGJkZjI0NTM1MzMwZWY4MThlYzg3NWRkOTAmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0JnJlc3BvbnNlLWNvbnRlbnQtdHlwZT12aWRlbyUyRm1wNCJ9.4QKShJeSr_gDRSrApB3X2iBMVIrmuxZQ5dsvjpa-34I)

## Get started

1. [Download Magnitude](https://magnitude.dev/download), install it, and open the app.
2. Choose a recommended model in **Discover** and download it.
3. Connect your agent in **Connections** and start using it.

The desktop app includes the `magnitude` CLI. No separate installation is needed.

## Why Magnitude?

- **Up to 2x faster than llama.cpp:** 92% faster decode on Metal, 19% on CUDA
- **Tuned on your device:** kernels are tuned on your hardware before a model runs
- **Built for the best models:** hand-optimized kernels for popular open-weight families
- **Memory that flexes:** 27% less memory per agent, freed when agents stop
- **Fast concurrent sessions:** sessions share prefix caches to prevent slowdown
- **Works with your agent:** one click to connect Pi, OpenCode, Hermes, Codex, and more
- **Free, private, open source:** no token costs, nothing leaves your machine, Apache 2.0

## Up to 2x faster than llama.cpp

Magnitude vs llama.cpp: 9% faster prefill and 92% faster decode on Metal, 23% faster prefill and 19% faster decode on CUDA

## FAQ

### What is Magnitude?

An open source inference engine that optimizes itself for your hardware. It ships as a desktop app that runs open models and connects them to the agent you already use.

### How is it faster than llama.cpp, Ollama, or LM Studio?

They ship kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your actual device before a model runs, so they fit your exact chip. [See the benchmarks against llama.cpp.](https://github.com/magnitudedev/magnitude#up-to-2x-faster-than-llamacpp)

### What hardware do I need?

Any Apple Silicon, NVIDIA, or AMD GPU, or nothing but a CPU. There is no fixed minimum. Smaller machines run smaller models, and more memory lets you run larger ones.

### What operating systems does it support?

macOS, Linux, and Windows.

### Which models does it support?

See the full list at [magnitude.dev/models](https://magnitude.dev/models). We write optimized kernels for the most popular open-weight families, which is how we beat generalist engines.

### Which agents work with it?

One click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works through the OpenAI-compatible API.

### Is it private?

Yes. Prompts, files, and models stay on your machine. No internet needed once a model is downloaded.
anerli19499
๐ŸŸง echo.github โญOpen source inference engine for agents that optimizes itself for your exact hardware; compiles and tunes kernels on-device so open models rmagnitudedev (Anders and Tom)โ€”โ€”
๐ŸŸ  redditOpen source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU.
LocalLLaMA
paranoidray9623
๐ŸŸ  redditStrata - RTX 3090 - 128 Ram - Qwen 3.8 Flash Next
LocalLLaMA
cezarducatti355

Interpretation history

Decision trace