2026-10-11 17:11 UTC

Independent reproduction will determine whether Cua’s GPU-passthrough approach gives Apple Silicon macOS virtual machines an 11–16× llama.cpp speedup and practically near-native local LLM inference.

state: expiredheat: lowuncertainty: highconvergesscott: lowapple-silicon-inference macos-virtualization local-inferenceCua

What is this?

Cua is presented as using a GPU-bridging approach to accelerate llama.cpp inside Apple Silicon macOS virtual machines, with a claimed 11–16× speedup and near-native inference performance. The supplied snippets support the broader mechanism: forwarding guest GPU/API calls to host Metal can substantially outperform CPU-only virtualization, and one independent report measured about 63 tokens/s in a container versus 78 tokens/s natively. However, the snippets describe API remoting or Vulkan-to-Metal para-virtualization rather than direct GPU passthrough, and they do not independently establish Cua’s exact 11–16× figures; those remain a claim awaiting reproducible benchmarks.

Why it matters to Scott

If independently reproduced, Cua’s result would operationally extend Scott’s hardware-aware local-inference position: accelerator access becomes an explicit substrate choice that could make isolated macOS VMs viable for near-native local model workloads. It could affect architecture choices for disposable agent environments, but the exact speedup and whether it generalises beyond this configuration remain unverified.
dev:concept.hardware-aware-local-inferenceip:source.give-the-agent-a-workshop-ebookip:concept.high-affordance-substrateradar:concept.local-inferenceradar:concept.llama-cppradar:concept.agent-sandboxing
queries asked of Scott's wikis
  • local LLM inference inside macOS virtual machines
  • Apple Silicon Metal acceleration and virtualization
  • GPU API remoting versus device passthrough
  • near-native inference in containers and agent sandboxes
  • reproducible llama.cpp benchmarking methodology
  • local model infrastructure for isolated coding agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cppfrabonacci30235

Interpretation history

Decision trace