2026-10-11 17:10 UTC

Independent benchmarks will determine whether Shoehorn can automatically quantize large language models to fit constrained Apple Silicon memory while preserving useful quality and inference speed.

state: expiredheat: lowuncertainty: highconvergesscott: lowlocal-inference quantization apple-silicon

What is this?

Shoehorn is presented as a library intended to take a BF16 language model plus an importance matrix and automatically quantize the model so it fits within a Mac’s Apple Silicon unified-memory constraints. The supplied results support the broader premise that 4-bit quantization can reduce model memory by roughly 4× while retaining much of its quality, but they also stress that fitting a model does not guarantee fast inference and that performance varies by model and hardware. None of the supplied snippets independently benchmarks Shoehorn itself, so its quality retention, automation reliability, and inference speed remain unverified here.

Why it matters to Scott

Shoehorn operationalizes Scott’s hardware-aware local-inference pattern by treating precision and Apple Silicon memory limits as runtime policy, while its need for independent quality and speed benchmarks matches his evaluation-driven approach. However, it is currently an unverified implementation of patterns Scott and the radar already cover extensively, so it adds little unless benchmarks demonstrate a material advantage in automatic model fitting.
dev:concept.hardware-aware-local-inferenceip:concept.evaluation-driven-developmentradar:concept.quantizationradar:concept.apple-silicon-inferenceradar:compressed-llm-fidelity-safety-gap
queries asked of Scott's wikis
  • automatic quantization and quality-aware model fitting
  • local inference memory budgets on Apple Silicon
  • quantization benchmarks and acceptable quality loss
  • unified memory capacity versus inference throughput
  • importance matrices for post-training quantization
  • local-model deployment and hardware-aware optimization

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Shoehorn, a library to quantize an LLM to fit your Mac's VRAMrhgraysonii50
🟧 echo.github ⭐The repository’s earliest recorded project goal contains the verbatim request: “LLM runtime that takes in a BF16 model file and an imatrix aBobby Grayson——

Interpretation history

Decision trace