Independent replication and adoption will determine whether the proposed intelligence-per-watt metric produces reproducible, decision-useful comparisons of local AI models and inference hardware.
state: expiredheat: lowuncertainty: highconvergesscott: mediuminference-economics local-inference model-evaluation
What is this?
Researchers associated with Stanford’s Hazy Research propose “intelligence per watt” (IPW), defined as mean task accuracy divided by mean inference power draw, for comparing local language-model and hardware combinations. They report an end-to-end, cross-platform profiling harness supporting NVIDIA, AMD, and Apple Silicon, tested across more than 20 local models, eight accelerators, and roughly one million queries. The supplied results report substantial efficiency gains, but they primarily describe the authors’ own study and related coverage; they do not establish independent replication or broad external adoption.
Why it matters to Scott
IPW independently operationalizes Scott’s model-plus-harness and AI unit-economics positions by evaluating useful accuracy against measured power for specific model–hardware combinations. If independently replicated, it could become a practical selection signal for his hardware-aware local inference and gamepc workloads; for now, the supplied evidence establishes only the authors’ benchmark, not external validation or adoption.
ip:concept.ai-unit-economicsip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.local-inferenceradar:concept.model-evaluationradar:concept.inference-economicsradar:concept.inference-efficiencyradar:concept.ai-hardware
queries asked of Scott's wikis
- local inference economics and hardware efficiency
- decision-useful evaluation metrics for local models
- cross-platform reproducible inference benchmarking
- accuracy versus energy cost tradeoffs
- local versus cloud workload routing economics
- evaluation harnesses for model-hardware combinations
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
| source | object | author | score | comments |
| 🟠 reddit | [2511.07885] Intelligence per Watt: Measuring Intelligence Efficiency of Local AI LocalLLaMA | pscoutou | 30 | 11 |
| 🟧 echo.paper ⭐ | The authors introduce intelligence per watt (IPW), defined as task accuracy per unit of power, and report evaluating 20+ local models, 8 acc | Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, J. Wes Griffin, Herumb Shandilya, Adrian Gamarra Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Kumar Guha, Shang Zhu, Ben Athiwaratkun, John Hennessy, Azalia Mirhoseini, Christopher Ré | — | — |
Interpretation history
2026-08-26T02:29:40Z
Repeated checks still show no external replication, reusable implementation, methodological revision, or adoption; the small Reddit drift is repetitive. This episode has faded, although IPW’s underlying utility remains unresolved rather than disproved.
2026-08-24T01:27:38Z
No new evidence has arrived since the methodological objections were recorded; IPW remains an unreplicated proposal with no implementation or adoption signal.
2026-08-22T00:23:50Z
The latest observation is only negligible Reddit engagement and adds no replication, implementation, adoption, or answer to the metric’s construct-validity concerns. The case remains a potentially useful but unvalidated benchmark proposal.
2026-08-19T23:42:48Z
The refreshed discussion adds substantive construct-validity concerns—power versus energy and throughput, plus idle and model-residency costs—rather than independent validation. IPW remains an unreplicated proposal and may require a broader workload-aware metric before it becomes decision-useful.
2026-08-19T09:34:11Z
Minor Reddit engagement adds only repetitive amplification; there is still no independent replication, implementation, or adoption showing that IPW is reproducible or decision-useful.
2026-08-19T09:27:35Z
grounded: converges/medium — IPW independently operationalizes Scott’s model-plus-harness and AI unit-economics positions by evaluating useful accuracy against measured power for specific m
2026-08-19T09:24:56Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vsh04u -> echo.paper.d86afb3210 by Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, J. Wes Griffin, Herumb Shandilya, Adrian Gamarra Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Kumar Guha, Shang Zhu, Ben Athiwaratkun, John Hennessy, Azalia Mirhoseini, Christopher Ré
2026-08-19T09:24:15Z
case created — The linked paper introduces a concrete evaluation metric for local inference efficiency that can be tested and adopted independently.
Decision trace
- 08-26 12:29expireRepeated checks still show no external replication, reusable implementation, methodological revision, or adoption; the small Reddit drift is repetitive. This episode has faded, although IPW’s underlyi
- 08-26 12:29alert_silentNo consequential event occurred; minor engagement does not warrant attention, and the case can be rediscovered if independent validation or adoption eventually appears.
- 08-26 12:29alert_routeNo consequential event occurred; minor engagement does not warrant attention, and the case can be rediscovered if independent validation or adoption eventually appears.
- 08-24 11:27repriceNo new evidence has arrived since the methodological objections were recorded; IPW remains an unreplicated proposal with no implementation or adoption signal.
- 08-24 11:27alert_silentThe staleness trigger contains no consequential delta; wait for an independent replication, reusable implementation, methodological revision, or adoption.
- 08-24 11:27alert_routeThe staleness trigger contains no consequential delta; wait for an independent replication, reusable implementation, methodological revision, or adoption.
- 08-22 10:23repriceThe latest observation is only negligible Reddit engagement and adds no replication, implementation, adoption, or answer to the metric’s construct-validity concerns. The case remains a potentially use
- 08-22 10:23alert_silentNo consequential new event occurred; independent reproduction, reusable tooling, methodological revision, or adoption would justify renewed attention.
- 08-22 10:23alert_routeNo consequential new event occurred; independent reproduction, reusable tooling, methodological revision, or adoption would justify renewed attention.
- 08-20 09:42repriceThe refreshed discussion adds substantive construct-validity concerns—power versus energy and throughput, plus idle and model-residency costs—rather than independent validation. IPW remains an unrepli
- 08-20 09:42alert_silentThe new comments sharpen known methodological questions but provide no replication, implementation, benchmark revision, or adoption that Scott needs before the next briefing.
- 08-20 09:42alert_routeThe new comments sharpen known methodological questions but provide no replication, implementation, benchmark revision, or adoption that Scott needs before the next briefing.
- 08-20 06:21sensor_dirtyengagement_update
- 08-20 03:21sensor_dirtycomment_update
- 08-20 00:21sensor_dirtycomment_update
- 08-19 21:21sensor_dirtyengagement_update
- 08-19 20:21sensor_dirtyengagement_update
- 08-19 19:34repriceMinor Reddit engagement adds only repetitive amplification; there is still no independent replication, implementation, or adoption showing that IPW is reproducible or decision-useful.
- 08-19 19:34alert_silentThe new delta is engagement-only and does not change the evidence base; independent benchmark reproduction, released reusable tooling, or adoption in hardware-selection workflows can wait for a normal
- 08-19 19:34alert_routeThe new delta is engagement-only and does not change the evidence base; independent benchmark reproduction, released reusable tooling, or adoption in hardware-selection workflows can wait for a normal
- 08-19 19:31alert_silentThe paper establishes a potentially useful accuracy-per-power benchmark across 20+ models and eight accelerators, but the only new delta is low-engagement resurfacing of the authors’ existing results.
- 08-19 19:31surface_candidateThe paper establishes a potentially useful accuracy-per-power benchmark across 20+ models and eight accelerators, but the only new delta is low-engagement resurfacing of the authors’ existing results.
- 08-19 19:31alert_routeThe paper establishes a potentially useful accuracy-per-power benchmark across 20+ models and eight accelerators, but the only new delta is low-engagement resurfacing of the authors’ existing results.
- 08-19 19:27groundIPW independently operationalizes Scott’s model-plus-harness and AI unit-economics positions by evaluating useful accuracy against measured power for specific model–hardware combinations. If independe
- 08-19 19:24promote_anchororigin walk conf 0.98
- 08-19 19:24createThe linked paper introduces a concrete evaluation metric for local inference efficiency that can be tested and adopted independently.