Production deployments will determine whether orchestration, retrieval, and tool-call overhead from agentic AI raises CPU demand enough to shift common CPU-to-GPU provisioning from roughly 1:4 toward 1:2 or 1:1.
state: expiredheat: lowuncertainty: highconvergesscott: mediumai-infrastructure agent-harnesses inference-economicsAMDArmMicrosoftOCP APAC
What is this?
The case concerns a hardware-capacity hypothesis: production AI agents may consume substantially more CPU per GPU than conventional inference because orchestration, retrieval, tool dispatch, response parsing, context rebuilding, and coordination occur around GPU-bound model execution. One supplied snippet supports CPU saturation when tool calls and coordination are added, but the results do not substantiate the specific shift from roughly 1:4 to 1:2 or 1:1, nor do they directly verify the attributed statements from AMD, Arm, Microsoft, Lisa Su, or OCP APAC. The proposed provisioning ratios should therefore be treated as an industry forecast awaiting production measurements, not as an established outcome.
Why it matters to Scott
The forecast extends Scott’s model-plus-harness position into infrastructure economics: orchestration and tool-use architecture, not model weights alone, may determine the CPU/GPU footprint of production agents. It creates a concrete trace-backed benchmarking and code-first architecture question, but the claimed 1:2–1:1 ratios remain unverified, so this is a measurement opportunity rather than an actionable provisioning conclusion.
ip:concept.model-plus-harness-benchmark-unitip:concept.agent-observabilityip:framework.code-first-architecturedev:concept.hardware-aware-local-inferenceradar:concept.ai-infrastructureradar:concept.inference-economicsradar:concept.agent-harnessesradar:backpressure-llm-serving-simulator
queries asked of Scott's wikis
- agent harness CPU overhead and bottlenecks
- tool-call and retrieval infrastructure economics
- CPU-to-GPU provisioning for production inference
- agent orchestration latency and resource accounting
- local inference hardware balance and utilization
- production benchmarks for multi-step agent workloads
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-08-15T13:30:01Z
No new evidence arrived within the monitoring window, and repeated discussion never advanced beyond conflicting anecdotes around AMD’s forecast. The episode has faded pending measured production workloads or concrete provisioning decisions.
2026-08-13T13:29:10Z
The refreshed comments remain conflicting systems anecdotes rather than measured production workloads or provisioning decisions. They neither validate the proposed 1:2–1:1 CPU/GPU shift nor materially weaken it, leaving the case as a vendor-led benchmarking hypothesis.
2026-08-13T02:31:33Z
Refreshed comments still provide only conflicting anecdotes, with no production measurements or independent provisioning evidence. The proposed CPU-to-GPU shift remains a vendor-led benchmarking hypothesis rather than an observed deployment trend.
2026-08-12T18:37:08Z
The refreshed discussion adds only competing anecdotes about harness CPU load, not workload measurements or deployment-level provisioning evidence. The ratio shift remains an unvalidated vendor forecast and benchmarking question.
2026-08-12T15:44:43Z
Refreshed discussion remains anecdotal amplification of the vendor forecast, without production measurements or independent provisioning evidence. The hypothesized CPU-to-GPU ratio shift is still a useful benchmark target, not an observed infrastructure trend.
2026-08-12T11:41:53Z
Refreshed comments add anecdotal agreement but no measured workload, independent deployment, or provisioning evidence. The case remains a vendor-led benchmarking hypothesis rather than evidence of an actual CPU-to-GPU ratio shift.
2026-08-12T10:30:59Z
The new activity is only modest amplification of the existing AMD forecast; no production measurements, independent implementation evidence, or workload definitions validate the proposed CPU-to-GPU ratio shift. The case remains a useful benchmarking hypothesis but no longer warrants near-term attention.
2026-08-12T10:29:11Z
grounded: converges/medium — The forecast extends Scott’s model-plus-harness position into infrastructure economics: orchestration and tool-use architecture, not model weights alone, may de
2026-08-12T10:25:31Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1vm9mh6 -> echo.other.0c0b5ee3fc by Advanced Micro Devices, Inc. (Lisa Su, Chair and CEO)
2026-08-12T10:24:23Z
case created — Two observations trace a concrete infrastructure forecast to statements by major compute vendors, but production deployment data is still needed to validate the claimed ratio shift.
Decision trace
- 08-15 23:30expireNo new evidence arrived within the monitoring window, and repeated discussion never advanced beyond conflicting anecdotes around AMD’s forecast. The episode has faded pending measured production workl
- 08-15 23:30alert_silentThe staleness trigger carries no substantive delta; Scott can wait until independent deployment measurements or procurement changes test the proposed CPU-to-GPU ratio shift.
- 08-15 23:30alert_routeThe staleness trigger carries no substantive delta; Scott can wait until independent deployment measurements or procurement changes test the proposed CPU-to-GPU ratio shift.
- 08-13 23:29repriceThe refreshed comments remain conflicting systems anecdotes rather than measured production workloads or provisioning decisions. They neither validate the proposed 1:2–1:1 CPU/GPU shift nor materially
- 08-13 23:29alert_silentOnly repetitive discussion changed; there is no new deployment measurement, independent implementation, or procurement evidence that Scott needs before the next briefing.
- 08-13 23:29alert_routeOnly repetitive discussion changed; there is no new deployment measurement, independent implementation, or procurement evidence that Scott needs before the next briefing.
- 08-13 23:21sensor_dirtycomment_update
- 08-13 19:21sensor_dirtyengagement_update
- 08-13 12:31repriceRefreshed comments still provide only conflicting anecdotes, with no production measurements or independent provisioning evidence. The proposed CPU-to-GPU shift remains a vendor-led benchmarking hypot
- 08-13 12:31alert_silentOnly discussion churn occurred; it neither validates nor materially weakens the forecast, so the case can wait for measured production workloads or independent provisioning changes.
- 08-13 12:31alert_routeOnly discussion churn occurred; it neither validates nor materially weakens the forecast, so the case can wait for measured production workloads or independent provisioning changes.
- 08-13 12:21sensor_dirtycomment_update
- 08-13 11:21sensor_dirtyengagement_update
- 08-13 07:21sensor_dirtyengagement_update
- 08-13 06:21sensor_dirtyengagement_update
- 08-13 04:37repriceThe refreshed discussion adds only competing anecdotes about harness CPU load, not workload measurements or deployment-level provisioning evidence. The ratio shift remains an unvalidated vendor foreca
- 08-13 04:37alert_silentNo consequential factual delta occurred; comment churn neither validates nor materially weakens the forecast, so this can wait for production benchmarks or independent deployment data.
- 08-13 04:37alert_routeNo consequential factual delta occurred; comment churn neither validates nor materially weakens the forecast, so this can wait for production benchmarks or independent deployment data.
- 08-13 04:21sensor_dirtycomment_update
- 08-13 02:21sensor_dirtyengagement_update
- 08-13 01:44repriceRefreshed discussion remains anecdotal amplification of the vendor forecast, without production measurements or independent provisioning evidence. The hypothesized CPU-to-GPU ratio shift is still a us
- 08-13 01:44alert_silentNo consequential factual delta occurred; refreshed comments and engagement do not validate the ratio forecast and can wait for production benchmarks or independent deployment data.
- 08-13 01:44alert_routeNo consequential factual delta occurred; refreshed comments and engagement do not validate the ratio forecast and can wait for production benchmarks or independent deployment data.
- 08-13 00:22sensor_dirtyengagement_update
- 08-13 00:22sensor_dirtycomment_update
- 08-12 22:21sensor_dirtyengagement_update
- 08-12 22:21sensor_dirtyengagement_update
- 08-12 21:41repriceRefreshed comments add anecdotal agreement but no measured workload, independent deployment, or provisioning evidence. The case remains a vendor-led benchmarking hypothesis rather than evidence of an
- 08-12 21:41alert_silentThe new delta is repetitive Reddit amplification and informal operator intuition, not a consequential factual change. It can wait until production benchmarks or independent provisioning data test the
- 08-12 21:41alert_routeThe new delta is repetitive Reddit amplification and informal operator intuition, not a consequential factual change. It can wait until production benchmarks or independent provisioning data test the
- 08-12 21:21sensor_dirtyengagement_update
- 08-12 21:21sensor_dirtycomment_update
- 08-12 20:30repriceThe new activity is only modest amplification of the existing AMD forecast; no production measurements, independent implementation evidence, or workload definitions validate the proposed CPU-to-GPU ra
- 08-12 20:30alert_silentOnly Reddit engagement changed, with no consequential factual delta beyond the already assessed vendor forecast. This can wait for concrete production benchmarks or independent operator provisioning d
- 08-12 20:30alert_routeOnly Reddit engagement changed, with no consequential factual delta beyond the already assessed vendor forecast. This can wait for concrete production benchmarks or independent operator provisioning d
- 08-12 20:29alert_silentAMD’s first-party comments establish that it is publicly forecasting materially higher CPU demand from agent workloads, including movement toward 1:1 CPU/GPU configurations. That is a useful infrastru
- 08-12 20:29surface_candidateAMD’s first-party comments establish that it is publicly forecasting materially higher CPU demand from agent workloads, including movement toward 1:1 CPU/GPU configurations. That is a useful infrastru
- 08-12 20:29alert_routeAMD’s first-party comments establish that it is publicly forecasting materially higher CPU demand from agent workloads, including movement toward 1:1 CPU/GPU configurations. That is a useful infrastru
- 08-12 20:29groundThe forecast extends Scott’s model-plus-harness position into infrastructure economics: orchestration and tool-use architecture, not model weights alone, may determine the CPU/GPU footprint of product
- 08-12 20:25promote_anchororigin walk conf 0.99