2026-10-11 18:01 UTC

Production deployments will determine whether orchestration, retrieval, and tool-call overhead from agentic AI raises CPU demand enough to shift common CPU-to-GPU provisioning from roughly 1:4 toward 1:2 or 1:1.

state: expiredheat: lowuncertainty: highconvergesscott: mediumai-infrastructure agent-harnesses inference-economicsAMDArmMicrosoftOCP APAC

What is this?

The case concerns a hardware-capacity hypothesis: production AI agents may consume substantially more CPU per GPU than conventional inference because orchestration, retrieval, tool dispatch, response parsing, context rebuilding, and coordination occur around GPU-bound model execution. One supplied snippet supports CPU saturation when tool calls and coordination are added, but the results do not substantiate the specific shift from roughly 1:4 to 1:2 or 1:1, nor do they directly verify the attributed statements from AMD, Arm, Microsoft, Lisa Su, or OCP APAC. The proposed provisioning ratios should therefore be treated as an industry forecast awaiting production measurements, not as an established outcome.

Why it matters to Scott

The forecast extends Scott’s model-plus-harness position into infrastructure economics: orchestration and tool-use architecture, not model weights alone, may determine the CPU/GPU footprint of production agents. It creates a concrete trace-backed benchmarking and code-first architecture question, but the claimed 1:2–1:1 ratios remain unverified, so this is a measurement opportunity rather than an actionable provisioning conclusion.
ip:concept.model-plus-harness-benchmark-unitip:concept.agent-observabilityip:framework.code-first-architecturedev:concept.hardware-aware-local-inferenceradar:concept.ai-infrastructureradar:concept.inference-economicsradar:concept.agent-harnessesradar:backpressure-llm-serving-simulator
queries asked of Scott's wikis
  • agent harness CPU overhead and bottlenecks
  • tool-call and retrieval infrastructure economics
  • CPU-to-GPU provisioning for production inference
  • agent orchestration latency and resource accounting
  • local inference hardware balance and utilization
  • production benchmarks for multi-step agent workloads

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAccording to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1
LocalLLaMA
ocean_protocol6023
🟠 redditAgentic AI could push CPU-to-GPU ratios from 1:4 toward 1:1 according to AMD
singularity
ocean_protocol5516
🟧 echo.other ⭐AMD’s Q1 FY2026 earnings-call Q&A said agentic AI is “largely additive” because agents “spawn more CPU tasks.” Lisa Su described the CPU/GPUAdvanced Micro Devices, Inc. (Lisa Su, Chair and CEO)——

Interpretation history

Decision trace