2026-10-11 16:38 UTC

Epoch AI (Jason Li) estimates the HBM shipped through 2027 could support only tens-to-hundreds of millions of concurrent frontier-model agents (up to ~1.9B on efficient open models), with even 20% utilization implying $2.6–5.3T/yr of API-equivalent spending against ~$1T projected developer revenue — and whether the figure becomes the standard reference for sizing agent demand against compute supply, or is credibly challenged as assumption-driven, resolves it.

state: watchingheat: lowuncertainty: mediumconvergesscott: highinference-economics ai-infrastructure agent-economicsJason LiEpoch AI

What is this?

Epoch AI researcher Jason Li published an analysis estimating that the high-bandwidth memory (HBM) shipped through 2027 could support only ~30–170M concurrent frontier-model agents (up to ~1.9B on efficient open models), because KV-cache memory — not raw FLOPs — bounds agent scale. Even at 20% utilization, serving implied demand would run $2.6–5.3T/yr of API-equivalent spending against ~$1T projected developer revenue, framing a structural gap between agent demand and compute supply. The estimate rests on contestable assumptions (per-agent-hour spend, GPU rental rates, KV-cache ratios) and third-party benchmarks. NOTE: the web search supplied for this grounding returned zero results, so nothing here could be externally corroborated — reception (citations, adoption, rebuttals) is attested only by the case's own tracked evidence, which one week post-publication shows minimal engagement.

Why it matters to Scott

Epoch AI independently quantifies the fleet-scale version of the constraint Scott has argued from the harness side since the VIC-20 framing: KV-cache memory — not FLOPs — bounds agent concurrency, making context engineering a supply-side lever. The analysis puts numbers on the structural gap (30–170M concurrent frontier agents; $2.6–5.3T/yr API-equivalent spend vs ~$1T developer revenue) that explains the flat-rate rationing Scott has tracked (OpenAI Pro pause, Kimi K3 halt). The load-bearing assumptions (S=$30/agent-hr, G=$5/GB300-hr, K=5–10x, u=2x) are exactly the parameters Scott's harness work operates on, giving him a dated-receipt position and a concrete target for audit or extension. If this becomes the standard reference for sizing agent demand against compute supply, Scott's prior framing gains external validation; if challenged, his harness-side evidence becomes relevant to the rebuttal.
ip:framework.context-engineeringip:concept.attention-budgetip:concept.ai-unit-economicsip:framework.the-mature-token-lawip:concept.economics-inversiondev:concept.cheap-model-front-doordev:project.llmreportwork:project.leverageairadar:concept.inference-economicsradar:concept.kv-cacheradar:concept.context-engineeringradar:concept.agent-economicsradar:concept.ai-infrastructureradar:openai-pro-200-usage-halvingradar:kimi-k3-subscription-capacityradar:2027-memory-capacity-selloutradar:epoch-price-of-thought
queries asked of Scott's wikis
  • KV-cache memory as the binding constraint on agent concurrency / harness memory footprint
  • context engineering as a supply-side lever for inference cost
  • inference economics: cost-per-agent-hour and capacity ceilings on agent fleets
  • flat-rate subscription rationing and revenue-vs-compute gap (OpenAI Pro pause, Kimi halt)
  • HBM supply chain and AI infrastructure buildout forecasts
  • agent fleet sizing: how many concurrent agents can current hardware support

Measured heat

now 0 pts/hpeak 7 pts/hcomments 0/hpeers p16momentum: steady2 platformsage 242h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-01 14:00⭐ origin echo-reconstructedEpoch AI report "How many AI agents could run on the AI chips shipped through 2027?" (subtitle: "Tens to hundreds of millions running today'
Jason Li (Epoch AI) on blog (echo) · attributed from hn.story.49944481
—
10-03 14:17first on hacker news · published · +48.3hHow many AI agents could run on the AI chips shipped through 2027?
iphonecorridor
—
10-03 14:17amplified on hacker news 👑hn.story.49944481
iphonecorridor
peak 4 · 2 comments · 67% of case engagement
10-08 17:36amplified on hacker newshn.story.50009029
gmays
peak 3 · 0 comments · 34% of case engagement
10-03 14:20our radar first saw it · +48.4hdiscovery anchor: hn.story.49944481—
pace: p41 vs 1188 stories at the 168h mark (now 242h old) — ahead of agentic-determinism-index (1.2x), behind acs-local-skill-risk-catalog (0.9x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnHow many AI agents could run on the AI chips shipped through 2027?
Retrieved article excerpt

Open article · Retrieved 2026-10-03T14:24:43.200416+00:00

Report

Oct. 2, 2026

# How many AI agents could run on the AI chips shipped through 2027?

Tens to hundreds of millions running today's most capable models, or billions running cheaper ones.

[Code and data](https://github.com/epoch-research/compute-to-agents)

Cite

[Jason Li's avatar](https://epoch.ai/about/team/jason-li)

By [Jason Li](https://epoch.ai/about/team/jason-li)

## Overview

AI companies are spending hundreds of billions of dollars a year on chips and data centers, on the premise that those chips will run AI agents to do work that people do today.
How many agents could this hardware buildout actually support?

- **AI chips shipped through 2027 could run tens to hundreds of millions of concurrent frontier-model agents.** Running nonstop, these agents would supply as many weekly working hours as about 140–720 million full-time employees.
- **More efficient models could potentially support billions of agents on the same hardware.** Applying DeepSeek V4 Pro serving benchmarks to the projected hardware supply yields approximately 1.9 billion concurrent agents supplying as many weekly working hours as 8 billion people each working 40 hours.
- **Even modest use of this capacity would require a massive increase in global demand for AI.** Using 20% of our central capacity estimate would imply $2.6–5.3 trillion a year in API-equivalent spending, against roughly $1 trillion in developer revenue by end-2027 at fivefold annual growth.
- **Hourly agent spending varies substantially across models and harnesses.** In our analysis of agent traces, Codex workloads averaged roughly $16–18 per hour of continuous agent activity, compared with $24–50 for Claude Code workloads.


## Potential agent capacity and spending

Anthropic’s Dario Amodei has described a future “[country of geniuses in a datacenter](https://darioamodei.com/essay/machines-of-loving-grace)”, but how many AI agents could future data centers actually support? That scale matters for AI’s potential impact on the economy and labor force.

We find that hardware using high-bandwidth memory (HBM) shipped during 2025–27 could eventually support tens to hundreds of millions of concurrent frontier-model agents, assuming full deployment and allocation to these workloads. HBM shipped during 2025–26 could support 16–56 million concurrent agents once deployed. Including shipments through 2027 raises that estimate to about 30–170 million.[1](https://epoch.ai/publications/estimating-the-agent-population#user-content-fn-1)

But unlike humans, an AI agent can work all 168 hours each week, 4.2 times the 40-hour workweek for a full-time employee. Therefore, these agents could work as many weekly hours as about 67–240 million people from hardware shipments through 2026, and about 140–720 million from shipments through 2027. For scale, the United States has a population of [342 million](https://www.census.gov/newsroom/press-releases/2026/population-growth-slows.html) and an estimated [100 million](https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2025/autonomous-generative-ai-agents-still-under-development.html) knowledge workers. These comparisons count working hours alone. Agents can also produce output much faster than humans, though the quality of that output varies.

Even if we use only 20% of the capacity from memory shipped through 2027, the implied spending at API prices would be $2.6–5.3 trillion a year once the hardware is deployed.[2](https://epoch.ai/publications/estimating-the-agent-population#user-content-fn-2)
For comparison, if model developers’ revenues keep growing fivefold each year, their combined annualized revenue would reach roughly $1 trillion by the end of 2027.

Demand could fall behind this potential supply, creating an overabundance of capacity.
The key uncertainty is whether sustained, rapid growth in demand for AI services will justify the investment.

Dot plot on a logarithmic axis showing estimated concurrent agents supportable by 2025–27 memory shipments for five models, ranging from 30–60 million for Claude Fable 5 to 1.9 billion for DeepSeek V4 Pro.


## How we estimate agent capacity

We estimate potential concurrent agents (\(A\)): how many agents could run at once on the memory shipped during 2025–27, assuming full deployment and allocation to the modeled workload.
We use these capacity estimates to derive working-hour and API-equivalent spending figures under the stated operating-time, pricing, allocation, and utilization assumptions.

Our estimate of potential concurrent agents is built from two terms:

\[\begin{aligned}
A &= E \times c, \\[6pt]
\text{where}\quad A &= \text{potential concurrent agents}, \\
E &= \text{effective hardware supply in GB300 equivalents}, \\
c &= \text{agents per GB300 equivalent}.
\end{aligned}\]

**Effective hardware supply (\(E\)):** measured in GB300-equivalent inference units, analogous to FLOP-based H100 equivalents.
For these workloads, memory capacity constrains concurrency and memory bandwidth constrains streaming speed.
We count high-bandwidth memory (HBM) shipped from 2025 onward, including HBM3E and newer generations, in units of 288 GB, matching a GB300 GPU.
We then adjust for the concurrency that newer hardware can support:

\[E = \frac{H\_3 + u \times H\_4}{288},\]

where \(H\_3\) and \(H\_4\) are cumulative HBM3E and HBM4/4E supply (GB).
\(u\) is the ratio of agent sessions per GB on HBM4/4E systems to agent sessions per GB on HBM3E systems.

**Serving capacity (\(c\)):** concurrent active agent sessions per GB300 equivalent.

We use

\[c = \frac{G \times K}{S}\]

for closed models.
\(S\) is API spending per active agent-hour ($/agent-hour), using durations adjusted to remove identified human waits and cap other idle gaps.
\(G\) is GPU rental cost ($/GB300-hour).
\(K\) is API-equivalent revenue divided by reference serving cost.

For open models, we use benchmarked concurrency.

**Main assumptions:** \(S = \$30/\text{agent-hour}\), \(G = \$5/\text{GB300-hour}\), \({K = 5\text{–}10\times}\), and \({u = 2\times}\). We also test \({u = 1\times}\) and \(4\times\).

Open-model benchmarks use P90 streaming-speed targets of 50 and 100 output tokens per second per user, with 200 as a sensitivity. We hold current model and workload requirements fixed.



## Estimating agent sessions per GPU today

### What counts as an agent?

We use “agent” as shorthand for an agentic workload running within a harness such as Codex or Claude Code. An agent session includes model calls and tool use, rather than continuous token generation.

We draw on two sources: SemiAnalysis’s AgentX serving benchmark for open models, and [TraceLab](https://github.com/uw-syfi/TraceLab), a public dataset of logged agent sessions, for closed models.
AgentX counts a main agent and its subagents as one session tree; TraceLab’s accounting groups do not always capture that complete tree. We use “agent” and “agent session” interchangeably when discussing capacity. A continuous agent session includes the time spent waiting for tool calls. We estimate the continuous working time by removing time spent waiting for human input and also capping unidentified idle time. One hour of this adjusted activity counts as one agent-hour.

For open models, serving benchmarks directly measure how many concurrent agent sessions the hardware supports at a given output speed. For closed models, we estimate concurrency from hourly spending and serving-cost assumptions.

### Open models: serving benchmarks measure concurrency directly

For open models, we use the benchmark data from [SemiAnalysis’s InferenceX AgentX](https://inferencex.semianalysis.com/inference). The AgentX benchmark uses a dataset of Claude Code agent session traces collected by SemiAnalysis. The benchmark replay uses synthetic text while preserving request lengths, shared context, and the timing and structure of model calls. Further details in the [AgentX methodology](https://inferencex.semianalysis.com/agentx/methodology).

Concurrency for AgentX is defined by the number of agent sessions launched. We divide the concurrency by the total number of GPUs to derive a metric of concurrent agent sessions per GPU. For prefill-decode disaggregated configurations, we count the number of combined GPUs. Figures 2–3 show the published configurations and the agent sessions per GPU.

**Speed targets: 50 and 100 tokens per second per user**

We use 50 and 100 TPS/user as round reference points for output speed in tokens per second. 200 TPS/user is also given in the appendix to test a more demanding scenario. For reference, both [OpenAI](https://artificialanalysis.ai/providers/openai) and [Anthropic](https://artificialanalysis.ai/providers/anthropic) tend to serve their frontier models around 50–70 TPS. In the benchmark data, P90 interactivity describes the output speed in TPS/user for the slower end of the distribution. This excludes the time to first token (TTFT) and does not measure end-to-end latency (see Appendix A).

The figures use the AgentX snapshot dated in Table 2 and cover seven models. We excluded any preview data not run directly through the InferenceX public repo.

Grid of seven log-log line charts plotting configured agent sessions per GPU against P90 streaming speed for seven models, with downward-sloping curves colored by Nvidia Blackwell, Hopper, and AMD GPU types.

**Concurrency per GPU at the main speed targets**

Heatmap table of agent sessions per GPU for seven models across seven GPU types, shown at streaming-speed targets of 50 and 100 tokens/s/user.

### Closed frontier models: concurrency inferred from API spending

Since closed frontier models lack the architecture details needed for transparent hardware benchmarks, we instead look at the API costs of a continuously running agent. We analyzed [TraceLab’s dataset](https://github.com/uw-syfi/TraceLab/releases/tag/v0.0.2) of Codex and Claude Code agent sessions. We adjusted for human delays (agent waiting on human response) and divided spending by agent working time to normalize to an hourly rate. Using the API prices listed in Table A3, we calculated the API-equivalent spending per agent-hour.

We then estimate how much of that spending covers serving costs, using an assumed markup ratio of API revenue to serving cost. Comparing the resulting cost per agent-hour with the rental cost of a GB300 GPU gives an estimate of how many concurrent agents each GPU could support.

Box plot comparing API-equivalent spending per agent-hour across Codex and Claude Code sessions for four models, plus high-context subsets, ranging from about $0 to $200.

**$30 per agent-hour sits between Codex and Claude Code costs**

Hourly spending varies across models and harnesses. Under Figure 4’s retained-cache[4](https://epoch.ai/publications/estimating-the-agent-population#user-content-fn-4) and five-minute gap-cap assumptions, TraceLab’s pooled rates are $18.2/hour for [GPT-5.5](https://developers.openai.com/api/docs/models/gpt-5.5), $15.5 for GPT-5.6 Sol, $24.3 for Opus 4.8, and $50.2 for Fable 5. We choose $30/hour as a round reference point which sits slightly towards the higher end. Future models may go up in price as with the jumps to Fable for Anthropic and GPT-6 Astra for OpenAI or they may go down due to competition and efficiency improvements. Figure 8 goes through a range of prices from $10 to $100 per agent-hour.

Let \(S\) be API spending per agent-hour and \(K\) the ratio of API-equivalent revenue to serving cost at our reference GPU rental price. \({K = 10\times}\) means $10 of API billing for every $1 of that cost.

The implied serving cost per agent-hour is \(S/K\). Dividing the GPU-hour rental price, \(G\), by this cost gives agent sessions per GB300 equivalent:

\[\begin{aligned}
\text{agent sessions per GB300 equivalent} &= \frac{G}{S/K} \\[4pt]
&= \frac{G \times K}{S}.
\end{aligned}\]

Table 1 uses \(S = \$30\) per agent-hour and \(G\) = [
iphonecorridor42
🟧 echo.blog ⭐Epoch AI report "How many AI agents could run on the AI chips shipped through 2027?" (subtitle: "Tens to hundreds of millions running today'Jason Li (Epoch AI)——
🟧 hnHow many AI agents could run on the AI chips shipped through 2027?gmays30

Interpretation history

Decision trace