2026-10-11 16:38 UTC

Epoch AI claims its published five-benchmark analysis finds fixed-performance inference costs fell about 47% per quarter over three years, implying substantially faster cost reductions than token-price comparisons alone capture.

state: corroboratedheat: highuncertainty: mediumconvergesscott: highinference-economics model-evaluationEpoch AILuke EmbersonDavid Roodman
Surfaced 2026-09-25T14:32:10Z — The report estimates approximately 47% quarterly cost declines at fixed benchmark performance, with public code and data and explicit limita — The episode's news phase is over: launch-window velocity decayed to zero on every platform and the last ~3 days produced no replication, no substantive methodological critique, and no new evidence objects — the 47%/qtr finding now reads as a settled, reproducible reference point rather than a moving story, corroborated at order-of-magnitude by independent prior work (Gundlach et al., Ihle, a16z) already converging with it. The magnitude-valve flag fired off the stale Sep 23-24 launch spike; current measured rates (0 pts/h, 13th percentile, hottest object far below cohort pace) say spread has stopped, not continuing, so heat cools to low even as the state advances to corroborated on the independent-lines convergence, not on engagement.

What is this?

Epoch AI has published 'The plunging price of thought', a research report (with public code and data) estimating that the cost of achieving a fixed level of AI benchmark performance fell roughly 47% per quarter — about 13x per year — across five benchmarks (math, hard sciences, games of skill) since 2023, a decline rate it claims exceeds any other transformative technology on record. The report's method compares cost at fixed benchmark performance rather than raw token prices, which the authors argue captures more of the true deflation. The snippets corroborate the headline finding from Epoch's own pages and a LinkedIn summary; an earlier Epoch analysis (March 2025) found per-benchmark price declines of 9x–900x per year, and an arXiv paper ('The Price of Progress') independently found large benchmark-price declines with algorithmic efficiency contributing ~3x/year — the snippets don't show independent replication of the specific 47% figure, and Epoch's own Substack note by JS Denain flags caveats (distilled models may be more brittle, trend may slow).

Why it matters to Scott

Epoch is a credible outside party quantifying the premise Scott's frameworks assume rather than measure — the 47%/quarter capability-adjusted deflation rate is a dated receipt for Model Perishability's 'models reprice faster than contracts' claim and the Re-Roll's collapsing recombination window. It also bears directly on his own pricing work: the token-price-vs-capability-adjusted measurement methodology is exactly the distinction underpinning the LLM Report pricing guide and AI Unit Economics, and no radar_hits page shows the radar already tracking this specific Epoch analysis (the 250-episode inference-economics concept is the topic, not this story).
ip:concept.model-perishabilityip:concept.cost-of-cognitionip:concept.ai-unit-economicsip:source.the-re-roll-ebookdev:project.llmreportradar:concept.inference-economicsradar:concept.inference-costsradar:concept.model-evaluationradar:vercel-september-open-weight-majority
queries asked of Scott's wikis
  • inference cost deflation economics local models
  • token price vs capability-adjusted cost measurement methodology
  • benchmark evaluation cost per dollar RAG pipeline budgeting
  • algorithmic efficiency gains distillation cheap frontier models
  • local open-weight inference vs API price trends
  • cost curve assumptions in agent harness and memory tooling design

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 482h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-21 14:00⭐ origin echo-reconstructedThe report estimates approximately 47% quarterly cost declines at fixed benchmark performance, with public code and data and explicit limita
Luke Emberson and David Roodman on blog (echo) · attributed from hn.story.49808743, reddit.post.1wnolng
—
09-22 21:58first on hacker news · published · +32.0hAI is getting cheaper more quickly than any other transformative tech in history
vitalnodo
—
09-22 22:49first on r/singularity · published · +32.8hThe plunging price of thought
Proper_Actuary2907
—
09-22 21:58amplified on hacker newshn.story.49808743
vitalnodo
peak 6 · 3 comments · 7% of case engagement
09-22 22:49amplified on r/singularity 👑reddit.post.1wnolng
Proper_Actuary2907
peak 170 · 40 comments · 93% of case engagement
09-22 22:20our radar first saw it · +32.4hdiscovery anchor: hn.story.49808743—
09-25 14:30reached heat=high · +96.5h · via queue+ledger——
pace: p76 vs 1032 stories at the 336h mark (now 482h old) — ahead of zed-agentic-xanadu (1.0x), behind coop-coding-agent-vm-isolation (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAI is getting cheaper more quickly than any other transformative tech in history
Retrieved article excerpt

Open article · Retrieved 2026-09-22T22:25:25.338074+00:00

Report

Sep. 22, 2026

# The plunging price of thought

Over the past three years, the cost of a given level of AI performance has fallen about 47% per quarter, faster than any other transformative technology in history.

[Code and data](https://github.com/droodman/inference-cost/)

Cite

[Interactive plots and tables](https://droodman.github.io/inference-cost/)

[Luke Emberson's avatar](https://epoch.ai/about/team/luke-emberson)David Roodman's avatar

By [Luke Emberson](https://epoch.ai/about/team/luke-emberson) and David Roodman

## Key Takeaways

- Over the past three years, the cost of a given level of AI performance has fallen an average of some 47% per quarter.
  That is a 13-fold drop every year – a faster rate than any other transformative technology in history.
- This rate is measured across the five benchmarks of AI capability for which we have the best data from the past three years — covering mathematics, hard sciences, and games of skill — as well as less sophisticated analysis with earlier data.
- We see somewhat slower cost drops on game-based puzzles, at 39–43% per quarter, and faster progress on math problems, at 50–52% per quarter.
- Costs have probably been falling this fast since the dawn of commercial LLM inference in November 2021, when OpenAI fully released GPT-3.
- The cost of a given level of performance often falls fastest right after that level is first achieved, that is, when it is state of the art (SOTA).
  Three of our five benchmarks exhibit this pattern.
  Averaging across all five, cost falls 66% per quarter (75x per year) for performance that has just debuted as SOTA.
  Two years later, prices fall half as fast, at 32% per quarter (4.7x per year).
- Despite falling prices, AI spending could remain high.
  If an important job for AI, like reviewing thousands of scientific papers for errors, demands as much cognition as running one of these benchmark tests a million times, then the spending would still add up even at a penny per run.
  Moreover, while prices fall, AI’s capability could keep rising.
  The price of passing a first-grade math test may now be trivial.
  The price of proving hard theorems is not.

Data and code are [on GitHub](https://github.com/droodman/inference-cost/).
An [overlay page](https://droodman.github.io/inference-cost/) has many plots and tables to explore.



## Overview

The “GPT” in “ChatGPT” stands for “Generative Pre-trained Transformer,” a technical description of how the AI inside it works.
Surely, though, the creators of OpenAI’s GPT models were nodding to an older meaning of the initialism: [general-purpose technology](https://doi.org/10.1016/0304-4076(94)01598-T).
They correctly foresaw that — like the steam engine, electricity, and the internet — large language models would someday touch every aspect of society.

It is now widely understood that the AI boom is a macroeconomic force powerful enough to raise prices for the inputs it demands: chips, power, even the labor of electricians.
Less well recognized is a paradoxical flip side: the price of the *output* from all those data centers is falling extraordinarily rapidly:

The next chart shows some examples.
On January 31, 2025, OpenAI released a new iteration in its series of “reasoning” models, called o3.
We estimate that for an average cost of 30 cents per question, it could achieve a 75% score on [GPQA Diamond](https://epoch.ai/benchmarks/gpqa-diamond), a multiple-choice exam covering PhD-level physics, chemistry, and biology.[1](https://epoch.ai/publications/the-plunging-price-of-thought#user-content-fn-1)
Just under 18 months later, OpenAI released GPT-5.6 Luna.
It scored just as well — for four hundredths of a penny per question ($0.0004).
That is a 725-fold drop in the price of thought in under 18 months.
It is like the sticker price on a new car falling from $50,000 to $69.
No other general purpose technology in history appears to have gotten so cheap so fast.

To measure these trends, we analyzed performance with a new and more comprehensive dataset that includes five AI performance benchmarks covering mathematics, hard sciences, and games of skill over the last three years.

Across that time, we find that the price for a given level of performance has fallen about 47% per quarter, or 13x per year.
We see slower drops on game-based puzzles, at about 39–43% per quarter (7–10x per year), and faster progress on math problems, at 50–52% per quarter (16–19x per year).

The cost decline for a given level of performance does tend to slow over time, though this pattern is not universal.
One possible explanation: when a performance level is first achieved, AI companies can briefly charge a premium for it, before competition and technological improvement quickly drive down the price.
In time, that dynamic slows.
Averaging across the five primary benchmarks, cost falls 66% per quarter (75x per year) at first.
Two years later, it falls half as fast, at a “mere” 32% per quarter (4.7x per year).

Our analysis comes with major caveats.
AI companies may be expressly training their models for some benchmarks (“benchmaxxing”), so that improvement on the benchmarks outstrips improvement for real-world tasks.
Even if they are not, doing well on a benchmark is not synonymous with useful work.
Because we focus on the frontier — the absolute cheapest model capable of any given level of performance — we implicitly posit an AI user who relentlessly searches for the most cost-effective model for each task, when real users do not switch models so often, and therefore do not reap quite the same savings.
Our data are incomplete and noisy: the timeframe is barely three years, and we do not include all combinations of AI model and benchmark.
Prices drop differently for different models, benchmarks, time periods, and performance ranges, and there are many reasonable ways to average over this variegated experience.
Overall, while we believe that our bottom-line numbers are reasonably representative of reality, they should not be read as exact.

The rest of this report details our analysis.
Parts of it are technical.



## Previous work

We are not the first to quantify how fast the price of AI is falling.
A [2024 post](https://a16z.com/llmflation-llm-inference-cost/) by Guido Appenzeller for Andreessen Horowitz documented how, in the three years following the general release of GPT-3, LLM costs fell by a factor of 1000, i.e., 10x per year.
A few months later, in March 2025, an [analysis by Epoch AI](https://epoch.ai/data-insights/llm-inference-price-trends) found 9–900x per year drops across six performance benchmarks.

Both of those early analyses measured *prices per token for models capable of achieving a given performance* as distinct from *actual cost to achieve given performance*.
With the advent of reasoning models — OpenAI released o1 in December 2024 — it has become more problematic to ignore this distinction.
Reasoning models can productively consume far more tokens, but as a result extract good performance from an underlying LLM that is smaller and cheaper to run per token.

In September 2025, Håvard Tveit Ihle [shared](https://www.lesswrong.com/posts/ifSBamvobbyB9KWjK/inference-costs-for-hard-coding-tasks-halve-roughly-every) an analysis on LessWrong that directly compared performance and cost on two suites of coding challenges.
The analysis finds that costs halved every 1.4–2 months (64–380x per year).

The most thorough analysis yet is the [March 2026 paper by Gundlach et al.](https://arxiv.org/abs/2511.23455v2) It, too, compares actual costs to performance.
As in the present analysis, it estimates trends both in the full body of data and in the subset of models defining the cost frontier at any given time.
It also disaggregates by level of performance — prices fall faster at the high end — and by model type (open or closed, dense or mixture of experts).
Overall they find declines of 5–10x per year.



## Data

The major novelty in the present analysis is to analyze trends in costs using a methodology that captures each model’s full expense-performance continuum.

If one LLM can score 80% on a benchmark, and costs $1 to do so, and another peaks out at 60%, for a price of $0.50, it is not obvious which is more cost-effective.
Perhaps the stronger model would, if run with fewer reasoning tokens, achieve 60% more cheaply than the weaker model.

More generally, our interest is in mapping the “Pareto” frontier of cost-effectiveness — the cheapest way to attain each performance level from the models available at any given time.
We would leave a lot of territory unmapped, and potentially a lot of frontier as well, if we only estimated cost and performance when LLMs are given an unlimited budget.
Any given modern LLM can produce a range of performance levels, depending on whether its reasoning is set to low, medium, high, or max, and depending on whether a budget limit is imposed.

One way to trace a model’s expense-performance curve is to run it many times against a benchmark, at various reasoning levels and token budgets.
That process would, however, be expensive and slow.
Instead, we follow a procedure developed by [the federal Center for AI Standards and Innovation (CAISI)](https://www.nist.gov/system/files/documents/2025/09/30/CAISI_Evaluation_of_DeepSeek_AI_Models.pdf#page=66).
It uses the *transcript* from a benchmark run with a high (or no) budget constraint to predict performance under tighter budgets.
The core idea is that since benchmarks consist of many questions, one can estimate how many an LLM would answer before consuming a given budget.
The transcript provides the needed information: how many tokens the model consumed in answering each question, and which it got right.[2](https://epoch.ai/publications/the-plunging-price-of-thought#user-content-fn-2)

For any arbitrary per-question budget of X tokens, we sum up the number of questions that were answered correctly in fewer than X output tokens.
Think of this as the maximum expense the model may incur before being forced to give up.
When a model does not answer in time for a particular budget threshold, that question is scored as the probability of guessing correctly, which could be, for example, 0.25 on four-way multiple choice and effectively 0 on other kinds of problems.
Repeating across a range of budgets produces a curve for performance as a function of hypothetical expenditure.

One concern about this procedure is that it could misestimate how LLMs would actually perform if *informed* of a budget.
If models were told that they should optimize for getting a good score at some lower threshold, they might do better than if the questions were silently truncated, as simulated in the CAISI methodology.

To test this concern, we modified our evaluations by explicitly telling models their per-question budgets.
Our results suggested that while high-effort versions of thinking models often improve their low-budget performance when those budgets are announced, they do not outperform lower-effort variants under standard, unannounced conditions.
That is, simply running thinking models at low thinking effort produces expense-performance curves similar to those from running at high thinking effort under announced limits.
(See Figure 3 for examples across several models.)
As a result, as long as we use the CAISI methodology across the range of thinking levels, we will probably cover most of a model’s cost-performance possibility domain.

Four-panel chart of GPQA Diamond accuracy against per-question output-token budget for GPT-5.2, o4-mini, Claude Sonnet 4.6 and Gemini 3.5 Flash. Curves from silent truncation at each reasoning effort closely track dots from runs where the model was told its budget.

Since historically our benchmarking efforts have tended to focus on capturing the upper limits of performance with minimal regard for cost efficiency, our data was systematically missing evaluations using lower reasoning 
vitalnodo63
🟠 redditThe plunging price of thought
singularity
Retrieved article excerpt

Open article · Retrieved 2026-09-22T23:21:42.319640+00:00

Report

Sep. 22, 2026

# The plunging price of thought

Over the past three years, the cost of a given level of AI performance has fallen about 47% per quarter, faster than any other transformative technology in history.

[Code and data](https://github.com/droodman/inference-cost/)

Cite

[Interactive plots and tables](https://droodman.github.io/inference-cost/)

[Luke Emberson's avatar](https://epoch.ai/about/team/luke-emberson)David Roodman's avatar

By [Luke Emberson](https://epoch.ai/about/team/luke-emberson) and David Roodman

## Key Takeaways

- Over the past three years, the cost of a given level of AI performance has fallen an average of some 47% per quarter.
  That is a 13-fold drop every year – a faster rate than any other transformative technology in history.
- This rate is measured across the five benchmarks of AI capability for which we have the best data from the past three years — covering mathematics, hard sciences, and games of skill — as well as less sophisticated analysis with earlier data.
- We see somewhat slower cost drops on game-based puzzles, at 39–43% per quarter, and faster progress on math problems, at 50–52% per quarter.
- Costs have probably been falling this fast since the dawn of commercial LLM inference in November 2021, when OpenAI fully released GPT-3.
- The cost of a given level of performance often falls fastest right after that level is first achieved, that is, when it is state of the art (SOTA).
  Three of our five benchmarks exhibit this pattern.
  Averaging across all five, cost falls 66% per quarter (75x per year) for performance that has just debuted as SOTA.
  Two years later, prices fall half as fast, at 32% per quarter (4.7x per year).
- Despite falling prices, AI spending could remain high.
  If an important job for AI, like reviewing thousands of scientific papers for errors, demands as much cognition as running one of these benchmark tests a million times, then the spending would still add up even at a penny per run.
  Moreover, while prices fall, AI’s capability could keep rising.
  The price of passing a first-grade math test may now be trivial.
  The price of proving hard theorems is not.

Data and code are [on GitHub](https://github.com/droodman/inference-cost/).
An [overlay page](https://droodman.github.io/inference-cost/) has many plots and tables to explore.



## Overview

The “GPT” in “ChatGPT” stands for “Generative Pre-trained Transformer,” a technical description of how the AI inside it works.
Surely, though, the creators of OpenAI’s GPT models were nodding to an older meaning of the initialism: [general-purpose technology](https://doi.org/10.1016/0304-4076(94)01598-T).
They correctly foresaw that — like the steam engine, electricity, and the internet — large language models would someday touch every aspect of society.

It is now widely understood that the AI boom is a macroeconomic force powerful enough to raise prices for the inputs it demands: chips, power, even the labor of electricians.
Less well recognized is a paradoxical flip side: the price of the *output* from all those data centers is falling extraordinarily rapidly:

The next chart shows some examples.
On January 31, 2025, OpenAI released a new iteration in its series of “reasoning” models, called o3.
We estimate that for an average cost of 30 cents per question, it could achieve a 75% score on [GPQA Diamond](https://epoch.ai/benchmarks/gpqa-diamond), a multiple-choice exam covering PhD-level physics, chemistry, and biology.[1](https://epoch.ai/publications/the-plunging-price-of-thought#user-content-fn-1)
Just under 18 months later, OpenAI released GPT-5.6 Luna.
It scored just as well — for four hundredths of a penny per question ($0.0004).
That is a 725-fold drop in the price of thought in under 18 months.
It is like the sticker price on a new car falling from $50,000 to $69.
No other general purpose technology in history appears to have gotten so cheap so fast.

To measure these trends, we analyzed performance with a new and more comprehensive dataset that includes five AI performance benchmarks covering mathematics, hard sciences, and games of skill over the last three years.

Across that time, we find that the price for a given level of performance has fallen about 47% per quarter, or 13x per year.
We see slower drops on game-based puzzles, at about 39–43% per quarter (7–10x per year), and faster progress on math problems, at 50–52% per quarter (16–19x per year).

The cost decline for a given level of performance does tend to slow over time, though this pattern is not universal.
One possible explanation: when a performance level is first achieved, AI companies can briefly charge a premium for it, before competition and technological improvement quickly drive down the price.
In time, that dynamic slows.
Averaging across the five primary benchmarks, cost falls 66% per quarter (75x per year) at first.
Two years later, it falls half as fast, at a “mere” 32% per quarter (4.7x per year).

Our analysis comes with major caveats.
AI companies may be expressly training their models for some benchmarks (“benchmaxxing”), so that improvement on the benchmarks outstrips improvement for real-world tasks.
Even if they are not, doing well on a benchmark is not synonymous with useful work.
Because we focus on the frontier — the absolute cheapest model capable of any given level of performance — we implicitly posit an AI user who relentlessly searches for the most cost-effective model for each task, when real users do not switch models so often, and therefore do not reap quite the same savings.
Our data are incomplete and noisy: the timeframe is barely three years, and we do not include all combinations of AI model and benchmark.
Prices drop differently for different models, benchmarks, time periods, and performance ranges, and there are many reasonable ways to average over this variegated experience.
Overall, while we believe that our bottom-line numbers are reasonably representative of reality, they should not be read as exact.

The rest of this report details our analysis.
Parts of it are technical.



## Previous work

We are not the first to quantify how fast the price of AI is falling.
A [2024 post](https://a16z.com/llmflation-llm-inference-cost/) by Guido Appenzeller for Andreessen Horowitz documented how, in the three years following the general release of GPT-3, LLM costs fell by a factor of 1000, i.e., 10x per year.
A few months later, in March 2025, an [analysis by Epoch AI](https://epoch.ai/data-insights/llm-inference-price-trends) found 9–900x per year drops across six performance benchmarks.

Both of those early analyses measured *prices per token for models capable of achieving a given performance* as distinct from *actual cost to achieve given performance*.
With the advent of reasoning models — OpenAI released o1 in December 2024 — it has become more problematic to ignore this distinction.
Reasoning models can productively consume far more tokens, but as a result extract good performance from an underlying LLM that is smaller and cheaper to run per token.

In September 2025, Håvard Tveit Ihle [shared](https://www.lesswrong.com/posts/ifSBamvobbyB9KWjK/inference-costs-for-hard-coding-tasks-halve-roughly-every) an analysis on LessWrong that directly compared performance and cost on two suites of coding challenges.
The analysis finds that costs halved every 1.4–2 months (64–380x per year).

The most thorough analysis yet is the [March 2026 paper by Gundlach et al.](https://arxiv.org/abs/2511.23455v2) It, too, compares actual costs to performance.
As in the present analysis, it estimates trends both in the full body of data and in the subset of models defining the cost frontier at any given time.
It also disaggregates by level of performance — prices fall faster at the high end — and by model type (open or closed, dense or mixture of experts).
Overall they find declines of 5–10x per year.



## Data

The major novelty in the present analysis is to analyze trends in costs using a methodology that captures each model’s full expense-performance continuum.

If one LLM can score 80% on a benchmark, and costs $1 to do so, and another peaks out at 60%, for a price of $0.50, it is not obvious which is more cost-effective.
Perhaps the stronger model would, if run with fewer reasoning tokens, achieve 60% more cheaply than the weaker model.

More generally, our interest is in mapping the “Pareto” frontier of cost-effectiveness — the cheapest way to attain each performance level from the models available at any given time.
We would leave a lot of territory unmapped, and potentially a lot of frontier as well, if we only estimated cost and performance when LLMs are given an unlimited budget.
Any given modern LLM can produce a range of performance levels, depending on whether its reasoning is set to low, medium, high, or max, and depending on whether a budget limit is imposed.

One way to trace a model’s expense-performance curve is to run it many times against a benchmark, at various reasoning levels and token budgets.
That process would, however, be expensive and slow.
Instead, we follow a procedure developed by [the federal Center for AI Standards and Innovation (CAISI)](https://www.nist.gov/system/files/documents/2025/09/30/CAISI_Evaluation_of_DeepSeek_AI_Models.pdf#page=66).
It uses the *transcript* from a benchmark run with a high (or no) budget constraint to predict performance under tighter budgets.
The core idea is that since benchmarks consist of many questions, one can estimate how many an LLM would answer before consuming a given budget.
The transcript provides the needed information: how many tokens the model consumed in answering each question, and which it got right.[2](https://epoch.ai/publications/the-plunging-price-of-thought#user-content-fn-2)

For any arbitrary per-question budget of X tokens, we sum up the number of questions that were answered correctly in fewer than X output tokens.
Think of this as the maximum expense the model may incur before being forced to give up.
When a model does not answer in time for a particular budget threshold, that question is scored as the probability of guessing correctly, which could be, for example, 0.25 on four-way multiple choice and effectively 0 on other kinds of problems.
Repeating across a range of budgets produces a curve for performance as a function of hypothetical expenditure.

One concern about this procedure is that it could misestimate how LLMs would actually perform if *informed* of a budget.
If models were told that they should optimize for getting a good score at some lower threshold, they might do better than if the questions were silently truncated, as simulated in the CAISI methodology.

To test this concern, we modified our evaluations by explicitly telling models their per-question budgets.
Our results suggested that while high-effort versions of thinking models often improve their low-budget performance when those budgets are announced, they do not outperform lower-effort variants under standard, unannounced conditions.
That is, simply running thinking models at low thinking effort produces expense-performance curves similar to those from running at high thinking effort under announced limits.
(See Figure 3 for examples across several models.)
As a result, as long as we use the CAISI methodology across the range of thinking levels, we will probably cover most of a model’s cost-performance possibility domain.

Four-panel chart of GPQA Diamond accuracy against per-question output-token budget for GPT-5.2, o4-mini, Claude Sonnet 4.6 and Gemini 3.5 Flash. Curves from silent truncation at each reasoning effort closely track dots from runs where the model was told its budget.

Since historically our benchmarking efforts have tended to focus on capturing the upper limits of performance with minimal regard for cost efficiency, our data was systematically missing evaluations using lower reasoning 
Proper_Actuary290717040
🟧 echo.blog ⭐The report estimates approximately 47% quarterly cost declines at fixed benchmark performance, with public code and data and explicit limitaLuke Emberson and David Roodman——

Interpretation history

Decision trace