2026-10-11 17:13 UTC

Vercel claims open-weight models reached 56% of its AI Gateway token volume but only 14% of spending in August 2026, helping lower average token prices by 23.2% and strengthening the economic case for workload-specific model routing.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumopen-weight-models inference-economics model-routingVercel

What is this?

Vercel, the cloud platform led by CEO Guillermo Rauch, operates AI Gateway — a routing layer between production applications and multiple model providers — and publishes a monthly Production Index from that traffic. The September 2026 index (primary blog post verified in these results) reports open-weight models processed 56% of August gateway tokens — their first majority, up from 7% in December 2025, 13% in April and 36% in July — while accounting for only 14% of estimated spend, as average price per token fell 23.2% in August (third consecutive monthly drop; closed-weight tokens cost ~7.8x open-weight ones per Techstrong's derivation). Anthropic nonetheless held 64% of gateway spend, and a September 19 daily snapshot from Rauch showed open-weight share touching 78.4% in a single day, with Moonshot AI and DeepSeek ranking third and fourth by estimated spend and their combined spend with Z.ai exceeding OpenAI's that day. Coverage is consistent across the primary source and multiple independent outlets, but all figures describe one gateway's self-selected customer population at list price — they measure Vercel's traffic, not the whole market.

Why it matters to Scott

Vercel's dated production index — open weights at 56% of tokens for 14% of spend, third straight monthly price drop — plus the FT's demand-side reporting show the world independently arriving at the task-risk/price barbell Scott already runs through his LiteLLM tiers, OpenRouter free-tier primary and Ollama bulk routes, making this a dated-receipts update for his Agent Token Manifesto/pricing guide rather than news that moves his position. The limits hold: gateway-population list-price figures still don't establish cost per successful task, which is exactly the divergence radar:hidden-reasoning-real-task-costs tracks and the open question for whether his actual routes should widen their open-weight share.
dev:concept.task-aware-model-routingdev:concept.cost-tiered-llm-routingdev:technology.litellmdev:technology.openrouterdev:project.llmreportradar:concept.model-routingradar:concept.inference-economicsradar:concept.open-weight-modelsradar:concept.llm-gatewaysradar:anthropic-openrouter-spend-premiumradar:hidden-reasoning-real-task-costs
queries asked of Scott's wikis
  • model barbell routing frontier vs cheap open-weight models by task risk
  • task-aware model routing LiteLLM implementation production
  • open-weight vs closed inference cost per token economics
  • token price deflation frontier lab pricing power and margins
  • inference gateway as control plane / model-swap infrastructure bet
  • agent workload cost optimization choosing models per task

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-18 17:43 (minted)⭐ origin echo-reconstructedThe September Production Index reports that open-weight models processed 56% of August gateway tokens and accounted for 14% of spend, while
Vercel on blog (echo) · attributed from hn.story.49757318 · published time unknown
—
09-18 17:13first on hacker news · published · lag ?Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend
eyehurtsme
—
09-18 17:13amplified on hacker newshn.story.49757318
eyehurtsme
peak 2 · 0 comments · 25% of case engagement
09-23 00:50amplified on hacker newshn.story.49810265
mgdo
peak 2 · 0 comments · 25% of case engagement
09-27 19:44amplified on hacker news 👑hn.story.49870126
macleginn
peak 3 · 0 comments · 38% of case engagement
09-28 11:49amplified on hacker newshn.story.49876558
ostenbom
peak 1 · 0 comments · 13% of case engagement
09-18 17:21our radar first saw it · lag ?discovery anchor: hn.story.49757318—

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOpen-weight models take 56% of token volume, Astra doubles Fable 5.1 spend
Retrieved article excerpt

Open article · Retrieved 2026-09-18T17:23:59.126596+00:00

*AI Gateway Production Index — September 2026*

Every month, [AI Gateway](https://vercel.com/ai-gateway) routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from [June](https://vercel.com/blog/ai-gateway-production-index-june-2026), [July](https://vercel.com/blog/ai-gateway-production-index-july-2026), and [August](https://vercel.com/blog/deepseek-overtakes-google-on-volume-cost-per-token-falls).

### [Copy link to heading](https://vercel.com/blog/ai-gateway-production-index-september-2026#september-2026-summary)**September 2026 summary**

The September index reports on AI Gateway data collected through August 2026.

- **Open-weight models ran the majority of gateway tokens for the first time,** up from 7% in December to 56% in August.
- **The average token costs less than half what it did five months ago.** Price per token fell 23.2% in August, the third straight monthly drop, and the median team paid 7.6% less.
- **Fable 5, Anthropic's most capable model, lost two-thirds of its share of gateway spend in one month.** Opus 5, at half the price, tripled its share. Anthropic kept 64% of all spend.
- **Gemini 3 Flash has lost 95% of its share of gateway tokens** since May, and more than three-quarters of the volume it lost went to models from other labs.

Special report: [OpenAI launches Astra](https://vercel.com/blog/ai-gateway-production-index-september-2026#special-report:-astra-took-a-third-of-openai-spend-within-48-hours-and-outpaced-fable-5.1-two-to-one-at-launch) on September 3

- **GPT-6 Astra took a third of OpenAI's spend within two days of launch and twice Fable 5.1's share of gateway spend.** Introduced two days apart at the same price, Astra took 7.7% of all gateway spend in its first twelve days, while Fable 5.1 took 3.7%.

## [Copy link to heading](https://vercel.com/blog/ai-gateway-production-index-september-2026#open-weight-models-take-a-majority-of-token-volume-for-the-first-time)Open-weight models take a majority of token volume for the first time

In August, open-weight models ran 56% of all tokens on AI Gateway, marking the first month they took the majority of volume.

In December 2025, they processed fewer than one in ten tokens, and only eight months later, they ran more token volume than all closed-weight models combined.

Open-weight model token share rose every month from April through August, rising from 13% to 56% of total volume. Open-weight model token share rose every month from April through August, rising from 13% to 56% of total volume. Open-weight model token share rose every month from April through August, rising from 13% to 56% of total volume. Open-weight model token share rose every month from April through August, rising from 13% to 56% of total volume. 

Open-weight model token share rose every month from April through August, rising from 13% to 56% of total volume.

Though the frontier kept the majority of spend, open-weight dollar share is accelerating. As open-weight models become more capable, customers are moving more production workloads over to them.

Open-weight models processed 56% of August’s gateway tokens and accounted for 14% of spend.Open-weight models processed 56% of August’s gateway tokens and accounted for 14% of spend.Open-weight models processed 56% of August’s gateway tokens and accounted for 14% of spend.Open-weight models processed 56% of August’s gateway tokens and accounted for 14% of spend.

Open-weight models processed 56% of August’s gateway tokens and accounted for 14% of spend.

Growth in open-weight model adoption helped push the average price per token across the gateway down 23.2% in August, its third consecutive monthly drop and the steepest since April. Among teams running more than ten million tokens in both months, the median team paid 7.6% less per token, more than double July's 2.9% decline.

Teams can now get more inference from the same budget and reserve frontier models only for the tasks that justify the premium.

## [Copy link to heading](https://vercel.com/blog/ai-gateway-production-index-september-2026#frontier-plateaus-as-fable-spend-goes-to-opus-5)Frontier plateaus as Fable spend goes to Opus 5

Production workloads that justify a frontier model don't always need the most expensive one. They need one that’s good enough.

Fable is the most capable model Anthropic sells. Opus is the tier below it and costs roughly half of Fable’s price per token. When the US export control on Fable 5 was lifted and its access restored on July 1, its gateway spend share surged to 13.2%. At the end of that same month, Opus 5 came online.

In August, Fable 5’s share of gateway spend fell to 4.9%, and Opus 5's share rose to 22.5%. Nine in ten of the teams that ran Fable cut their usage, and more of them moved their workloads to Opus 5 than any other model. Fable’s extra capability wasn’t worth double the price.

Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's spend share rose to 22.5%. Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's spend share rose to 22.5%. Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's spend share rose to 22.5%. Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's spend share rose to 22.5%. 

Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's spend share rose to 22.5%.

Teams left Fable, Anthropic's most expensive and capable model, but the lab retained the lion’s share of gateway spend because those workloads stepped down to Opus 5, not a different lab.

Anthropic has taken at least 61 cents of every dollar spent through AI Gateway every month since December, and 64 cents in August. Its models have held the top two spots by spend every month since December, even as the models in those spots changed.

Anthropic has held the top two spots by spend every month since December. Google and OpenAI have each cracked third twice.Anthropic has held the top two spots by spend every month since December. Google and OpenAI have each cracked third twice.Anthropic has held the top two spots by spend every month since December. Google and OpenAI have each cracked third twice.Anthropic has held the top two spots by spend every month since December. Google and OpenAI have each cracked third twice.

Anthropic has held the top two spots by spend every month since December. Google and OpenAI have each cracked third twice.

## [Copy link to heading](https://vercel.com/blog/ai-gateway-production-index-september-2026#customer-loyalty-follows-the-model-profile,-not-the-lab)Customer loyalty follows the model profile, not the lab

Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins.

When a new model preserves what users valued in its predecessor, the lab retains its customers. When it doesn’t, those customers fill the need through other providers.

When Claude Opus 5 launched, it gained almost twice what Fable lost, because it handled the same workloads at half the price. And within five days of Z.ai launching GLM-5.3-Flash, it was running three times GLM-5.2's daily volume.

GLM-5.3-Flash overtook GLM-5.2 one day after appearing on AI Gateway and processed two-thirds of Z.ai’s tokens by August 31.GLM-5.3-Flash overtook GLM-5.2 one day after appearing on AI Gateway and processed two-thirds of Z.ai’s tokens by August 31.GLM-5.3-Flash overtook GLM-5.2 one day after appearing on AI Gateway and processed two-thirds of Z.ai’s tokens by August 31.GLM-5.3-Flash overtook GLM-5.2 one day after appearing on AI Gateway and processed two-thirds of Z.ai’s tokens by August 31.

GLM-5.3-Flash overtook GLM-5.2 one day after appearing on AI Gateway and processed two-thirds of Z.ai’s tokens by August 31.

Google struggled to retain customers with its new models. Because the new offerings didn’t provide a relative advantage on capability or price, a majority of Gemini 3 Flash’s workloads moved to OpenAI, Anthropic, and DeepSeek.

More than three-quarters of the volume Gemini 3 Flash lost went to other labs, and Google's own full-size Flash successors took under a tenth of it.More than three-quarters of the volume Gemini 3 Flash lost went to other labs, and Google's own full-size Flash successors took under a tenth of it.More than three-quarters of the volume Gemini 3 Flash lost went to other labs, and Google's own full-size Flash successors took under a tenth of it.More than three-quarters of the volume Gemini 3 Flash lost went to other labs, and Google's own full-size Flash successors took under a tenth of it.

More than three-quarters of the volume Gemini 3 Flash lost went to other labs, and Google's own full-size Flash successors took under a tenth of it.

About half of the volume that left Gemini 3 Flash went to cheaper models, led by GPT-5.6 Luna, which costs less than half as much per token. Most of the other half went to higher-priced models, led by Claude Opus 5 and Sonnet 5, which cost roughly nine and three times as much as Gemini 3 Flash, respectively.

The flight to better-fit models meant that over the same period, Google’s share of gateway token volume fell from 30% to 5%, with Gemini 3 Flash accounting for 22 of the 25 percentage points lost.

## [Copy link to heading](https://vercel.com/blog/ai-gateway-production-index-september-2026#special-report:-astra-took-a-third-of-openai-spend-within-48-hours-and-outpaced-fable-5.1-two-to-one-at-launch)Special report: Astra took a third of OpenAI spend within 48 hours and outpaced Fable 5.1 two to one at launch

GPT-6 Astra launched on the AI Gateway on September 3 at the same price as Fable 5.1 and two and a half times the price of GPT-5.6 Sol. Two days later, it accounted for one in every three dollars spent on OpenAI models through the gateway. Its share of spend has held, hovering between 28% and 39% since.

Within OpenAI’s model lineup, Astra and Sol processed 27% of OpenAI’s tokens but accounted for 71% of its spending from September 4 through 16. Luna and Nano processed more than twice as many tokens for about one-ninth as much spending.

Luna processed more than eight times as many tokens as Astra, but Astra accounted for more than four times as much spend.Luna processed more than eight times as many tokens as Astra, but Astra accounted for more than four times as much spend.Luna processed more than eight times as many tokens as Astra, but Astra accounted for more than four times as much spend.Luna processed more than eight times as many tokens as Astra, but Astra accounted for more than four times as much spend.

Luna processed more than eight times as many tokens as Astra, but Astra accounted for more than four times as much spend.

Anthropic launched Fable 5.1 on September 1, two days before Astra. Over each model's first twelve days on the gateway, Astra took 7.7% of all gateway spend, more than twice Fable 5.1's share of 3.7%, and was used by twice as many teams.

Astra passed Fable 5.1’s cumulative gateway spend on day four and reached twice Fable’s 12-day total by day 12.Astra passed Fable 5.1’s cumulative gateway spend on day four and reached twice Fable’s 12-day total by day 12.Astra passed Fable 5.1’s cumulative gateway spend on day four and reached twice Fable’s 12-day total by day 12.Astra passed Fable 5.1’s cumulative gateway spend on day four and reached twice Fable’s 12-day total by day 12.

Astra passed Fable 5.1’s cumulative gateway spend on day four and reached twice Fable’s 12-day total by day 12.

OpenAI’s cheaper models carry its volume, while Astra’s early lead over Fable shows it can also attract teams at the highest price point. Together, they let OpenAI compete with other frontier labs for both scale and premium spend.

Stay tuned for more in next month's report.

## [Copy link to heading]
eyehurtsme20
🟧 echo.blog ⭐The September Production Index reports that open-weight models processed 56% of August gateway tokens and accounted for 14% of spend, while Vercel——
🟧 hnOpen-weight models take 56% of token volumemgdo20
🟧 hnCorporate America embraces cheaper 'open' AI modelsmacleginn30
🟧 hnCorporate America embraces cheaper 'open' AI modelsostenbom10

Interpretation history

Decision trace