2026-10-11 16:38 UTC

Ornn Data claims open-weight models can deliver comparable intelligence at roughly one-fifth the cost of closed models and that self-hosted sparse inference can favor older A100 GPUs, potentially extending the economic life of existing accelerator fleets.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highopen-weight-inference inference-economics ai-infrastructureOrnn DataNVIDIA
Surfaced 2026-09-28T12:30:47Z — The Economics of Open-Weight Inference — The FT reporting on enterprises rejecting frontier pricing for open models is an independent adoption-side line, distinct from Ornn's self-interested paper and from the toy-scale P100 demo — that satisfies the two-independent-lines bar and promotes the case to corroborated. The contested core (the A100-beats-H100 sparse reversal) remains single-source and unreplicated, so this is corroboration of the economics thesis, not of Ornn's flagship number.

What is this?

Ornn Data — a commercial firm whose own materials (per the case's disclosure note) sell GPU-market data and rent GPUs — published 'The Economics of Open-Weight Inference' on 7 Sep 2026, arguing that once models are standardized for intelligence on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model completes a task for roughly one-fifth the cost of the cheapest comparable closed model ($0.09 vs $0.43), that self-hosting on rented hardware computes to $0.12–$0.35 per million output tokens at full utilization, and that on the sparse model gpt-oss-120b an older A100 produces output more cheaply than an H100 — a hardware-ranking reversal implying existing accelerator fleets keep economic life longer than assumed. The paper itself flags its soft spots: open weights are cheaper only at some sampled intelligence thresholds (the closed frontier still tops the range), the sparse A100/H100 rows combine third-party serving measurements rather than MLPerf results, the dense A100 row is estimated, and all costs are compute-only. The supplied web record corroborates the headline figures mainly through Ornn's own channels (its site, its X account, and a team member's LinkedIn announcement), with adjacent independent work pointing the same direction — an arXiv on-premise cost-benefit study and MIT Sloan's estimate that open-model inference is ~87% cheaper — while a visible counter-debate (self-hosting rarely beats hosted APIs in practice; 'open-weight' licenses are restrictive) stands unaddressed, and neither the paper's A100 rental-market data nor its citation of a reported NVIDIA acquisition of Hugging Face is corroborated anywhere in these snippets. The one independent evidence object attached to the case is a hobbyist demo running a 27B model on two ~$80 decade-old Tesla P100s at ~60 t/s after custom kernel work — genuine support for the old-hardware mechanism, but toy-scale and not a replication of Ornn's economics.

Why it matters to Scott

Ornn Data independently prices what Scott's canon argues: ip:concept.ai-unit-economics' self-hosting break-even and vintage-GPU queries get market-grade numbers (A100 5-year terms retaining 80.2% of price at 90% occupancy), ip:concept.model-perishability gets its best dated receipt yet — a quantified hardware-ranking reversal showing old fleets keep economic life — and the $0.09/task qualifying open model is a concrete price point for the model-barbell's cheap end and cost-tiered routing on gamepc. The hobbyist P100 demo adds independent (toy-scale) implementation evidence for the same mechanism the radar already tracks as DumpsterCluster / v100-skinny / CMP-170HX bets. Caveats — the source sells GPU data and rents GPUs, the A100 sparse reversal is single-source and unreplicated — make this dated-receipts ammunition for the LeverageAI pitch and a watch-item, not a settled claim, but nothing here contradicts the canon: it corroborates it from outside.
ip:concept.ai-unit-economicsip:concept.model-perishabilityip:concept.model-barbelldev:project.gamepcdev:concept.cost-tiered-llm-routingwork:concept.ai-consulting-practiceradar:concept.inference-economicsradar:concept.open-modelsradar:concept.self-hostingradar:concept.gpu-infrastructureradar:concept.sparse-moeradar:dumpstercluster-low-cost-70b-servingradar:dumpstercluster-retired-gpu-inferenceradar:v100-nvfp4-inferenceradar:nvidia-cmp-vram-unlockradar:expertcache-gpt-oss-120b-m1radar:vercel-september-open-weight-majority
queries asked of Scott's wikis
  • model perishability open-weight gap closing lifecycle
  • self-hosting break-even inference vs API cost
  • vintage GPU inference economics old accelerators
  • cost-tiered model routing intelligence per dollar
  • sparse MoE local inference quantization
  • DumpsterCluster used GPU engineering bets

Measured heat

now 0 pts/hpeak 31 pts/hcomments 0/hpeers p16momentum: steady2 platformsage 458h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-22 13:45⭐ origin directly observedThe Economics of Open-Weight Inference
marinesebastian on hacker news
—
09-27 20:42first on r/LocalLLaMA · published · +126.9h2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0
Kmic68
—
09-28 10:59first on r/singularity · published · +141.2hFT: Corporate America rejects overpriced frontier, embraces open models
chocolateUI
—
10-09 11:51first on hacker news · published · +406.1hThe Case for Small Specialized Models
julesbelveze
—
09-22 13:45amplified on hacker news 👑hn.story.49801218
marinesebastian
peak 58 · 27 comments · 60% of case engagement
09-27 20:42amplified on r/LocalLLaMAreddit.post.1wrv2l3
Kmic68
peak 21 · 24 comments · 18% of case engagement
09-28 10:59amplified on r/singularityreddit.post.1wsbi4o
chocolateUI
peak 37 · 4 comments · 16% of case engagement
10-09 11:51amplified on hacker newshn.story.50019194
julesbelveze
peak 1 · 0 comments · 1% of case engagement
10-09 20:55amplified on hacker newshn.story.50026463
yogthos
peak 5 · 2 comments · 5% of case engagement
09-22 16:21our radar first saw it · +2.6hdiscovery anchor: hn.story.49801218—
09-28 12:25reached heat=high · +142.7h · via ledger——
pace: p73 vs 1032 stories at the 336h mark (now 458h old) — ahead of sakana-pc-alm-local-training (1.0x), behind perplexity-numeric-citation-audit (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐The Economics of Open-Weight Inference
Retrieved article excerpt

Open article · Retrieved 2026-09-23T16:27:34.508475+00:00

[Publications](https://data.ornn.com/publications)

# The Economics of Open-Weight Inference

How open-weight demand can support the useful life of NVIDIA GPU families

Ornn Data · 7 September 2026

## Abstract

GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.

## Key findings

- Hosted open-weight models were cheaper at several sampled common score thresholds on the Artificial Analysis Intelligence Index, not at every threshold. Closed models remain the cheaper qualifying option at some scores, and the closed frontier exceeds the open sample at the top of the range.
- Self-hosting on rented hardware produced compute-only costs of $0.12 to $0.35 per million output tokens at full utilization in the printed sample.
- On gpt-oss-120b, the A100 produced output more cheaply than the H100 at spot and at the three- and five-year term prices. The A100/H100 sparse result combines different third-party serving setups. The dense A100 row is estimated.
- The five-year A100 term price retains 80.2% of the one-month price, versus 43.7% to 59.8% for Hopper and 53.8% for Blackwell. Forward marks are analyst-produced indicators, not executable quotes.
- A100 occupancy rose from 74% to 90% as listed capacity rose 13% and the spot index rose 20%. The paper does not establish that open-weight demand caused A100 occupancy or rental-price behavior.

## Selected figures and tables

$0.00$0.50$1.000.750.29A100 SXM40.680.64H100 SXM0.810.98H2000.390.47B200Dense (Llama-2-70B)Sparse (gpt-oss-120b)

Compute-only cost per million output tokens at full use and the base case

| GPU | Dense full use | Dense base case | Sparse full use | Sparse base case |
| --- | --- | --- | --- | --- |
| A100 SXM4 | $0.28 | $0.75 | $0.12 | $0.29 |
| H100 SXM | $0.26 | $0.68 | $0.27 | $0.64 |
| H200 | $0.32 | $0.81 | $0.35 | $0.98 |
| B200 | $0.16 | $0.39 | $0.15 | $0.47 |

Source: Ornn Data calculation from published throughput and 1 September 2026 spot rents (paper Table 7). Full use takes Offline throughput with no headroom. The base case takes Server throughput where available and otherwise rescales the GPUStack baseline; the latter does not establish a Server service level. Dense A100 is estimated. The A100/H100 sparse rows are third-party measurements, not MLPerf results. Costs are compute-only USD per million output tokens.

A100 occupancy, listed-capacity change, and spot-price change

| Family | Occupancy, 1 Mar 2026 | Occupancy, 1 Sep 2026 | Occupancy, Mar–Sep mean | Listed capacity, Mar–Sep | Spot, Mar–Sep |
| --- | --- | --- | --- | --- | --- |
| A100 SXM4 | 74% | 90% | 80% | +13% | +20% |

Source: Ornn occupancy and listed-capacity series and the settled spot index (paper Table 8), retrieved 2 September 2026. Occupancy is rented capacity divided by listed capacity across tracked on-demand providers in the global region. Listed capacity measures tracked on-demand supply, not the installed base.

406080100806044541M6M1Y3Y5YTerm price as % of the 1-month markA100 SXM4H100 SXMH200B200B300

Forward term-price retention by GPU family

| Family | 6M / 1M | 1Y / 1M | 3Y / 1M | 5Y / 1M | Fwd 37–60 / 1M | Age at 5Y end (yrs) |
| --- | --- | --- | --- | --- | --- | --- |
| A100 SXM4 | 95.0% | 93.1% | 84.2% | 80.2% | 74% | 11.3 |
| H100 SXM | 95.0% | 86.2% | 68.2% | 59.8% | 47% | 9.4 |
| H200 | 92.4% | 80.2% | 51.0% | 43.7% | 33% | 7.8 |
| B200 | 99.1% | 97.9% | 69.1% | 53.8% | 31% | 7.5 |
| B300 | 93.8% | 81.3% | 60.9% | 53.8% | 43% | 6.5 |

Source: Ornn reported term marks (paper Table 9). A100, H100, H200, and B200 marks were published 13 August 2026; B300 on 1 September 2026. Retention and implied-forward ratios are calculations from those marks. Ages use an illustrative 1 September 2026 start measured from family announcement dates. Forward marks are analyst-produced indicators, not executable quotes.

## Open-weight portability

Closed models are available only through deployments authorized by their developers. Open weights allow independent deployment and, where compatible endpoints exist, a choice of providers serving the same checkpoint. Reported September 2026 OpenRouter snapshots listed eighteen to twenty-two providers for several widely served open models, with highest-to-lowest output-price ratios between 1.8 and 5.6. Those figures illustrate price variation, not interchangeable service: quantization, capacity, reliability, and migration costs can limit substitution.

The same choice extends to hardware. Memory, numerical format, software, and licensing determine which deployments are feasible. gpt-oss-120b uses MXFP4 expert weights and has been served on a single 80 GB A100. Operators can therefore consider older hardware where the model fits, the software supports it efficiently, and the license permits the intended use. The economic distinction is who can make the deployment decision.

## GPU useful life

Ornn publishes a settled daily rental index and term-price curves at one-month, six-month, one-year, three-year, and five-year tenors. A term price is a flat GPU-hour rate over a stated contract tenor. An implied forward differences cumulative term costs to obtain a rate for the interval between two tenors. The marks price a rental service; an individual device’s resale value and operating life are separate quantities.

The A100’s five-year term price retains four-fifths of the one-month price for a contract that would run to September 2031, when the family will be more than eleven years old. The Hopper and Blackwell curves are steeper. Rising occupancy on an expanding A100 listing base means the family’s spot strength cannot be attributed solely to fewer tracked GPUs being available for rent. A new generation does not, by itself, make its predecessors economically obsolete.

## Self-hosting economics

Self-hosting has no posted per-token price. Its compute-only cost is the GPU-hour rent on Ornn’s spot index divided by the throughput the operator achieves, adjusted for productive utilization and reserved headroom. The paper constructs that cost for two workloads: dense Llama-2-70B and sparse gpt-oss-120b.

The ordering reverses across those workloads. The dense comparison favors newer hardware. For gpt-oss-120b, the A100’s base-case cost is $0.29 per million output tokens, less than half the H100’s $0.64. Those A100 and H100 sparse inputs come from different third-party serving setups. The dense A100 row is estimated from a vendor throughput ratio in the absence of a matched MLPerf submission. Electricity at nameplate GPU power is a few percent of spot rent and does not reverse the printed rankings.

## Interpretation

Long-running agents, batch evaluation, and parts of reinforcement learning offer flexibility over when and where work runs. Workloads that tolerate latency can route to any cost-efficient compatible hardware. Older hardware does not need to win every workload to retain an economic role. Newer GPUs can offer lower costs on demanding tasks while older capacity serves work suited to its memory, software, and rental price.

The right measure of hardware usefulness is the cost of the work it can serve and the demand for that work, not simply the age of the chip. Evaluating an individual investment still requires acquisition costs, operating cash flows, and the alternatives available to the owner. Financing and contract terms belong in that assessment because they help determine the price of a long commitment.

## Limitations and disclosures

Forward marks are analyst-produced indicators, not executable quotes or verified averages of comparable executed contracts at every tenor. Differencing term marks does not identify expected future spot rent. August term marks and September spot observations are different vintages.

Occupancy records rental status across tracked on-demand providers, not the installed base, tenant workloads, or realized token throughput. Family ages run from NVIDIA’s announcement, not a device’s commissioning date. Strong percentage retention can also follow earlier declines in the starting rent. The marks do not establish individual-device survival, profitable operating life, residual value, or the appropriateness of an owner’s accounting depreciation schedule.

Dense A100 throughput is estimated; alternative precision assumptions can reverse its cost ranking. The A100/H100 sparse result combines different third-party serving setups. A common Index threshold screens models but does not establish equal task success. The two self-hosted examples do not establish the performance of the higher-scoring models in the hosted comparison.

The rental observations do not identify the contribution of open-weight demand. The paper does not establish that open-weight demand caused A100 occupancy or rental-price behavior. Device retirements, dense inference, non-LLM work, operating withdrawal costs, and financing or scarcity effects remain alternative explanations.

Ornn Data publishes this paper and the rental index, occupancy series, and term marks used in it. It licenses data commercially and also offers GPU rentals through Ornn Compute. These activities create a commercial interest in both the interpretation and adoption of its data. There is no external funding to report. Nothing in this paper is an offer to trade or investment, accounting, tax, or legal advice.

## Are open-weight models cheaper than closed models?

In the paper sample, the cheapest qualifying open-weight model, standardized by intelligence on the Artificial Analysis Intelligence Index, completes a task at roughly one fifth the cost of a comparable closed model. Open-weight models are cheaper only at several sampled common score thresholds, not at every threshold. Closed models remain the cheaper qualifying option at some scores, and the closed frontier exceeds the open sample at the top of the range.

## Is the A100 cheaper than the H100 for self-hosted open-weight inference?

On gpt-oss-120b, a sparse mixture-of-experts model with 5.1 billion active parameters, the computed A100 cost is $0.12 per million output tokens at full use and $0.29 in the base case, below the H100 figures of $0.27 and $0
marinesebastian5827
🟠 reddit2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0
LocalLLaMA
Kmic682124
🟠 redditFT: Corporate America rejects overpriced frontier, embraces open models
singularity
chocolateUI334
🟧 hnThe Case for Small Specialized Modelsjulesbelveze10
🟧 hnEcosia switches from Mistral to open-weight AI models including Qwen, GLM, Kimiyogthos52

Interpretation history

Decision trace