2026-10-11 17:11 UTC

Independent replications will determine whether hidden reasoning tokens and failed agent attempts cause cross-provider API costs on realistic tasks to diverge far more than published token prices imply.

state: resolvedheat: lowuncertainty: highconvergesscott: mediumapi-costs reasoning-tokens agent-economicsOpenAIAnthropicGoogleMoonshot AI

What is this?

The case concerns a purported Weckr benchmark comparing the real cost of completing 40 tasks across APIs from OpenAI, Anthropic, Google, and Moonshot AI, reporting a 10.6× cost spread despite only a 2× difference in published prices. The supplied snippets support the underlying concern that hidden reasoning tokens, retries, failed attempts, and output-heavy workflows can make per-token pricing a poor proxy for agent costs, with “cost per successful outcome” proposed as a better metric. However, the search results do not independently document or replicate Weckr’s benchmark, methodology, or raw results, so the headline comparison remains unverified by the supplied web evidence.

Why it matters to Scott

The purported benchmark independently operationalizes Scott’s existing AI unit-economics position that providers should be compared by total cost per successful outcome—including reasoning, retries, and failures—not headline token rates. It also directly extends his provider-execution benchmark and LLM pricing work, but the publishing opportunity remains conditional because the supplied evidence does not independently validate Weckr’s methodology or 10.6× result.
ip:concept.ai-unit-economicsip:concept.agent-observabilityip:concept.evaluation-driven-developmentdev:project.remote-execdev:project.llmreportradar:concept.ai-benchmarksradar:concept.benchmark-integrityradar:rtk-coding-agent-cost-regressionradar:claude-phantom-token-billing-bug
queries asked of Scott's wikis
  • cost per successful outcome for coding agents
  • agent retries and failed-attempt economics
  • reasoning-token observability and API billing audits
  • model routing by task success and total cost
  • cross-provider agent benchmark methodology
  • eval harnesses for realistic end-to-end API cost

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (33) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditReal task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]
MachineLearning
pixelo232330
🟧 echo.github ⭐The primary artifact is Weckr’s published benchmark suite and raw results. Its commit says: “First run 2026-07-23: all 4 providers, 40 of 40Ghiles Asmani (Ghiles3232)——
🟠 redditjust another benchmark: $0.34 vs $27.60 for the same tasks solved
LocalLLaMA
arsenyinfo255
🟠 reddit5.0 v 4.8 - credit usage
ClaudeAI
pgpnw36
🟠 redditEnterprise PSA: Check your default speed settings - Fast (2.5x usage) is set automatically
OpenAI
mawhii111
🟠 redditTask that used 5% of weekly now using 50%?
OpenAI
PM__me_sth147
🟠 redditI open-sourced a privacy-safe benchmark for coding-agent token experiments
OpenAI
bestofdesp75
🟠 redditPSA if you're on the API: the `thinking` default flipped between Opus 4.8 and Opus 5
ClaudeAI
BFitch8521
🟧 hnAmazon accidentally spent $1.8M using Claude for a menial coding task, wentsbulaev11
🟠 redditThe most expensive prompt I ever sent was two words
ClaudeAI
pyjuunu11
🟠 redditClaude Code burned my entire five-hour limit in 6 minutes 32 seconds: 10.26M tokens, zero lines of code
ClaudeAI
cosmintrica029
🟧 hnAmazon spent $1.8M using Claude for menial coding task, went 860% over budgetPLenz70
🟠 redditOpenAI uses 10 X the tokens for the same prompt. why?
OpenAI
michael_g_williams23
🟠 redditIntelligence density went up a lot this year and my bill didn't move. The $/M number is not where the money goes.
singularity
truecakesnake203
🟧 hnKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoningsbulaev20
🟧 hnReal cost running LLMs productionjaviermanzano10
🟠 redditDeepSeek tops AI models in affordability, new study says
artificial
LinkedInNews122
🟠 redditA cheaper AI model is not necessarily cheaper once retries are counted
artificial
ExtremeAdmirable409705
🟠 redditBaseline token usage across models
ClaudeAI
jerryadc14
🟠 redditSonnet 5's adaptive thinking quietly ate my JSON budget, and four other things that broke shipping a Claude-powered iOS app
ClaudeAI
haytchsquared01
🟧 hnLLM intelligence vs. cost per task, Dec 2024–Aug 2026theanonymousone20
🟠 redditI benchmarked 5 token saving tools across Codex and Claude code. The 60-90% token saving claims didnt hold up
ClaudeAI
Obvious_Gap_57682313
🟠 redditFable practically unusable with credits
ClaudeAI
Emergency-Bobcat64851923
🟧 hnQwen 3.8 and Claude Opus 5 show why raw benchmark scores don't predict the billashurandi20
🟠 redditAPI equivalent cost
ClaudeAI
IllustriousWedding94214
🟠 redditExtended thinking silently consumed 674k tokens — UI showed only 1.1k used
ClaudeAI
forchat412
🟧 hnI spent $200 in API credits asking AI agents to scaffold a Next.js startervladzoff10
🟧 hnShow HN: Claude Code and Codex usage screen for TRMNL X e-inkstared20
🟧 hnCan Claude Code in a loop improve an enterprise AI agent with $10,745 of budget?jeremytian54
🟠 redditIs more reasoning necessarily better?
artificial
CoVegGirl18
🟠 redditClaude Code Costs: Max 20x vs. Teams or Enterprise -- 60 Times Cheaper
ClaudeAI
Wsz2020010
🟠 redditHow much does your Claude Code usage cost at API rates?
ClaudeAI
entelligenceai17616
🟠 redditWhat is going on with my Claude usage?
ClaudeAI
Many-Teach981308

Interpretation history

Decision trace