2026-10-11 15:03 UTC

GitHub user tongroy's AgentMeasure audit reports 45 verified token-accounting bugs in AI cost dashboards, suggesting inference-cost measurements are unreliable enough to distort spending decisions.

state: acceleratingheat: lowuncertainty: mediumconvergesscott: highinference-economics

What is this?

AgentMeasure is an open-source GitHub project by user roy-tong (HN: tongroy) billing itself as 'open measurement infrastructure for AI agents'; its flagship artifact, issue #18 'The Token-Accounting Bug Report' (opened Sep 10, 2026), claims an audit of ~110–124 AI usage/cost tools using synthetic JSONL fixtures, yielding 45+ 'verified' billing bugs across five recurring classes — per-block re-summation, re-emitted events, cache-pricing semantics, price-table drift, and resume/fork loss. The project's own numbers drift across surfaces (10 merged fixes in the issue body vs. 19 in the README; 45+ bugs vs. 65+ findings; ~110 vs. 124 tools), and the supplied material is entirely self-description — no exposed tool inventory or independent validation — with traction limited to a 218-star repo and a 1-point self-submitted Show HN. Around it, commercial actors are monetizing the same anxiety: Vaudit launched a TokenAudit bill-recovery product via press release, and observability vendors (Larridin, CloudZero, Mavvrik) cite Bloomberg/Axios reporting on enterprise AI budget blowouts (Uber capped employee agentic-coding usage after exhausting its 2026 budget in four months; one enterprise reportedly spent ~$500M in a single month on Claude licenses), though vendor blog citations are secondhand and none of the supplied snippets independently verifies AgentMeasure's specific bug counts.

Why it matters to Scott

Converges and is now operationally live: the independently corroborated Claude Code ~2x /usage overcount sits in exactly the token counters gating his LiteLLM/Langfuse cost-tiered routing and autonomy budgets, and the corroborated mechanism — raw JSONL receipts recount correctly while the aggregation layer double-counts — vindicates his agent-receipts-over-dashboards doctrine rather than merely illustrating it. Meanwhile the WSJ unbudgetable-spend piece and Vaudit's $1.7M recovery claim are dated mainstream receipts for the observability ebook and bounded-budget token-economics arguments, and his own indexed Claude Code JSONLs (dev:project.search) make a cheap local recount immediately actionable — with AgentMeasure's specific 45-bug list still self-described and best treated as an unverified checklist, not a finding.
ip:source.observability-for-agentic-systems-what-to-log-how-to-redact-how-to-debug-ebookip:concept.agent-observabilityip:concept.token-economicsip:concept.autonomy-budgetip:concept.agent-receiptsdev:concept.cost-tiered-llm-routingdev:technology.litellmdev:technology.langfusedev:technology.claude-coderadar:concept.inference-economicsradar:concept.agent-observabilityradar:concept.token-economicsradar:provider-token-inflation-auditradar:claude-phantom-token-billing-bugradar:claude-code-headless-metering-3xradar:replay-prompt-cache-miss-audit
queries asked of Scott's wikis
  • cost-tiered model routing gateway provider-switch thresholds
  • LiteLLM Langfuse token counters usage stats
  • bounded budget token economics observability ebook
  • agent observability measurement trust dashboards
  • Claude Code JSONL usage logs audit overcount
  • billing audit conformance fixtures verification

Measured heat

now 0 pts/hpeak 24 pts/hcomments 0/hpeers p7momentum: steady3 platformsage 752h
points/hour across evidence · reading as of 2026-10-12 01:07:16.990013+11:00 · deterministic, not a model opinion

How the heat travelled

09-10 07:33 (minted)⭐ origin echo-reconstructedAI cost dashboard is probably wrong - 45 verified token-accounting bugs
tongroy on github (echo) · attributed from hn.story.49639249 · published time unknown
—
09-10 06:30first on hacker news · published · lag ?AI cost dashboard is probably wrong – 45 verified token-accounting bugs
tongroy
—
09-16 11:44first on r/ClaudeAI · published · lag ?Claude Code's /usage Stats tab overstates your tokens ~2x. Reported on GitHub since Aug 2025, still unfixed.
Background-Basis-672
—
09-16 18:27first on r/OpenAI · published · lag ?SillyTavern v1.19.0 released
Fcking_Chuck
—
09-10 06:30amplified on hacker newshn.story.49639249
tongroy
peak 1 · 0 comments · 0% of case engagement
09-11 11:15amplified on hacker news 👑hn.story.49656471
michalwarda
peak 170 · 84 comments · 62% of case engagement
09-14 17:44amplified on hacker newshn.story.49700895
practicalsystem
peak 1 · 0 comments · 0% of case engagement
09-14 21:14amplified on hacker newshn.story.49704131
lognebudo
peak 1 · 0 comments · 0% of case engagement
09-16 11:44amplified on r/ClaudeAIreddit.post.1whuvcf
Background-Basis-672
peak 7 · 2 comments · 1% of case engagement
09-16 18:27amplified on r/OpenAIreddit.post.1wi5h3r
Fcking_Chuck
peak 9 · 0 comments · 1% of case engagement
4 more amplifiers in ainews.case_chain
09-10 07:21our radar first saw it · lag ?discovery anchor: hn.story.49639249—
pace: p82 vs 520 stories at the 720h mark (now 752h old) — ahead of coop-coding-agent-vm-isolation (1.0x), behind deathray-macos-webgpu-hang (1.0x)

Evidence (11) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAI cost dashboard is probably wrong – 45 verified token-accounting bugstongroy10
🟧 echo.github ⭐AI cost dashboard is probably wrong - 45 verified token-accounting bugstongroy——
🟧 hnRTK reports token savings, but our cost benchmarks disagreemichalwarda17084
🟧 hnShow HN: CostClaw, a local audit of your Claude Code logspracticalsystem10
🟧 hnOllama 0.33.3 changed what prompt_eval_duration measureslognebudo10
🟠 redditClaude Code's /usage Stats tab overstates your tokens ~2x. Reported on GitHub since Aug 2025, still unfixed.
ClaudeAI
Background-Basis-67272
🟠 redditSillyTavern v1.19.0 released
OpenAI
Fcking_Chuck80
🟧 hnShow HN: AgentMeasure – healthchecks and settlement statements for AI billstongroy20
🟧 hnScan your LLM agent's tools for cost-amplification bugs (no API key)sarojas12310
🟧 hnSpending on AI Is Becoming Almost Impossible for Businesses to Budgetswolpers5779
🟧 hnShow HN: Tallyhook – price Claude Code and Codex logs per client you billvitaljudge21

Interpretation history

Decision trace