AgentMeasure is an open-source GitHub project by user roy-tong (HN: tongroy) billing itself as 'open measurement infrastructure for AI agents'; its flagship artifact, issue #18 'The Token-Accounting Bug Report' (opened Sep 10, 2026), claims an audit of ~110–124 AI usage/cost tools using synthetic JSONL fixtures, yielding 45+ 'verified' billing bugs across five recurring classes — per-block re-summation, re-emitted events, cache-pricing semantics, price-table drift, and resume/fork loss. The project's own numbers drift across surfaces (10 merged fixes in the issue body vs. 19 in the README; 45+ bugs vs. 65+ findings; ~110 vs. 124 tools), and the supplied material is entirely self-description — no exposed tool inventory or independent validation — with traction limited to a 218-star repo and a 1-point self-submitted Show HN. Around it, commercial actors are monetizing the same anxiety: Vaudit launched a TokenAudit bill-recovery product via press release, and observability vendors (Larridin, CloudZero, Mavvrik) cite Bloomberg/Axios reporting on enterprise AI budget blowouts (Uber capped employee agentic-coding usage after exhausting its 2026 budget in four months; one enterprise reportedly spent ~$500M in a single month on Claude licenses), though vendor blog citations are secondhand and none of the supplied snippets independently verifies AgentMeasure's specific bug counts.
Converges and is now operationally live: the independently corroborated Claude Code ~2x /usage overcount sits in exactly the token counters gating his LiteLLM/Langfuse cost-tiered routing and autonomy budgets, and the corroborated mechanism — raw JSONL receipts recount correctly while the aggregation layer double-counts — vindicates his agent-receipts-over-dashboards doctrine rather than merely illustrating it. Meanwhile the WSJ unbudgetable-spend piece and Vaudit's $1.7M recovery claim are dated mainstream receipts for the observability ebook and bounded-budget token-economics arguments, and his own indexed Claude Code JSONLs (dev:project.search) make a cheap local recount immediately actionable — with AgentMeasure's specific 45-bug list still self-described and best treated as an unverified checklist, not a finding.
ip:source.observability-for-agentic-systems-what-to-log-how-to-redact-how-to-debug-ebookip:concept.agent-observabilityip:concept.token-economicsip:concept.autonomy-budgetip:concept.agent-receiptsdev:concept.cost-tiered-llm-routingdev:technology.litellmdev:technology.langfusedev:technology.claude-coderadar:concept.inference-economicsradar:concept.agent-observabilityradar:concept.token-economicsradar:provider-token-inflation-auditradar:claude-phantom-token-billing-bugradar:claude-code-headless-metering-3xradar:replay-prompt-cache-miss-audit
queries asked of Scott's wikis
- cost-tiered model routing gateway provider-switch thresholds
- LiteLLM Langfuse token counters usage stats
- bounded budget token economics observability ebook
- agent observability measurement trust dashboards
- Claude Code JSONL usage logs audit overcount
- billing audit conformance fixtures verification
2026-10-08T21:24:54Z
New periphery implementation (Tallyhook per-client billing tool) attached but adds no new facts to the corroborated core — Claude Code ~2× overcount, SillyTavern fix, Ollama semantics change remain the only independently verified defects. AgentMeasure's 45-bug list still unverified. Case-specific engagement has cooled to 0.33 pts/hr (steady momentum) though the inference-economics topic stays hot with 119 open episodes. Periphery expansion already priced at last look; this is another data point in the same direction.
2026-10-08T17:49:20Z
evidence attached: hn.story.50005611 — Per-client billing tool for Claude Code/Codex logs directly addresses the token-accounting reliability problem the case tracks.
2026-10-08T03:52:51Z
The WSJ story (hn.story.49964537) triggered a velocity spike (51× baseline) and the magnitude valve fired, confirming cross-platform spread at top-decile engagement — but this is renewed attention on an already-attached evidence object, not new facts. The corroborated core (Claude Code ~2× overcount, SillyTavern fix, Ollama semantics change) and the unverified status of AgentMeasure's 45-bug list are unchanged. The periphery expansion into mainstream press and commercial audit entrants was already priced in at the last look.
2026-10-05T16:22:31Z
grounded: converges/high — Converges and is now operationally live: the independently corroborated Claude Code ~2x /usage overcount sits in exactly the token counters gating his LiteLLM/L
2026-10-05T16:12:33Z
The case graduates to accelerating: the periphery has crossed from dev communities into mainstream business press (WSJ on unbudgetable AI spend, moving at top-decile velocity for its cohort) and commercial audit entrants (Vaudit's ~$1.7M recovery claim), matching the hypothesis's second clause that cost-measurement unreliability now distorts spending at scale. Caveat kept sharp: the WSJ piece documents unpredictable spend, not proven accounting error, so it corroborates the stakes, not the defect — and AgentMeasure's specific 45-bug list remains unverified.
2026-10-05T15:30:48Z
evidence attached: hn.story.49964537 — WSJ's high-engagement piece on enterprises unable to budget AI token spend is mainstream macro corroboration that cost measurement and predictability are now distorting spending decisions.
2026-10-02T00:21:31Z
grounded: converges/high — AgentMeasure is substantiated well beyond the headline: a released conformance suite testing exactly the semantics Scott's observability work prescribes (a retr
2026-10-02T00:12:41Z
The case's center of gravity has shifted from tongroy's unverified 45-bug audit to a corroborated pattern: Claude Code's ~2x Stats overcount (user recount with a concrete duplication mechanism), SillyTavern's shipped fix for inflated /tokens output, and Ollama's prompt_eval_duration semantics change are independent lines that token/metric accounting in shipping tools is defective — none of them validates AgentMeasure, and no distorted spending decision is demonstrated. The new cost-amplification scanner is adjacent periphery (inflated real costs are a different defect class from misreported costs) and does not by itself change the assessment; the promotion reweights the accumulated independent instances across the two-line bar.
2026-10-01T23:31:30Z
evidence attached: hn.story.49927585 — Independent tooling that scans agent stacks for cost-amplification bugs reinforces the open case's claim that agent cost accounting is unreliable enough to distort spending.
2026-09-19T09:22:00Z
The new same-author Show HN announcement positions AgentMeasure as a billing-audit product, but supplies neither an inspectable implementation nor independent validation of its 45-bug claim. The attachment rationale overstates the evidence: this is an HN headline, not a supplied GitHub artifact or demonstrated accounting result.
2026-09-19T09:21:37Z
evidence attached: hn.story.49764767 — This first-party GitHub artifact is direct corroboration that AgentMeasure is released as a health-check and billing-audit tool for unreliable AI cost accounting.
2026-09-16T19:30:46Z
The SillyTavern release summary alleges a fix to inflated tokenizer-command output, not demonstrated cost-dashboard or billing errors. Without primary release details or a connection to AgentMeasure, it adds an adjacent verification lead but does not materially strengthen this audit case.
2026-09-16T19:22:34Z
evidence attached: reddit.post.1wi5h3r — The release fixes inflated tokenizer counts, providing a concrete corroborating example that AI cost dashboards can misreport token usage.
2026-09-16T12:25:40Z
A separate Claude Code user now reports a concrete duplication mechanism and a raw-log recount, making dashboard overcounting a testable risk rather than just an audit headline. This supports the broader accounting concern, but does not verify AgentMeasure’s 45 bugs or establish incorrect billing; the same report says the adjacent Usage tab deduplicates correctly.
2026-09-16T12:21:48Z
evidence attached: reddit.post.1whuvcf — Detailed independent recounting reports the same kind of token-accounting duplication that AgentMeasure flags, strengthening the measurement-reliability case.
2026-09-14T21:33:39Z
The Ollama headline alleges a change in timing-metric semantics, not a token-counting or billing error, and provides no supporting implementation details. It adds a separate observability caution but does not corroborate AgentMeasure’s audit or establish a threat to Scott’s budget controls.
2026-09-14T21:22:03Z
evidence attached: hn.story.49704131 — Independent evidence that inference metrics can change across runtime versions materially supports the case that cost accounting is unreliable.
2026-09-14T18:44:25Z
CostClaw’s headline introduces a potential local auditing tool, but supplies neither implementation results nor accounting discrepancies and does not corroborate AgentMeasure. The case remains an unverified audit claim rather than demonstrated risk to Scott’s budget controls.
2026-09-14T18:22:33Z
evidence attached: hn.story.49700895 — A local Claude Code log-audit tool bears directly on the reliability and observability of coding-agent cost measurements.
2026-09-11T12:36:01Z
The RTK discussion concerns whether reported token savings translate into lower task costs, not whether dashboards miscount tokens; it does not independently corroborate AgentMeasure’s claimed 45 bugs. The attachment adds a related measurement caution but leaves this audit unverified.
2026-09-11T12:22:45Z
evidence attached: hn.story.49656471 — Independent disagreement over reported coding-agent token savings supports the open hypothesis that inference-cost measurements and optimization claims are unreliable.
2026-09-10T07:39:08Z
No substantive evidence has arrived: the GitHub echo repeats the submitter’s headline rather than independently substantiating the audit, so the initial first-party characterization was too strong. Accounting errors could affect Scott’s budget-based routing, but neither reproducible bugs nor affected tools have been identified in the supplied evidence.
2026-09-10T07:36:01Z
grounded: converges/medium — The reported audit converges with Scott’s Agent Observability and bounded-budget Token Economics positions, with a concrete verification target: his cost-tiered
2026-09-10T07:33:42Z
case created — First-party audit report with a concrete bug count relevant to inference-cost measurement reliability.