Reddit user Bitter-Truck1049's controlled six-run /usage measurement claims headless Claude Code (Agent SDK and `claude -p`) consumes roughly 3x more of the 5-hour limit per dollar of API-equivalent tokens than interactive use, and Anthropic documenting, confirming, or correcting that undocumented differential decides who actually bears the usage-limit cut for headless workloads.
state: resolvedheat: lowuncertainty: mediumnovelscott: highagent-harnesses inference-economics claude-codeAnthropic
What is this?
A Reddit user (Bitter-Truck1049, Sep 28, 2026) posted a controlled six-run measurement claiming headless Claude Code entrypoints โ the Agent SDK and `claude -p` โ burn ~3x more of Claude's 5-hour subscription window per dollar of API-equivalent tokens than interactive use (~5.2% vs ~1.7% per $1), with in-thread corroboration citing transcript entrypoint fields (`sdk-py`/`sdk-cli` vs `cli`/`claude-vscode`). A second independent thread documents subagents defaulting to a 5-minute prompt-cache TTL vs 1 hour for the main session โ corroborating a family of undocumented limit-drain mechanics while also offering a mundane mechanism (cold-cache misses inflating real token cost) for part or all of the 3x; no supplied source confirms or corrects the entrypoint-metering claim itself, and Anthropic has said nothing about it. The instrumentation opacity is documented: `/usage` is TUI-only, and the SDK's `rate_limit_event` messages omit utilization fields in normal operation (anthropics/claude-code#50518). The policy backdrop is conflicting in the supplied material: most sources describe the announced June 15, 2026 split of programmatic usage onto a separate API-priced credit as in effect, one guide (checked June 26) reports it paused โ and the Sep 28 measurement's premise (headless burn still landing on the 5-hour window) implies the subscription path was still live for that user.
Why it matters to Scott
Directly actionable on Scott's live headless scheduling stack: the router project's hourly/weekly Claude Code agents and the cron-ebook's cache-window-tuned cadence are exactly the `claude -p` pattern the claimed ~3x burn taxes, and if the 5-min subagent cache TTL is the mechanism it hands fresh economic support to prompt-interrupt's warm-kernel design over cold fresh runs (while conditioning the headless applicability of his prefix-caching-economics near-linearity claim). The contested entrypoint-vs-cache attribution is also precisely the controlled, same-fixture comparison his Langfuse / trace-backed-agent-comparison tooling was built to run โ a first-to-settle measurement opportunity, not just a topic match.
dev:project.routerip:source.ask-yourself-if-you-re-finished-cron-as-the-poor-man-s-orchestrator-ebookip:framework.prompt-interrupt-architectureip:concept.prefix-caching-economicsdev:technology.claude-codedev:technology.langfusedev:concept.trace-backed-agent-comparisonradar:claude-code-usage-limit-cutradar:claude-code-cache-ttl-analyzerradar:cache-tax-idle-session-warmingradar:replay-prompt-cache-miss-auditradar:frontierharness-17x-cost-variationradar:codepress-subscription-cloud-agentsradar:hermes-claude-directsdkradar:underclass-sticky-subscription-poolradar:claude-phantom-token-billing-bug
queries asked of Scott's wikis
- claude -p headless cron scheduled orchestration runs
- subscription vs API arbitrage pricing case
- prompt cache TTL cache-miss token cost inflation
- Langfuse trace token attribution comparison tooling
- metering opacity undocumented rate-limit changes platform trust
- agent harness token overhead context plumbing
Measured heat
now 4 pts/hpeak 18 pts/hcomments 1/hpeers p91momentum: steady1 platformsage 221h
points/hour across evidence ยท reading as of 2026-10-08 08:06:52.397066+11:00 ยท deterministic, not a model opinion
How the heat travelled
Evidence (3) โ โญ canonical anchor
Interpretation history
2026-10-07T21:30:01Z
Anthropic's Oct 7 Help Center update dissolves the case's premise rather than answering it: headless Agent SDK usage is moved off subscription limits onto monthly API credits, so the claimed 3x differential no longer decides who bears the headless cut โ headless no longer draws on the 5-hour window at all. The measurement's factual questions (is the 3x real; entrypoint metering vs subagent cache-TTL mechanism) were never settled and now turn historical; resolved as superseded by the pricing-structure change.
2026-10-07T20:59:03Z
evidence attached: reddit.post.1x04j1x โ The October 7 coverage change โ headless Agent SDK/claude -p moved off subscription limits onto monthly API credits โ directly bears on the metering case's question of who actually bears the headless usage-limit cut.
2026-10-07T01:51:59Z
grounded: novel/high โ Directly actionable on Scott's live headless scheduling stack: the router project's hourly/weekly Claude Code agents and the cron-ebook's cache-window-tuned cad
2026-10-07T01:43:35Z
The second thread (different author) independently documents subagents defaulting to a 5-min cache TTL vs the main session's 1h โ corroborating the family of undocumented limit-drain mechanics in headless/subagent workloads, but also supplying a mundane candidate mechanism (cold-cache misses) for the 3x figure, so the phenomenon earns corroborated while its attribution as deliberate entrypoint metering becomes contested. Attention cools: the episode peaked on one platform days ago (~0.5 pts/h now, 63rd percentile), the periphery is not expanding, and Anthropic remains silent.
2026-10-06T23:36:37Z
evidence attached: reddit.post.1wze3rr โ Subagent 5-min cache default draining limits is the same family of undocumented Claude Code consumption mechanics as the headless metering differential and materially contextualises it.
2026-09-28T18:50:47Z
grounded: novel/high โ No Scott wiki page carries the headless-metering claim itself, so this is new information rather than a held or echoed position โ but it reprices load-bearing p
2026-09-28T18:39:53Z
case created โ A reproducible, specific metering claim with real spread that directly reweights who absorbs the significant usage-limit-cut episode, and it has no existing case.
Decision trace
- 10-08 08:30resolveAnthropic's Oct 7 Help Center update dissolves the case's premise rather than answering it: headless Agent SDK usage is moved off subscription limits onto monthly API credits, so the claimed
- 10-08 08:02attention_routeQuiet hours with the briefing two hours out and the credit rollout only beginning 'this week' โ no scheduled job breaks before 10am, so there is no concrete cost of waiting and no justificat
- 10-08 07:59attention_candidateattach
- 10-08 07:59attachThe October 7 coverage change โ headless Agent SDK/claude -p moved off subscription limits onto monthly API credits โ directly bears on the metering case's question of who actually bears the head
- 10-08 00:59attention_routeFirst look at a novel economic fact about the stack he runs daily, with a same-day mitigation available. No overnight interrupt is justified โ nothing degrades by morning โ so it goes into the 10am br
- 10-07 12:51repriceThe second thread (different author) independently documents subagents defaulting to a 5-min cache TTL vs the main session's 1h โ corroborating the family of undocumented limit-drain mechanics in
- 10-07 12:51groundDirectly actionable on Scott's live headless scheduling stack: the router project's hourly/weekly Claude Code agents and the cron-ebook's cache-window-tuned cadence are exactly the `cla
- 10-07 10:36attachSubagent 5-min cache default draining limits is the same family of undocumented Claude Code consumption mechanics as the headless metering differential and materially contextualises it.
- 10-07 10:32propose_attachSubagent 5-min cache default draining limits is the same family of undocumented Claude Code consumption mechanics as the headless metering differential and materially contextualises it.
- 09-29 04:50groundNo Scott wiki page carries the headless-metering claim itself, so this is new information rather than a held or echoed position โ but it reprices load-bearing practice: his router project runs hourly/
- 09-29 04:39createA reproducible, specific metering claim with real spread that directly reweights who absorbs the significant usage-limit-cut episode, and it has no existing case.