Cross-provider testing will determine whether changes to tool schemas routinely invalidate prompt caches and materially raise the cost and latency of tool-using agent workloads.
state: expiredheat: lowuncertainty: highconvergesscott: highprompt-caching llm-apis inference-economics agent-harnessesOpenAI
What is this?
Chirag reports a replicated API experiment in which removing one tool from the available schema invalidated the prompt cache for OpenAI’s GPT-5.5, while GPT-5.2 retained its cache. The reported result suggests that tool-schema changes may affect the cost and latency of tool-using agents differently across model versions, motivating broader cross-provider testing. The supplied web snippets discuss prompt caching, agent latency, and regression testing generally, but do not independently document Chirag’s experiment or establish that cache invalidation is routine across providers.
Why it matters to Scott
The experiment independently substantiates Scott’s prefix-caching economics and evaluation-driven approach while adding a consequential caveat: cache behavior can change between model versions when tool schemas change. This directly affects his LiteLLM-routed `ask` harness and multi-provider routing decisions, warranting version-bound cache regression fixtures; the broader claim that this routinely occurs across providers remains unproven.
ip:concept.prefix-caching-economicsip:concept.evaluation-driven-developmentdev:project.askdev:concept.task-aware-model-routingdev:concept.trace-backed-agent-comparisonradar:cache-hunter-prompt-cache-debuggingradar:concept.inference-economicsradar:concept.llm-apisradar:concept.agent-harnesses
queries asked of Scott's wikis
- prompt-cache invalidation from tool schema changes
- agent harness tool-schema stability and versioning
- cross-provider LLM API regression testing
- cached-token economics for tool-using agents
- tool registry changes and agent latency benchmarks
- provider-specific prompt caching semantics
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-14T15:48:14Z
No independent replication or cross-provider testing emerged within the active horizon, so the finding remains a useful one-off regression fixture rather than evidence of a routine industry pattern.
2026-08-12T14:57:24Z
No new evidence extends the single model-version experiment into cross-provider corroboration; this remains an actionable regression-test lead rather than an established pattern. The unchanged reobservation adds no alert-worthy delta.
2026-08-12T14:49:36Z
grounded: converges/high — The experiment independently substantiates Scott’s prefix-caching economics and evaluation-driven approach while adding a consequential caveat: cache behavior c
2026-08-12T14:47:21Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49271916 -> echo.blog.09b1256b62 by Chirag Rathod (Srce Cde)
2026-08-12T14:46:08Z
case created — The reported version-dependent cache behavior is directly actionable for agent-harness design and inference-cost control.
Decision trace
- 08-15 01:48expireNo independent replication or cross-provider testing emerged within the active horizon, so the finding remains a useful one-off regression fixture rather than evidence of a routine industry pattern.
- 08-15 01:48alert_silentThis reobservation adds no evidence beyond the already-routed GPT-5.2 versus GPT-5.5 experiment; any future provider replication can reopen the case as a new material delta.
- 08-15 01:48alert_routeThis reobservation adds no evidence beyond the already-routed GPT-5.2 versus GPT-5.5 experiment; any future provider replication can reopen the case as a new material delta.
- 08-13 00:57repriceNo new evidence extends the single model-version experiment into cross-provider corroboration; this remains an actionable regression-test lead rather than an established pattern. The unchanged reobser
- 08-13 00:57alert_silentThe implementation-relevant GPT-5.2 versus GPT-5.5 result was already routed, and this look adds neither replication nor provider coverage.
- 08-13 00:57alert_routeThe implementation-relevant GPT-5.2 versus GPT-5.5 result was already routed, and this look adds neither replication nor provider coverage.
- 08-13 00:51alert_shadowThe published rig and raw data establish an implementation-relevant model-version regression: dropping one tool left 2,560 cached tokens on GPT-5.2 but zero on GPT-5.5. Scott is likely to value adding
- 08-13 00:51alert_routeThe published rig and raw data establish an implementation-relevant model-version regression: dropping one tool left 2,560 cached tokens on GPT-5.2 but zero on GPT-5.5. Scott is likely to value adding
- 08-13 00:49groundThe experiment independently substantiates Scott’s prefix-caching economics and evaluation-driven approach while adding a consequential caveat: cache behavior can change between model versions when to
- 08-13 00:47promote_anchororigin walk conf 0.99
- 08-13 00:46createThe reported version-dependent cache behavior is directly actionable for agent-harness design and inference-cost control.