2026-10-11 18:03 UTC

Cross-provider testing will determine whether changes to tool schemas routinely invalidate prompt caches and materially raise the cost and latency of tool-using agent workloads.

state: expiredheat: lowuncertainty: highconvergesscott: highprompt-caching llm-apis inference-economics agent-harnessesOpenAI

What is this?

Chirag reports a replicated API experiment in which removing one tool from the available schema invalidated the prompt cache for OpenAI’s GPT-5.5, while GPT-5.2 retained its cache. The reported result suggests that tool-schema changes may affect the cost and latency of tool-using agents differently across model versions, motivating broader cross-provider testing. The supplied web snippets discuss prompt caching, agent latency, and regression testing generally, but do not independently document Chirag’s experiment or establish that cache invalidation is routine across providers.

Why it matters to Scott

The experiment independently substantiates Scott’s prefix-caching economics and evaluation-driven approach while adding a consequential caveat: cache behavior can change between model versions when tool schemas change. This directly affects his LiteLLM-routed `ask` harness and multi-provider routing decisions, warranting version-bound cache regression fixtures; the broader claim that this routinely occurs across providers remains unproven.
ip:concept.prefix-caching-economicsip:concept.evaluation-driven-developmentdev:project.askdev:concept.task-aware-model-routingdev:concept.trace-backed-agent-comparisonradar:cache-hunter-prompt-cache-debuggingradar:concept.inference-economicsradar:concept.llm-apisradar:concept.agent-harnesses
queries asked of Scott's wikis
  • prompt-cache invalidation from tool schema changes
  • agent harness tool-schema stability and versioning
  • cross-provider LLM API regression testing
  • cached-token economics for tool-using agents
  • tool registry changes and agent latency benchmarks
  • provider-specific prompt caching semantics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDropping one tool invalidated the prompt cache on GPT-5.5 but not on 5.2srcecde20
🟧 echo.blog ⭐This is the primary source: Chirag reports his own replicated API experiment, stating, “This post is the experiment,” with GPT-5.2 retainingChirag Rathod (Srce Cde)——

Interpretation history

Decision trace