2026-10-11 17:10 UTC

prompt-caching

band: hotmomentum: stable score: 0.671
temperature history

Episodes (11)

Cache-hunter will gain adoption among LLM harness builders as a reproducible way to detect prompt changes that invalidate prefill caches.
resolvedconvergesscott: medium
Cross-provider testing will determine whether changes to tool schemas routinely invalidate prompt caches and materially raise the cost and latency of tool-using agent workloads.
expiredconvergesscott: high
Cache Analyzer’s creator claims analyzing Claude Code sessions can show when five-minute versus one-hour prompt-cache retention offers a better cost and responsiveness tradeoff for coding-agent workloads.
expiredknownscott: low
Replay maintainer Daniel Saito claims its released local CLI identifies the turns, causes, and token costs of prompt-cache breaks in agent transcripts, exposing 42.9 million re-billed tokens in his 119-session corpus and enabling turn-level inference-cost diagnosis.
seedconvergesscott: medium
karanb192 claims the released cache-tax tool prevents idle Claude Code context from expiring through scheduled warming requests, potentially lowering resumed-session costs when avoided cache writes outweigh warming charges.
corroboratedconvergesscott: medium
ghuntley claims Underclass's released local proxy pools ChatGPT/Codex and GitHub Copilot subscriptions behind an OpenAI-compatible endpoint with persistent session affinity, automatic quota cooldowns, and fail-fast saturation, reducing account-management overhead and prompt-cache disruption for agent workloads.
watchingconvergesscott: medium
OpenAI claims its GPT-6 caching update preserves eligible prefixes for 30 minutes and adds explicit breakpoints, diagnostics, and cache-preserving reasoning changes, reducing latency and input costs for persistent agents.
corroboratedconvergesscott: high
OpenAI claims its released GPT-6 Sol and Luna improve coding and professional-agent performance while cutting API prices roughly in half versus GPT-5.6 promotional rates, materially lowering sustained agent-work costs.
resolvedconvergesscott: medium
Anthropic's Opus 5.5 repricing cut cache-read rates 60% (to $0.20/M) versus 20% for input/output tokens, materially changing the economics of cache-heavy long-context agent workloads.
resolvedconvergesscott: high
Asana reports a 76x cost reduction (from $36.21 to $0.47 per run) and 5.6x speedup for a browser-agent workflow by stabilizing page history for prompt caching and batch-pruning screenshots, establishing a referenced cost-control pattern for long-running agent workflows.
seedconvergesscott: high
A builder releases `/dehistorize`, a reusable agent skill that strips edit-history leakage from model outputs β€” preventing models from oversharing deleted content or anchoring to prior versions β€” as a practical harness-level mitigation for history-contamination in agent workflows.
seedconvergesscott: high

Trajectory notes