2026-10-11 18:00 UTC

Cache-hunter will gain adoption among LLM harness builders as a reproducible way to detect prompt changes that invalidate prefill caches.

state: resolvedheat: lowuncertainty: mediumconvergesscott: mediumllm-harness prompt-caching cache-invalidation developer-tooling

What is this?

Cache-hunter is presented in a Reddit post as a simple tool for harness builders to detect when prompt changes invalidate LLM prefill caches, a problem its unnamed author encountered while building a local-first harness. The surrounding results establish that cache invalidation can affect inference cost and latency, and that related harness-level prompt-caching tools exist. However, the supplied snippets do not identify Cache-hunter’s creator, explain its implementation, or provide evidence that it has gained—or will gain—adoption.

Why it matters to Scott

Cache-hunter operationalizes Scott’s prefix-caching economics and evaluation-driven development positions by treating prompt-induced cache invalidation as a reproducible harness regression. It could directly inform testing of the active `ask` harness, whose append-only tool loop and context compaction can alter prompt prefixes, although the supplied evidence does not establish the tool’s implementation quality or adoption.
ip:concept.prefix-caching-economicsip:concept.evaluation-driven-developmentdev:project.ask
queries asked of Scott's wikis
  • prompt-prefix stability in agent harnesses
  • prefill cache observability and regression testing
  • cache-aware context compaction strategies
  • local inference economics and KV-cache reuse
  • prompt mutation and cache invalidation
  • LLM harness cost and latency instrumentation

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐If you're building a harness, here is a simple tool to catch cache invalidation in your calls to LLMs
LocalLLaMA
t4a894514147
🟧 hnAgentic Workflow's Cache Keepalive Costs 8x Too Muchmempko21
🟠 redditI burned 246M tokens in 22 hours on Claude Code and measured exactly where every one went. The answer surprised me.
ClaudeAI
IndividualEngine8579024
🟠 redditI Built a Proxy to See What Claude Code Is Really Doing (243 Sessions Later, Here's What I Found)
ClaudeAI
Muttawakkil117
🟠 redditA prompt-cache benchmark for Anthropic-compatible /v1/messages endpoints that report cache usage
LocalLLaMA
Appropriate_Heat_86603
🟧 hnI tested whether verifying LLM cache hits in real time helps (weak yes)ChengyouXin10
🟧 hnShow HN: A plugin to monitor OpenCode cache hitsnmdra10

Interpretation history

Decision trace