2026-10-11 17:10 UTC

Independent benchmarks will determine whether CachyLlama’s SSD-backed multi-tier persistent KV cache materially reduces repeated prompt-processing latency in long local-agent sessions on slower hardware without unacceptable storage or correctness tradeoffs.

state: expiredheat: lowuncertainty: highnovelscott: nonekv-cache local-inference coding-agents llama-cppCachyLlama

What is this?

CachyLlama is presented in the case as a llama.cpp fork whose maintainer introduced an SSD-backed persistent KV cache with hot, warm, and cold tiers, intended to reduce repeated prompt processing during long local-agent sessions. The supplied web results are unrelated to the project and provide no independent evidence about its implementation, benchmarks, storage costs, correctness, or performance on slower hardware. Consequently, the claimed benefit remains an unverified hypothesis pending relevant primary artifacts and reproducible third-party testing.

Why it matters to Scott

No intersection found in Scott’s wikis, and no radar pages already track CachyLlama or this development. The case remains an unverified local-inference performance hypothesis, so the supplied material does not establish a Scott-specific consequence or actionable connection.
queries asked of Scott's wikis
  • persistent KV cache for long-running agents
  • tiered agent memory hot warm cold storage
  • local inference latency on constrained hardware
  • llama.cpp forks and local coding agents
  • KV-cache persistence correctness and invalidation
  • SSD-backed inference cache economics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditCachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful
LocalLLaMA
UsualResult5626
🟧 echo.github ⭐The earliest primary artifact is the maintainer’s commit introducing the feature: “squash: SSD-backed KV cache with hot/warm/cold tiering.” fewtarius——
🟠 redditCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows
LocalLLaMA
UsedMorning9886134

Interpretation history

Decision trace