2026-10-11 16:38 UTC

Spomin creator wgaca2 claims the released router and llama.cpp fork replace context with summaries directly in the live KV cache for experimental Qwen sessions, potentially sustaining long-running local agents without repeatedly reprocessing retained context.

state: seedheat: lowuncertainty: highconvergesscott: mediumkv-cache context-compaction local-inference agent-harnesseswgaca2alekk89

What is this?

The supplied case describes Spomin as an experimental Qwen router and companion llama.cpp fork whose creator, identified as wgaca2, claims it replaces context with summaries directly in the live KV cache. None of the returned web snippets identifies Spomin or verifies its repositories, implementation, or performance; alekk89’s role is also unestablished. Related snippets discuss costly prompt reprocessing, context-shifting limitations, and recurrent-state checkpoint constraints in llama.cpp, establishing a relevant engineering problem but not demonstrating that Spomin solves it or sustains long-running agents.

Why it matters to Scott

Spomin’s claimed live-cache summary replacement extends Scott’s agent-authored context compaction with a potentially testable way to avoid retained-context reprocessing, relevant to Ask’s lossy --compact path and his Prefix-Caching Economics claim that append-only transcript retention is cost-optimal. This remains an unverified creator claim—not a demonstrated economic challenge or substitute for his durable external state—and the supplied radar pages track related cache-runtime approaches, not Spomin itself.
dev:concept.agent-authored-context-compactiondev:project.askip:concept.prefix-caching-economicsip:framework.long-running-agentsradar:yandex-kv-cache-agent-runtimeradar:cachyllama-persistent-kv-cacheradar:anthropic-context-compaction-cost-reversalradar:concept.kv-cacheradar:concept.context-compression
queries asked of Scott's wikis
  • Agent harness context compaction and summary fidelity
  • Local inference KV cache reuse and prefill latency
  • Long-running agents bounded context and persistent memory
  • Qwen llama.cpp coding-agent projects
  • Recurrent model state checkpoints and context editing

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 727h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-11 09:23 (minted)⭐ origin echo-reconstructedThe creator links this Spomin router repository and a companion llama.cpp fork as the implementation that replaces context with summaries di
alekk89 on github (echo) · attributed from reddit.post.1wdaq1v · published time unknown
—
09-11 08:55first on r/LocalLLaMA · published · lag ?Spomin - Live KV cache compaction (Experimental for Qwen)
wgaca2
—
09-15 16:20first on hacker news · published · lag ?Ask HN: Why aren't LLM token CDN cache's a thing?
devrob
—
09-11 08:55amplified on r/LocalLLaMA 👑reddit.post.1wdaq1v
wgaca2
peak 24 · 2 comments · 88% of case engagement
09-15 16:20amplified on hacker newshn.story.49714867
devrob
peak 1 · 1 comments · 12% of case engagement
09-11 09:20our radar first saw it · lag ?discovery anchor: reddit.post.1wdaq1v—
pace: p55 vs 519 stories at the 720h mark (now 727h old) — ahead of jasper-from-scratch-t2i-kit (1.1x), behind grith-syscall-agent-supervision (0.9x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditSpomin - Live KV cache compaction (Experimental for Qwen)
LocalLLaMA
wgaca2242
🟧 echo.github ⭐The creator links this Spomin router repository and a companion llama.cpp fork as the implementation that replaces context with summaries dialekk89——
🟧 hnAsk HN: Why aren't LLM token CDN cache's a thing?devrob11

Interpretation history

Decision trace