2026-10-11 16:37 UTC

The authors of arXiv:2610.10845 claim a practical long-term memory mechanism with a 50M token window for LLMs โ€” if validated, it advances agent-memory architectures beyond current context limits.

state: seedheat: lowuncertainty: highconvergesscott: highagent-memory long-context research-artifact

What is this?

arXiv:2610.10845 introduces 'galahad-kv', an open-source KV-cache persistence layer that stores transformer key-value states in 16K-token blocks on encrypted local NVMe, enabling 50-million-token effective context windows by loading cached KV states byte-exact instead of recomputing. The mechanism targets long-horizon agent workloads where repeated context recomputation is the bottleneck. The paper was posted to arXiv on 2026-10-07 and appeared on Hacker News the same week; independent replication or benchmark results beyond the authors' claims are not yet visible in the supplied snippets.

Why it matters to Scott

The galahad-kv claim (50M-token effective context via byte-exact KV-cache persistence on encrypted NVMe) independently arrives at a mechanism Scott's frameworks already argue for: durable external state that survives context degradation without rereading full transcripts (long-running-agents), and a memory hierarchy where cold storage is paged just-in-time rather than preloaded (context-engineering). It bears directly on his attention-budget thesis โ€” the paper treats capacity as the constraint; Scott treats attention quality as the constraint โ€” so validation would force a reevaluation of whether raw KV persistence complements or competes with selective context loading. It also moves the local-inference economics frontier: NVMe-backed KV offload vs. recompute at 50M-token scale is a concrete tradeoff his hardware-aware local inference work and token-economics doctrine price explicitly.
ip:framework.context-engineeringip:framework.long-running-agentsip:source.breaking-the-1hr-barrierip:concept.durable-external-stateip:concept.attention-budgetdev:concept.hardware-aware-local-inferenceradar:adaptive-kv-cache-streamingradar:agent-memory-leaderboard-validationradar:awareness-agent-memory-longmemevalradar:beellama-kvarn-kv-cache-validationradar:afm3-prompt-conditioned-pruning
queries asked of Scott's wikis
  • agent-memory architecture patterns: KV-cache persistence vs. retrieval-augmented approaches
  • local inference economics: NVMe-backed KV offload vs. recompute tradeoffs at 50M token scale
  • open-weights tooling: galahad-kv integration path for local model runtimes (llama.cpp, vLLM, MLC)
  • model sovereignty: encrypted local KV storage as a privacy/control primitive for agent memory
  • long-context evaluation: whether 50M-token windows change agent benchmark design (LoCoMo, LongMemEval)

Measured heat

now 0 pts/hpeak 3 pts/hcomments 0/hpeers p37momentum: steady2 platformsage 54h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-09 09:35โญ origin directly observedShow HN: Long term Memory and 50M token window for LLM
Corbenic on hacker news
โ€”
10-10 11:44first on r/LocalLLaMA ยท published ยท +26.1hReal long-term memory for local LLMs: saved KV state on NVMe, reloaded byte-exact instead of recompute (I built it, free on 1 GPU)
MindPsychological140
โ€”
10-10 21:36first on hacker news ยท published ยท +36.0hShow HN: Try Free Long-Term Memory for AI 50M-Token Window
Corbenic
โ€”
10-09 09:35amplified on hacker news ๐Ÿ‘‘hn.story.50018178
Corbenic
peak 6 ยท 1 comments ยท 54% of case engagement
10-10 11:44amplified on r/LocalLLaMAreddit.post.1x2d9wb
MindPsychological140
peak 0 ยท 5 comments ยท 21% of case engagement
10-10 21:36amplified on hacker newshn.story.50037428
Corbenic
peak 3 ยท 0 comments ยท 24% of case engagement
10-09 13:35our radar first saw it ยท +4.0hdiscovery anchor: hn.story.50018178โ€”
pace: p55 vs 1204 stories at the 48h mark (now 54h old) โ€” ahead of agentgit-accountless-agent-handoffs (1.1x), behind anthropic-usage-policy-election-interference-ban (0.9x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hn โญShow HN: Long term Memory and 50M token window for LLMCorbenic61
๐ŸŸ  redditReal long-term memory for local LLMs: saved KV state on NVMe, reloaded byte-exact instead of recompute (I built it, free on 1 GPU)
LocalLLaMA
MindPsychological14005
๐ŸŸง hnShow HN: Try Free Long-Term Memory for AI 50M-Token WindowCorbenic30

Interpretation history

Decision trace