2026-10-11 18:01 UTC

Independent evaluations will determine whether Memory Bench reliably identifies when dedicated agent-memory layers outperform full-chat-history baselines.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-memory memory-benchmark context-managementLeapMemory

What is this?

The case describes Memory Bench, reportedly an open benchmark from LeapMemory comparing dedicated agent-memory layers against simply supplying full chat history for agent tasks. The supplied snippets do not directly document Memory Bench or LeapMemory, so its design, release status, and results remain unverified here. They do establish a broader field of open agent-memory evaluations measuring capabilities such as remembering, reasoning, recommendation, retention, poisoning resistance, task success, and user experience—and report known failure modes such as reusing invalid memories or failing to reconcile updates.

Why it matters to Scott

Memory Bench independently operationalizes Scott’s core comparison between ever-growing chat history and selective external memory/context compaction, potentially providing a falsifiable test of his Long-Running Agents and Context Engineering claims. No results or independent validation are supplied yet, so this is a meaningful evaluation opportunity rather than confirmation of those claims.
ip:framework.long-running-agentsip:framework.context-engineeringip:concept.evaluation-driven-developmentdev:concept.agent-authored-context-compactionradar:concept.agent-memoryradar:concept.ai-benchmarksradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • dedicated agent memory vs full context history
  • agent memory benchmark design and evaluation
  • context management economics for long-running agents
  • memory reliability stale facts and contradiction handling
  • agent-maintained wikis as external memory
  • memory-layer poisoning and trust boundaries

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Open benchmark for memory layers vs. pasting full chat historycanu2110
🟧 echo.github ⭐An open benchmark comparing dedicated memory layers with pasting full chat history for agent tasks.LeapMemory——

Interpretation history

Decision trace