2026-10-11 17:10 UTC

Independent audits and repeated evaluations will determine whether the Agent Memory Leaderboard produces reproducible, decision-useful comparisons across open-source and commercial agent-memory systems.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-memory memory-evaluation agent-harnessesAgent Memory Leaderboard

What is this?

The Agent Memory Leaderboard is a public evaluation platform intended to compare agent-memory systems under shared conditions, with participants exposing Add/Search APIs and being scored on textual-memory and coding-agent-memory tracks. Its stated motivation is that vendors currently report results using different datasets, answer models, and judges, making scores difficult to compare. The supplied snippets include secondary claims that Hindsight leads several benchmarks, but they do not establish independent audits, repeated reproductions, or enough methodology and primary-result detail to conclude that the leaderboard’s comparisons are yet decision-useful.

Why it matters to Scott

The leaderboard independently moves toward Scott’s vendor-neutral, model-plus-harness evaluation position and directly overlaps his trace-backed agent-comparison work. It creates a practical comparison and publishing opportunity, but the supplied evidence does not yet show the audits, replayable traces, controlled harness disclosure, or repeated evaluations needed to establish decision-useful results.
ip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditdev:concept.trace-backed-agent-comparisondev:project.thinkerradar:memory-bench-layer-baseline-validityradar:concept.agent-memoryradar:concept.agent-evaluationradar:concept.agent-benchmarks
queries asked of Scott's wikis
  • agent memory evaluation methodology and reproducibility
  • memory benchmarks for coding agents and harnesses
  • file vs vector vs graph memory architectures
  • agent-maintained wiki memory evaluation
  • retrieval quality versus downstream agent performance
  • vendor-neutral benchmarks and evaluation harnesses

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Agent Memory Leaderboard – first public results for AI memory systemsIreneAI20
🟧 echo.other ⭐This is the primary results publication, not a report quoting an earlier announcement. The site publishes the academic textual leaderboard aAgent Memory Leaderboard Organizers——
🟧 hnShow HN: I evaluated file, vector, graph and RL based memory frameworkspinglin100
🟠 redditA scale for your AI's notes. Notes that pay their rent stay. Notes that don't get thrown out. And it refuses to claim a saving it can't prove.
ClaudeAI
tvuk12
🟧 hnAI Memory for Claude: An Honest 4-Way ComparisonlabyrinthAC10

Interpretation history

Decision trace