The Agent Memory Leaderboard is a public evaluation platform intended to compare agent-memory systems under shared conditions, with participants exposing Add/Search APIs and being scored on textual-memory and coding-agent-memory tracks. Its stated motivation is that vendors currently report results using different datasets, answer models, and judges, making scores difficult to compare. The supplied snippets include secondary claims that Hindsight leads several benchmarks, but they do not establish independent audits, repeated reproductions, or enough methodology and primary-result detail to conclude that the leaderboard’s comparisons are yet decision-useful.
The leaderboard independently moves toward Scott’s vendor-neutral, model-plus-harness evaluation position and directly overlaps his trace-backed agent-comparison work. It creates a practical comparison and publishing opportunity, but the supplied evidence does not yet show the audits, replayable traces, controlled harness disclosure, or repeated evaluations needed to establish decision-useful results.
ip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditdev:concept.trace-backed-agent-comparisondev:project.thinkerradar:memory-bench-layer-baseline-validityradar:concept.agent-memoryradar:concept.agent-evaluationradar:concept.agent-benchmarks
queries asked of Scott's wikis
- agent memory evaluation methodology and reproducibility
- memory benchmarks for coding agents and harnesses
- file vs vector vs graph memory architectures
- agent-maintained wiki memory evaluation
- retrieval quality versus downstream agent performance
- vendor-neutral benchmarks and evaluation harnesses
2026-08-22T17:27:58Z
Repeated reviews have produced only adjacent evaluation examples and minor amplification, with no audit, replayable traces, or versioned rerun testing the inaugural rankings. The episode has faded; a direct reproduction or new leaderboard run should open a fresh case.
2026-08-20T16:41:15Z
Minor engagement on an adjacent comparison adds no methodology, replay materials, audit, or repeated leaderboard run. The case remains a cold, prospective benchmark-validation question despite the broader topic’s heat.
2026-08-18T15:46:09Z
The four-way Claude memory comparison remains a title-only adjacent evaluation with no disclosed methods, results, traces, or replay materials. It neither reproduces nor audits the Agent Memory Leaderboard, so the case remains an unvalidated first run awaiting direct independent testing.
2026-08-18T14:23:50Z
evidence attached: hn.story.49345976 — This hunted comparison is independent evaluation evidence bearing directly on whether Claude memory approaches can be compared in a decision-useful way.
2026-08-17T17:41:03Z
No audit, reproduction, disclosed trace set, or repeated versioned run has appeared; the attached items remain adjacent evaluation examples rather than validation of the leaderboard. The case stays open but cold until evidence directly tests the inaugural rankings.
2026-08-15T16:39:27Z
Token Warden adds a concrete A/B harness pattern for testing whether persistent memory rules justify their context cost, but it neither audits nor reproduces the Agent Memory Leaderboard. The case remains an unvalidated first leaderboard run awaiting independent replay, disclosed traces, or a repeated versioned evaluation.
2026-08-15T16:23:04Z
evidence attached: reddit.post.1vp6178 — The released token-warden plugin introduces a concrete task-based method for measuring whether persistent coding-agent memory rules repay their context cost.
2026-08-15T15:36:55Z
grounded: converges/medium — The leaderboard independently moves toward Scott’s vendor-neutral, model-plus-harness evaluation position and directly overlaps his trace-backed agent-compariso
2026-08-15T15:34:13Z
The newly attached comparison is only a bare title and does not disclose methods, versions, results, or any audit of the leaderboard, so it does not yet provide an independent line of validation. The case remains about whether the released inaugural rankings can withstand reproduction and repeated versioned runs.
2026-08-15T15:23:04Z
evidence attached: hn.story.49310834 — A comparative evaluation of file, vector, graph, and RL memory frameworks materially informs whether agent-memory comparisons are decision-useful.
2026-08-15T04:25:04Z
Only minor engagement arrived; there is still no independent audit, reproduction, or second versioned run to validate the leaderboard’s rankings. The case remains a prospective benchmark-integrity question rather than established comparative infrastructure.
2026-08-13T03:35:45Z
grounded: known/medium — Scott already holds the core position in “Model-Plus-Harness Benchmark Unit” and “Trace-backed agent comparison”: comparisons must bind results to a disclosed,
2026-08-13T03:33:16Z
No independent audit, repeated run, or implementation evidence has arrived; this is unchanged first-party publication rather than validation of reproducibility or decision usefulness. The cached grounding is now misleading because the inaugural results have been released.
2026-08-13T03:31:51Z
grounded: known/low — Scott already holds the core position in Model-Plus-Harness Benchmark Unit and Trace-backed agent comparison, while the radar tracks substantially the same vali
2026-08-13T03:29:22Z
origin walked (codex/luna, conf 0.9): anchor hn.story.49281370 -> echo.other.a705d6dc35 by Agent Memory Leaderboard Organizers
2026-08-13T03:27:27Z
case created — The first public evaluation covering 69 memory frameworks is a substantive artifact that could become useful comparative infrastructure if its methodology holds up.