The case attributes to IreneAI a Show HN announcement of an open Add/Search framework for comparing agent-memory systems while keeping answer generation and evaluation under framework control. A Zylos Research snippet describes a matching protocol called AML: participants implement history ingestion (Add) and memory retrieval (Search), while AML controls generation, evaluation, aggregation, and orchestration. That snippet says comparability requires matching versioned benchmarks, pipelines, model configurations, and scoring rules; it supports the proposed mechanism, not proof that memory quality is fully isolated. The supplied web results do not directly establish IreneAI’s authorship or verify the announced release.
The described Add/Search protocol converges with Scott’s Trace-backed Agent Comparison and Reflexive Agent Design: controlled comparisons rather than conflating backend quality with evaluation setup, offering a concrete candidate protocol for testing his RAG/Wiki Substrate Rule. This could inform his memory-backend evaluations, but IreneAI’s authorship and release remain unverified, and fixed generation settings do not establish model-independent memory quality; the radar’s Agent Memory Leaderboard validation page tracks a related concern, not demonstrably this release.
dev:concept.trace-backed-agent-comparisonip:framework.reflexive-agent-designip:framework.rag-wiki-substrate-ruleradar:agent-memory-leaderboard-validationradar:ship-harness-benchradar:concept.memory-evaluation
queries asked of Scott's wikis
- agent memory retrieval evaluation harness
- benchmark confounds model versus system performance
- reproducible evaluation versioned scoring contracts
- agent-maintained wiki memory quality testing
- memory backend comparison fixed answer model
now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 565h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
2026-10-05T21:43:56Z
First non-IreneAI evidence enters the case: independent academic work on evaluating memory structures in LLM agents. It does not corroborate the Add/Search isolation claim — release, adoption and reproducibility remain unverified — but it reframes the case from one team's self-promotional loop to a contested evaluation space, modestly lowering the odds Add/Search becomes the reference protocol while independently confirming the problem's importance to Scott's own comparison work.
2026-10-05T20:39:23Z
evidence attached: hn.story.49968612 — Independent academic work on evaluating agent-memory structures is direct competitive context for whether Add/Search becomes the reference memory-evaluation framework.
2026-09-22T09:22:24Z
The latest attachment is another same-author submission sharing an existing link, not an independent adoption report or new evaluation result. The challenge remains a relevant candidate protocol, but repeated promotion without methodological or implementation evidence does not warrant continued elevated attention.
2026-09-22T07:23:23Z
evidence attached: hn.story.49797488 — shared external link with case evidence
2026-09-21T14:13:37Z
The organizer’s Cycle 2 launch turns the sparse framework claim into a concrete, currently open challenge with a shared Add/Search interface, standardized answer/evaluation pipeline, three tracks, public results, and a prize pool. This substantiates availability and intended controls, but not yet reproducibility, participation, or whether the protocol isolates memory quality in practice.
2026-09-21T14:12:15Z
anchor promoted to claim owner's artifact: echo.blog.3c37e05a38 -> hn.story.49783671 — The organizer's Cycle 2 announcement is a new phase of the same Add/Search agent-memory evaluation effort.
2026-09-21T14:12:15Z
evidence attached: hn.story.49783671 — The organizer's Cycle 2 announcement is a new phase of the same Add/Search agent-memory evaluation effort.
2026-09-19T15:25:39Z
The same author's new question identifies memory freshness as an evaluation concern, but supplies no method, implementation, or result connecting it to the announced framework. It does not substantiate the release or independently corroborate the comparability claim; the earlier characterization of a concrete released artifact remains unverified.
2026-09-19T15:21:43Z
evidence attached: hn.story.49766795 — The observation directly raises freshness evaluation for agent memory, a material dimension for judging memory systems.
2026-09-18T03:25:31Z
grounded: converges/medium — The described Add/Search protocol converges with Scott’s Trace-backed Agent Comparison and Reflexive Agent Design: controlled comparisons rather than conflating
2026-09-18T03:22:24Z
case created — A distinct evaluation-framework release provides a concrete artifact relevant to memory selection, but the single sparse announcement does not yet establish implementation details or comparative results.