2026-10-11 16:37 UTC

IreneAI claims the released open Add/Search evaluation framework compares agent-memory systems without team-selected answer models or evaluation pipelines, potentially separating memory quality from evaluation-setup advantages.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumagent-memory agent-evaluation memory-benchmarksIreneAI

What is this?

The case attributes to IreneAI a Show HN announcement of an open Add/Search framework for comparing agent-memory systems while keeping answer generation and evaluation under framework control. A Zylos Research snippet describes a matching protocol called AML: participants implement history ingestion (Add) and memory retrieval (Search), while AML controls generation, evaluation, aggregation, and orchestration. That snippet says comparability requires matching versioned benchmarks, pipelines, model configurations, and scoring rules; it supports the proposed mechanism, not proof that memory quality is fully isolated. The supplied web results do not directly establish IreneAI’s authorship or verify the announced release.

Why it matters to Scott

The described Add/Search protocol converges with Scott’s Trace-backed Agent Comparison and Reflexive Agent Design: controlled comparisons rather than conflating backend quality with evaluation setup, offering a concrete candidate protocol for testing his RAG/Wiki Substrate Rule. This could inform his memory-backend evaluations, but IreneAI’s authorship and release remain unverified, and fixed generation settings do not establish model-independent memory quality; the radar’s Agent Memory Leaderboard validation page tracks a related concern, not demonstrably this release.
dev:concept.trace-backed-agent-comparisonip:framework.reflexive-agent-designip:framework.rag-wiki-substrate-ruleradar:agent-memory-leaderboard-validationradar:ship-harness-benchradar:concept.memory-evaluation
queries asked of Scott's wikis
  • agent memory retrieval evaluation harness
  • benchmark confounds model versus system performance
  • reproducible evaluation versioned scoring contracts
  • agent-maintained wiki memory quality testing
  • memory backend comparison fixed answer model

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 565h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-21 06:24⭐ origin directly observedEvaluating Long-Term Memory for AI Agents: AML Cycle 2 Is Now Open
IreneAI on hacker news
—
09-18 02:52first on hacker news · published · +-75.5hShow HN: An open Add/Search evaluation framework for agent memory
IreneAI
—
09-18 03:22first on blog (echo) · first seen by us · +-75.0hThe team's Show HN announcement presents an open Add/Search evaluation framework aimed at making agent-memory systems comparable “without le
Framework team represented by IreneAI
—
09-18 02:52amplified on hacker news 👑hn.story.49749689
IreneAI
peak 2 · 0 comments · 25% of case engagement
09-19 14:20amplified on hacker newshn.story.49766795
IreneAI
peak 2 · 0 comments · 25% of case engagement
09-21 06:24amplified on hacker newshn.story.49783671
IreneAI
peak 2 · 0 comments · 25% of case engagement
09-22 06:40amplified on hacker newshn.story.49797488
IreneAI
peak 1 · 0 comments · 13% of case engagement
10-05 18:30amplified on hacker newshn.story.49968612
matt_d
peak 1 · 0 comments · 13% of case engagement
09-18 03:20our radar first saw it · +-75.1hdiscovery anchor: hn.story.49749689—
pace: p39 vs 1032 stories at the 336h mark (now 565h old) — ahead of agentsec-static-config-auditing (1.2x), behind anthropic-meta-lawsuit (0.8x)

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: An open Add/Search evaluation framework for agent memoryIreneAI20
🟧 echo.blogThe team's Show HN announcement presents an open Add/Search evaluation framework aimed at making agent-memory systems comparable “without leFramework team represented by IreneAI——
🟧 hnHow should we evaluate whether an AI agent's memory is still current?IreneAI20
🟧 hn ⭐Evaluating Long-Term Memory for AI Agents: AML Cycle 2 Is Now Open
Retrieved article excerpt

Open article · Retrieved 2026-09-21T07:23:12.876366+00:00

[@AgentMemoryL](https://x.com/AgentMemoryL)

[Agent Memory Leaderboard](https://x.com/AgentMemoryL)[@AgentMemoryL](https://x.com/AgentMemoryL)

Agent Memory Challenge 2026 Cycle 2 is now open.
Long-term memory is not just about storing more history. It is about retrieving the right evidence, recognizing what has changed, and avoiding stale context when an agent needs to act.
Three tracks: Textual · Coding · Multimodal
Open-source Methods · Commercial Products
Over USD 22,000 prize pool for eligible open-source teams.
A shared Add/Search interface. Standardized Answer/Eval. Public, comparable results.
Join: [agentmemoryleaderboard.ai/evaluation](https://agentmemoryleaderboard.ai/evaluation)

[3:35 AM · Sep 20, 2026](https://x.com/AgentMemoryL/status/2101515447816663222)·[2,807

Views](https://x.com/AgentMemoryL/status/2101515447816663222)

[7](https://x.com/AgentMemoryL/status/2101515447816663222)

3

15

2
IreneAI20
🟧 hnBeyond Context Windows: Evaluating Long-Term Memory for AI AgentsIreneAI10
🟧 hnEvaluating Memory Structure in LLM Agentsmatt_d10

Interpretation history

Decision trace