2026-10-11 17:11 UTC

Achiral AI claims its released Cognoscenti benchmark can measure whether agent-memory systems retain and retrieve information accurately, securely, and consistently, potentially establishing a common evaluation for trustworthy long-running memory.

state: expiredheat: lowuncertainty: highknownscott: lowagent-memory memory-evaluation trustworthy-agentsAchiral AI

What is this?

Achiral AI has released Cognoscenti, which it claims evaluates whether agent-memory systems retain and retrieve information accurately, securely, and consistently, using established benchmarks including LoCoMo, LongMemEval, and BEAM. The supplied results confirm that these benchmarks cover long-term, temporal, and multi-session recall, while production memory evaluation may also test latency, token use, and cross-agent isolation. However, the snippets do not independently document Cognoscenti’s methodology or adoption, so its potential to become a common evaluation standard remains an unverified claim.

Why it matters to Scott

Scott already holds the relevant position in Evaluation-Driven Development and Model-Plus-Harness Benchmark Unit, while the radar already tracks validation of Memory Bench and the Agent Memory Leaderboard. Cognoscenti is another claimed evaluation in that established territory, but without supplied methodology, reproducibility evidence, or adoption, it does not yet extend or change what Scott builds or argues.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitip:framework.long-running-agentsradar:memory-bench-layer-baseline-validityradar:agent-memory-leaderboard-validationradar:concept.agent-memoryradar:concept.agent-evaluation
queries asked of Scott's wikis
  • agent memory evaluation and benchmark design
  • long-running agent memory architecture
  • memory security isolation and leakage
  • temporal recall and knowledge-update handling
  • agent-maintained wikis memory reliability
  • RAG versus persistent agent memory

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnCognoscenti: A Benchmark for Trustworthy AI Memorymarvindanig11
🟧 echo.github ⭐Releases Cognoscenti as a benchmark for trustworthy AI memory.Achiral AIβ€”β€”

Interpretation history

Decision trace