Achiral AI claims its released Cognoscenti benchmark can measure whether agent-memory systems retain and retrieve information accurately, securely, and consistently, potentially establishing a common evaluation for trustworthy long-running memory.
state: expiredheat: lowuncertainty: highknownscott: lowagent-memory memory-evaluation trustworthy-agentsAchiral AI
What is this?
Achiral AI has released Cognoscenti, which it claims evaluates whether agent-memory systems retain and retrieve information accurately, securely, and consistently, using established benchmarks including LoCoMo, LongMemEval, and BEAM. The supplied results confirm that these benchmarks cover long-term, temporal, and multi-session recall, while production memory evaluation may also test latency, token use, and cross-agent isolation. However, the snippets do not independently document Cognoscentiβs methodology or adoption, so its potential to become a common evaluation standard remains an unverified claim.
Why it matters to Scott
Scott already holds the relevant position in Evaluation-Driven Development and Model-Plus-Harness Benchmark Unit, while the radar already tracks validation of Memory Bench and the Agent Memory Leaderboard. Cognoscenti is another claimed evaluation in that established territory, but without supplied methodology, reproducibility evidence, or adoption, it does not yet extend or change what Scott builds or argues.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitip:framework.long-running-agentsradar:memory-bench-layer-baseline-validityradar:agent-memory-leaderboard-validationradar:concept.agent-memoryradar:concept.agent-evaluation
queries asked of Scott's wikis
- agent memory evaluation and benchmark design
- long-running agent memory architecture
- memory security isolation and leakage
- temporal recall and knowledge-update handling
- agent-maintained wikis memory reliability
- RAG versus persistent agent memory
Measured heat
no measured readings yet β the hourly heat pass fills this in
How the heat travelled
no chain yet β the hourly chain pass fills this in
Evidence (2) β β canonical anchor
Interpretation history
2026-09-04T15:51:37Z
No methodology, results, replication, or adoption emerged within the observation horizon, and the repeated 404 does not establish withdrawal. The artifact has not earned continued tracking unless it resurfaces with accessible evidence or implementation uptake.
2026-09-02T15:48:33Z
No substantive evidence has arrived beyond the original release claim; the failed reobservation neither validates the benchmark nor establishes that it was withdrawn. Without methodology, reproducibility, results, or adoption, this remains an unverified artifact in an already crowded evaluation space.
2026-09-02T15:43:14Z
grounded: known/low β Scott already holds the relevant position in Evaluation-Driven Development and Model-Plus-Harness Benchmark Unit, while the radar already tracks validation of M
2026-09-02T15:39:33Z
case created β A dedicated open benchmark is a concrete new artifact in the active agent-memory evaluation space.
Decision trace
- 09-05 01:51expireNo methodology, results, replication, or adoption emerged within the observation horizon, and the repeated 404 does not establish withdrawal. The artifact has not earned continued tracking unless it r
- 09-05 01:51alert_silentThe staleness trigger and failed reobservation add no consequential evidence; Scott can wait for an accessible benchmark, independent validation, comparative results, or adoption.
- 09-05 01:51alert_routeThe staleness trigger and failed reobservation add no consequential evidence; Scott can wait for an accessible benchmark, independent validation, comparative results, or adoption.
- 09-03 01:48repriceNo substantive evidence has arrived beyond the original release claim; the failed reobservation neither validates the benchmark nor establishes that it was withdrawn. Without methodology, reproducibil
- 09-03 01:48alert_silentThe only new delta is a 404 during reobservation, which is insufficient to infer a consequential change. The benchmark can wait for evidence of accessible methodology, independent replication, compara
- 09-03 01:48alert_routeThe only new delta is a 404 during reobservation, which is insufficient to infer a consequential change. The benchmark can wait for evidence of accessible methodology, independent replication, compara
- 09-03 01:43alert_silentAchiral AI appears to have released Cognoscenti, but the supplied evidence provides no methodology, reproducibility results, comparative findings, or adoption showing that it changes agent-memory eval
- 09-03 01:43alert_routeAchiral AI appears to have released Cognoscenti, but the supplied evidence provides no methodology, reproducibility results, comparative findings, or adoption showing that it changes agent-memory eval
- 09-03 01:43groundScott already holds the relevant position in Evaluation-Driven Development and Model-Plus-Harness Benchmark Unit, while the radar already tracks validation of Memory Bench and the Agent Memory Leaderb
- 09-03 01:39createA dedicated open benchmark is a concrete new artifact in the active agent-memory evaluation space.