2026-10-11 17:09 UTC

BrainAPI’s developer claims its externally managed memory layer outperforms Mem0, Zep, and Letta on LoCoMo and BEAM1M, which would make BrainAPI a competitive backend for persistent agent memory if the reported results hold.

state: expiredheat: lowuncertainty: highknownscott: lowagent-memory memory-benchmarksBrainAPI

What is this?

BrainAPI is presented as an externally managed persistent-memory backend for AI agents whose developer claims it outperforms Mem0, Zep, and Letta on the LoCoMo and BEAM1M benchmarks. The supplied snippets establish that agent-memory benchmark results are heavily disputed: Zep challenged Mem0’s LoCoMo methodology, while even a simple file-and-grep setup reportedly scored 74.0% versus Mem0’s published 68.5% in one comparison. The material does not provide BrainAPI’s actual scores, methodology, paper, or independent reproduction, and the separate Memanto result is too thin here to establish its relationship to BrainAPI, so the competitive claim remains unverified testimony rather than a grounded ranking.

Why it matters to Scott

Scott already holds the substantive positions in “Long-Running Agents” and “Model-Plus-Harness Benchmark Unit”: durable agent continuity belongs in external state, while memory scores are meaningful only with a disclosed, reproducible harness. BrainAPI’s unsupported ranking adds no validated result or architectural detail that would yet change his builds or arguments, and the radar already tracks this benchmark-validity problem in “Agent Memory Leaderboard Validation” and “Memory Bench Layer Baseline Validity.”
ip:framework.long-running-agentsip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentradar:agent-memory-leaderboard-validationradar:memory-bench-layer-baseline-validityradar:concept.agent-memory
queries asked of Scott's wikis
  • external memory services vs agent-owned memory
  • agent memory benchmark validity and evaluation harnesses
  • persistent memory retrieval architecture for coding agents
  • files and grep vs vector or graph memory
  • memory backend portability and vendor lock-in
  • long-horizon agent memory failure modes

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditWe beat Mem0, Zep and Letta on two memory benchmarks. The score isn't the interesting part
LocalLLaMA
shbong015
🟧 echo.paper ⭐The original paper introduces Memanto and reports: “Memanto achieves state of the art accuracy scores of 89.8 percent and 87.1 percent” on LSeyed Moein Abtahi, Rasa Rahnema, Hetkumar Patel, Neel Patel, Majid Fekri, Tara Khani——

Interpretation history

Decision trace