2026-10-11 16:37 UTC

Open Codenames benchmark gains traction as a cited reference for evaluating LLM reasoning and communication in multi-agent settings.

state: seedheat: lowuncertainty: mediumnovelscott: lowagent-evaluation llm-benchmarks multi-agentAttilaT

What is this?

The web results show an academic paper (arXiv:2412.11373) proposing the board game Codenames as a benchmark for LLM reasoning, theory of mind, and epistemic reasoning, plus a Codenames AI Competition at CoG 2025 organized by researchers from Flinders University and University of Melbourne. However, the supplied snippets do not surface a specific 'Open Codenames' GitHub release or benchmark by 'AttilaT' โ€” the GitHub result returned is a general benchmark aggregation repo (awaresome_LLM_eval_benchmark) that does not mention this project. The case's central claim (a new GitHub-released benchmark by AttilaT gaining traction as a cited reference) is not substantiated in the provided material.

Why it matters to Scott

The case claims a specific 'Open Codenames' GitHub release by 'AttilaT' gaining traction as a cited reference, but the grounding notes this claim is not substantiated in the supplied material โ€” the web snippets show only an academic paper and a competition, not the claimed GitHub benchmark. Scott's canon has deep frameworks on agent evaluation (model-plus-harness, evaluation-driven-development, trace-backed-agent-comparison), multi-agent coordination (micro-agents, scatter-gather, dialectical-tree-search), and benchmark validity (version-bound-assessment, proof-of-read-citation-gate), but none reference this specific benchmark or author. The radar tracks the general topics (agent-benchmarks, multi-agent-coordination, benchmark-validity) but has no episode for this case. Since the central entity is unverified in the supplied evidence and has no intersection with Scott's specific positions or projects, there is no meaningful connection โ€” merely a topical overlap with areas he works in.
queries asked of Scott's wikis
  • agent-evaluation frameworks Scott has built or argued for (e.g., multi-agent harnesses, benchmark design principles)
  • open-weights model sovereignty and local inference economics โ€” does this benchmark assume API models or support local eval?
  • multi-agent communication/coordination benchmarks โ€” how does Codenames compare to AgentBench, PerspectiveGap, HumanEvalComm in Scott's view?
  • benchmark reproducibility and citation dynamics โ€” what makes a benchmark become a 'cited reference' vs. a one-off release?

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p16momentum: steady1 platformsage 48h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-09 16:27โญ origin directly observedOpen Codenames: a new way to benchmark LLMs
AttilaT on hacker news
โ€”
10-09 16:27amplified on hacker news ๐Ÿ‘‘hn.story.50022892
AttilaT
peak 1 ยท 1 comments ยท 98% of case engagement
10-09 19:34our radar first saw it ยท +3.1hdiscovery anchor: hn.story.50022892โ€”
pace: p33 vs 968 stories at the 24h mark (now 48h old) โ€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentlane-git-native-coordination (0.7x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hn โญOpen Codenames: a new way to benchmark LLMsAttilaT11

Interpretation history

Decision trace