2026-10-11 17:12 UTC

Independent replication will determine whether frontier LLM factual errors are primarily caused by failures to recall stored parametric knowledge rather than by absence of that knowledge.

state: expiredheat: lowuncertainty: highconvergesscott: mediumfactuality retrieval model-evaluation ragGoogle Research

What is this?

Google Research reports that several frontier LLMs encode substantially more factual knowledge than they can reliably produce: its study claims 95–98% encoding for Gemini-3-Pro and GPT-5, despite direct-recall failures on 26–34% of facts and residual failures of 11–12% with thinking. The work frames factual errors as a “lost keys” problem—stored parametric knowledge being inaccessible—rather than solely “empty shelves,” or absent knowledge. The supplied material establishes an original Google Research report and arXiv paper, but does not establish that an independent replication has occurred; the broader snippets provide related evaluation work, not a replication of these specific findings.

Why it matters to Scott

Google Research’s reported distinction between absent knowledge and inaccessible encoded knowledge converges with Scott’s separation of answer failures by stage and his argument that residual model recall is often the wrong evaluation unit. If independently replicated, it could sharpen when retrieval or inference-time search should recover latent knowledge versus supply external evidence, but the current material remains an unreplicated report and does not weaken Scott’s provenance-focused case for external memory.
ip:concept.answer-failure-classesip:concept.benchmarking-the-wrong-unitip:concept.retrieval-augmented-generationip:concept.inference-time-scalingradar:concept.model-evaluationradar:concept.ragradar:concept.knowledge-systems
queries asked of Scott's wikis
  • parametric memory versus retrieval failure
  • RAG when knowledge is encoded but inaccessible
  • agent memory recall and retrieval design
  • factuality evaluation: knowledge versus access
  • reasoning as a factual-recall mechanism
  • external memory versus model-weight knowledge

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnFrontier LLMs know more facts than they can recallMarcoDewey92
🟧 echo.paper ⭐The original study is the authors’ arXiv paper, which distinguishes missing knowledge (“empty shelves”) from inaccessible encoded knowledge Nitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek, and Gal Yona——
🟠 redditI was trying to improve factual recall with prompting and found near-identical confidence can hide very different perturbation responses
LocalLLaMA
Any-Chipmunk548043

Interpretation history

Decision trace