Google Research reports that several frontier LLMs encode substantially more factual knowledge than they can reliably produce: its study claims 95–98% encoding for Gemini-3-Pro and GPT-5, despite direct-recall failures on 26–34% of facts and residual failures of 11–12% with thinking. The work frames factual errors as a “lost keys” problem—stored parametric knowledge being inaccessible—rather than solely “empty shelves,” or absent knowledge. The supplied material establishes an original Google Research report and arXiv paper, but does not establish that an independent replication has occurred; the broader snippets provide related evaluation work, not a replication of these specific findings.
Google Research’s reported distinction between absent knowledge and inaccessible encoded knowledge converges with Scott’s separation of answer failures by stage and his argument that residual model recall is often the wrong evaluation unit. If independently replicated, it could sharpen when retrieval or inference-time search should recover latent knowledge versus supply external evidence, but the current material remains an unreplicated report and does not weaken Scott’s provenance-focused case for external memory.
ip:concept.answer-failure-classesip:concept.benchmarking-the-wrong-unitip:concept.retrieval-augmented-generationip:concept.inference-time-scalingradar:concept.model-evaluationradar:concept.ragradar:concept.knowledge-systems
queries asked of Scott's wikis
- parametric memory versus retrieval failure
- RAG when knowledge is encoded but inaccessible
- agent memory recall and retrieval design
- factuality evaluation: knowledge versus access
- reasoning as a factual-recall mechanism
- external memory versus model-weight knowledge
2026-09-01T04:28:53Z
After a full monitoring ladder, no independent replication, comparable critique, or implementation has emerged; retire the active episode without treating the underlying claim as disproved. Reopen if a genuinely independent test appears.
2026-08-30T04:25:55Z
No independent replication, comparable methodological critique, or implementation has appeared, leaving the bottleneck hypothesis an unverified single-team result. The 48-hour staleness trigger is too short for this research question; retain it on a weekly cadence.
2026-08-28T03:29:51Z
No independent replication, comparable critique, or implementation has appeared; the engagement decline in the adjacent Reddit experiment adds no evidentiary meaning. Keep this as a low-temperature single-team claim and avoid further checks until the weekly research cadence.
2026-08-26T02:30:01Z
No independent replication or comparable methodological test has appeared, so the encoded-but-inaccessible knowledge claim remains a single-team finding. Further unchanged checks add no meaning; move to a weekly research cadence.
2026-08-24T02:23:13Z
No independent replication, comparable critique, or implementation has appeared; repeated staleness does not further weaken or corroborate the single-team finding. Shift monitoring to a weekly research cadence rather than treating each stale interval as meaningful.
2026-08-22T01:28:30Z
Another unchanged observation provides no independent replication, critique, or implementation evidence; the adjacent Reddit experiment also remains methodologically non-corroborative. This is still a single-team research claim best monitored on a longer replication timescale.
2026-08-20T01:24:59Z
A methodological comment plausibly explains the adjacent Reddit result as serial-position sensitivity rather than evidence of inaccessible encoded knowledge, weakening its already limited corroborative value. The Google claim remains a single-team finding awaiting a comparable independent replication.
2026-08-19T12:32:05Z
The attached perturbation experiment adds adjacent evidence that similar apparent confidence can mask unstable recall behavior, but it does not test whether the underlying facts are encoded or replicate the Google study’s methodology. The central bottleneck claim therefore remains a single-team result awaiting an independent, comparable test.
2026-08-19T12:23:23Z
evidence attached: reddit.post.1vske9c — The controlled perturbation experiment materially probes whether factual failures reflect recall dynamics rather than simply missing knowledge.
2026-08-17T17:41:45Z
A further stale interval without replication, critique, or implementation leaves the claim unchanged as a single-team result. The case remains worth monitoring on a research timescale, but repeated absence of evidence adds no momentum or urgency.
2026-08-15T17:31:01Z
No replication, critique, or implementation evidence has emerged, so the claim remains a single-team result rather than a broader finding. The replication horizon is longer than this stale interval, warranting continued low-temperature monitoring rather than expiry.
2026-08-13T16:37:08Z
No independent replication or methodological challenge has appeared; this remains a single-team, testable result rather than evidence that recall failure is the dominant factuality bottleneck. The unchanged reobservation adds no momentum, so the case cools while awaiting a genuinely independent test.
2026-08-13T16:33:03Z
grounded: converges/medium — Google Research’s reported distinction between absent knowledge and inaccessible encoded knowledge converges with Scott’s separation of answer failures by stage
2026-08-13T16:30:33Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49288011 -> echo.paper.53c6fea35e by Nitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek, and Gal Yona
2026-08-13T16:29:38Z
case created — Google Research presents a testable finding with direct implications for retrieval augmentation and factuality evaluation.