2026-10-11 17:20 UTC

Independent replication will determine whether poisoned or stale retrieved documents reliably steer local open-weight models to assert planted false values and whether practical retrieval defenses prevent the failure.

state: expiredheat: lowuncertainty: highnovelscott: nonerag-poisoning local-models knowledge-systems

What is this?

The evidence title describes a proposed or reported experiment ranking 11 local open-weight models by how often poisoned or stale retrieved sources cause them to assert a planted false value. The supplied snippets establish adjacent risks—Anthropic reports reliable training-time backdoors from as few as 250 poisoned documents, while a separate study finds stale repository retrieval can steer code models toward incompatible output and that current evidence can largely rescue the failure. However, the snippets do not identify who conducted the 11-model test, provide its results, or establish whether practical retrieval defenses were evaluated successfully, so the case’s specific replication claim remains unverified here.

Why it matters to Scott

No intersection found: the supplied Scott wiki and radar hits are empty, so there is no grounded basis to connect this replication claim to Scott’s existing positions, projects, or tracked cases.
queries asked of Scott's wikis
  • RAG source poisoning and trust boundaries
  • stale knowledge detection in agent memory
  • retrieval provenance and document freshness
  • local model susceptibility to retrieved context
  • RAG defenses against planted falsehoods
  • knowledge-system source validation and conflict resolution

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Poisoned and stale sources: 11 local models ranked by how often they assert a planted false value
LocalLLaMA
arcandor14

Interpretation history

Decision trace