Independent replication will determine whether poisoned or stale retrieved documents reliably steer local open-weight models to assert planted false values and whether practical retrieval defenses prevent the failure.
state: expiredheat: lowuncertainty: highnovelscott: nonerag-poisoning local-models knowledge-systems
What is this?
The evidence title describes a proposed or reported experiment ranking 11 local open-weight models by how often poisoned or stale retrieved sources cause them to assert a planted false value. The supplied snippets establish adjacent risks—Anthropic reports reliable training-time backdoors from as few as 250 poisoned documents, while a separate study finds stale repository retrieval can steer code models toward incompatible output and that current evidence can largely rescue the failure. However, the snippets do not identify who conducted the 11-model test, provide its results, or establish whether practical retrieval defenses were evaluated successfully, so the case’s specific replication claim remains unverified here.
Why it matters to Scott
No intersection found: the supplied Scott wiki and radar hits are empty, so there is no grounded basis to connect this replication claim to Scott’s existing positions, projects, or tracked cases.
queries asked of Scott's wikis
- RAG source poisoning and trust boundaries
- stale knowledge detection in agent memory
- retrieval provenance and document freshness
- local model susceptibility to retrieved context
- RAG defenses against planted falsehoods
- knowledge-system source validation and conflict resolution
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-07T17:47:38Z
No replication, scripts, results, or defense testing appeared within the observation horizon. The lone self-report has faded without enough evidence to sustain an active case, though the underlying failure mode remains testable.
2026-07-30T01:22:41Z
The attached evidence adds no independent replication, scripts, or defense testing; it is the same self-reported claim already assessed. The case remains testable but uncorroborated, with no reason for near-term attention.
2026-07-30T00:23:39Z
The new discussion challenges reproducibility and posting quality rather than independently validating the reported failure mode. Without scripts, results, or a second implementation, the case remains an uncorroborated self-report and has cooled slightly.
2026-07-29T17:24:16Z
grounded: novel/none — No intersection found: the supplied Scott wiki and radar hits are empty, so there is no grounded basis to connect this replication claim to Scott’s existing pos
2026-07-29T17:22:01Z
case created — The post reports a bounded, testable failure mode across 11 local models, but currently has only one self-reported observation and minimal engagement.
Decision trace
- 08-08 03:47expireNo replication, scripts, results, or defense testing appeared within the observation horizon. The lone self-report has faded without enough evidence to sustain an active case, though the underlying fa
- 08-08 03:47alert_silentThe only change is elapsed time without new evidence; there is no consequential delta for Scott and no confirming fact expected imminently.
- 08-08 03:47alert_routeThe only change is elapsed time without new evidence; there is no consequential delta for Scott and no confirming fact expected imminently.
- 07-30 11:22repriceThe attached evidence adds no independent replication, scripts, or defense testing; it is the same self-reported claim already assessed. The case remains testable but uncorroborated, with no reason fo
- 07-30 11:20mark_dirtycomment_update
- 07-30 10:23repriceThe new discussion challenges reproducibility and posting quality rather than independently validating the reported failure mode. Without scripts, results, or a second implementation, the case remains
- 07-30 03:24groundNo intersection found: the supplied Scott wiki and radar hits are empty, so there is no grounded basis to connect this replication claim to Scott’s existing positions, projects, or tracked cases.
- 07-30 03:22createThe post reports a bounded, testable failure mode across 11 local models, but currently has only one self-reported observation and minimal engagement.