Independent evaluations will determine whether 1Password's SCAM benchmark realistically measures agents' susceptibility to scams and social engineering and supports effective defenses.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security agent-evaluation prompt-injection1Password
What is this?
1Password has released SCAM (Security Comprehension Awareness Measure), an open-source benchmark for testing whether AI agents safely handle scams and social-engineering attacks embedded in realistic, multi-turn workplace tasks involving email, links, forms, and credentials. Its reported evaluations of eight models found substantial variation and critical failures, while 1Password acknowledges that the 30 scenarios omit multi-agent workflows, real browser environments, long conversation histories, and evolving attacks. The supplied evidence is primarily from 1Password and launch coverage; it does not include independent evaluations establishing the benchmark’s realism, validity, or usefulness for developing defenses.
Why it matters to Scott
1Password’s task-based benchmark independently moves toward Scott’s model-plus-harness and representative-evaluation position, while its acknowledged lack of real browsers and fuller workflows creates a dated-receipts opportunity to test whether it benchmarks the right unit. It could also distinguish probabilistic scam resistance from Scott’s structural provenance, credential-isolation, and containment defenses, but no independent validation yet shows that SCAM should change those designs.
ip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditip:concept.evaluation-driven-developmentip:framework.agent-provenance-stackip:concept.guardrail-illusiondev:project.silo-osradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.agentic-securityradar:concept.prompt-injectionradar:concept.benchmark-integrityradar:onepassword-claude-secret-injection
queries asked of Scott's wikis
- agent security evaluation harnesses
- realistic benchmark design for tool-using agents
- prompt injection versus social engineering defenses
- credential handling and least privilege for agents
- multi-turn adversarial testing of agent workflows
- benchmark validity and defense transfer
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-17T14:41:10Z
No independent reproduction, extension, or demonstrated defense transfer emerged during the observation window. Retire the active episode and reopen only if external testing establishes SCAM’s realism or practical security value.
2026-08-15T14:36:07Z
The reobservation adds no independent evaluation, implementation, or defense-transfer evidence, so the benchmark remains a promising first-party artifact rather than a validated measure. Cool the case pending substantive external testing.
2026-08-15T14:31:40Z
grounded: converges/medium — 1Password’s task-based benchmark independently moves toward Scott’s model-plus-harness and representative-evaluation position, while its acknowledged lack of re
2026-08-15T14:28:54Z
origin walked (codex/luna, conf 0.97): anchor hn.story.49310309 -> echo.github.83da53cc19 by 1Password (initial commit by Jason Meller)
2026-08-15T14:28:00Z
case created — A first-party benchmark introduces a concrete evaluation artifact for a consequential and undermeasured agent-security failure mode.
Decision trace
- 08-18 00:41expireNo independent reproduction, extension, or demonstrated defense transfer emerged during the observation window. Retire the active episode and reopen only if external testing establishes SCAM’s realism
- 08-18 00:41alert_silentThe delta is only another unchanged reobservation of an old first-party artifact; there is no new consequential fact for Scott before the next briefing.
- 08-18 00:41alert_routeThe delta is only another unchanged reobservation of an old first-party artifact; there is no new consequential fact for Scott before the next briefing.
- 08-16 00:36repriceThe reobservation adds no independent evaluation, implementation, or defense-transfer evidence, so the benchmark remains a promising first-party artifact rather than a validated measure. Cool the case
- 08-16 00:36alert_silentThis is only an unchanged legacy-state recheck; Scott can wait for an independent reproduction, benchmark extension, or demonstrated security improvement before revisiting it.
- 08-16 00:36alert_routeThis is only an unchanged legacy-state recheck; Scott can wait for an independent reproduction, benchmark extension, or demonstrated security improvement before revisiting it.
- 08-16 00:32alert_silent1Password’s open-source SCAM benchmark is an established, relevant release—30 workplace scenarios across nine threat categories—but the visible delta adds no independent validation, comparative result
- 08-16 00:32surface_candidate1Password’s open-source SCAM benchmark is an established, relevant release—30 workplace scenarios across nine threat categories—but the visible delta adds no independent validation, comparative result
- 08-16 00:32alert_route1Password’s open-source SCAM benchmark is an established, relevant release—30 workplace scenarios across nine threat categories—but the visible delta adds no independent validation, comparative result
- 08-16 00:31ground1Password’s task-based benchmark independently moves toward Scott’s model-plus-harness and representative-evaluation position, while its acknowledged lack of real browsers and fuller workflows creates
- 08-16 00:28promote_anchororigin walk conf 0.97
- 08-16 00:28createA first-party benchmark introduces a concrete evaluation artifact for a consequential and undermeasured agent-security failure mode.