2026-10-11 17:11 UTC

Independent evaluations will determine whether 1Password's SCAM benchmark realistically measures agents' susceptibility to scams and social engineering and supports effective defenses.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security agent-evaluation prompt-injection1Password

What is this?

1Password has released SCAM (Security Comprehension Awareness Measure), an open-source benchmark for testing whether AI agents safely handle scams and social-engineering attacks embedded in realistic, multi-turn workplace tasks involving email, links, forms, and credentials. Its reported evaluations of eight models found substantial variation and critical failures, while 1Password acknowledges that the 30 scenarios omit multi-agent workflows, real browser environments, long conversation histories, and evolving attacks. The supplied evidence is primarily from 1Password and launch coverage; it does not include independent evaluations establishing the benchmark’s realism, validity, or usefulness for developing defenses.

Why it matters to Scott

1Password’s task-based benchmark independently moves toward Scott’s model-plus-harness and representative-evaluation position, while its acknowledged lack of real browsers and fuller workflows creates a dated-receipts opportunity to test whether it benchmarks the right unit. It could also distinguish probabilistic scam resistance from Scott’s structural provenance, credential-isolation, and containment defenses, but no independent validation yet shows that SCAM should change those designs.
ip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditip:concept.evaluation-driven-developmentip:framework.agent-provenance-stackip:concept.guardrail-illusiondev:project.silo-osradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.agentic-securityradar:concept.prompt-injectionradar:concept.benchmark-integrityradar:onepassword-claude-secret-injection
queries asked of Scott's wikis
  • agent security evaluation harnesses
  • realistic benchmark design for tool-using agents
  • prompt injection versus social engineering defenses
  • credential handling and least privilege for agents
  • multi-turn adversarial testing of agent workflows
  • benchmark validity and defense transfer

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn1Password's new benchmark teaches AI agents how not to get scammedmrkd40
🟧 echo.github ⭐The initial SCAM repository commit is the furthest-back primary artifact. Its README says: “SCAM is an open-source benchmark that measures t1Password (initial commit by Jason Meller)——

Interpretation history

Decision trace