2026-10-11 17:10 UTC

Independent reproduction will determine whether Claude Code can autonomously discover and exploit consequential SAML implementation flaws in realistic applications.

state: expiredheat: lowuncertainty: highknownscott: mediumagentic-security coding-agents claude-codeEric ChiangOblique SecurityAnthropic

What is this?

The case concerns a reported experiment by Eric Chiang of Oblique Security using Anthropic’s Claude Code to search SAML implementations for exploitable security flaws. The supplied search results do not independently document that specific experiment or its findings, but broader testing shows Claude Code can find real vulnerabilities while producing high false-positive rates; source-only analysis also struggles to establish exploitability in deployed systems. Independent reproduction would therefore need to verify both the claimed SAML flaws and Claude Code’s ability to discover and exploit them autonomously under realistic conditions.

Why it matters to Scott

Scott already holds the case’s central position in Security Reviewer Method: agent-found vulnerabilities remain conditional until dangerous primitives are closed into reachable, independently checkable exploit paths. Reproducing the SAML work would directly exercise his active security-review harness and Claude Code practice, but the supplied evidence does not yet establish a new result; the radar also already tracks closely related Claude Code security validation in “Claude Security plugin beta.”
ip:source.security-reviewer-method-ebookip:concept.claim-bounded-adversarial-verificationip:concept.mechanically-different-verifiersdev:project.wordpress-security-reviewdev:technology.claude-coderadar:claude-security-plugin-betaradar:exploitgym-agent-exploitation-validationradar:concept.coding-agent-securityradar:concept.agentic-security
queries asked of Scott's wikis
  • autonomous coding agents for vulnerability research
  • agent harnesses for exploit validation and false-positive reduction
  • coding-agent evaluation on realistic end-to-end tasks
  • AI security research reproducibility and independent verification
  • agentic testing of authentication and identity protocols
  • source analysis versus runtime exploitability

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnHacking SAML with Claude Codeericchiang10
🟧 echo.blog ⭐This is the original account of Eric Chiang’s SAML research using Claude Code. He says he “attempted to hack every SAML implementation I couOblique——

Interpretation history

Decision trace