The case concerns a reported experiment by Eric Chiang of Oblique Security using Anthropic’s Claude Code to search SAML implementations for exploitable security flaws. The supplied search results do not independently document that specific experiment or its findings, but broader testing shows Claude Code can find real vulnerabilities while producing high false-positive rates; source-only analysis also struggles to establish exploitability in deployed systems. Independent reproduction would therefore need to verify both the claimed SAML flaws and Claude Code’s ability to discover and exploit them autonomously under realistic conditions.
Scott already holds the case’s central position in Security Reviewer Method: agent-found vulnerabilities remain conditional until dangerous primitives are closed into reachable, independently checkable exploit paths. Reproducing the SAML work would directly exercise his active security-review harness and Claude Code practice, but the supplied evidence does not yet establish a new result; the radar also already tracks closely related Claude Code security validation in “Claude Security plugin beta.”
ip:source.security-reviewer-method-ebookip:concept.claim-bounded-adversarial-verificationip:concept.mechanically-different-verifiersdev:project.wordpress-security-reviewdev:technology.claude-coderadar:claude-security-plugin-betaradar:exploitgym-agent-exploitation-validationradar:concept.coding-agent-securityradar:concept.agentic-security
queries asked of Scott's wikis
- autonomous coding agents for vulnerability research
- agent harnesses for exploit validation and false-positive reduction
- coding-agent evaluation on realistic end-to-end tasks
- AI security research reproducibility and independent verification
- agentic testing of authentication and identity protocols
- source analysis versus runtime exploitability
2026-08-22T00:23:19Z
After the monitoring horizon, the account still has no independent reproduction, named affected implementation, concrete flaw, or reachable exploit path. The episode has faded as a standalone signal and should be reopened only if a reproducible technical artifact emerges.
2026-08-19T23:39:02Z
No independent reproduction, concrete vulnerability, affected implementation, or reachable exploit path has emerged; this remains an unvalidated primary-account claim rather than a developing result. The unchanged observation adds no new meaning, so the case cools while awaiting substantive technical evidence.
2026-08-19T23:32:13Z
grounded: known/medium — Scott already holds the case’s central position in Security Reviewer Method: agent-found vulnerabilities remain conditional until dangerous primitives are close
2026-08-19T23:29:33Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49368038 -> echo.blog.73c8f84e29 by Oblique
2026-08-19T23:28:24Z
case created — The original security write-up presents a concrete coding-agent vulnerability-research episode with transferable control and evaluation lessons.