2026-10-11 18:04 UTC

Aisle claims its AI-assisted security review found six curl CVEs after assessments associated with OpenAI and Anthropic found none, suggesting audit methodology and harness design materially affect AI vulnerability-discovery results.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security vulnerability-research software-securityAislecurlOpenAIAnthropic

What is this?

AISLE, an AI security research firm, says its AI-assisted review of curl uncovered six CVEs, including authentication, memory-safety, host-validation, and protocol-related flaws. It contrasts that result with earlier assessments associated with OpenAI and Anthropic that reportedly found none, implying that the surrounding audit methodology, agent harness, and human verification process may matter as much as model capability. The supplied snippets support AISLE’s discoveries and a broader verification/triage bottleneck, but the “OpenAI and Anthropic found zero” comparison is only asserted in the summary and evidence titles, not independently detailed in the search excerpts.

Why it matters to Scott

AISLE’s claimed six-versus-zero result is a concrete dated receipt for Scott’s model-plus-harness benchmark unit and his hypothesis-first, independently verified security-review method; it also bears directly on his active WordPress security-review system. It creates a publishing and evaluation-design opportunity, though the vendor comparison remains insufficiently substantiated in the supplied evidence.
ip:source.security-reviewer-method-ebookip:concept.model-plus-harness-benchmark-unitip:concept.mechanically-different-verifiersdev:project.wordpress-security-reviewdev:concept.trace-backed-agent-comparisonradar:concept.agent-harnessesradar:concept.agentic-securityradar:concept.agent-evaluationradar:concept.security-agents
queries asked of Scott's wikis
  • security-agent harness design and evaluation
  • LLM coding agents for vulnerability discovery
  • agentic workflows versus base-model capability
  • verification and triage of AI-generated findings
  • AI security benchmark methodology and false negatives
  • human-in-the-loop software security research

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnSix curl CVEs after OpenAI and Anthropic came back with zerogoobreee17865
🟧 echo.blog ⭐Aisle says it discovered six curl CVEs after OpenAI and Anthropic found zero.Aisle——

Interpretation history

Decision trace