AVDH (Agentic Vulnerability Discovery Harness) is a source-code-review pipeline published by Google Cloud/Mandiant on August 19, 2026 (authors Alex Tselevich and Michael Maturi): an Explorer agent maps the target codebase and dispatches Specialist Explorer subagents, followed by hypothesis generation, skeptical validation, and human exploit verification, run as a deterministic pipeline across proactive reviews, pentests, red team, and incident-response engagements. A Google reseller summary (Softprom, Sept 7, 2026) credits the harness with environments spanning tens of millions of lines of code, thousands of pipelines executed, tens of thousands of findings, 12 assigned CVEs (e.g. CVE-2026-13242, CVE-2026-55803) plus about a dozen more in active disclosure, and positions AVDH alongside CodeMender scanning as a two-layer defense β but those performance figures are secondhand and Google's own deep-dive promised at the Cyber Defense Summit (Sept 15β16, 2026) has no results in the supplied material. The case also carries Trail of Bits' independent judgment that AI-assisted security review is practically 'good enough,' and AVDH sits inside a broader Google Cloud agentic-security push (AI Threat Defense autonomous remediation, Threat Hunting and Detection Engineering agents). Measured accuracy, comparison with conventional review tools, and remediation outcomes thus remain unvalidated first-party claims.
Google/Mandiant's AVDH independently ships the deterministic, hypothesis-first, verification-gated multi-agent review architecture Scott's Security Reviewer Method already documents β skeptical claim-bounded validation, mechanically distinct verifiers, findings conditional until exploit paths close β and Trail of Bits' independent 'good enough' judgment adds a second consequential party, making this a dated-receipts publishing opportunity rather than mere repetition. Mantis is additionally the concrete, evaluable harness artifact to benchmark against his live WordPress security-review pipeline, and the case's unresolved accuracy/remediation questions are exactly the evaluation-gates he argues such claims must pass before counting.
ip:source.security-reviewer-method-ebookip:concept.claim-bounded-adversarial-verificationip:concept.mechanically-different-verifiersip:concept.model-plus-harness-benchmark-unitdev:project.wordpress-security-reviewradar:cloudflare-security-audit-skillradar:harnesseval-code-review-gainsradar:cross-model-code-review-validationradar:concept.agent-harnessesradar:concept.coding-agent-security
queries asked of Scott's wikis
- security reviewer method multi-agent code review pipeline
- mechanically different verifiers security findings validation
- evaluation-driven development agent harness benchmarking
- WordPress automated security review workflow
- automated patch remediation agent CodeMender
- coding agent harness prompt injection adversarial robustness
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1322h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
2026-09-25T22:53:56Z
The flagged velocity spike was the Trail of Bits thread's tail, not re-ignition: it finished its cycle at 93/16 and now sits at ~0 pts/h and 12.5th percentile, with only generic 'good enough / assume breach' debate already priced in. More meaningfully, the Cyber Defense Summit (Sept 15β16) has now closed without the promised AVDH deep-dive results surfacing in evidence, so the expectation of near-term vendor validation lapses; the direction stays corroborated but attention-worthy pressure is gone β cool to low.
2026-09-24T21:30:02Z
grounded: converges/high β Google/Mandiant's AVDH independently ships the deterministic, hypothesis-first, verification-gated multi-agent review architecture Scott's Security Reviewer Met
2026-09-24T21:23:50Z
Trail of Bits β an independent, credentialed first-party auditor β now publicly judges AI-assisted security review practically adequate ('good enough'), converting the case from a single Google vendor claim into a corroborated direction, though AVDH's specific effectiveness numbers and the promised Cyber Defense Summit deep-dive remain unvalidated. Heat moves to medium on steady multi-week periphery expansion (new independent actors, hottest thread at ~84th percentile for its cohort) rather than raw engagement, which stays thin.
2026-09-24T20:37:40Z
evidence attached: hn.story.49793957 β Trail of Bits, a credible first-party auditor, independently addresses whether AI-driven security review is practically adequate β material context for the agent-driven-review case.
2026-09-10T18:04:00Z
The Monero item adds a potential implementation comparison, but its headline alone establishes neither agentic operation nor deployment outcomes; the attachment's characterization as independent corroboration is premature. Google's defensive-effectiveness hypothesis remains unvalidated, with no new detail on the separately surfaced coding-agent threat claims.
2026-09-10T15:26:00Z
evidence attached: hn.story.49645017 β A deployed automated security-review workflow provides independent corroboration for agent-driven source-code security analysis.
2026-09-09T15:38:10Z
The new headline attributes attacker use of prompt injection against coding agents to Google, adding a directly relevant operational threat claim without validating AVDH's defensive effectiveness. That claim has already been routed for attention; the next useful evidence is the underlying report's attack context, affected tools and outcomes, not further headlines linking the same publication.
2026-09-09T15:23:42Z
evidence attached: hn.story.49628034 β shared external link with case evidence
2026-09-09T06:23:12Z
The credential-harvesting headline adds an offensive claim attributed to Google, not corroboration that AVDH is an effective defensive control. Without source details distinguishing an observed attack from a controlled demonstration and specifying starting access, it does not materially strengthen the workflow hypothesis or establish a new operational risk for Scott.
2026-09-09T06:22:36Z
anchor promoted to claim owner's artifact: echo.blog.2b253af153 -> hn.story.49621461 β The linked Google Threat Intelligence artifact connects credential harvesting to the existing adversarial-AI reporting episode, adding an offensive finding rather than warranting a separate case and supplying the claim owner's source as anchor.
2026-09-09T06:22:36Z
evidence attached: hn.story.49621461 β The linked Google Threat Intelligence artifact connects credential harvesting to the existing adversarial-AI reporting episode, adding an offensive finding rather than warranting a separate case and supplying the claim owner's source as anchor.
2026-09-07T18:24:38Z
The Mantis getting-started headline adds a potentially useful implementation reference, but the supplied evidence does not establish its relationship to AVDH or independently validate the defensive workflow. The earlier AVDH architecture and performance claims remain reconstructed testimony; the newly attached guide has already been routed for attention.
2026-09-07T18:22:39Z
evidence attached: hn.story.49601113 β Google's Mantis harness is first-party corroborating evidence for agent-assisted security bug discovery and remediation.
2026-09-06T23:49:20Z
No new evidence since prior look; staleness trigger only. AVDH remains an unvalidated first-party claim pending the September Cyber Defense Summit demonstration or independent reproduction.
2026-09-04T23:30:51Z
The minor engagement increase is repetitive amplification, not validation; AVDH remains a consequential first-party implementation whose performance claims await technical results, demonstration, or independent reproduction.
2026-09-02T22:44:59Z
No new evidence changes the case: AVDH remains a consequential first-party implementation with unvalidated performance claims. Keep watching for the announced summit demonstration, benchmarks, or independent reproduction rather than treating staleness as disproof.
2026-08-31T21:48:05Z
The first-party disclosure establishes a real Mandiant workflow and reported deployment, moving this beyond a bare concept, but no independent validation or new evidence supports its effectiveness claims. The latest observation is unchanged, so attention should cool pending the promised summit demonstration or technical results.
2026-08-31T21:44:14Z
grounded: converges/high β Google/Mandiant is a consequential independent party advancing the same core direction Scott has already documented and built: deterministic, multi-agent source
2026-08-31T21:41:10Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49514820 -> echo.blog.2b253af153 by Mandiant (Google Cloud)
2026-08-31T21:40:08Z
case created β Googleβs first-party security publication introduces a concrete agentic review workflow with direct relevance to secure software operations.