2026-10-11 16:37 UTC

Google Threat Intelligence claims its agentic source-code review workflow can help defenders identify and remediate security weaknesses associated with adversarial AI, potentially making agent-driven review a practical defensive control.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagentic-security coding-agents automated-code-reviewGoogle Threat IntelligenceGoogle Cloud

What is this?

AVDH (Agentic Vulnerability Discovery Harness) is a source-code-review pipeline published by Google Cloud/Mandiant on August 19, 2026 (authors Alex Tselevich and Michael Maturi): an Explorer agent maps the target codebase and dispatches Specialist Explorer subagents, followed by hypothesis generation, skeptical validation, and human exploit verification, run as a deterministic pipeline across proactive reviews, pentests, red team, and incident-response engagements. A Google reseller summary (Softprom, Sept 7, 2026) credits the harness with environments spanning tens of millions of lines of code, thousands of pipelines executed, tens of thousands of findings, 12 assigned CVEs (e.g. CVE-2026-13242, CVE-2026-55803) plus about a dozen more in active disclosure, and positions AVDH alongside CodeMender scanning as a two-layer defense β€” but those performance figures are secondhand and Google's own deep-dive promised at the Cyber Defense Summit (Sept 15–16, 2026) has no results in the supplied material. The case also carries Trail of Bits' independent judgment that AI-assisted security review is practically 'good enough,' and AVDH sits inside a broader Google Cloud agentic-security push (AI Threat Defense autonomous remediation, Threat Hunting and Detection Engineering agents). Measured accuracy, comparison with conventional review tools, and remediation outcomes thus remain unvalidated first-party claims.

Why it matters to Scott

Google/Mandiant's AVDH independently ships the deterministic, hypothesis-first, verification-gated multi-agent review architecture Scott's Security Reviewer Method already documents β€” skeptical claim-bounded validation, mechanically distinct verifiers, findings conditional until exploit paths close β€” and Trail of Bits' independent 'good enough' judgment adds a second consequential party, making this a dated-receipts publishing opportunity rather than mere repetition. Mantis is additionally the concrete, evaluable harness artifact to benchmark against his live WordPress security-review pipeline, and the case's unresolved accuracy/remediation questions are exactly the evaluation-gates he argues such claims must pass before counting.
ip:source.security-reviewer-method-ebookip:concept.claim-bounded-adversarial-verificationip:concept.mechanically-different-verifiersip:concept.model-plus-harness-benchmark-unitdev:project.wordpress-security-reviewradar:cloudflare-security-audit-skillradar:harnesseval-code-review-gainsradar:cross-model-code-review-validationradar:concept.agent-harnessesradar:concept.coding-agent-security
queries asked of Scott's wikis
  • security reviewer method multi-agent code review pipeline
  • mechanically different verifiers security findings validation
  • evaluation-driven development agent harness benchmarking
  • WordPress automated security review workflow
  • automated patch remediation agent CodeMender
  • coding agent harness prompt injection adversarial robustness

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1322h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-09 05:29⭐ origin directly observedGoogle: AI agents harvested credentials in under six hours
lbw1215 on hacker news
β€”
08-17 14:00first on blog (echo) Β· published Β· +-543.5hThis is the primary publication: Mandiant introduces its Agentic Vulnerability Discovery Harness (AVDH), describing a deterministic multi-ag
Mandiant (Google Cloud)
β€”
08-31 21:04first on hacker news Β· published Β· +-200.4hStaying Ahead of Adversarial AI Through Agentic Source Code Review
wslh
β€”
08-31 21:04amplified on hacker newshn.story.49514820
wslh
peak 1 Β· 0 comments Β· 1% of case engagement
09-07 18:02amplified on hacker newshn.story.49601113
joshcsimmons
peak 2 Β· 0 comments Β· 1% of case engagement
09-09 05:29amplified on hacker newshn.story.49621461
lbw1215
peak 3 Β· 0 comments Β· 2% of case engagement
09-09 15:19amplified on hacker newshn.story.49628034
fourfire
peak 13 Β· 1 comments Β· 10% of case engagement
09-10 15:05amplified on hacker newshn.story.49645017
Cider9986
peak 3 Β· 0 comments Β· 2% of case engagement
09-21 21:53amplified on hacker news πŸ‘‘hn.story.49793957
aray07
peak 96 Β· 16 comments Β· 83% of case engagement
08-31 21:21our radar first saw it Β· +-200.1hdiscovery anchor: hn.story.49514820β€”

Evidence (7) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnStaying Ahead of Adversarial AI Through Agentic Source Code Reviewwslh10
🟧 echo.blogThis is the primary publication: Mandiant introduces its Agentic Vulnerability Discovery Harness (AVDH), describing a deterministic multi-agMandiant (Google Cloud)β€”β€”
🟧 hnGetting started with the Mantis harness to find and fix bugsjoshcsimmons20
🟧 hn ⭐Google: AI agents harvested credentials in under six hourslbw121530
🟧 hnGoogle: Attackers are using prompt injection against coding agentsfourfire131
🟧 hnAutomated security review of Monero-project/Monero pull requestsCider998630
🟧 hnSecurity auditing in the age of (good enough) AIaray079616

Interpretation history

Decision trace