2026-10-11 16:36 UTC

Microsoft claims its validation-first framework provides a practical way to verify agent outputs and tool actions before they are trusted or executed, potentially making validation gates a reusable control in agent harnesses.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumagent-verification agent-harnesses agentic-securityMicrosoft

What is this?

The case concerns “Only believe what you can validate,” identified in the supplied evidence titles as a Microsoft article presenting Julia Kordick’s opinionated verification framework for agentic-AI outputs. The web results do not include that article: secondary sources instead describe Microsoft-related policy gates, human approval before tool execution, and middleware for input validation. These establish related control patterns, not the specific article’s mechanisms or demonstrated practicality; one source explicitly distinguishes authorization checks from validating business data, so approval should not be equated with output correctness.

Why it matters to Scott

The attributed Microsoft article offers a potential publishing comparison with Scott’s Verification Loops and Decision Authority Infrastructure: checking observable evidence before trusting claims, and independently gating consequential actions rather than trusting model proposals. This is provisional convergence, not demonstrated architectural equivalence or practicality—the primary article is missing from the supplied grounding, output correctness must remain distinct from authorization, and no radar hit establishes that this same development is already tracked.
ip:concept.verification-loopsip:framework.decision-authority-infrastructuredev:concept.deterministic-agent-control-planeradar:concept.agent-verificationradar:conduct-tool-call-guardrailsradar:dmx-gated-coding-agent-loops
queries asked of Scott's wikis
  • agent harness verification gates before tool execution
  • coding agent claims evidence actual test lint execution
  • authorization versus output correctness validation
  • deterministic validators versus LLM adversarial review
  • agent tool middleware human approval trust boundaries

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady3 platformsage 1058h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-28 14:00⭐ origin echo-reconstructedThis Microsoft article is the primary source. It presents Julia Kordick’s explicitly opinionated verification framework for agentic-AI outpu
Microsoft on blog (echo) · attributed from hn.story.49497440
—
08-30 10:36first on hacker news · published · +44.6hOnly believe what you can validate: a verification framework for agentic AI
yani__
—
08-30 14:05first on r/ClaudeAI · published · +48.1hWorkflow: stopping Claude Code from reporting a lint pass it did not actually run
Goldziher
—
09-18 04:58first on r/MachineLearning · published · +495.0hA pre-execution gate that stops LLMs from filling in missing tool arguments [P]
Jay299792458
—
08-30 10:36amplified on hacker newshn.story.49497440
yani__
peak 2 · 0 comments · 0% of case engagement
08-30 14:05amplified on r/ClaudeAIreddit.post.1w2ihiz
Goldziher
peak 1 · 4 comments · 0% of case engagement
08-31 02:44amplified on hacker newshn.story.49505043
Mkld27
peak 12 · 8 comments · 3% of case engagement
08-31 14:05amplified on r/ClaudeAIreddit.post.1w3etje
Massive-Zucchini2560
peak 1 · 12 comments · 1% of case engagement
08-31 21:55amplified on r/ClaudeAIreddit.post.1w3sdu7
gryswynd
peak 1 · 1 comments · 0% of case engagement
09-01 12:00amplified on r/ClaudeAIreddit.post.1w49vmz
Jazzlike-Echidna-670
peak 0 · 1 comments · 0% of case engagement
17 more amplifiers in ainews.case_chain
08-30 11:20our radar first saw it · +45.4hdiscovery anchor: hn.story.49497440—
pace: p89 vs 519 stories at the 720h mark (now 1058h old) — ahead of docs-first-agent-continuity-protocol (1.0x), behind amd-threadripper-halo-station (1.0x)

Evidence (24) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnOnly believe what you can validate: a verification framework for agentic AIyani__20
🟧 echo.blog ⭐This Microsoft article is the primary source. It presents Julia Kordick’s explicitly opinionated verification framework for agentic-AI outpuMicrosoft——
🟠 redditWorkflow: stopping Claude Code from reporting a lint pass it did not actually run
ClaudeAI
Goldziher14
🟧 hnShow HN: Prove your code produced your claims without making reviewers rerun itMkld27128
🟠 redditresearch-graph: a CLI that checks whether a multi-agent run actually held together (MIT, built with Claude Code)
ClaudeAI
Massive-Zucchini2560012
🟠 redditLaunched an app today where Claude is the content engine: Opus writes daily Japanese word puzzles, Sonnet adversarially reviews them
ClaudeAI
gryswynd11
🟠 redditClaude builds videos in my editor over MCP, and the fix that made it work was moving one rule into a tool description
ClaudeAI
Jazzlike-Echidna-67001
🟧 hnShow HN: MC/DC coverage so coding agents could work overnightNedomas10
🟠 redditI measured how often Claude Code told me "tests pass" on evidence that was already stale. 26 percent. So I wrote a hook that checks.
ClaudeAI
SmiLePLSSS67
🟠 redditI spent months getting my Claude Code pipeline to stop grading its own homework, then found the plugin I’d already published had the same flaw
ClaudeAI
mshadmanrahman06
🟧 hnShow HN: Local audit trail for Claude Code tool callszaghaghi10
🟧 hnPublic beta: a decision-governance runtime for AI agentslca1310
🟠 redditOne Claude Code feature I was underusing: hooks
ClaudeAI
Pretend_Sell659235069
🟧 hnHow well do agents use test/verification techniques?vinhnx19169
🟠 redditWatch Skill: let Claude inspect a screen recording and check whether a fix worked
ClaudeAI
Fearless-Role-270711
🟧 hnShow HN: MaruCheck – Independent QA for AI-generated codeKidusMT30
🟧 hnI stopped reading Cursor's diffs, so I built a gate that refuses themdarylprodev10
🟧 hnGuardrail Your Agents with Leanbarthelomew50
🟧 hnShow HN: EthersFlow-Adversarial Multi-Model Consensus Gate for AI Agent ActionEthersFlow10
🟧 hnShow HN: Weftgate, a local verification gate for coding agentsAvinashamudala10
🟠 redditmy ai coding agent said something was "tested" and it very much was not
ClaudeAI
chandra_kumar020
🟧 hnShow HN: Pinocchio: Harness for Verifiable Worklazaruslong10
🟧 hnShow HN: ClientCoded – QA Platform for AI Agentstravishcronin10
🟠 redditA pre-execution gate that stops LLMs from filling in missing tool arguments [P]
MachineLearning
Jay29979245801

Interpretation history

Decision trace