2026-10-11 17:11 UTC

Ingot claims four reproducible vLLM parser failures can return HTTP 200 responses containing incorrect tool calls, creating a silent correctness risk for agents unless serving or caller-side validation is hardened.

state: expiredheat: lowuncertainty: highconvergesscott: mediumllm-serving tool-calling agent-reliability ai-infrastructureIngotvLLM

What is this?

Ingot reports four reproducible failure modes in vLLM’s tool-call parsers where an HTTP 200 response can contain an incorrectly extracted tool call, allowing an agent to treat a semantic failure as success. vLLM’s documentation confirms that its parser layer extracts tool calls from model-generated text, and a separate PydanticAI issue reports broken vLLM tool-calling behavior, but the supplied snippets do not independently document Ingot’s four reproductions. The search summary says the issue was fixed in vLLM 0.17.0, though the visible snippet mentioning that version concerns an unrelated SSRF vulnerability, so the fix claim is not established here.

Why it matters to Scott

The reported HTTP-success/semantic-failure mode independently supports Scott’s position that parsed tool calls require deterministic validation before dispatch, directly bearing on his validation-gated extraction and multi-format tool-call parsing work. It could justify new adversarial fixtures and fail-closed caller checks, but the four reproductions and claimed fix are not independently established by the supplied material.
ip:concept.verification-loopsip:concept.runtime-governancedev:concept.validation-gated-llm-extractiondev:concept.multi-format-tool-call-parsingdev:concept.deterministic-agent-control-planeradar:concept.tool-callingradar:concept.agent-reliabilityradar:concept.vllmradar:frontier-api-zero-output-voids
queries asked of Scott's wikis
  • semantic validation of agent tool calls
  • HTTP success versus agent task correctness
  • structured-output parser failure modes
  • caller-side validation for tool execution
  • agent harness invariants and fail-closed design
  • vLLM serving reliability and tool calling

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnFour reproducible vLLM parser failures that return 200 with the wrong tool callglitch00310
🟧 echo.blog ⭐A report describing four reproducible vLLM parser failures that return successful HTTP responses with the wrong tool call.Ingot——

Interpretation history

Decision trace