2026-10-11 16:37 UTC

Ruiyang Wang and coauthors claim GAVEL's explicit graph world model raises Qwen3-8B task success on BEHAVIOR-1K from 41.2% to 91.8% for single tasks and 19.9% to 92.6% for multi-task instructions, potentially making compact-model embodied planning reliable through external verification and repair.

state: seedheat: mediumuncertainty: mediumconvergesscott: highagent-planning world-models agent-evaluationRuiyang WangMiroslav Pajic

What is this?

GAVEL (Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning) is a 2026 arXiv preprint (arXiv:2609.19315) by Ruiyang Wang, Miroslav Pajic, and colleagues. The paper introduces an explicit graph world model that wraps an LLM (tested with Qwen3-8B) to maintain object relations (inside, on top, near, holding) and state flags, enabling external verification and repair of generated plans. On the BEHAVIOR-1K embodied benchmark, they report dramatic gains: single-task success from 41.2% to 91.8%, and multi-task from 19.9% to 92.6%. The result positions graph-structured world models as a practical verification layer that makes compact-model embodied planning reliable.

Why it matters to Scott

GAVEL is a direct, benchmarked instantiation of Scott's core architectural thesis: an explicit graph world model acts as an independent verification layer outside the LLM (Qwen3-8B), turning a compact open model into a reliable embodied planner via external verification and repair — exactly the 'two-leashes' / 'verification loops' / 'executable worldview' pattern he argues for. The 41%→92% and 20%→93% jumps on BEHAVIOR-1K validate the model-barbell claim (cheap model + verification harness beats raw model power) and the usable-mass-over-unusable-power principle. This is not merely an example of his pattern; it is a consequential independent arrival that creates a dated-receipts publishing opportunity.
ip:concept.verification-loopsip:framework.executable-worldviewip:framework.decision-authority-infrastructureip:framework.two-leashesip:concept.guardrail-illusionip:concept.zero-trust-for-decisionsip:framework.siloosip:concept.world-loop-closureip:concept.runtime-governanceip:concept.model-barbellip:concept.usable-mass-over-unusable-powerip:concept.model-plus-harness-benchmark-unitip:concept.durable-external-stateip:framework.long-running-agentsip:framework.agent-provenance-stackip:concept.cognitive-provenanceip:concept.forensic-artefact-problemip:concept.deterministic-agent-control-planeip:concept.validation-gated-llm-extractionip:concept.propose-finalise-gateip:concept.claim-bounded-adversarial-verificationip:framework.agent-loopip:concept.high-not-maxip:concept.model-dividendip:dev:technology.ollamaip:dev:project.gamepcip:dev:concept.hierarchical-task-decompositionip:dev:concept.dialectical-tree-searchip:dev:concept.metacognitive-resolution-controlip:dev:concept.progressive-resolution-orchestrationip:dev:concept.cheap-model-front-doorip:dev:concept.hardware-aware-local-inferenceip:dev:concept.cost-tiered-llm-routingradar:grapharc-runtime-agent-graph-gatesradar:certora-autoprover-agent-verificationradar:procedural-graphs-agent-executionradar:aph-agent-notarization-protocolradar:qwen38-27b-local-agent-capabilityradar:civbench-long-horizon-planningradar:laya-vega-decision-model-releaseradar:longhorizon-harness-validationradar:intern-decision-one-pass-decisionsradar:canary-agent-code-verificationradar:world-labs-atlas-spatial-modelradar:world-model-optimizer-agent-routing
queries asked of Scott's wikis
  • graph world models explicit verification repair planning
  • compact model embodied planning reliability 8B 7B
  • external verification layer LLM planning agent
  • BEHAVIOR-1K benchmark embodied agent evaluation
  • Qwen3 local open model planning performance
  • long-horizon task planning LLM world model

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady1 platformsage 502h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-20 18:04⭐ origin directly observedGraph World Models for Verified and Efficient Long-Horizon LLM Task Planning
gmays on hacker news
—
09-20 18:04amplified on hacker news 👑hn.story.49778275
gmays
peak 1 · 0 comments · 106% of case engagement
09-20 18:20our radar first saw it · +0.3hdiscovery anchor: hn.story.49778275—
pace: p9 vs 1032 stories at the 336h mark (now 502h old) — behind addom-local-coding-harness (0.5x)

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
Retrieved article excerpt

Open article · Retrieved 2026-09-20T18:22:54.954878+00:00

Agents · Robotics

# GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic

[Chat with Paper](https://academy.dair.ai/dashboard/paper-chat/gavel-graph-world-models-for-verified-and-efficient-long-horizon-llm-task-planni-2609.19315)

First page

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

The curator’s take

Ruiyang Wang and colleagues present GAVEL, which verifies and repairs long-horizon LLM robot plans against an explicit graph world model holding object relations, action preconditions and effects, and probabilistic beliefs over unobserved locations.

## Ask this paper

Key points

01

Single-task success goes from 41.2 to 91.8 percent. With Qwen3-8B on BEHAVIOR-1K across 100 long-horizon tasks; multi-task success rises from 19.9 to 92.6 percent across 500 multi-task instructions.

02

The graph predicts consequences before execution. Violations are detected and repaired directly when the correction follows from the world model, and LLM replanning is reserved for errors that need semantic reasoning.

03

Belief reasoning reorders subtasks. Reasoning over distributions of possible object locations reduces expected search cost and cuts travel distance about 5.4 percent against a static variant.

04

The gain holds for a compact model. Most of the improvement comes from the harness rather than model capability, which is the argument for putting the world model outside the LLM.

Abstract

Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model. The graph represents relevant object-relations, action pre-conditions and effects, and probabilistic beliefs over unobserved object locations. This model can predict the consequences of LLM-generated actions before execution, detect violations, and repair those whose corrections follow directly from the world model. This method also reserves LLM replanning solely for errors requiring semantic reasoning. For multi-task instructions, GAVEL reasons over distributions of possible object locations to reorder remaining subtasks and minimize expected search cost. We evaluate GAVEL on BEHAVIOR-1K across 100 single long-horizon tasks and 500 multi-task instructions. With Qwen3-8B, GAVEL improves single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%. Distributional belief reasoning also reduces travel distance by approximately 5.4% compared with a static variant. These improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

Every Monday

##### Get next week’s papers.

[Subscribe on Substack](https://nlp.elvissaravia.com/subscribe)
gmays10

Interpretation history

Decision trace