2026-10-11 17:11 UTC

Multinet AI claims current frontier reasoning agents systematically fail interactive 2D mazes, revealing a material gap in spatial planning and tool-mediated environment control.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-evaluation reasoning-agents interactive-environmentsMultinet AI

What is this?

Multinet AI’s MultiNet 2.0 Preview evaluates frontier reasoning agents in interactive 2D maze environments. It reports that Claude Opus 4.8, Kimi K2.6, and Qwen 3.6-27B solved only 6 of 150 runs across 50 mazes, with 45 mazes unsolved by every model, which Multinet presents as evidence that performance degrades over interactive task horizons. The supplied snippets do not establish the evaluation methodology, scaffolding, controls, or whether the failures specifically arise from spatial planning, visual perception, tool use, latency, or another component.

Why it matters to Scott

The claimed collapse on interactive horizons converges with Scott’s view that agent capability belongs to the model-plus-harness system, including its hands, eyes, state, and feedback loop—not the weights alone. It creates a concrete benchmark-analysis opportunity, but the missing harness, trace, and control details prevent attributing the failures specifically to spatial reasoning rather than perception or tool control; the radar already tracks the closely related ActiveVision gap, though not this same evaluation.
ip:concept.model-plus-harness-benchmark-unitip:concept.agent-hands-and-eyesip:concept.perceptual-engineeringdev:concept.trace-backed-agent-comparisonradar:activevision-repeated-perception-gap
queries asked of Scott's wikis
  • interactive environment benchmarks for coding agents
  • agent harness failures over long task horizons
  • spatial planning versus tool-control failures
  • closed-loop agent evaluation and observability
  • benchmark validity for scaffolded agents
  • state tracking and recovery in interactive agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnFrontier Reasoning Agents Fail on Interactive 2D Mazesjiggle123112
🟧 echo.blog ⭐Multinet AI reports that frontier reasoning agents fail on interactive 2D maze tasks.Multinet AI——

Interpretation history

Decision trace