Multinet AI claims current frontier reasoning agents systematically fail interactive 2D mazes, revealing a material gap in spatial planning and tool-mediated environment control.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-evaluation reasoning-agents interactive-environmentsMultinet AI
What is this?
Multinet AI’s MultiNet 2.0 Preview evaluates frontier reasoning agents in interactive 2D maze environments. It reports that Claude Opus 4.8, Kimi K2.6, and Qwen 3.6-27B solved only 6 of 150 runs across 50 mazes, with 45 mazes unsolved by every model, which Multinet presents as evidence that performance degrades over interactive task horizons. The supplied snippets do not establish the evaluation methodology, scaffolding, controls, or whether the failures specifically arise from spatial planning, visual perception, tool use, latency, or another component.
Why it matters to Scott
The claimed collapse on interactive horizons converges with Scott’s view that agent capability belongs to the model-plus-harness system, including its hands, eyes, state, and feedback loop—not the weights alone. It creates a concrete benchmark-analysis opportunity, but the missing harness, trace, and control details prevent attributing the failures specifically to spatial reasoning rather than perception or tool control; the radar already tracks the closely related ActiveVision gap, though not this same evaluation.
ip:concept.model-plus-harness-benchmark-unitip:concept.agent-hands-and-eyesip:concept.perceptual-engineeringdev:concept.trace-backed-agent-comparisonradar:activevision-repeated-perception-gap
queries asked of Scott's wikis
- interactive environment benchmarks for coding agents
- agent harness failures over long task horizons
- spatial planning versus tool-control failures
- closed-loop agent evaluation and observability
- benchmark validity for scaffolded agents
- state tracking and recovery in interactive agents
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-28T15:38:37Z
The claim has gone stale without methodology, traces, replication, or independent corroboration, leaving its causal attribution unresolved. The episode can fade unless substantive benchmark materials or a replication later reopen it.
2026-08-26T14:39:20Z
No new methodology, traces, replication, or independent evidence has appeared, so the benchmark claim remains an uncorroborated lead rather than a demonstrated frontier-agent limitation. Unchanged engagement adds no urgency.
2026-08-26T14:35:07Z
grounded: converges/medium — The claimed collapse on interactive horizons converges with Scott’s view that agent capability belongs to the model-plus-harness system, including its hands, ey
2026-08-26T14:33:35Z
case created — The claim identifies a specific, reproducible capability gap with implications for evaluating interactive and long-horizon agents.
Decision trace
- 08-29 01:38expireThe claim has gone stale without methodology, traces, replication, or independent corroboration, leaving its causal attribution unresolved. The episode can fade unless substantive benchmark materials
- 08-29 01:38alert_silentThe only delta is a staleness trigger and negligible engagement movement; no new fact changes the benchmark’s credibility or implications for Scott.
- 08-29 01:38alert_routeThe only delta is a staleness trigger and negligible engagement movement; no new fact changes the benchmark’s credibility or implications for Scott.
- 08-27 00:39repriceNo new methodology, traces, replication, or independent evidence has appeared, so the benchmark claim remains an uncorroborated lead rather than a demonstrated frontier-agent limitation. Unchanged eng
- 08-27 00:39alert_silentThis is only a legacy-state re-evaluation with no substantive delta; the unresolved attribution between planning, perception, and tool-control failures can wait for a normal briefing or methodological
- 08-27 00:39alert_routeThis is only a legacy-state re-evaluation with no substantive delta; the unresolved attribution between planning, perception, and tool-control failures can wait for a normal briefing or methodological
- 08-27 00:35alert_silentThe first-party claim is relevant to agent-harness evaluation, but the visible evidence provides no methodology, traces, model list, controls, or results separating spatial reasoning failures from per
- 08-27 00:35surface_candidateThe first-party claim is relevant to agent-harness evaluation, but the visible evidence provides no methodology, traces, model list, controls, or results separating spatial reasoning failures from per
- 08-27 00:35alert_routeThe first-party claim is relevant to agent-harness evaluation, but the visible evidence provides no methodology, traces, model list, controls, or results separating spatial reasoning failures from per
- 08-27 00:35groundThe claimed collapse on interactive horizons converges with Scott’s view that agent capability belongs to the model-plus-harness system, including its hands, eyes, state, and feedback loop—not the wei
- 08-27 00:33createThe claim identifies a specific, reproducible capability gap with implications for evaluating interactive and long-horizon agents.