2026-10-11 17:21 UTC

Independent evaluation will determine whether Orivael’s non-LLM reasoning system can reproduce its perfect ft09 result and generalize across additional ARC-AGI-3 task families.

state: expiredheat: lowuncertainty: highknownscott: lowagent-benchmarks symbolic-reasoning arc-agiOrivael

What is this?

Orivael claims a non-LLM reasoning system achieved 100% on the ARC-AGI-3 ft09 task without any model calls, with independent reproduction and broader generalization still unestablished. ARC-AGI-3 is an interactive benchmark that tests agents’ ability to explore unfamiliar environments, infer goals and mechanics, form world models, and plan efficiently through action and feedback. The supplied web results describe the benchmark and its competition infrastructure, but they do not independently document Orivael’s system, methods, result, or planned evaluation.

Why it matters to Scott

The radar already tracks this same ARC-AGI-3 claim-validation pattern in the Schema and Seed IQ cases, including the need for independent reproduction and cross-environment generalization. It aligns with Scott’s evidence-ceiling and replayable-evaluation doctrines, but without methods, evaluator results, or broader task performance, it is currently another unverified benchmark claim rather than something that changes what he builds or argues.
ip:framework.discussed-is-not-deployedip:framework.challenger-never-arbiterradar:schema-arc-agi-3-claimradar:seed-iq-arc-agi-3d-doomradar:concept.arc-agiradar:concept.benchmark-integrityradar:concept.agent-benchmarks
queries asked of Scott's wikis
  • non-LLM agent architectures and symbolic reasoning
  • benchmark reproduction and independent evaluation
  • ARC-style abstraction and interactive reasoning
  • agent world models from action and feedback
  • benchmark overfitting versus cross-task generalization
  • model-free agents and inference economics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐We got 100% on ARC-AGI-3 ft09 with zero model calls. The failures are more interesting.
artificial
Living_Substance1274013

Interpretation history

Decision trace