Anubis maintainer robbe1912 claims roughly 100 agent-hours exposed enough failure modes to make its coding-agent hallucination detector unreliable as an execution gate, suggesting detector-only safeguards are brittle.
state: expiredheat: lowuncertainty: highconvergesscott: mediumcoding-agents agent-evaluation hallucination-detectionrobbe1912Anubis
What is this?
The case alleges that an Anubis maintainer, robbe1912, tested a hallucination detector for coding agents for roughly 100 agent-hours and found enough failure modes that it could not reliably serve as an execution gate. The supplied search snippets do not directly substantiate the named maintainer, test duration, detector design, or results; most refer to unrelated projects also called Anubis or to coding-agent risk generally. The strongest broader support is only that executed code can be checked empirically and that agent failures may have serious consequences, so the central claim remains thinly grounded here.
Why it matters to Scott
The claimed 100-agent-hour failure study directly converges with Scott’s position that probabilistic detectors are defence-in-depth, not binding execution gates, and supports his use of deterministic controls and mechanically different verifiers. It could provide a useful empirical receipt for his agent-control architectures, but the supplied evidence does not substantiate the experiment strongly enough for high relevance.
ip:concept.guardrail-illusionip:framework.architecture-not-vibesip:concept.mechanically-different-verifiersip:concept.deterministic-coredev:concept.deterministic-agent-control-planeradar:concept.agent-evaluationradar:concept.agent-verificationradar:opencode-guardians-tool-call-verificationradar:tracelint-deterministic-agent-trace-checks
queries asked of Scott's wikis
- coding-agent execution gates and layered safeguards
- detector-only guardrails versus deterministic validation
- coding-agent hallucination evaluation over long horizons
- fail-closed harness design for autonomous coding agents
- tests sandboxes permissions and rollback as agent controls
- false positives and false negatives in agent evaluators
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-09-05T18:32:15Z
After the 48-hour review horizon, the evidence remains a title-level negative-result claim, with no detector design, failure examples, artifact, or independent evaluation supplied. This has not matured into an empirical receipt for Scott’s execution-gate decisions; expire it without treating the underlying claim as disproved.
2026-09-03T18:00:56Z
No methodological detail, artifact, or independent corroboration has emerged; the negligible engagement change does not strengthen the title-level self-report. The case remains a potentially useful anecdote rather than an empirical receipt against detector-only execution gates.
2026-09-03T17:53:09Z
grounded: converges/medium — The claimed 100-agent-hour failure study directly converges with Scott’s position that probabilistic detectors are defence-in-depth, not binding execution gates
2026-09-03T17:49:51Z
case created — The released artifact and reported negative result form a specific engineering episode about the limits of coding-agent safety gates.
Decision trace
- 09-06 04:32expireAfter the 48-hour review horizon, the evidence remains a title-level negative-result claim, with no detector design, failure examples, artifact, or independent evaluation supplied. This has not mature
- 09-06 04:32alert_silentThere is no new substantive delta or expected confirming event that warrants Scott’s attention. Concrete failure data or a reproducible evaluation could reopen the case; the engagement change does not
- 09-06 04:32alert_routeThere is no new substantive delta or expected confirming event that warrants Scott’s attention. Concrete failure data or a reproducible evaluation could reopen the case; the engagement change does not
- 09-04 04:00repriceNo methodological detail, artifact, or independent corroboration has emerged; the negligible engagement change does not strengthen the title-level self-report. The case remains a potentially useful an
- 09-04 04:00alert_silentThere is no consequential new delta to surface; wait for concrete failure data, a reproducible artifact, or independent evaluation.
- 09-04 04:00alert_routeThere is no consequential new delta to surface; wait for concrete failure data, a reproducible artifact, or independent evaluation.
- 09-04 03:54alert_silentThe only visible evidence is the maintainer’s title-level self-report; it provides no detector design, failure modes, measurements, methodology, or reproducible artifacts. The claimed lesson is releva
- 09-04 03:54surface_candidateThe only visible evidence is the maintainer’s title-level self-report; it provides no detector design, failure modes, measurements, methodology, or reproducible artifacts. The claimed lesson is releva
- 09-04 03:54alert_routeThe only visible evidence is the maintainer’s title-level self-report; it provides no detector design, failure modes, measurements, methodology, or reproducible artifacts. The claimed lesson is releva
- 09-04 03:53groundThe claimed 100-agent-hour failure study directly converges with Scott’s position that probabilistic detectors are defence-in-depth, not binding execution gates, and supports his use of deterministic
- 09-04 03:49createThe released artifact and reported negative result form a specific engineering episode about the limits of coding-agent safety gates.