Independent evaluation will determine whether Donely AI's autonomous security agent can obtain root access on realistic targets with approximately 90% success.
state: expiredheat: lowuncertainty: highknownscott: lowagentic-security autonomous-hacking security-harnessesDonely AI
What is this?
Donely AI is presented as the company behind a security-agent harness that it claims can obtain root access in nine out of ten attempts. The supplied search snippets discuss general agent autonomy, evaluation, identity, and access controls, but provide no direct information about Donely AI, its harness, target realism, testing protocol, or results. Despite the web answer’s assertion, the supplied evidence does not establish that an independent evaluation occurred or verified an approximately 90% success rate.
Why it matters to Scott
This is already covered by Scott’s Capability Audit position and the radar’s open ExploitGym agent-exploitation validation case: security-agent success claims require representative targets, calibrated controls, reproducible traces, and independently checkable attack paths. Donely AI’s unverified “9/10 root” claim adds no results or methodology that would yet change what Scott builds or argues.
ip:concept.capability-auditip:source.security-reviewer-method-ebookradar:exploitgym-agent-exploitation-validationradar:concept.agent-evaluationradar:concept.autonomous-hacking
queries asked of Scott's wikis
- autonomous security-agent evaluation harnesses
- agentic coding harnesses for offensive security
- benchmark realism and reproducibility for autonomous agents
- permission boundaries and sandboxing for tool-using agents
- autonomy levels and human assistance in agent evaluations
- dual-use risks of autonomous hacking agents
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-22T09:25:21Z
The claim has produced no methodology, reproducible evidence, independent evaluation, or follow-on interest within its observation horizon, and no confirming event is currently expected. It remains an unvalidated first-party capability assertion, so continued active tracking is not justified.
2026-08-20T08:36:41Z
No new methodology, independent evaluation, reproducible traces, or target details have appeared; the case remains a single-source capability claim, and unchanged engagement adds no substance.
2026-08-20T08:32:59Z
grounded: known/low — This is already covered by Scott’s Capability Audit position and the radar’s open ExploitGym agent-exploitation validation case: security-agent success claims r
2026-08-20T08:30:31Z
case created — The consequential and readily testable root-compromise claim warrants tracking despite currently resting on a single first-party report.
Decision trace
- 08-22 19:25expireThe claim has produced no methodology, reproducible evidence, independent evaluation, or follow-on interest within its observation horizon, and no confirming event is currently expected. It remains an
- 08-22 19:25alert_silentThe only new delta is elapsed staleness with unchanged evidence and engagement; there is no event or validation Scott needs before the next briefing.
- 08-22 19:25alert_routeThe only new delta is elapsed staleness with unchanged evidence and engagement; there is no event or validation Scott needs before the next briefing.
- 08-20 18:36repriceNo new methodology, independent evaluation, reproducible traces, or target details have appeared; the case remains a single-source capability claim, and unchanged engagement adds no substance.
- 08-20 18:36alert_silentThe reobservation contains no consequential delta and does not validate the claimed root-compromise rate; it can wait for independent results or inspectable methodology.
- 08-20 18:36alert_routeThe reobservation contains no consequential delta and does not validate the claimed root-compromise rate; it can wait for independent results or inspectable methodology.
- 08-20 18:33alert_silentDonely AI’s self-reported “9/10 root” result provides no target set, controls, assistance level, reproducible traces, or independent evaluation. It is an unvalidated capability claim rather than a con
- 08-20 18:33alert_routeDonely AI’s self-reported “9/10 root” result provides no target set, controls, assistance level, reproducible traces, or independent evaluation. It is an unvalidated capability claim rather than a con
- 08-20 18:32groundThis is already covered by Scott’s Capability Audit position and the radar’s open ExploitGym agent-exploitation validation case: security-agent success claims require representative targets, calibrate
- 08-20 18:30createThe consequential and readily testable root-compromise claim warrants tracking despite currently resting on a single first-party report.