2026-10-11 17:12 UTC

URML-MARS claims its released URML harness can reproducibly evaluate safety failures in AI agents controlling laboratory and factory hardware, extending agent-security testing into consequential physical environments.

state: expiredheat: lowuncertainty: highknownscott: mediumphysical-ai agentic-security safety-evaluationURML-MARS

What is this?

URML-MARS is presented as the creator of a released agent-safety evaluation harness for laboratory and factory hardware. The claimed system maps model outputs to executable physical behaviors while using deterministic checks, reproducible conditions, metrics, and containment controls to identify unsafe or unpredictable actions. The supplied search results support the broader harness-engineering approach, but they do not directly document the URML repository, its maintainers, implementation, or demonstrated results, so the release’s specific capabilities remain only lightly substantiated here.

Why it matters to Scott

SiloOS and Decision Authority Infrastructure already establish Scott’s position that untrusted action-taking agents require structural containment and independent deterministic gates. URML’s claimed contribution is a potentially useful physical-hardware testbed for those architectures, but the supplied evidence does not yet establish implementation quality or results, and the radar already tracks adjacent physical-agent failures and reproducible agent-evaluation harnesses.
ip:framework.siloosip:framework.decision-authority-infrastructureip:concept.evaluation-driven-developmentdev:project.silo-osradar:kinetic-prompt-injection-physical-agentsradar:understudy-agent-scenario-testingradar:concept.agent-evaluationradar:concept.embodied-agents
queries asked of Scott's wikis
  • physical-agent safety and hardware control
  • agent harness containment and executable policy
  • reproducible evals for tool-using agents
  • deterministic checks for nondeterministic agents
  • sandboxing, interlocks, and fail-safe actions
  • physical-AI failure testing and observability

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: URML – safety-eval harness for AI agents on lab and factory hardwareYahalomi20
🟧 echo.github ⭐The repository releases a physical-AI safety-evaluation example for agents operating laboratory and factory hardware.URML-MARS——
🟧 hnPLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physicalsbulaev10

Interpretation history

Decision trace