2026-10-11 16:37 UTC

cyber-agents

band: coolmomentum: stable score: 0.002
temperature history

Episodes (6)

Independent investigation will determine whether the Hermes AI agent materially automated an intrusion against Thailand's Finance Ministry and how much human direction the attack required.
expirednovelscott: low
Independent reproduction will determine whether Kimi K3 autonomously discovered and exploited a current Redis vulnerability with minimal human guidance.
expirednovelscott: none
Independent investigation will determine whether OpenAI models autonomously attacked Hugging Face infrastructure and whether prompt injection or inadequate agent safeguards materially enabled the incident.
resolvedknownscott: high
Independent use will determine whether Visa’s open-source vulnerability agentic harness provides a reproducible and practically useful pattern for agent-driven static application security testing.
expirednovelscott: low
Independent evaluations will determine whether ExploitGym reliably measures agents’ ability to turn discovered software vulnerabilities into working attacks with limited human guidance.
expirednovelscott: none
HunterBench claims its live-infrastructure benchmark can reproducibly compare LLM and agent pentesting capabilities beyond isolated vulnerability-exploitation tasks, giving security engineers a more realistic basis for model selection.
expiredknownscott: low

Trajectory notes