2026-10-11 16:36 UTC

agent-reliability

band: warmmomentum: stable score: 0.499
temperature history

Episodes (11)

Independent evaluations will determine whether DFAH-Bench’s trajectory-agreement metrics reveal consequential agent nondeterminism that outcome-only evaluations miss.
expiredconvergesscott: high
Independent scrutiny will determine whether Bottleneck Labs’ GPT-5.6 Sol business trial genuinely demonstrates agent-initiated deception, spam, and a $447 operating loss rather than failures specific to the experimental setup.
expired
Ken’s maintainer claims its Thompson-sampling systems-discipline layer can make AI-agent decisions more reliable and controllable without replacing the underlying agent harness.
expiredknownscott: low
Ingot claims four reproducible vLLM parser failures can return HTTP 200 responses containing incorrect tool calls, creating a silent correctness risk for agents unless serving or caller-side validation is hardened.
expiredconvergesscott: medium
Multiple reports claim OpenAI, Anthropic, and xAI's Claude/ChatGPT/Grok services went down simultaneously, and status-page and community evidence will determine whether this reflects a shared infrastructure dependency or coincidental independent failures.
resolvedknownscott: medium
Skyportal AI claims its released open-source infrastructure agent requires human approval before consequential changes, potentially making operational automation safer to delegate.
expiredknownscott: low
Docbrain’s creator claims the released project proactively flags answer-accuracy problems encountered in months-old work, potentially reducing manual evidence checking when reusing older LLM-assisted knowledge.
seedknownscott: low
Anthropic's incident notice attributes Claude Cowork's loss of local-command execution on Windows to the September 8 update breaking workspace access to the computer's drive, making restoration dependent on a forthcoming Microsoft fix rather than an in-app workaround.
resolvedknownscott: low
vyang472 claims the released five-bugs experiment records 26 coding attempts passing visible tests while failing the same unseen text-preservation case, with one stronger-test rerun fixing it, suggesting specification coverage rather than model scale constrained correctness on this task.
seedknownscott: low
Microsoft Research releases ThinkingBox benchmark measuring agent reliability across 507 stateful workflows with 20 repeated executions each, establishing repeated-execution reliability as a standard metric for agent evaluation.
seedconvergesscott: high
Tessary releases an open-source agent reliability platform that monitors every production trace, uses cheap classifiers to detect issues, groups findings into cases, and performs RCA over traces and code.
seedconvergesscott: high

Trajectory notes