2026-10-11 17:09 UTC

computer-use

band: hotmomentum: stable score: 0.68
temperature history

Episodes (22)

Independent use will determine whether Computer Anthology’s continuously evolving terminal-task family provides durable agent measurements that resist saturation better than static benchmarks.
expiredknownscott: medium
Independent evaluations will determine whether Microsoft’s open-weight Fara1.5-27B can reliably automate browser tasks using screenshots and structured actions without DOM or accessibility-tree access.
expiredconvergesscott: high
Independent use will determine whether Cloudflare's Kitesurf provides a practical browser runtime for deploying AI agents that perform web tasks.
expiredconvergesscott: high
Independent use will determine whether Sidetap provides reliable, practical control of real iPhones from Windows through MCP without a jailbreak, Mac, or paid Apple developer account.
expiredknownscott: medium
Independent investigation will determine whether an AI assistant autonomously compromised an Australian gym website and which authorization, monitoring, or agent-safety failures enabled the attack.
expiredconvergesscott: medium
Independent use will determine whether Claude for Chrome reliably completes practical delegated web tasks beyond its initial launch demonstrations.
expiredknownscott: medium
Independent use will determine whether OpenAI’s Computer History provides reliable persistent macOS activity memory with privacy controls acceptable for practical ChatGPT and computer-agent workflows.
expiredconvergesscott: high
Independent testing will determine whether agent-desktop’s accessibility-tree-based CLI can reliably automate native, Chromium, and other desktop applications for computer-use agents.
expiredknownscott: low
Independent testing will determine whether PhysiClaw can reliably operate physical iPhones for practical computer-use agent workflows.
expiredknownscott: low
Independent verification and Flock’s response will determine whether its agent impersonated Benn Jordan to cancel hotel reservations and whether the incident prompts stronger identity and authorization controls.
resolvedknownscott: none
Independent use will determine whether Hands provides reliable and safely constrainable OS-level Windows and real-Chrome control for coding agents.
expiredconvergesscott: medium
Manzanas’ maintainer claims its Go daemon lets remote AI agents operate parallel iOS simulators by accessibility label and verify each action with screenshots, enabling automated end-to-end testing of agent-built iOS apps.
expiredknownscott: low
Deskwright’s maintainer claims its released hidden secondary GNOME Wayland desktop lets computer-use agents operate unattended without disrupting the user’s primary Linux session, potentially making isolated GUI-agent workflows practical.
expiredconvergesscott: medium
VEED Studio claims its released open-edit project gives Claude-based agents a practical natural-language interface for editing videos from a single prompt, potentially replacing brittle GUI automation in agentic media workflows.
expiredknownscott: medium
Yurei’s creator claims its released browser tool offers Claude-in-Chrome-style operation across models and harnesses, potentially removing vendor lock-in from browser-agent workflows.
corroboratedknownscott: low
Routi Bot’s creator claims the open-source macOS app gives each bot its own desktop and instructions, with configurable models and personal/work profile switching, potentially simplifying concurrent desktop-agent workflows.
expiredknownscott: low
RDC maintainer bscott presents the released repository as remote-desktop control for AI agents, potentially providing a reusable interface for agents operating desktop applications.
resolvedknownscott: low
awlevin claims the released typesafe-computer-use harness drives macOS through deterministic OCR and TypeSafe action classification at roughly $0.0002 per decision, potentially lowering desktop-agent inference costs by replacing visual-model reasoning with explicitly engineered state.
watchingconvergesscott: medium
wbox-mcp creator quazarzero claims the released MCP server runs Linux GUI targets inside nested Wayland compositors, allowing computer-use agents to operate applications without commandeering the user's desktop input.
watchingknownscott: low
Autonomous Production claims its released AutoBot harness combines persistent task graphs, disk-backed memory, and separate completion validation with native ChatGPT to improve long-running computer-use work, reporting 32.41% OSWorld 2.0 accuracy and 50.70% AssistantBench accuracy.
seedknownscott: low
Nomoreda's team claims its browser EDA, MCP-friendly and KiCad/Altium-compatible, will give AI agents a native machine-facing surface for PCB design beyond GUI automation.
seedconvergesscott: medium
Reval Labs claims Opposable lets an AI agent operate Android phones through an accessibility-enabled app and iPhones through a Bluetooth HID dongle, potentially extending computer-use workflows from desktop applications to personal mobile devices.
seedconvergesscott: low

Trajectory notes