2026-10-11 17:15 UTC

Simular claims Sai tops OSWorld 2.0 against leading computer-use models while operating at roughly two-thirds their cost, establishing a potentially stronger cost-quality frontier for computer-use agents.

state: expiredheat: lowuncertainty: highknownscott: lowcomputer-use-agents agent-evaluation inference-economicsSimularSai

What is this?

Simular describes itself as an autonomous-computer company and says its open agentic framework, Agent S, achieved a 72.6% success rate on the original OSWorld benchmark. The supplied evidence titles attribute a separate claim to Simular that a system called Sai scored 73% on the long-horizon OSWorld 2.0 benchmark at about two-thirds the cost of GPT-5. However, the provided public leaderboard snippets do not list Sai and show different leaders and scores, so neither the top-ranking claim nor the cost comparison is independently established by these materials.

Why it matters to Scott

Scott already treats agent performance as a model-plus-harness property and evaluates cost against validated capability, as captured in Model-Plus-Harness Benchmark Unit and Inference-Time Scaling; the radar also already tracks computer-use evaluation and inference economics. Sai is therefore an unverified new datapoint in an established territory, not yet evidence that would change his designs or claims because neither its leaderboard position nor cost comparison is independently established.
ip:concept.model-plus-harness-benchmark-unitip:concept.inference-time-scalingdev:concept.trace-backed-agent-comparisonradar:concept.computer-use-agentsradar:concept.agent-evaluationradar:concept.inference-economicsradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • computer-use agent reliability and long-horizon task execution
  • benchmark validity for coding and GUI agents
  • cost-quality frontiers in agent inference
  • agent harness versus base-model performance
  • economics of retries, planning, and tool-use loops
  • open agent frameworks for desktop automation

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnSimular's Sai tops OSWorld 2.0, beats GPT and Opus at 2/3 the costtaro66620
🟧 echo.blog ⭐Simular’s own announcement says: “Sai ... has achieved a 73% success rate on OSWorld 2.0” and did so at “about two-thirds the cost” of GPT-5Simular Team——

Interpretation history

Decision trace