Simular claims Sai tops OSWorld 2.0 against leading computer-use models while operating at roughly two-thirds their cost, establishing a potentially stronger cost-quality frontier for computer-use agents.
state: expiredheat: lowuncertainty: highknownscott: lowcomputer-use-agents agent-evaluation inference-economicsSimularSai
What is this?
Simular describes itself as an autonomous-computer company and says its open agentic framework, Agent S, achieved a 72.6% success rate on the original OSWorld benchmark. The supplied evidence titles attribute a separate claim to Simular that a system called Sai scored 73% on the long-horizon OSWorld 2.0 benchmark at about two-thirds the cost of GPT-5. However, the provided public leaderboard snippets do not list Sai and show different leaders and scores, so neither the top-ranking claim nor the cost comparison is independently established by these materials.
Why it matters to Scott
Scott already treats agent performance as a model-plus-harness property and evaluates cost against validated capability, as captured in Model-Plus-Harness Benchmark Unit and Inference-Time Scaling; the radar also already tracks computer-use evaluation and inference economics. Sai is therefore an unverified new datapoint in an established territory, not yet evidence that would change his designs or claims because neither its leaderboard position nor cost comparison is independently established.
ip:concept.model-plus-harness-benchmark-unitip:concept.inference-time-scalingdev:concept.trace-backed-agent-comparisonradar:concept.computer-use-agentsradar:concept.agent-evaluationradar:concept.inference-economicsradar:concept.benchmark-integrity
queries asked of Scott's wikis
- computer-use agent reliability and long-horizon task execution
- benchmark validity for coding and GUI agents
- cost-quality frontiers in agent inference
- agent harness versus base-model performance
- economics of retries, planning, and tool-use loops
- open agent frameworks for desktop automation
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-31T23:35:48Z
Repeated observation windows have produced no leaderboard inclusion, methodology, or independent validation, so this is no longer an active developing episode. Archive the self-reported claim and reopen only if substantive verification appears.
2026-08-29T22:38:04Z
The claim remains an isolated, self-reported benchmark and cost result; another observation window produced no official leaderboard entry, methodology, or independent validation. Keep the case dormant pending a substantive verification event rather than continued routine attention.
2026-08-27T21:41:42Z
No new validation or discussion has appeared; Sai remains a first-party benchmark and cost claim awaiting official leaderboard inclusion and methodological scrutiny. The unchanged reobservation lowers immediate attention temperature without weakening or confirming the underlying hypothesis.
2026-08-27T21:39:25Z
grounded: known/low — Scott already treats agent performance as a model-plus-harness property and evaluates cost against validated capability, as captured in Model-Plus-Harness Bench
2026-08-27T21:36:40Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49471349 -> echo.blog.33b588fefb by Simular Team
2026-08-27T21:34:50Z
case created — The first-party benchmark claim combines a reported capability lead with a material cost advantage relevant to production agent economics.
Decision trace
- 09-01 09:35expireRepeated observation windows have produced no leaderboard inclusion, methodology, or independent validation, so this is no longer an active developing episode. Archive the self-reported claim and reop
- 09-01 09:35alert_silentThe staleness trigger adds no consequential evidence, and the isolated claim still does not affect Scott’s designs or conclusions. Any future official listing, reproducible methodology, or independent
- 09-01 09:35alert_routeThe staleness trigger adds no consequential evidence, and the isolated claim still does not affect Scott’s designs or conclusions. Any future official listing, reproducible methodology, or independent
- 08-30 08:38repriceThe claim remains an isolated, self-reported benchmark and cost result; another observation window produced no official leaderboard entry, methodology, or independent validation. Keep the case dormant
- 08-30 08:38alert_silentNo consequential new delta occurred; the 48-hour staleness trigger adds no evidence. Wait for official OSWorld 2.0 inclusion, reproducible evaluation details, or independent cost validation.
- 08-30 08:38alert_routeNo consequential new delta occurred; the 48-hour staleness trigger adds no evidence. Wait for official OSWorld 2.0 inclusion, reproducible evaluation details, or independent cost validation.
- 08-28 07:41repriceNo new validation or discussion has appeared; Sai remains a first-party benchmark and cost claim awaiting official leaderboard inclusion and methodological scrutiny. The unchanged reobservation lowers
- 08-28 07:41alert_silentThis is only a legacy-state re-evaluation with no consequential new delta. Wait for official OSWorld 2.0 listing, reproducible evaluation details, or independent cost validation.
- 08-28 07:41alert_routeThis is only a legacy-state re-evaluation with no consequential new delta. Wait for official OSWorld 2.0 listing, reproducible evaluation details, or independent cost validation.
- 08-28 07:39alert_silentSimular’s announcement establishes the benchmark and cost claim as a new self-reported datapoint, but the result is still awaiting official leaderboard inclusion and lacks enough cost-methodology or i
- 08-28 07:39surface_candidateSimular’s announcement establishes the benchmark and cost claim as a new self-reported datapoint, but the result is still awaiting official leaderboard inclusion and lacks enough cost-methodology or i
- 08-28 07:39alert_routeSimular’s announcement establishes the benchmark and cost claim as a new self-reported datapoint, but the result is still awaiting official leaderboard inclusion and lacks enough cost-methodology or i
- 08-28 07:39groundScott already treats agent performance as a model-plus-harness property and evaluates cost against validated capability, as captured in Model-Plus-Harness Benchmark Unit and Inference-Time Scaling; th
- 08-28 07:36promote_anchororigin walk conf 0.98
- 08-28 07:34createThe first-party benchmark claim combines a reported capability lead with a material cost advantage relevant to production agent economics.