ClaudeAI user sebasmtl claims his open-source OpenPhysicsAI physics lab โ 13 flags scored against sealed, previously-unpredicted experimental measurements plus 4 starter trials โ lets anyone's AI compete as a solver, and real external solver attempts would establish it as a used machine-verified benchmark for AI scientific capability.
state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation ai-for-physics open-benchmarks ai-researchsebasmtlOpenPhysicsAI
What is this?
OpenPhysicsAI is an open-source 'physics lab' announced on a ClaudeAI community forum by user sebasmtl: 13 experimental flags scored against sealed, previously-unpredicted real measurements plus 4 starter trials, built so anyone's AI can compete as a solver, with machine-scoring against the sealed data serving as verification; the creator's stated aim is that genuine external solver attempts would establish it as a used, machine-verified benchmark for AI scientific capability. The supplied web results contain no independent coverage of OpenPhysicsAI or sebasmtl โ the case rests entirely on the creator's own announcement, and its (already small) engagement is unconfirmed. The snippets do show the surrounding territory is active: Apodex's TRACES benchmark grades AI solvers on unresolved research problems against hidden ground truth using outcome and process verifiers, and a recent arXiv benchmark anchors frontier coding agents against an external Connect Four solver โ so machine-scored, contamination-resistant scientific-capability evaluation is an emerging pattern, but whether this grassroots instance gains uptake is unknown.
Why it matters to Scott
An unknown grassroots builder independently assembled Scott's own evaluation canon in a new domain: sealed previously-unpredicted measurements are the Future-Leakage Rule (freeze the world before the window) and Hidden Gates held-out-gate discipline applied to AI-for-physics, and 'will external solvers come?' is exactly his publishing-as-active-sensor uptake question โ a dated receipt that blind-machine-scored evaluation is emerging organically. The sharpest angle for him: by his own model-plus-harness-benchmark-unit doctrine, 'whose AI can beat it' is underspecified โ unless the solver harness is pinned, the leaderboard scores harness variance, not model capability, a confound he could engage, test with his own blind-review harness patterns, or publish against (and it answers the retained-question leakage worry in the Sous physics-regrading episode). Medium rather than high because the entity is one zero-engagement forum post with no independent coverage; the convergence receipt is real but its consequence depends on uptake that hasn't happened.
ip:framework.hidden-gates-frameworkip:concept.future-leakage-ruleip:concept.model-plus-harness-benchmark-unitip:framework.publishing-is-an-active-sensordev:concept.blind-workbook-to-wiki-reconciliationradar:concept.benchmark-integrityradar:concept.ai-for-scienceradar:sous-physics-benchmark-regradingradar:ship-harness-benchradar:concept.agent-benchmarks
queries asked of Scott's wikis
- machine-scored benchmark sealed hidden ground truth
- eval contamination training-data leakage prevention position
- agent harness as eval unit model versus tools-and-loop
- open benchmark cold-start adoption community uptake
- AI-for-science physics solver benchmark territory
- verifier gaming reward hacking machine scoring
Measured heat
now 0 pts/hpeak 1 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 362h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p32 vs 1032 stories at the 336h mark (now 362h old) โ ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-09-27T21:38:07Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wruzwt -> echo.github.84b070ad35 by molanocortes (GitHub) = u/sebasmtl (Reddit)
2026-09-27T21:32:52Z
grounded: converges/medium โ An unknown grassroots builder independently assembled Scott's own evaluation canon in a new domain: sealed previously-unpredicted measurements are the Future-Le
2026-09-27T21:26:10Z
case created โ The creator's own announcement of a concrete, resolvable challenge (sealed real measurements, machine-scored) is a plausible new evaluation episode worth tracking for external uptake, despite currently tiny engagement.
Decision trace
- 09-28 07:38promote_anchororigin walk conf 0.85
- 09-28 07:32groundAn unknown grassroots builder independently assembled Scott's own evaluation canon in a new domain: sealed previously-unpredicted measurements are the Future-Leakage Rule (freeze the world before
- 09-28 07:26createThe creator's own announcement of a concrete, resolvable challenge (sealed real measurements, machine-scored) is a plausible new evaluation episode worth tracking for external uptake, despite cur