2026-10-11 16:37 UTC

ClaudeAI user sebasmtl claims his open-source OpenPhysicsAI physics lab โ€” 13 flags scored against sealed, previously-unpredicted experimental measurements plus 4 starter trials โ€” lets anyone's AI compete as a solver, and real external solver attempts would establish it as a used machine-verified benchmark for AI scientific capability.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation ai-for-physics open-benchmarks ai-researchsebasmtlOpenPhysicsAI

What is this?

OpenPhysicsAI is an open-source 'physics lab' announced on a ClaudeAI community forum by user sebasmtl: 13 experimental flags scored against sealed, previously-unpredicted real measurements plus 4 starter trials, built so anyone's AI can compete as a solver, with machine-scoring against the sealed data serving as verification; the creator's stated aim is that genuine external solver attempts would establish it as a used, machine-verified benchmark for AI scientific capability. The supplied web results contain no independent coverage of OpenPhysicsAI or sebasmtl โ€” the case rests entirely on the creator's own announcement, and its (already small) engagement is unconfirmed. The snippets do show the surrounding territory is active: Apodex's TRACES benchmark grades AI solvers on unresolved research problems against hidden ground truth using outcome and process verifiers, and a recent arXiv benchmark anchors frontier coding agents against an external Connect Four solver โ€” so machine-scored, contamination-resistant scientific-capability evaluation is an emerging pattern, but whether this grassroots instance gains uptake is unknown.

Why it matters to Scott

An unknown grassroots builder independently assembled Scott's own evaluation canon in a new domain: sealed previously-unpredicted measurements are the Future-Leakage Rule (freeze the world before the window) and Hidden Gates held-out-gate discipline applied to AI-for-physics, and 'will external solvers come?' is exactly his publishing-as-active-sensor uptake question โ€” a dated receipt that blind-machine-scored evaluation is emerging organically. The sharpest angle for him: by his own model-plus-harness-benchmark-unit doctrine, 'whose AI can beat it' is underspecified โ€” unless the solver harness is pinned, the leaderboard scores harness variance, not model capability, a confound he could engage, test with his own blind-review harness patterns, or publish against (and it answers the retained-question leakage worry in the Sous physics-regrading episode). Medium rather than high because the entity is one zero-engagement forum post with no independent coverage; the convergence receipt is real but its consequence depends on uptake that hasn't happened.
ip:framework.hidden-gates-frameworkip:concept.future-leakage-ruleip:concept.model-plus-harness-benchmark-unitip:framework.publishing-is-an-active-sensordev:concept.blind-workbook-to-wiki-reconciliationradar:concept.benchmark-integrityradar:concept.ai-for-scienceradar:sous-physics-benchmark-regradingradar:ship-harness-benchradar:concept.agent-benchmarks
queries asked of Scott's wikis
  • machine-scored benchmark sealed hidden ground truth
  • eval contamination training-data leakage prevention position
  • agent harness as eval unit model versus tools-and-loop
  • open benchmark cold-start adoption community uptake
  • AI-for-science physics solver benchmark territory
  • verifier gaming reward hacking machine scoring

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 362h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-26 14:00โญ origin echo-reconstructedThe Reddit post (17s self-uploaded video, selftext removed by automod) announces the author's own project; its surviving sibling post in r/C
molanocortes (GitHub) = u/sebasmtl (Reddit) on github (echo) ยท attributed from reddit.post.1wruzwt
โ€”
09-27 20:39first on r/ClaudeAI ยท published ยท +30.6hI built a physics lab with Claude agents. Now I want to see whose AI can beat it
sebasmtl
โ€”
09-27 20:39amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wruzwt
sebasmtl
peak 3 ยท 2 comments ยท 98% of case engagement
09-27 21:20our radar first saw it ยท +31.3hdiscovery anchor: reddit.post.1wruzwtโ€”
pace: p32 vs 1032 stories at the 336h mark (now 362h old) โ€” ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditI built a physics lab with Claude agents. Now I want to see whose AI can beat it
ClaudeAI
sebasmtl32
๐ŸŸง echo.github โญThe Reddit post (17s self-uploaded video, selftext removed by automod) announces the author's own project; its surviving sibling post in r/Cmolanocortes (GitHub) = u/sebasmtl (Reddit)โ€”โ€”

Interpretation history

Decision trace