2026-10-11 16:37 UTC

John Sous and coauthors claim expert repairs and regrading reveal near-saturation of retained physics benchmark questions by frontier models, undermining low leaderboard scores as evidence of weak closed-form physics capability.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediummodel-evaluation benchmarks research-agentsJohn SousArman CohanYale

What is this?

The case describes a claimed expert audit of physics evaluations, attributed to John Sous and coauthors: repairing benchmark questions and regrading answers reportedly raises frontier-model scores to near-saturation on the retained questions. This would concern performance on those closed-form questions, not establish broader research capability. All supplied web snippets are unrelated to the study, so they do not verify its existence, authorship, Yale or Arman Cohan's involvement, methodology, or score changes.

Why it matters to Scott

The claimed expert repairs converge with Scott’s separate answer-key audit in “LLM rubric grading with independent key audit” and his requirement for mechanically different verifiers: if verified, the reported score changes would make key validity a consequential control in his grading pipeline. The supplied material does not verify the study or its results; radar pages already track benchmark cleaning and saturation, but their snippets do not establish coverage of this specific physics audit.
dev:concept.llm-rubric-gradingip:concept.mechanically-different-verifiersradar:cleaned-benchmarks-frontier-rankingsradar:ai-benchmark-saturation-distortionradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • benchmark validity and ground-truth answer errors
  • leaderboard scores versus actual model capability
  • evaluation harnesses expert review and grader reliability
  • benchmark saturation versus open-ended research agents
  • model selection based on task-specific evaluations

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 616h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 00:22 (minted)⭐ origin echo-reconstructedExpert regrading and repairs reveal broken physics evaluations and near-saturation of leading benchmarks; Sous's accompanying account notes
John Sous and coauthors on paper (echo) · attributed from reddit.post.1whgoci · published time unknown
—
09-15 23:37first on r/singularity · published · lag ?Turns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.
Profanion
—
09-16 19:19first on hacker news · published · lag ?How good are frontier models at physics?
qt31415926
—
09-15 23:37amplified on r/singularityreddit.post.1whgoci
Profanion
peak 225 · 24 comments · 48% of case engagement
09-16 19:19amplified on hacker news 👑hn.story.49731620
qt31415926
peak 100 · 50 comments · 52% of case engagement
09-17 04:14amplified on hacker newshn.story.49736325
teleforce
peak 2 · 0 comments · 1% of case engagement
09-16 00:20our radar first saw it · lag ?discovery anchor: reddit.post.1whgoci—
pace: p82 vs 1032 stories at the 336h mark (now 616h old) — ahead of astra-zerobench-human-baseline (1.0x), behind oracle-new-mexico-force-majeure (1.0x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditTurns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.
singularity
Retrieved article excerpt

Open article · Retrieved 2026-09-16T00:22:20.704113+00:00

# Is Physics Dead?

Broken benchmarks, and re-evaluating the capabilities of frontier models in physics

By John Sous ·
September 14, 2026

*Based on [How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks](https://arxiv.org/abs/2609.13009), arXiv:2609.13009. Thanks to my co-authors Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Lucas Baker, and Arman Cohan, and to the Yale physics faculty and graduate researchers who carried out the audits.*

> *Physics is dead.*
> (not Friedrich Nietzsche)

AI has already come for mathematics. What of physics?

Anyone who has spent time working with frontier AI models knows that they are capable of astonishing feats of reason, but also of equally shocking mistakes and naïveté. Within the context of a research project they still drift, make inconsistent choices, forget previously established work, and often struggle to identify the next meaningful step despite knowing the mechanics. To spend a week applying an agent to a difficult problem is to experience both sides of this duality: the somewhat terrifying breadth of its knowledge, and its seeming inability to put the knowledge to practical use.

In the wake of Navier-Stokes, any question of capability on the mathematics front seems settled. Physics, however, seems to offer some additional barriers that favor humans, such as fuzzier notions of proof and more interactions with the real world. Most concretely, the benchmarks would suggest that frontier AI has yet to approach even graduate level, so physics should be safe for a while.

Is it?

## Looking at the benchmarks

Judging from the leaderboards on popular sites such as Artificial Analysis, physics appears to be one of the last holdouts against the inexorable march of AI capabilities. On CritPt, a benchmark of research-level challenges, GPT-5.6 Sol scored 32% (the latest release, GPT-6 Astra, does no better). On the physics portion of Humanity’s Last Exam, it scored 47%. Apparently, even a model good enough to be declared artificial general intelligence gets roughly half of graduate-level physics wrong.

On the surface, this kind of result is plausible even if the latest models are strong enough to answer essentially any analytical question. Physics is not primarily about executing calculations. Mostly, the hardest part is to decide which calculations matter and which form of the question makes sense to tackle.

According to an old joke, if you hire a physicist to help a dairy farm, they will start by telling you to “assume a spherical cow.” The joke captures the deeper truth that physicists are concerned mostly with finding the idealization that discards almost everything about a problem while preserving exactly what matters. For example, in 1983, Robert Laughlin set out to explain the fractional quantum Hall effect. The direct approach would have required an approach for a vast number of strongly interacting electrons, beyond any computer. Instead, Laughlin guessed a many-electron wavefunction with the right symmetry and limiting behavior. His guess was correct, and for it he received the Nobel Prize. Another great example is how Kenneth Wilson (who was also awarded the Nobel Prize) and Michael Fisher implemented a brilliant idea for studying phase transitions, where they treated a problem in three spatial dimensions through an expansion around four dimensions, solved it in 4 − ε dimensions, and then set ε = 1. These are not just calculations. They may omit a certain form of rigor, nonetheless they have the quality of being determined by an intuition about which calculation might expose the underlying physics.

After all, perhaps the low leaderboard scores meant that physics remained beyond the reach of the best models. Perhaps we are safe?

## Measuring carefully

In an effort to explain the gap between frontier model performance on physics and mathematics benchmarks, we looked at why the models failed. Unlike most benchmark analyses, which seek to understand gaps in AI ability at scale and parse differences in statistics, we decided it was necessary to interrogate the questions in detail. We collected questions where the AI answers had largely been rejected, handed them to highly qualified physicists (my amazing students!), and asked them to check the grading.

To put it lightly, we were surprised.

Pre-audit and corrected accuracy on six physics benchmarks


Light bars are pre-audit scores, solid bars are scores after correction.

Across all benchmarks, we found that a large majority of questions where frontier model answers were rejected were false negatives, either because the evaluation pipeline rejected equivalent forms of correct answers or because the benchmark items themselves were defective. We found incorrect arithmetic in reference answers, omissions of key information leading to underspecified problem statements, and problems whose results depended on unstated conventions. Most surprisingly, these issues were also prevalent in the expert-curated benchmarks used to evaluate frontier physics capabilities, including CritPt and the physics component of Humanity’s Last Exam. [1](https://jsous.github.io/blogs/is-physics-dead/#fn:1)

We worked with faculty members and their students at Yale (and a few external physicists) to scale the audit process, assigning every individual question in their area of expertise and attempting repairs wherever possible. We found defects in 30 of 50 CMT-Benchmark questions and 21 of 56 CritPt questions. After correcting graders and repairing or excluding flawed questions, we found that not only were the results enormously different pre- and post-audit, but the strongest models available at the time of audit approached saturation on even the hardest benchmarks.

| Benchmark | Pre-audit mean@4 | Corrected mean@4 | Corrected pass@4 |
| --- | --- | --- | --- |
| HLE-Physics | 47.3% | 78.7% | 91.4% |
| CMT-Benchmark | 61.0% | 87.2% | 98.0% |
| CritPt | 32.3%\* | 87.5% | 94.4% |

*The pre-audit CritPt figure is reported mean@5 on all 70 challenges according to Artificial Analysis. Our corrected results use the 54 retained challenges.*

There are important qualifications: corrected scores are computed on retained or repaired questions, some runs used tools while others did not, and our HLE audit only covered questions where GPT-5.6 Sol’s answers had been initially rejected. In addition, even the corrected evaluator had an error rate of about 4%. However, none of these factors affect the central conclusion that all physics benchmarks have been dramatically understating frontier AI capabilities in physics and continue to do so.

To restate: even the prior generation of frontier models (Fable 5 and GPT-5.6 Sol) now solve nearly all well-posed, closed-form physics problems in every available benchmark, from undergraduate mechanics to research-level condensed matter theory. The belief that benchmarks imply some key difference between physics and mathematics, programming, and other areas where AI capabilities now approach superhuman levels is no more than a comforting fiction.

None of this implies that physics is dead, that agents will replace human physicists, or even that everything important about physics can be reduced to a form that is cleanly solvable by AI. However, it strongly suggests a rude surprise awaits anyone who believes the progression that has so far advanced through chess, Go, poker, protein folding, programming, and now mathematics is about to stop at physics.

## Physics is not yet solved

Given the evidence that physics and mathematics are trending strongly in the same direction, we were also curious whether they would make similar progress on open questions. Our results in this area, although not reported in the paper, look a bit more optimistic for the physicists.

We tested a GPT-based agentic system, which previously succeeded in resolving several open mathematical conjectures, on open problems in theoretical physics. To our temporary relief and encouragement, it has not managed to fully resolve even one autonomously. We also noted that these agents made considerably less progress on the partially resolved physics problems than on open mathematics problems of comparable difficulty, both by our analysis and independent agent-based analysis of partial results in each domain.

Our preliminary conclusion is that the difference comes down exactly to the gaps in high-level intuition that one notices working with these systems in person, and that these gaps may prove more of an obstacle to physics research than to mathematics because problem formulation is a more critical component of what makes a physics question “open”. In other words, frontier models are demonstrably excellent at the portions of any research problem that resemble a problem set. Calculations, numerics, limit-checking, and efficient program formulation are no problem: what they lack is the meta-awareness to step back and ask not only what can be solved but what should be. So far, the traces demonstrate ample evidence of understanding the problem as posed, but not the fundamental will to reshape the question that made approaches such as Laughlin’s, Wilson’s, and Fisher’s possible.

## Redefining frontier physics for AI

Given that all existing benchmarks are approaching saturation, what do we need in order to get a true picture of AI capabilities in physics research?

**Harder tasks.** The next generation of benchmarks should focus on tasks that emphasize the unique difficulties of physics, favoring those where choosing a favorable representation is part of the problem and there may not be a single reference answer. This will also require a more careful effort to create standards of evaluation flexible enough to admit any answer a qualified physicist would deem correct without relying purely on agentic self-judgment.

**Better agents and harnesses for physics.** AI physics currently lacks the infrastructure of AI mathematics. Besides the simple fact that far more effort has been devoted to the latter, the prominence of Lean has contributed critically to the progression of AI mathematical capabilities. So far, physics has no clear parallel. We need verification tools built around our own culture: limiting cases, dimensional analysis, agreement with known results, and simply better ways of asking the questions that used to exist only on a blackboard. Before tackling the most famous grand challenges, we can develop these tools to address open problems with a more well-defined space of plausible idealizations.

**Open challenges.** We must establish a process of cooperation between humans and AI on open physics questions, including publicly documenting attempts, analyzing partial progress and instructive failures, and creating models for interpretable AI-driven physics research that contributes to human knowledge. Mathematics, as stated by Toffoli and Duede, is not merely the production of solutions, and nor is physics. Results should be comprehensible to humans, and we should have a definite idea of what positive collaboration looks like.

It is always tempting to stick our heads in the sand. What makes this moment unique for physicists is that it is equally easy to do so in several ways: naïve faith in the benchmark gaps, an emphasis on the intuition and judgment supposedly unique to humans, or a refusal to furnish the tools that have helped power the recent progress of AI in mathematics. We propose that the braver and more tenable choice is to face facts directly and help steer the evolution of cooperation between humans and AI. What would it look like to have a process where both sides contribute, results are pursued not for bragging rights or benchmarks but for the evolution of knowledge, and accessibility to human minds remains a core tenet and design component of the research process?

That part is still up to us.

1. We found some or all
Profanion22522
🟧 echo.paper ⭐Expert regrading and repairs reveal broken physics evaluations and near-saturation of leading benchmarks; Sous's accompanying account notes John Sous and coauthors——
🟧 hnHow good are frontier models at physics?
Retrieved article excerpt

Open article · Retrieved 2026-09-16T20:22:40.795592+00:00

# Computer Science > Artificial Intelligence

**arXiv:2609.13009** (cs)

[Submitted on 11 Sep 2026]

# Title:How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Authors:[Ali Ansari](https://arxiv.org/search/cs?searchtype=author&query=Ansari,+A), [Haoran Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+H), [Andy Zeyi Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+A+Z), [Mark Jabbour](https://arxiv.org/search/cs?searchtype=author&query=Jabbour,+M), [Yongshan Ding](https://arxiv.org/search/cs?searchtype=author&query=Ding,+Y), [Steven Girvin](https://arxiv.org/search/cs?searchtype=author&query=Girvin,+S), [Yu He](https://arxiv.org/search/cs?searchtype=author&query=He,+Y), [Sohrab Ismail-Beigi](https://arxiv.org/search/cs?searchtype=author&query=Ismail-Beigi,+S), [Aleksander Kubica](https://arxiv.org/search/cs?searchtype=author&query=Kubica,+A), [Owen D. Miller](https://arxiv.org/search/cs?searchtype=author&query=Miller,+O+D), [Corey O'Hern](https://arxiv.org/search/cs?searchtype=author&query=O'Hern,+C), [Vidvuds Ozolins](https://arxiv.org/search/cs?searchtype=author&query=Ozolins,+V), [David Poland](https://arxiv.org/search/cs?searchtype=author&query=Poland,+D), [A. Douglas Stone](https://arxiv.org/search/cs?searchtype=author&query=Stone,+A+D), [Frank C. van den Bosch](https://arxiv.org/search/cs?searchtype=author&query=van+den+Bosch,+F+C), [Logan Wright](https://arxiv.org/search/cs?searchtype=author&query=Wright,+L), [Navid Akbari](https://arxiv.org/search/cs?searchtype=author&query=Akbari,+N), [Santanu Antu](https://arxiv.org/search/cs?searchtype=author&query=Antu,+S), [Kangle Cai](https://arxiv.org/search/cs?searchtype=author&query=Cai,+K), [Andrew Calabrese-Day](https://arxiv.org/search/cs?searchtype=author&query=Calabrese-Day,+A), [Mateo Cárdenes Wuttig](https://arxiv.org/search/cs?searchtype=author&query=Wuttig,+M+C), [Meng Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+M), [Barry T. Chiang](https://arxiv.org/search/cs?searchtype=author&query=Chiang,+B+T), [Ali Ghorashi](https://arxiv.org/search/cs?searchtype=author&query=Ghorashi,+A), [Shouzhen Gu](https://arxiv.org/search/cs?searchtype=author&query=Gu,+S), [Haoyang Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+H), [Zhibo Kang](https://arxiv.org/search/cs?searchtype=author&query=Kang,+Z), [Lukas Kienesberger](https://arxiv.org/search/cs?searchtype=author&query=Kienesberger,+L), [Hantian Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+H), [Charles Lomba](https://arxiv.org/search/cs?searchtype=author&query=Lomba,+C), [Zhongling Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+Z), [Wenchao Ma](https://arxiv.org/search/cs?searchtype=author&query=Ma,+W), [Rohin E. McIntosh](https://arxiv.org/search/cs?searchtype=author&query=McIntosh,+R+E), [Evan McKinney](https://arxiv.org/search/cs?searchtype=author&query=McKinney,+E), [Ivan Rojkov](https://arxiv.org/search/cs?searchtype=author&query=Rojkov,+I), [Xulei Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+X), [Yarone Meir Tokayer](https://arxiv.org/search/cs?searchtype=author&query=Tokayer,+Y+M), [Naveen Balaji Umasankar](https://arxiv.org/search/cs?searchtype=author&query=Umasankar,+N+B), [Mira Varma](https://arxiv.org/search/cs?searchtype=author&query=Varma,+M), [Leda Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+L), [Qimin Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Q), [Tyler Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+T), [Haoyu Wei](https://arxiv.org/search/cs?searchtype=author&query=Wei,+H), [Jinming Yang](https://arxiv.org/search/cs?searchtype=author&query=Yang,+J), [Jinchen Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+J), [Sherlock Tingrui Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+S+T), [Qinyuan Zheng](https://arxiv.org/search/cs?searchtype=author&query=Zheng,+Q), [Jay S. Zou](https://arxiv.org/search/cs?searchtype=author&query=Zou,+J+S), [Lucas Baker](https://arxiv.org/search/cs?searchtype=author&query=Baker,+L), [Arman Cohan](https://arxiv.org/search/cs?searchtype=author&query=Cohan,+A), [John Sous](https://arxiv.org/search/cs?searchtype=author&query=Sous,+J)

View a PDF of the paper titled How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks, by Ali Ansari and 50 other authors

[View PDF](https://arxiv.org/pdf/2609.13009)
[HTML (experimental)](https://arxiv.org/html/2609.13009v1)
> Abstract:Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

|  |  |
| --- | --- |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2609.13009](https://arxiv.org/abs/2609.13009) [cs.AI] |
|  | (or  [arXiv:2609.13009v1](https://arxiv.org/abs/2609.13009v1) [cs.AI] for this version) |
|  | <https://doi.org/10.48550/arXiv.2609.13009> Focus to learn more  arXiv-issued DOI via DataCite (pending registration) |

## Submission history

From: Ali Ansari [[view email](https://arxiv.org/show-email/9f732102/2609.13009)]   
 **[v1]**
Fri, 11 Sep 2026 16:06:50 UTC (178 KB)

Full-text links:

## Access Paper:

View a PDF of the paper titled How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks, by Ali Ansari and 50 other authors

- [View PDF](https://arxiv.org/pdf/2609.13009)
- [HTML (experimental)](https://arxiv.org/html/2609.13009v1)
- [TeX Source](https://arxiv.org/src/2609.13009)

[license icon](http://creativecommons.org/licenses/by/4.0/ "Rights to this article")

### Current browse context:

cs.AI

[< prev](https://arxiv.org/prevnext?id=2609.13009&function=prev&context=cs.AI "previous in cs.AI (accesskey p)")
  |   
[next >](https://arxiv.org/prevnext?id=2609.13009&function=next&context=cs.AI "next in cs.AI (accesskey n)")

[new](https://arxiv.org/list/cs.AI/new)
 | 
[recent](https://arxiv.org/list/cs.AI/recent)
 | [2026-09](https://arxiv.org/list/cs.AI/2026-09)

Change to browse by:

[cs](https://arxiv.org/abs/2609.13009?context=cs)

### References & Citations

- [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2609.13009)
- [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2609.13009)
- [Semantic Scholar](https://api.semanticscholar.org/arXiv:2609.13009)

export BibTeX citation
Loading...

## BibTeX formatted citation

×

loading...

Data provided by:

### Bookmark

[BibSonomy](http://www.bibsonomy.org/BibtexHandler?requTask=upload&url=https://arxiv.org/abs/2609.13009&description=How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks "Bookmark on BibSonomy")
[Reddit](https://reddit.com/submit?url=https://arxiv.org/abs/2609.13009&title=How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks "Bookmark on Reddit")



Bibliographic Tools

# Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer *([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))*

Connected Papers Toggle

Connected Papers *([What is Connected Papers?](https://www.connectedpapers.com/about))*

Litmaps Toggle

Litmaps *([What is Litmaps?](https://www.litmaps.co/))*

scite.ai Toggle

scite Smart Citations *([What are Smart Citations?](https://www.scite.ai/))*

Code, Data, Media

# Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv *([What is alphaXiv?](https://alphaxiv.org/))*

Links to Code Toggle

CatalyzeX Code Finder for Papers *([What is CatalyzeX?](https://www.catalyzex.com))*

DagsHub Toggle

DagsHub *([What is DagsHub?](https://dagshub.com/))*

GotitPub Toggle

Gotit.pub *([What is GotitPub?](http://gotit.pub/faq))*

Huggingface Toggle

Hugging Face *([What is Huggingface?](https://huggingface.co/huggingface))*

ScienceCast Toggle

ScienceCast *([What is ScienceCast?](https://sciencecast.org/welcome))*

Demos

# Demos

Replicate Toggle

Replicate *([What is Replicate?](https://replicate.com/docs/arxiv/about))*

Spaces Toggle

Hugging Face Spaces *([What is Spaces?](https://huggingface.co/docs/hub/spaces))*

Spaces Toggle

TXYZ.AI *([What is TXYZ.AI?](https://txyz.ai))*

Related Papers

# Recommenders and Search Tools

Link to Influence Flower

Influence Flower *([What are Influence Flowers?](https://influencemap.cmlab.dev/))*

Core recommender toggle

CORE Recommender *([What is CORE?](https://core.ac.uk/services/recommender))*

- Author
- Venue
- Institution
- Topic


About arXivLabs

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.13009) |
Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
qt3141592610050
🟧 hnIs Physics Dead: Broken benchmarks and re-evaluating frontier models in physicsteleforce20

Interpretation history

Decision trace