2026-10-11 17:09 UTC

Genghan Zhang, Yixin Dong, Kunle Olukotun and coauthors claim PTXBench is an auditable benchmark and adaptation testbed for LLM-written architecture-specific PTX β€” finding no evaluated model consistently matches frontier libraries on H100/B200 GEMM and attention workloads, with repair-conditioned SFT helping unevenly β€” and adoption by agent-driven kernel-engineering work (cf. the open TIRx case) would make it the reference testbed while fading citations close it as a quiet benchmark paper.

state: seedheat: lowuncertainty: mediumnovelscott: mediumagent-evaluation gpu-kernel-optimization ai-assisted-systemsGenghan ZhangYixin DongKunle Olukotun

What is this?

PTXBench is an open-source benchmark from a Stanford-led team (Genghan Zhang, Yixin Dong, Kunle Olukotun and coauthors, with CMU and RadixArk affiliations also listed) that tests whether LLMs can write architecture-specific PTX β€” NVIDIA's low-level intermediate assembly β€” for GPU kernel optimization on H100 and B200. Released as arXiv:2608.17379 in August 2026, it scores functional correctness, whether the chosen target instructions actually execute at runtime, and speedup over frontier libraries on GEMM and attention workloads, finding the capability deeply uneven: no evaluated model consistently matches frontier libraries, success falls substantially on complex attention-backward workloads, and executing the target instructions does not translate into competitive performance. The authors also adapt Qwen3.6-27B via repair-conditioned supervised fine-tuning, which improves several tasks but generalizes unevenly, implicating data coverage, balance, and reasoning-teacher quality rather than dataset size alone. Code is open-sourced at github.com/zhang677/PTXBench; the supplied snippets show no external adoption yet, don't name the frontier libraries used as the reference bar, and the hypothesis's agent-driven kernel-engineering adoption question (cf. TIRx) is forward-looking rather than grounded in the material.

Why it matters to Scott

No canon position is challenged or arrived at β€” Scott's wikis hold no stance on LLM-written PTX β€” so this is new ground, but it bears on him twice over: it is a first-party counterweight to the agentic kernel-optimization claims he already tracks (Codex autoresearch's 232Γ—, CUDA Agent, the open TIRx substrate, for which PTXBench is a natural adoption testbed), and its repair-conditioned SFT of a ~27B open model β€” uneven generalization blamed on data coverage, balance and teacher quality rather than size β€” lands directly on his own fine-tuning-data-factory practice and CUDA/local-inference substrate.
dev:technology.cudadev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.synthetic-finetuning-datasetradar:mlc-tirx-agentic-gpu-harnessradar:cuda-agent-kernel-generation-validationradar:codex-autoresearch-gpu-kernel-speedupradar:concept.gpu-kernelsradar:concept.gpu-optimizationradar:concept.agent-benchmarksradar:concept.post-training
queries asked of Scott's wikis
  • agent evaluation harness design
  • code-generation benchmark critique
  • repair loop self-correction runtime vs weights
  • GPU kernel performance engineering
  • local inference cost and performance
  • open-weight model post-training adaptation

Measured heat

now 0 pts/hpeak 3 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1322h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

08-17 14:00⭐ origin echo-reconstructed"PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries acro
Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng, Yueming Yuan, Shaowei Zhu, Kunle Olukotun on paper (echo) Β· attributed from hn.story.49939792
β€”
10-02 23:21first on hacker news Β· published Β· +1113.4hPTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization
matt_d
β€”
10-02 23:21amplified on hacker news πŸ‘‘hn.story.49939792
matt_d
peak 3 Β· 0 comments Β· 101% of case engagement
10-03 00:21our radar first saw it Β· +1114.3hdiscovery anchor: hn.story.49939792β€”

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnPTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization
Retrieved article excerpt

Open article Β· Retrieved 2026-10-03T00:26:03.212924+00:00

# Computer Science > Computation and Language

**arXiv:2608.17379** (cs)

[Submitted on 18 Aug 2026 ([v1](https://arxiv.org/abs/2608.17379v1)), last revised 28 Sep 2026 (this version, v3)]

# Title:PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX

Authors:[Genghan Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+G), [Yixin Dong](https://arxiv.org/search/cs?searchtype=author&query=Dong,+Y), [Chengze Fan](https://arxiv.org/search/cs?searchtype=author&query=Fan,+C), [Zhichen Zeng](https://arxiv.org/search/cs?searchtype=author&query=Zeng,+Z), [Yueming Yuan](https://arxiv.org/search/cs?searchtype=author&query=Yuan,+Y), [Shaowei Zhu](https://arxiv.org/search/cs?searchtype=author&query=Zhu,+S), [Kunle Olukotun](https://arxiv.org/search/cs?searchtype=author&query=Olukotun,+K)

View a PDF of the paper titled PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX, by Genghan Zhang and 6 other authors

[View PDF](https://arxiv.org/pdf/2608.17379)
[HTML (experimental)](https://arxiv.org/html/2608.17379v3)
> Abstract:We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.

|  |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2608.17379](https://arxiv.org/abs/2608.17379) [cs.CL] |
|  | (or  [arXiv:2608.17379v3](https://arxiv.org/abs/2608.17379v3) [cs.CL] for this version) |
|  | <https://doi.org/10.48550/arXiv.2608.17379> Focus to learn more  arXiv-issued DOI via DataCite |

## Submission history

From: Genghan Zhang [[view email](https://arxiv.org/show-email/c22d7b9d/2608.17379)]   
 **[[v1]](https://arxiv.org/abs/2608.17379v1)**
Tue, 18 Aug 2026 05:14:41 UTC (5,480 KB)  
**[[v2]](https://arxiv.org/abs/2608.17379v2)**
Wed, 19 Aug 2026 20:36:11 UTC (5,480 KB)  
**[v3]**
Mon, 28 Sep 2026 05:21:04 UTC (7,272 KB)

Full-text links:

## Access Paper:

View a PDF of the paper titled PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX, by Genghan Zhang and 6 other authors

- [View PDF](https://arxiv.org/pdf/2608.17379)
- [HTML (experimental)](https://arxiv.org/html/2608.17379v3)
- [TeX Source](https://arxiv.org/src/2608.17379)

[license icon](http://creativecommons.org/licenses/by/4.0/ "Rights to this article")

### Current browse context:

cs.CL

[<Β prev](https://arxiv.org/prevnext?id=2608.17379&function=prev&context=cs.CL "previous in cs.CL (accesskey p)")
Β  | Β  
[nextΒ >](https://arxiv.org/prevnext?id=2608.17379&function=next&context=cs.CL "next in cs.CL (accesskey n)")

[new](https://arxiv.org/list/cs.CL/new)
 | 
[recent](https://arxiv.org/list/cs.CL/recent)
 | [2026-08](https://arxiv.org/list/cs.CL/2026-08)

Change to browse by:

[cs](https://arxiv.org/abs/2608.17379?context=cs)  
[cs.AI](https://arxiv.org/abs/2608.17379?context=cs.AI)

### References & Citations

- [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2608.17379)
- [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2608.17379)
- [Semantic Scholar](https://api.semanticscholar.org/arXiv:2608.17379)

export BibTeX citation
Loading...

## BibTeX formatted citation

Γ—

loading...

Data provided by:

### Bookmark

[BibSonomy](http://www.bibsonomy.org/BibtexHandler?requTask=upload&url=https://arxiv.org/abs/2608.17379&description=PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX "Bookmark on BibSonomy")
[Reddit](https://reddit.com/submit?url=https://arxiv.org/abs/2608.17379&title=PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX "Bookmark on Reddit")



Bibliographic Tools

# Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer *([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))*

Connected Papers Toggle

Connected Papers *([What is Connected Papers?](https://www.connectedpapers.com/about))*

Litmaps Toggle

Litmaps *([What is Litmaps?](https://www.litmaps.co/))*

scite.ai Toggle

scite Smart Citations *([What are Smart Citations?](https://www.scite.ai/))*

Code, Data, Media

# Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv *([What is alphaXiv?](https://alphaxiv.org/))*

Links to Code Toggle

CatalyzeX Code Finder for Papers *([What is CatalyzeX?](https://www.catalyzex.com))*

DagsHub Toggle

DagsHub *([What is DagsHub?](https://dagshub.com/))*

GotitPub Toggle

Gotit.pub *([What is GotitPub?](http://gotit.pub/faq))*

Huggingface Toggle

Hugging Face *([What is Huggingface?](https://huggingface.co/huggingface))*

ScienceCast Toggle

ScienceCast *([What is ScienceCast?](https://sciencecast.org/welcome))*

Demos

# Demos

Replicate Toggle

Replicate *([What is Replicate?](https://replicate.com/docs/arxiv/about))*

Spaces Toggle

Hugging Face Spaces *([What is Spaces?](https://huggingface.co/docs/hub/spaces))*

Spaces Toggle

TXYZ.AI *([What is TXYZ.AI?](https://txyz.ai))*

Related Papers

# Recommenders and Search Tools

Link to Influence Flower

Influence Flower *([What are Influence Flowers?](https://influencemap.cmlab.dev/))*

Core recommender toggle

CORE Recommender *([What is CORE?](https://core.ac.uk/services/recommender))*

- Author
- Venue
- Institution
- Topic


About arXivLabs

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.17379) |
Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
matt_d30
🟧 echo.paper ⭐"PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries acroGenghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng, Yueming Yuan, Shaowei Zhu, Kunle Olukotunβ€”β€”

Interpretation history

Decision trace