2026-10-11 16:37 UTC

OpenGEMM's maintainer claims its released CUDA library provides autotuned dense and block-scaled GEMM plus standalone kernel emission for NVIDIA B200, enabling developers to deploy shape-specific kernels outside the library.

state: seedheat: mediumuncertainty: mediumnovelscott: lowgpu-kernels ai-infrastructure open-gemmaramesh10OpenGEMM

What is this?

The case describes OpenGEMM as an open-source CUDA matrix-multiplication (GEMM) library targeting NVIDIA B200/sm_100a, attributed to maintainer aramesh10. Its reported release offers dense and block-scaled GEMM, autotuning for unoptimized shapes, and standalone CUDA source emission so developers can deploy shape-specific kernels outside the library. None of the supplied web results directly documents OpenGEMM or its maintainer: NVIDIA CUTLASS documentation and DeepSeek's separate DeepGEMM repository establish related GPU-kernel tooling, but do not corroborate this release or its claimed capabilities.

Why it matters to Scott

Scottโ€™s CUDA use and hardware-aware local inference policy provide adjacent context, but the hits establish neither B200 use nor custom GEMM integration that this release would change; no meaningful claim-level intersection is established. The radar tracks related Blackwell GEMM optimization, not this OpenGEMM development, whose capabilities remain uncorroborated maintainer claims in the supplied material.
radar:blackwell-nvfp4-gemm-optimizationradar:concept.gpu-kernels
queries asked of Scott's wikis
  • CUDA kernel autotuning shape-specific inference optimization
  • standalone code generation runtime dependency removal
  • B200 Blackwell inference deployment projects
  • low-precision block-scaled GEMM inference economics
  • CUTLASS DeepGEMM custom kernel integration

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 671h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-13 17:32 (minted)โญ origin echo-reconstructedOpenGEMM supplies CUDA GEMM kernels for B200 sm_100a, autotunes unoptimized shapes, and emits standalone CUDA source for dense and block-sca
aramesh10 on github (echo) ยท attributed from hn.story.49686236 ยท published time unknown
โ€”
09-13 17:19first on hacker news ยท published ยท lag ?OpenGEMM: Open-source B200 GEMM kernels
adityar2
โ€”
09-13 17:19amplified on hacker news ๐Ÿ‘‘hn.story.49686236
adityar2
peak 2 ยท 1 comments ยท 101% of case engagement
09-13 17:22our radar first saw it ยท lag ?discovery anchor: hn.story.49686236โ€”
pace: p32 vs 1032 stories at the 336h mark (now 671h old) โ€” ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnOpenGEMM: Open-source B200 GEMM kernels
Retrieved article excerpt

Open article ยท Retrieved 2026-09-13T17:23:30.183661+00:00

# OpenGEMM

GEMM kernels for NVIDIA B200 (sm\_100a) in CUDA.

```
import opengemm as og

c = og.gemm(a, b)                    # C[M, N] = A[M, K] @ B[N, K].T
c = og.gemm(a, b, sfa, sfb)          # block-scaled: nvfp4, mxfp8, mxfp4

og.emit_kernel(a, b, file="k.cu")    # emits .cu/.cuh for this shape
c = og.run_kernel("k.cu", a, b)      # compiles emitted kernel and runs it
```

Check [API.md](https://github.com/aramesh10/OpenGEMM/blob/main/API.md) for documentation.

## Install

From PyPI

```
pip install opengemm
```

From a clone:

```
git clone https://github.com/aramesh10/OpenGEMM.git
cd OpenGEMM
pip install -e .
```

Requirements:

- sm\_100a
- PyTorch 2.8+
- CUDA 12.9+

The kernels are compiled into two libraries on the first `gemm()` call and takes ~25s. Run `python -m opengemm` or from python `og.prebuild()` to pay the cost at install time instead.

## Agent Quickstart

Give your agent this prompt to use OpenGEMM as a tool:

```
OpenGEMM emits standalone CUDA GEMM kernels for B200 (sm_100a), no GPU
needed to emit:

python -c "
import opengemm as og
S = dict(m=1024, n=1024, k=1024)

og.emit_kernel(**S, atype='bf16', file='k')         # writes k.cu and k.cuh
og.emit_kernel(**S, atype='e4m3', btype='e5m2')     # mixed, names itself
og.emit_kernel(**S, atype='e2m1', sftype='ue4m3')   # block-scaled (nvfp4)
src, hdr = og.emit_kernel(**S, atype='bf16')        # the text, always returned
print(src, hdr)
"
atype / btype: bf16 f16 tf32 s8 u8 e4m3 e5m2 e3m2 e2m3 e2m1
sftype (block-scaled): ue4m3 (nvfp4) or ue8m0 (mxfp8, mxfp4)
```

## Dense and block-scaled

`C[M, N] = A[M, K] @ B[N, K].T`. Both operands are row-major with K innermost.

| GEMM | `atype` / `btype` | `sftype` | output | `torch.dtype` (in โ†’ out) |
| --- | --- | --- | --- | --- |
| bfloat16 | bf16 | โ€” | f32 | `bfloat16` โ†’ `float32` |
| float16 | f16 | โ€” | f32 | `float16` โ†’ `float32` |
| tf32 | tf32 | โ€” | f32 | `float32` โ†’ `float32` |
| int8 | s8 | โ€” | s32 | `int8` โ†’ `int32` |
| uint8 | u8 | โ€” | s32 | `uint8` โ†’ `int32` |
| fp8 | e4m3 | โ€” | f32 | `float8_e4m3fn` โ†’ `float32` |
| fp8 | e5m2 | โ€” | f32 | `float8_e5m2` โ†’ `float32` |
| mixed fp8 | e4m3, e5m2 | โ€” | f32 | `float8_e4m3fn`, `float8_e5m2` โ†’ `float32` |
| fp6 | e3m2 | โ€” | f32 | `uint8` โ†’ `float32` |
| fp6 | e2m3 | โ€” | f32 | `uint8` โ†’ `float32` |
| fp4 | e2m1 | โ€” | f32 | `uint8` โ†’ `float32` |
| nvfp4 | e2m1 | ue4m3 (per 16) | bf16 | `float4_e2m1fn_x2`, `float8_e4m3fn` โ†’ `bfloat16` |
| mxfp8 | e4m3 | ue8m0 (per 32) | bf16 | `float8_e4m3fn`, `float8_e8m0fnu` โ†’ `bfloat16` |
| mxfp4 | e2m1 | ue8m0 (per 32) | bf16 | `float4_e2m1fn_x2`, `float8_e8m0fnu` โ†’ `bfloat16` |

Note: fp6 and fp4 have no torch dtype. They arrive densely packed in `uint8` and are named - `gemm(a, b, atype="e2m1")`
Use `btype=` when the two operands differ.

Output is `[M, N]`, row-major, like `torch.mm`.

## Tuning and performance

There is no heursitic to choose the config. Optimized configs are stored in `configs.json`.
If a particular shape has not been optimized, the library autotunes and returns and saves the best config locally to `./opengemm-configs/tuned_configs.json` or to `OPENGEMM_CONFIGS` env variable.

```
CUDA_VISIBLE_DEVICES=0 python scripts/tune.py --dtype f16 --shape 4096 4096 4096
CUDA_VISIBLE_DEVICES=0 python scripts/benchmark.py --dtype bf16 e4m3    # vs cuBLAS
CUDA_VISIBLE_DEVICES=0 python scripts/test.py                           # correctness
```

`tune.py` ablates every compiled configuration for a shape and records the best performing config to `configs.json`

## Standalone kernels

```
python scripts/emit_kernel.py --dtype e4m3 --shape 4096 4096 4096 --file emitted/e4m3_4k.cu
python scripts/run_kernel.py emitted/e4m3_4k.cu       # correctness, then timing vs cuBLAS
```

OpenGEMM can also emit the optimized CUDA files for a kernel given a shape and dtype. It can be ran with `scripts/run_kernel.py` or built with `nvcc`:

```
nvcc -O3 -std=c++20 -gencode=arch=compute_100a,code=sm_100a --expt-relaxed-constexpr -shared -Xcompiler -fPIC -lcuda <KERNEL_FILE>.cu -o <KERNEL_FILE>.so
```

The entry point is `extern "C" void mm_<dtype>_<M>_<N>_<K>(a, b, c, stream)`,
or `smm_<dtype>_<M>_<N>_<K>(a, b, sfa, sfb, c, stream)`

`emit_kernel` reads only shapes and dtypes, so meta tensors work:
`emit_kernel(torch.empty(4096, 4096, dtype=torch.bfloat16, device="meta"), ...)`.
adityar221
๐ŸŸง echo.github โญOpenGEMM supplies CUDA GEMM kernels for B200 sm_100a, autotunes unoptimized shapes, and emits standalone CUDA source for dense and block-scaaramesh10โ€”โ€”

Interpretation history

Decision trace