2026-10-11 18:04 UTC

The maintainer claims its open-source FlashMLA build adds sm_120 support for consumer Blackwell GPUs and delivers 2–3× the attention-kernel performance of PyTorch SDPA, potentially accelerating local LLM training and inference.

state: expiredheat: lowuncertainty: highknownscott: lowflashmla local-inference inference-optimizationsmashedshanky

What is this?

FlashMLA is described as DeepSeek’s open-source CUDA kernel library for accelerating Multi-head Latent Attention in large-language-model inference. Maintainer smashedshanky claims a build adds `sm_120` kernels for consumer NVIDIA Blackwell GPUs and achieves 2–3× the attention-kernel performance of PyTorch SDPA, potentially benefiting local training and inference. The supplied results establish that `sm_120` is the consumer-Blackwell target and that related MLA kernels exist, but they do not independently verify this build’s benchmark methodology or performance claim.

Why it matters to Scott

Scott already treats accelerator placement, precision, compilation, and hardware-specific optimization as explicit local-inference policy on the “Hardware-aware local inference” page. This is a new example of that established approach rather than a change to it; the supplied evidence neither verifies the 2–3× claim nor establishes that Scott’s current hardware and models can use these Blackwell-specific MLA kernels.
dev:concept.hardware-aware-local-inferencedev:technology.cudaradar:concept.inference-optimizationradar:concept.gpu-kernelsradar:concept.local-inferenceradar:concept.cuda
queries asked of Scott's wikis
  • local LLM inference economics and GPU bottlenecks
  • custom CUDA kernels versus framework-native attention
  • consumer GPU support for local AI workloads
  • open-source inference optimization strategy
  • FP8 KV cache and memory-bandwidth optimization
  • hardware-specific kernels and portability tradeoffs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditFlashMLA sm_120 kernel build with 2-3x performance increase from SPDA
LocalLLaMA
smashedshanky87
🟧 echo.github ⭐This is the earliest primary artifact underlying the Reddit claim: a code commit implementing the SM120 dense/sparse kernels and SDPA comparsmurthy96 (GitHub account: SuperLuminalHelper; commit email identifies Shashank Murthy)——

Interpretation history

Decision trace