2026-10-11 16:38 UTC

Nunchux AI claims VC-Attention accelerates MiniMax-H3 attention kernels by 1.51–1.59× over BF16 FlashAttention-4 on B300 and B200 without retraining, with better B200 output fidelity than SageAttention2, potentially reducing video-generation inference costs.

state: seedheat: lowuncertainty: mediumnovelscott: lowlow-bit-attention inference-economics inference-optimization video-generationNunchux AI

What is this?

The case attributes VC-Attention to Nunchux AI: a training-free, low-bit attention method claiming 1.59× B200 and 1.51× B300 kernel speedups over BF16 FlashAttention-4 on MiniMax-H3, plus better B200 output fidelity than SageAttention2. None of the supplied web results directly documents Nunchux AI or VC-Attention, so these claims remain uncorroborated here. The vLLM recipe establishes that MiniMax-H3 is an open-weight audiovisual generation model and reports FlashAttention-4 consuming roughly 76% of diffusion device time in one profiled workload, making attention optimization a plausible performance lever. That recipe also warns of fidelity trade-offs from other quantized and sparse attention methods; the supplied material does not establish VC-Attention’s end-to-end speedup or cost savings.

Why it matters to Scott

The hits establish Scott’s hardware-aware local inference work, but not use of MiniMax-H3 or B200/B300; this uncorroborated kernel claim does not yet change a documented build or position. The radar already tracks H3 deployment and session costs, but not VC-Attention, and the supplied material establishes neither end-to-end savings nor applicability to Scott’s systems.
radar:minimax-h3-comfyui-local-validationradar:qrun-h3-session-cost-measurements
queries asked of Scott's wikis
  • inference economics kernel speedups versus end-to-end cost
  • quantization precision trade-offs output fidelity evaluation
  • training-free inference optimization open-weight deployment
  • video generation pipelines GPU attention bottlenecks
  • vLLM Blackwell inference serving projects

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 626h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedIntroduces training-free VC-Attention using V-Smooth and ExpCast-FP8, reporting 1.59× B200 and 1.51× B300 attention-kernel speedups on MiniM
Nunchux AI Team on blog (echo) · attributed from hn.story.49736020
—
09-17 03:20first on hacker news · published · +37.3hVC-Attention: Faster Low-Bit Attention Without Retraining
lmxyy
—
09-17 03:20amplified on hacker newshn.story.49736020
lmxyy
peak 4 · 0 comments · 28% of case engagement
09-23 16:34amplified on hacker news 👑hn.story.49818681
lmxyy
peak 7 · 3 comments · 71% of case engagement
09-17 03:21our radar first saw it · +37.4hdiscovery anchor: hn.story.49736020—
pace: p48 vs 1032 stories at the 336h mark (now 626h old) — ahead of acs-local-skill-risk-catalog (1.1x), behind agentgit-accountless-agent-handoffs (0.9x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnVC-Attention: Faster Low-Bit Attention Without Retraining
Retrieved article excerpt

Open article · Retrieved 2026-09-17T03:22:13.124334+00:00

[← Back to blog](https://www.nunchux.ai/blog)Research

# VC-Attention: Faster Low-Bit Attention Without Retraining

Nunchux AI Team·September 16, 2026·5 min read

[![](/blog/attention-is-the-video-bottleneck/demo.webp)](https://jdbvvg5ymgludm5a.public.blob.vercel-storage.com/blog-assets/02-VC-Attention/demo-v28.mp4)

VC-Attention in one minute.

#### NVIDIA B200

FlashAttention-4

1.00×

Attn-QAT

1.40×

VC-Attention

1.59×

**Nunchux Attention**

1.91×

#### NVIDIA B300

FlashAttention-4

1.00×

Attn-QAT

1.34×

VC-Attention

1.51×

**Nunchux Attention**

1.83×

**Attention speedup over BF16 FlashAttention-4 [[1]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-1) on B200 and B300.** We benchmark the attention workload in MiniMax-H3 when generating 243 frames at 1344×768. Attn-QAT [[3]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-3) uses QK4 · PV8; VC-Attention uses 8-bit attention with ExpCast-FP8; the B200 bar also runs V-Smooth, time-weighted by the deployed schedule (on for the first quarter of the denoising steps, off for the rest). Nunchux Attention is our proprietary extension of VC-Attention.

Show as a table

Attention speedup by kernel and GPU, relative to each card’s own BF16 reference.

| Kernel | B200 speedup | B300 speedup |
| --- | --- | --- |
| FlashAttention-4 | 1.00× | 1.00× |
| Attn-QAT | 1.40× | 1.34× |
| VC-Attention | 1.59× | 1.51× |
| Nunchux Attention | **1.91×** | **1.83×** |

[![](/blog/attention-is-the-video-bottleneck/fa4-107.webp)](/blog/attention-is-the-video-bottleneck/fa4-107.mp4)

**FlashAttention-4** · BF16  
1.00× attention speedup

[![](/blog/attention-is-the-video-bottleneck/sage2-107.webp)](/blog/attention-is-the-video-bottleneck/sage2-107.mp4)

**SageAttention2** · 8-bit

[![](/blog/attention-is-the-video-bottleneck/vc-107.webp)](/blog/attention-is-the-video-bottleneck/vc-107.mp4)

**VC-Attention** · 8-bit  
1.59× attention speedup

Play all0:00 / 0:00

[![](/blog/attention-is-the-video-bottleneck/fa4-103.webp)](/blog/attention-is-the-video-bottleneck/fa4-103.mp4)

**FlashAttention-4** · BF16  
1.00× attention speedup

[![](/blog/attention-is-the-video-bottleneck/sage2-103.webp)](/blog/attention-is-the-video-bottleneck/sage2-103.mp4)

**SageAttention2** · 8-bit

[![](/blog/attention-is-the-video-bottleneck/vc-103.webp)](/blog/attention-is-the-video-bottleneck/vc-103.mp4)

**VC-Attention** · 8-bit  
1.59× attention speedup

Play all0:00 / 0:00

Two MiniMax-H3 examples generated on an NVIDIA B200 at 1344×768, 243 frames. Speedups are for the attention kernel relative to BF16 FlashAttention-4. Across 100 prompts, VC-Attention achieves 20.2 dB mean PSNR against the BF16 reference, compared with 19.9 dB for SageAttention2. Higher PSNR means the output is closer to the BF16 reference.

Nunchux is building the frontier of multimodal inference: the fastest, cheapest, and highest quality inference for image, video, and world models. For video, the operator that matters most is attention.

The 10 second clips above were generated with MiniMax-H3. On one B200 with BF16 FlashAttention-4, about two thirds of every denoising step is attention. Attention cost grows with the square of the token count, so longer clips make it worse. Low-bit attention offers a path to faster video generation, but on B200 and B300, realizing those gains while preserving fidelity requires addressing both quantization error and the softmax bottleneck.

Today we introduce **VC-Attention**, a faster and more accurate training-free low-bit attention. VC-Attention combines two innovations, V-Smooth and ExpCast-FP8. V-Smooth reduces the quantization error; ExpCast-FP8 removes the softmax bottleneck. On B200, VC-Attention runs the attention kernel 1.6× faster than BF16 FlashAttention-4 [[1]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-1) on MiniMax-H3, and it is more faithful to the BF16 output than SageAttention2 [[2]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-2).

## V-Smooth

Existing methods already keep the query and key product accurate at low precision, for example with a Hadamard transform that spreads outliers across channels. With that in place, value quantization becomes a major source of error. V-Smooth groups the value tokens with a lightweight k-means so each hardware block holds similar tokens, then subtracts the block mean and quantizes only the residual. Across the four models in the report [[4]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-4), this alone adds 1.1 to 2.8 dB of PSNR over SageAttention2.

rows: value tokens t1…t8cells: channels, darker is largerbrace: one hardware block, one scale

t1

t2

t3

t4

t5

t6

t7

t8

block 1  
one scale

+μ1

block 2  
one scale

+μ2

Easy to Quantize

Each block gives up its mean μ, stored exactly. What is left is small and evenly spread, and only that goes to the 8-bit quantizer.

1 · Sequence order2 · Group with k-means3 · Subtract the block mean

**V-Smooth, step by step.** Each row is one value token, and a hardware block of tokens shares one quantization scale. In sequence order every block mixes large and small tokens. V-Smooth groups similar tokens into the same block, then subtracts each block’s mean and stores it exactly. The residual that remains is small and evenly spread, so it quantizes with far less error.

## ExpCast-FP8

Low-bit compute speeds up the two matrix products, but the softmax between them still runs in high precision and becomes the longest stage of the kernel. ExpCast-FP8 replaces that stage with a linear approximation that maps each score directly to its FP8 probability code.

Three attention pipelines. BF16 runs QK-transpose in BF16, softmax in FP32 and PV in BF16. SageAttention2 runs QK-transpose in INT8, softmax still in FP32, and PV in FP8. VC-Attention runs QK-transpose in INT8, softmax with ExpCast-FP8, and PV in FP8. A dashed box around the two FP32 softmax stages is labelled: softmax stays in FP32.softmax stays in FP32BF16QK⊤BF16SoftmaxFP32PVBF16SageAttention2QK⊤INT8SoftmaxFP32PVFP8VC-AttentionQK⊤INT8SoftmaxExpCast-FP8PVFP8

**The stage that low-bit Tensor Cores do not touch.** Once both matrix products run in 8-bit, the softmax between them still runs in FP32. ExpCast-FP8 replaces it with a linear approximation that writes the FP8 probability code directly.

Given the log-domain score zzz, the standard path evaluates p=2zp = 2^{z}p=2z in FP32 and then casts ppp to E4M3. An E4M3 byte stores roughly 8log⁡2p+568\log\_2 p + 568log2​p+56, so ExpCast-FP8 computes the byte directly with one linear map, 8z+56+β8z + 56 + \beta8z+56+β, rounded to an integer.

##### Standard: exponentiate, then cast

```
fp32  p     = exp2(z);     // slow FP32 exp
e4m3  p_fp8 = to_e4m3(p);  // cast to FP8
```

##### ExpCast-FP8: write the E4M3 byte

```
fp32  c     = 8*z + 56 + beta;    // fast FMA
uint8 code  = round(clip(c, 0, 120));
e4m3  p_fp8 = view_as_e4m3(code); // no-op
```

no exponentialno cast to E4M3

**ExpCast-FP8 skips the slow FP32 exp and the cast to E4M3.**

## From VC-Attention to Nunchux Attention

**Nunchux Attention** is our proprietary extension of VC-Attention. It adds algorithm and kernel optimizations developed for our inference stack. On the MiniMax-H3 attention workload, it runs 1.91× faster than BF16 FlashAttention-4 on B200 and 1.83× faster on B300.

## What this makes possible

VC-Attention speeds up attention in existing video models without retraining. It can be combined with sparse attention [[5](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-5),[6](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-6),[7](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-7)], few-step distillation, and multi-GPU execution to further reduce the cost of video generation. Together, these improvements make higher-resolution and longer videos more practical to generate.

---

Read the report: [*VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention*](https://arxiv.org/pdf/2609.15810).

Free access to MiniMax-H3 arrives very soon. Join the [Nunchux Modelverse](https://www.nunchux.ai/waitlist) waiting list today to be among the first to use it.

If you run visual models at scale, [contact sales](https://www.nunchux.ai/enterprise). We are hiring. Please visit our [careers page](https://www.nunchux.ai/careers) for more details.

## References

1. [1]Ted Zadouri, Markus Hoehnerbach, Jay Shah, Vijay Thakkar, and Tri Dao. [FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling](https://proceedings.mlsys.org/paper_files/paper/2026/hash/ae8b0b5838ba510daff1198474e7b984-Abstract-Conference.html). MLSys 2026.
2. [2]Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, and Jianfei Chen. [SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization](https://arxiv.org/abs/2411.10958). ICML 2025.
3. [3]Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, and Hao Zhang. [Attn-QAT: 4-Bit Attention with Quantization-Aware Training](https://arxiv.org/abs/2603.00040). arXiv:2603.00040, 2026.
4. [4]Xingyang Li, Dongyun Zou, Shining Zhang, Jiacheng Chen, Haocheng Xi, Lvmin Zhang, Jun-Yan Zhu, Song Han, Zhekai Zhang, Yujun Lin, and Muyang Li. [VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention](https://arxiv.org/abs/2609.15810). arXiv:2609.15810, 2026.
5. [5]Haocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu, Muyang Li, Xiuyu Li, Yujun Lin, Han Cai, Jintao Zhang, Dacheng Li, Jianfei Chen, Ion Stoica, Kurt Keutzer, and Song Han. [Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity](https://arxiv.org/abs/2502.01776). ICML 2025.
6. [6]Xingyang Li, Muyang Li, Tianle Cai, Haocheng Xi, Shuo Yang, Yujun Lin, Lvmin Zhang, Songlin Yang, Jinbo Hu, Kelly Peng, Maneesh Agrawala, Ion Stoica, Kurt Keutzer, and Song Han. [Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation](https://arxiv.org/abs/2506.19852). NeurIPS 2025.
7. [7]Jintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei, Haocheng Xi, Jun Zhu, and Jianfei Chen. [SpargeAttn: Accurate Sparse Attention Accelerating Any Model Inference](https://arxiv.org/abs/2502.18137). ICML 2025.
lmxyy40
🟧 echo.blog ⭐Introduces training-free VC-Attention using V-Smooth and ExpCast-FP8, reporting 1.59× B200 and 1.51× B300 attention-kernel speedups on MiniMNunchux AI Team——
🟧 hnNunchux on AMD MI355X: 5s MiniMax-H3 Videos in 1.3slmxyy73

Interpretation history

Decision trace