Retrieved article excerpt
Open article · Retrieved 2026-09-17T03:22:13.124334+00:00
[← Back to blog](https://www.nunchux.ai/blog)Research
# VC-Attention: Faster Low-Bit Attention Without Retraining
Nunchux AI Team·September 16, 2026·5 min read
[](https://jdbvvg5ymgludm5a.public.blob.vercel-storage.com/blog-assets/02-VC-Attention/demo-v28.mp4)
VC-Attention in one minute.
#### NVIDIA B200
FlashAttention-4
1.00×
Attn-QAT
1.40×
VC-Attention
1.59×
**Nunchux Attention**
1.91×
#### NVIDIA B300
FlashAttention-4
1.00×
Attn-QAT
1.34×
VC-Attention
1.51×
**Nunchux Attention**
1.83×
**Attention speedup over BF16 FlashAttention-4 [[1]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-1) on B200 and B300.** We benchmark the attention workload in MiniMax-H3 when generating 243 frames at 1344×768. Attn-QAT [[3]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-3) uses QK4 · PV8; VC-Attention uses 8-bit attention with ExpCast-FP8; the B200 bar also runs V-Smooth, time-weighted by the deployed schedule (on for the first quarter of the denoising steps, off for the rest). Nunchux Attention is our proprietary extension of VC-Attention.
Show as a table
Attention speedup by kernel and GPU, relative to each card’s own BF16 reference.
| Kernel | B200 speedup | B300 speedup |
| --- | --- | --- |
| FlashAttention-4 | 1.00× | 1.00× |
| Attn-QAT | 1.40× | 1.34× |
| VC-Attention | 1.59× | 1.51× |
| Nunchux Attention | **1.91×** | **1.83×** |
[](/blog/attention-is-the-video-bottleneck/fa4-107.mp4)
**FlashAttention-4** · BF16
1.00× attention speedup
[](/blog/attention-is-the-video-bottleneck/sage2-107.mp4)
**SageAttention2** · 8-bit
[](/blog/attention-is-the-video-bottleneck/vc-107.mp4)
**VC-Attention** · 8-bit
1.59× attention speedup
Play all0:00 / 0:00
[](/blog/attention-is-the-video-bottleneck/fa4-103.mp4)
**FlashAttention-4** · BF16
1.00× attention speedup
[](/blog/attention-is-the-video-bottleneck/sage2-103.mp4)
**SageAttention2** · 8-bit
[](/blog/attention-is-the-video-bottleneck/vc-103.mp4)
**VC-Attention** · 8-bit
1.59× attention speedup
Play all0:00 / 0:00
Two MiniMax-H3 examples generated on an NVIDIA B200 at 1344×768, 243 frames. Speedups are for the attention kernel relative to BF16 FlashAttention-4. Across 100 prompts, VC-Attention achieves 20.2 dB mean PSNR against the BF16 reference, compared with 19.9 dB for SageAttention2. Higher PSNR means the output is closer to the BF16 reference.
Nunchux is building the frontier of multimodal inference: the fastest, cheapest, and highest quality inference for image, video, and world models. For video, the operator that matters most is attention.
The 10 second clips above were generated with MiniMax-H3. On one B200 with BF16 FlashAttention-4, about two thirds of every denoising step is attention. Attention cost grows with the square of the token count, so longer clips make it worse. Low-bit attention offers a path to faster video generation, but on B200 and B300, realizing those gains while preserving fidelity requires addressing both quantization error and the softmax bottleneck.
Today we introduce **VC-Attention**, a faster and more accurate training-free low-bit attention. VC-Attention combines two innovations, V-Smooth and ExpCast-FP8. V-Smooth reduces the quantization error; ExpCast-FP8 removes the softmax bottleneck. On B200, VC-Attention runs the attention kernel 1.6× faster than BF16 FlashAttention-4 [[1]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-1) on MiniMax-H3, and it is more faithful to the BF16 output than SageAttention2 [[2]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-2).
## V-Smooth
Existing methods already keep the query and key product accurate at low precision, for example with a Hadamard transform that spreads outliers across channels. With that in place, value quantization becomes a major source of error. V-Smooth groups the value tokens with a lightweight k-means so each hardware block holds similar tokens, then subtracts the block mean and quantizes only the residual. Across the four models in the report [[4]](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-4), this alone adds 1.1 to 2.8 dB of PSNR over SageAttention2.
rows: value tokens t1…t8cells: channels, darker is largerbrace: one hardware block, one scale
t1
t2
t3
t4
t5
t6
t7
t8
block 1
one scale
+μ1
block 2
one scale
+μ2
Easy to Quantize
Each block gives up its mean μ, stored exactly. What is left is small and evenly spread, and only that goes to the 8-bit quantizer.
1 · Sequence order2 · Group with k-means3 · Subtract the block mean
**V-Smooth, step by step.** Each row is one value token, and a hardware block of tokens shares one quantization scale. In sequence order every block mixes large and small tokens. V-Smooth groups similar tokens into the same block, then subtracts each block’s mean and stores it exactly. The residual that remains is small and evenly spread, so it quantizes with far less error.
## ExpCast-FP8
Low-bit compute speeds up the two matrix products, but the softmax between them still runs in high precision and becomes the longest stage of the kernel. ExpCast-FP8 replaces that stage with a linear approximation that maps each score directly to its FP8 probability code.
Three attention pipelines. BF16 runs QK-transpose in BF16, softmax in FP32 and PV in BF16. SageAttention2 runs QK-transpose in INT8, softmax still in FP32, and PV in FP8. VC-Attention runs QK-transpose in INT8, softmax with ExpCast-FP8, and PV in FP8. A dashed box around the two FP32 softmax stages is labelled: softmax stays in FP32.softmax stays in FP32BF16QK⊤BF16SoftmaxFP32PVBF16SageAttention2QK⊤INT8SoftmaxFP32PVFP8VC-AttentionQK⊤INT8SoftmaxExpCast-FP8PVFP8
**The stage that low-bit Tensor Cores do not touch.** Once both matrix products run in 8-bit, the softmax between them still runs in FP32. ExpCast-FP8 replaces it with a linear approximation that writes the FP8 probability code directly.
Given the log-domain score zzz, the standard path evaluates p=2zp = 2^{z}p=2z in FP32 and then casts ppp to E4M3. An E4M3 byte stores roughly 8log2p+568\log\_2 p + 568log2p+56, so ExpCast-FP8 computes the byte directly with one linear map, 8z+56+β8z + 56 + \beta8z+56+β, rounded to an integer.
##### Standard: exponentiate, then cast
```
fp32 p = exp2(z); // slow FP32 exp
e4m3 p_fp8 = to_e4m3(p); // cast to FP8
```
##### ExpCast-FP8: write the E4M3 byte
```
fp32 c = 8*z + 56 + beta; // fast FMA
uint8 code = round(clip(c, 0, 120));
e4m3 p_fp8 = view_as_e4m3(code); // no-op
```
no exponentialno cast to E4M3
**ExpCast-FP8 skips the slow FP32 exp and the cast to E4M3.**
## From VC-Attention to Nunchux Attention
**Nunchux Attention** is our proprietary extension of VC-Attention. It adds algorithm and kernel optimizations developed for our inference stack. On the MiniMax-H3 attention workload, it runs 1.91× faster than BF16 FlashAttention-4 on B200 and 1.83× faster on B300.
## What this makes possible
VC-Attention speeds up attention in existing video models without retraining. It can be combined with sparse attention [[5](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-5),[6](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-6),[7](https://www.nunchux.ai/blog/attention-is-the-video-bottleneck#ref-7)], few-step distillation, and multi-GPU execution to further reduce the cost of video generation. Together, these improvements make higher-resolution and longer videos more practical to generate.
---
Read the report: [*VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention*](https://arxiv.org/pdf/2609.15810).
Free access to MiniMax-H3 arrives very soon. Join the [Nunchux Modelverse](https://www.nunchux.ai/waitlist) waiting list today to be among the first to use it.
If you run visual models at scale, [contact sales](https://www.nunchux.ai/enterprise). We are hiring. Please visit our [careers page](https://www.nunchux.ai/careers) for more details.
## References
1. [1]Ted Zadouri, Markus Hoehnerbach, Jay Shah, Vijay Thakkar, and Tri Dao. [FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling](https://proceedings.mlsys.org/paper_files/paper/2026/hash/ae8b0b5838ba510daff1198474e7b984-Abstract-Conference.html). MLSys 2026.
2. [2]Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, and Jianfei Chen. [SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization](https://arxiv.org/abs/2411.10958). ICML 2025.
3. [3]Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, and Hao Zhang. [Attn-QAT: 4-Bit Attention with Quantization-Aware Training](https://arxiv.org/abs/2603.00040). arXiv:2603.00040, 2026.
4. [4]Xingyang Li, Dongyun Zou, Shining Zhang, Jiacheng Chen, Haocheng Xi, Lvmin Zhang, Jun-Yan Zhu, Song Han, Zhekai Zhang, Yujun Lin, and Muyang Li. [VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention](https://arxiv.org/abs/2609.15810). arXiv:2609.15810, 2026.
5. [5]Haocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu, Muyang Li, Xiuyu Li, Yujun Lin, Han Cai, Jintao Zhang, Dacheng Li, Jianfei Chen, Ion Stoica, Kurt Keutzer, and Song Han. [Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity](https://arxiv.org/abs/2502.01776). ICML 2025.
6. [6]Xingyang Li, Muyang Li, Tianle Cai, Haocheng Xi, Shuo Yang, Yujun Lin, Lvmin Zhang, Songlin Yang, Jinbo Hu, Kelly Peng, Maneesh Agrawala, Ion Stoica, Kurt Keutzer, and Song Han. [Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation](https://arxiv.org/abs/2506.19852). NeurIPS 2025.
7. [7]Jintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei, Haocheng Xi, Jun Zhu, and Jianfei Chen. [SpargeAttn: Accurate Sparse Attention Accelerating Any Model Inference](https://arxiv.org/abs/2502.18137). ICML 2025.