Nunchux AI released VC-Attention, a training-free low-bit attention kernel for video diffusion transformers. It pairs V-Smooth, which clusters value tokens and quantizes residuals after block-mean subtraction, with ExpCast-FP8, which replaces the FP32 exponential with a single multiply-add. Benchmarks report 6.02× speed over SageAttention2 on B200 and 1.60× over BF16 FlashAttention-4 on MiniMax-H3, with PSNR gains on Wan2.2-14B. No public kernel release yet.
No score is assigned. Sources and their independence are shown in the citation chain below.