videoJul 17, 2026 · 20 min
Fixing block-sparse INT8 attention for Wan 2.2 — and a 1.78× two-expert cache
The published SOTA sparse-attention kernel gives ~1× on Wan 2.2 out of the box, because its block selector marks most video-latent blocks 'unpredictable.' Swapping only the selector recovers 1.3–1.76×. Plus a step cache tuned to the model's two experts, and a free sparsity signal already in the kernel's registers.
video generationattentionsparsityquantization
1.78×two-expert step cacheno cache ▸ per-expert