For causal attention, skipping blocks wholly above the causal diagonal is reported to make the kernel around 1.7–1.8× faster than attention without a causal mask; only one square block per row needs elementwise causal masking. · CiteArk
For causal attention, skipping blocks wholly above the causal diagonal is reported to make the kernel around 1.7–1.8× faster than attention without a causal mask; only one square block per row needs elementwise causal masking.