正在加载页面…
The CUDA-graph-captured per-layer attention primitive was faster than FP16-SDPA through rank 64 at all tested lengths and bit widths, but slower at full rank 128. · CiteArk