正在加载页面…
The H100 measurements use the unmodified implementation and no H100-specific TMA or fourth-generation Tensor Core instructions. The authors forecast, rather than measure, a further 1.5–2× speedup from those features and leave that optimization to future work. · CiteArk