正在加载页面…
For 1.3B- and 2.7B-parameter GPT-style models at 2k and 8k context, FlashAttention-2 is summarized as up to 1.3× faster than FlashAttention and 2.8× faster than a baseline without FlashAttention, reaching 225 TFLOPs/s/GPU and 72% model-FLOPs utilization. · CiteArk