Loading page…
Table 1 reports per-GPU training throughput for GPT3-2.7B at 8k context on 8×A100 GPUs for the baseline without FlashAttention, FlashAttention, and FlashAttention-2. · CiteArk