Loading page…
Figure 5b reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 128. · CiteArk