Loading page…
Figure 5a reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 64. · CiteArk