Loading page…
Figure 5c reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 64. · CiteArk