正在加载页面…
The reported end-to-end TFLOPs/s uses 6 × sequence_length × parameter_count + 12 × layer_count × hidden_dimension × sequence_length². The paper notes that the attention term could be halved for causal masking, but deliberately does not halve it to remain consistent with Megatron-LM and prior reporting. · CiteArk