Loading page…
For both prompt-token concentration measures, the text and multimodal models’ per-dataset values do not overlap, and the text model remains more concentrated in the same direction across all 48 MoE layers. · CiteArk