For both prompt-token concentration measures, the text and multimodal models’ per-dataset values do not overlap, and the text model remains more concentrated in the same direction across all 48 MoE layers. · CiteArk
For both prompt-token concentration measures, the text and multimodal models’ per-dataset values do not overlap, and the text model remains more concentrated in the same direction across all 48 MoE layers.
来源:fixed-paper:PDF p. 11, Section 6.3 and Figure 4