Loading page…
The compared models use matched sparse-expert and hybrid-attention patterns but different widths, heads, and body sizes; in SST only the Prefiller contains K/V projections and K normalization, while both towers have independent Transformer-block parameters. · CiteArk