正在加载页面…
The Based and ReBased kernel-approximation variants do not improve over the corresponding simple activation or normalization baselines in the reported 130M and 380M ablations. The paper notes that the negative result may reflect the difference between SSD's 1-semiseparable mask and vanilla linear-attention methods, and that the expanded-feature variants use smaller B/C projections and adjusted layer counts. · CiteArk