Under the same SPG training budget, the IM reward curve remains above an entropy-scored masking variant throughout training. The paper concludes that marginal token predictability alone does not preserve contextual structure as effectively as the rollout-derived neighbor-confidence-shift score. · CiteArk
Under the same SPG training budget, the IM reward curve remains above an entropy-scored masking variant throughout training. The paper concludes that marginal token predictability alone does not preserve contextual structure as effectively as the rollout-derived neighbor-confidence-shift score.
来源:paper-fixed:PDF pp. 8–9, Section 4.2 Q5 and Figure 3(b)