The priority score is a cheap trajectory-derived proxy rather than an exact token-influence measure. When multiple tokens are revealed together, a confidence shift cannot be uniquely attributed to one token; the score also depends on the sampled decoding order. The study uses one base model, leaves train/inference masking alignment open, and reports the entropy comparison only with SPG rather than GDPO or ESPO. · CiteArk
The priority score is a cheap trajectory-derived proxy rather than an exact token-influence measure. When multiple tokens are revealed together, a confidence shift cannot be uniquely attributed to one token; the score also depends on the sampled decoding order. The study uses one base model, leaves train/inference masking alignment open, and reports the entropy comparison only with SPG rather than GDPO or ESPO.
来源:paper-fixed:PDF p. 9, Section 6; PDF p. 16, Appendices C.2–C.3