Loading page…
The priority score is a cheap trajectory-derived proxy rather than an exact token-influence measure. When multiple tokens are revealed together, a confidence shift cannot be uniquely attributed to one token; the score also depends on the sampled decoding order. The study uses one base model, leaves train/inference masking alignment open, and reports the entropy comparison only with SPG rather than GDPO or ESPO. · CiteArk