正在加载页面…
At fixed mask ratios on generated completions from all four benchmarks, downstream-biased masking yields higher average target-token log-probability on masked positions than uniform random masking, whereas upstream-biased masking yields substantially lower confidence. The separation grows with stronger priority bias; the paper interprets downstream subproblems as better posed because upstream reasoning context remains visible. · CiteArk