Explore ArkGraph and select the steps to run.
This paper studies how masked reconstruction subproblems are selected when reinforcement-learning objectives for diffusion language models can sample only a few perturbations per rollout. It identifies an upstream/downstream structure from confidence changes along denoising trajectories and defines a token priority score. Informed Masking uses that score to bias perturbations toward downstream tokens while leaving the host method’s reward, mask budget, and estimator intact. Experiments with LLaDA-8B-Instruct across GSM8K, MATH-500, Countdown, and Sudoku integrate the method into ESPO, GDPO, and SPG. Reported results show generally higher accuracy, better-conditioned reconstruction estimates, reduced gradient dispersion, faster convergence, and improved stability, alongside limitations in attribution and decoding-order dependence.
In plain language, the paper argues that reinforcement learning for diffusion language models benefits from hiding tokens the visible reasoning can still determine, instead of choosing hidden tokens uniformly. Its masking rule uses information already produced during generation and improves the reported results across several math and planning tasks. The evidence is limited to one 8-billion-parameter base model and a small set of host algorithms; the score is only a trajectory-dependent proxy for influence, not a causal or order-invariant measure.
The paper’s claims are available in Research claims.