Loading page…
The experiments use veRL with vLLM rollouts and FSDP actor training; OPD-based methods use one response per prompt, batch and PPO minibatch size 1024, maximum prompt length 2048, maximum response length 16384, RADAR with constant 1×10^-5 learning rate, temperature/top-p 1.0/1.0, and tensor parallel size 4. Math uses a DAPO-style boxed-answer verifier and code uses execution tests; binary outcomes map to ±1. · CiteArk