The experiments use veRL with vLLM rollouts and FSDP actor training; OPD-based methods use one response per prompt, batch and PPO minibatch size 1024, maximum prompt length 2048, maximum response length 16384, RADAR with constant 1×10^-5 learning rate, temperature/top-p 1.0/1.0, and tensor parallel size 4. Math uses a DAPO-style boxed-answer verifier and code uses execution tests; binary outcomes map to ±1. · CiteArk
Not assessedNo independent reproduction scheduledMethodclaim-experimental-configuration
The experiments use veRL with vLLM rollouts and FSDP actor training; OPD-based methods use one response per prompt, batch and PPO minibatch size 1024, maximum prompt length 2048, maximum response length 16384, RADAR with constant 1×10^-5 learning rate, temperature/top-p 1.0/1.0, and tensor parallel size 4. Math uses a DAPO-style boxed-answer verifier and code uses execution tests; binary outcomes map to ±1.
Source: paper:PDF pp. 16–17, Table 4 and Appendix E
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.