Across 100 GRPO steps on ALFWorld and WebShop, the plotted training and validation success rates generally rise while executor turns fall; the paper describes training as stable under task reward alone, without auxiliary content-quality reward, task grouping, or return shaping. · CiteArk
Not assessedPlan blockedFindingclaim-training-curves-017
Across 100 GRPO steps on ALFWorld and WebShop, the plotted training and validation success rates generally rise while executor turns fall; the paper describes training as stable under task reward alone, without auxiliary content-quality reward, task grouping, or return shaping.
Source: paper:PDF pp. 24-25, Figures 5 and 6 and accompanying paragraph
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.