Across 100 GRPO steps on ALFWorld and WebShop, the plotted training and validation success rates generally rise while executor turns fall; the paper describes training as stable under task reward alone, without auxiliary content-quality reward, task grouping, or return shaping. · CiteArk
Across 100 GRPO steps on ALFWorld and WebShop, the plotted training and validation success rates generally rise while executor turns fall; the paper describes training as stable under task reward alone, without auxiliary content-quality reward, task grouping, or return shaping.
来源:paper:PDF pp. 24-25, Figures 5 and 6 and accompanying paragraph