Loading page…
On WebShop, rebuilding the training bank after 100 GRPO steps and training 50 more raises success by 2.8 points for Qwen3-8B and 0.9 for GPT-5.4 but leaves Gemini-2.5-Pro unchanged; warm-starting the test bank with 100 training trajectories changes success by at most 1.3 points and remains within standard deviation. · CiteArk