正在加载页面…
Removing the current task from the curator input lowers untrained JITMEM-base by up to 3.1 ALFWorld and 4.6 WebShop success-rate points; after RL training, the maximum drops widen to 11.4 and 10.4 points, which the paper interprets as evidence that RL learns to exploit the task signal. · CiteArk