Removing the current task from the curator input lowers untrained JITMEM-base by up to 3.1 ALFWorld and 4.6 WebShop success-rate points; after RL training, the maximum drops widen to 11.4 and 10.4 points, which the paper interprets as evidence that RL learns to exploit the task signal.
Training-free JITMEM design ablations on WebShop
实验方案已经存在,但当前处理任务没有对应的执行目标。
RL-trained curator task/retrieval ablations on WebShop
实验方案已经存在,但当前处理任务没有对应的执行目标。
Training-free JITMEM design ablations on ALFWorld
实验方案已经存在,但当前处理任务没有对应的执行目标。
RL-trained curator task/retrieval ablations on ALFWorld
实验方案已经存在,但当前处理任务没有对应的执行目标。