Loading page…
Forcing the RL-trained curator to receive no retrieved trajectories degrades it to or below untrained JITMEM-base across all executors, with maximum reported success-rate drops of 14.8 points on ALFWorld and 15.2 on WebShop; the paper takes this as evidence that RL learns to distill retrieved experience rather than only emit hints from parametric knowledge. · CiteArk