正在加载页面…
For untrained JITMEM-base, storing all trajectories with correctness labels rather than filtering to successful trajectories lowers success by 1.5-2.9 points on ALFWorld and 2.3-3.4 on WebShop across executors; replacing raw traces with write-time ReasoningBank-style distillations lowers success by 1.7-2.9 on ALFWorld and 6.8-8.2 on WebShop. · CiteArk