Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz, Shafiq Joty
Why it is worth reading
The paper argues that an agent can reuse experience more effectively by deciding what matters when the next task arrives, rather than permanently compressing experience when it is stored. This produces sizable reported gains and shorter executor interactions, and a trained curator transfers between executor models. The evidence is limited to three benchmarks, uses a simple BM25 retriever, requires an extra model call, relies on benchmark-specific hand-designed payload formats, and does not train on tau2-bench.
Core research claims
- Removing the current task from the curator input lowers untrained JITMEM-base by up to 3.1 ALFWorld and 4.6 WebShop success-rate points; after RL training, the maximum drops widen to 11.4 and 10.4 points, which the paper interprets as evidence that RL learns to exploit the task signal.
- Appendix examples show task-specific payload synthesis on all three benchmarks: combining partially relevant episodes into an ALFWorld clean-and-place procedure, converting unrelated product searches into WebShop search and attribute-selection advice, and distilling MMS cases into an ordered tau2-bench diagnostic that retains the policy requirement for explicit approval before a mutating call.
- On tau2-bench with GPT-5.4 executor, prompted JITMEM-gpt reports 75.6 micro-average success rate, 3.9 points above the strongest baseline, driven by 72.6 on Telecom; on Airline and Retail, the paper says no memory method improves over no memory beyond variance, and JITMEM remains on par with baselines.