Explore ArkGraph and select the steps to run.
This paper introduces JITMEM, an agent-memory framework that stores complete trajectories and postpones their distillation until a new task is known. A task-conditioned curator converts retrieved raw traces into a compact payload for a frozen executor; the curator is optimized with GRPO using immediate task reward. Experiments on ALFWorld, WebShop, and tau2-bench compare prompted and trained read-time curation with no-memory, heuristic, and learned write-time systems across several executors. Reported results favor read-time curation in success, context efficiency, and cross-executor transfer. Ablations attribute the gains to task conditioning, successful-trajectory filtering, retention of raw traces, and actual use of retrieved experience.
The paper argues that an agent can reuse experience more effectively by deciding what matters when the next task arrives, rather than permanently compressing experience when it is stored. This produces sizable reported gains and shorter executor interactions, and a trained curator transfers between executor models. The evidence is limited to three benchmarks, uses a simple BM25 retriever, requires an extra model call, relies on benchmark-specific hand-designed payload formats, and does not train on tau2-bench.
The paper’s claims are available in Research claims.