正在加载页面…
On tau2-bench with GPT-5.4 executor, prompted JITMEM-gpt reports 75.6 micro-average success rate, 3.9 points above the strongest baseline, driven by 72.6 on Telecom; on Airline and Retail, the paper says no memory method improves over no memory beyond variance, and JITMEM remains on par with baselines. · CiteArk