正在加载页面…
JAM has the highest LLM-judge accuracy among systems evaluated with the same GPT-4o-mini protocol on LoCoMo, NarrativeQA and HotpotQA. Its gains over the strongest baseline in each benchmark are 12.54, 20.00 and 11.20 percentage points. Memory-R1 scores are quoted separately and are not directly comparable under this judging pipeline. · CiteArk