学科 信息检索 · 2 个仓库
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
0/29This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.
Just-In-Time Agent Memory with Runtime Agentic Research
0/18This paper introduces Just-In-Time Agent Memory (JAM), which constructs query-specific context by exploring complete stored histories at runtime. A fixed Memorizer organizes raw sessions into a hierarchical workspace with navigational summaries, while a trainable Researcher searches, opens directories, and browses evidence. Memory-Gym synthesizes evidence-grounded tasks across nine task types and six domains. Researcher training combines verified-trajectory supervised fine-tuning with hint-guided reinforcement learning. Evaluations compare JAM with retrieval, ahead-of-time memory systems, and trained memory agents on conversational memory, long-document reasoning, and multi-hop tasks. Additional analyses examine code-domain transfer, component ablations, runtime scaling, human data-quality audits, inference variability, and efficiency.