Explore ArkGraph and select the steps to run.
This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.
Finding the right document is not the same as finding the fact needed to answer a reasoning step. The paper argues that standard document-recall metrics hide this difference, so improving retrieval alone can leave many multi-hop failures untouched. A lightweight predictor can target extra retrieval more efficiently on some datasets, but the effect depends on the dataset, question type, retriever, and reasoning model; the work does not solve cases where the corpus passage itself omits the needed relation.
The paper’s claims are available in Research claims.
Past research reports
Saved reports remain available. New report generation is paused.