The study evaluates a Self-Ask-style multi-hop QA pipeline using GPT-4.1-mini as the reasoning model, BM25 retrieval of three passages per generated sub-question, accumulated content-deduplicated passages, and three benchmarks; Contriever is additionally evaluated on MuSiQue and HotpotQA. · CiteArk
The study evaluates a Self-Ask-style multi-hop QA pipeline using GPT-4.1-mini as the reasoning model, BM25 retrieval of three passages per generated sub-question, accumulated content-deduplicated passages, and three benchmarks; Contriever is additionally evaluated on MuSiQue and HotpotQA.
来源:paper:PDF p.3, Figure 2 and §3, paragraphs 'Datasets' and 'QA system'