自然语言处理实证方法会议 · 3 篇论文
Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
citeark/finding-where-the-buck-stops-an-automated-failure-attribution-based-reflection-f--j878vv
The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
citeark/when-and-how-should-an-agent-clarify-cigask-teaching-llms-to-clarify-via-counter--iisl0l
The paper presents CIGAsk, a reinforcement-learning recipe for teaching language models both when to request clarification and how to formulate an informative question. In multi-turn GRPO, Counterfactual Information Gain rewards clarification turns according to how much a frozen reference model's likelihood of the gold answer increases after the simulated user's response, while an asymmetric ambiguity bonus rewards selective asking. Experiments use Qwen2.5 3B and 7B policies on PACIFIC, AbgCoQA, and AmbigNQ. The authors report gains over prompting, supervised warmstarts, and contextual external baselines, component and reference-model ablations, cross-backbone transfer, simulator sensitivity, and retention on two out-of-domain closed-book QA datasets.
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
citeark/diagnosing-the-fact-grounding-gap-in-multi-hop-question-answering--arkkn9
This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.