浏览 ArkGraph,选择本次要执行的步骤。
The paper presents CIGAsk, a reinforcement-learning recipe for teaching language models both when to request clarification and how to formulate an informative question. In multi-turn GRPO, Counterfactual Information Gain rewards clarification turns according to how much a frozen reference model's likelihood of the gold answer increases after the simulated user's response, while an asymmetric ambiguity bonus rewards selective asking. Experiments use Qwen2.5 3B and 7B policies on PACIFIC, AbgCoQA, and AmbigNQ. The authors report gains over prompting, supervised warmstarts, and contextual external baselines, component and reference-model ablations, cross-backbone transfer, simulator sensitivity, and retention on two out-of-domain closed-book QA datasets.
CIGAsk is evidence that a model can learn selective, useful clarification from ordinary answer supervision plus ambiguity labels, rather than merely being prompted to ask. Its frozen-reference reward provides local feedback about whether a user's reply actually helps answer the question. The evidence is promising but narrow: results are single-seed, use cooperative simulators with latent-intent access, cover English benchmarks and 3B-9B models, and some baseline comparisons use different evaluation protocols. The approach also still needs ambiguity labels during training.
论文中的结论可在「研究结论」中查看。