When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
Yunxiang Li, Xixin Wu, Helen Meng
为什么值得读
CIGAsk is evidence that a model can learn selective, useful clarification from ordinary answer supervision plus ambiguity labels, rather than merely being prompted to ask. Its frozen-reference reward provides local feedback about whether a user's reply actually helps answer the question. The evidence is promising but narrow: results are single-seed, use cooperative simulators with latent-intent access, cover English benchmarks and 3B-9B models, and some baseline comparisons use different evaluation protocols. The approach also still needs ambiguity labels during training.
核心研究结论
- Sensitivity to the ambiguity-bonus weight is non-monotonic on PACIFIC: among rho values 0.1, 0.3, and 0.5, rho=0.3 has the highest overall F1, ambiguous-query clarification rate, and reported selectivity. Raising rho to 0.5 reduces F1 and recall and does not improve selectivity over the main setting.
- On unseen closed-book TriviaQA and NQ Open subsets, forced-direct CIGAsk 7B slightly exceeds the Qwen2.5-7B base F1, while under a multi-action prompt it clarifies only 5.6% and 9.4% of queries, respectively.
- Replacing GPT-4o with Qwen2.5-7B-Instruct as the fixed latent-intent-aware simulator leaves PACIFIC performance close to the main setting and slightly raises selectivity, whereas Qwen2.5-3B-Instruct substantially lowers post-clarification F1 and ambiguous-query recall.