CIGAsk trains a multi-turn GRPO policy with an outcome reward, per-clarification Counterfactual Information Gain computed by a frozen reference model, and a terminal Asymmetric Ambiguity Bonus keyed to the gold ambiguity label. · CiteArk
CIGAsk trains a multi-turn GRPO policy with an outcome reward, per-clarification Counterfactual Information Gain computed by a frozen reference model, and a terminal Asymmetric Ambiguity Bonus keyed to the gold ambiguity label.
来源:paper:PDF pp. 3-5, Section 3, especially Equations 2-8; PDF p. 11, Table 7