In a matched-seed 3B comparison, CIGAsk and outcome-only GRPO learned to clarify at similar rates, but only the CIG-equipped model consistently converted those turns into improved post-clarification F1; the outcome-only variant fell to near zero. · CiteArk
In a matched-seed 3B comparison, CIGAsk and outcome-only GRPO learned to clarify at similar rates, but only the CIG-equipped model consistently converted those turns into improved post-clarification F1; the outcome-only variant fell to near zero.
来源:paper:PDF p. 7, Section 5.1; PDF p. 11, Appendix A.1