正在加载页面…
On PACIFIC fullval, CIGAsk 7B reports the best overall F1, post-clarification F1, and ambiguous-query clarification rate in Table 2; CIGAsk 3B also exceeds all prompting and SFT baselines. Scaling CIGAsk from 3B to 7B raises F1, post-clarification F1, and ambiguous-query recall while also raising the clear-query false-positive rate from 0.179 to 0.260. Prompting and SFT expose negative behaviors: Direct and SFT never clarify, FATA clarifies every input and has low F1, and ReAct clarifies more often on clear than ambiguous inputs. · CiteArk