浏览 ArkGraph,选择本次要执行的步骤。
论文中的结论可在「研究结论」中查看。
0 / 26 条结论已通过验证
其余仍在验证中
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
rho: 0.1 | 0.702 score (0-1) | — | 尚未评估 |
rho: 0.1 | 0.681 proportion | — | 尚未评估 |
rho: 0.1 | 0.175 proportion | — | 尚未评估 |
rho: 0.1 | 0.506 proportion-point difference | — | 尚未评估 |
rho: 0.3 | 0.795 score (0-1) | — | 尚未评估 |
rho: 0.3 | 0.873 proportion | — | 尚未评估 |
rho: 0.3 | 0.26 proportion | — | 尚未评估 |
rho: 0.3 | 0.613 proportion-point difference | — | 尚未评估 |
rho: 0.5 | 0.703 score (0-1) | — | 尚未评估 |
rho: 0.5 | 0.789 proportion | — | 尚未评估 |
rho: 0.5 | 0.188 proportion | — | 尚未评估 |
rho: 0.5 | 0.601 proportion-point difference | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: CIGAsk 7B · prompt: forced-direct · dataset: TriviaQA | 56.2 F1 points | — | 尚未评估 |
model: Qwen2.5-7B base · prompt: forced-direct · dataset: TriviaQA | 54 F1 points | — | 尚未评估 |
model: CIGAsk 7B · prompt: multi-action · dataset: TriviaQA | 5.6% | — | 尚未评估 |
model: CIGAsk 7B · prompt: forced-direct · dataset: NQ Open | 25.9 F1 points | — | 尚未评估 |
model: Qwen2.5-7B base · prompt: forced-direct · dataset: NQ Open | 25.1 F1 points | — | 尚未评估 |
model: CIGAsk 7B · prompt: multi-action · dataset: NQ Open | 9.4% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
simulator: GPT-4o | 0.795 score (0-1) | — | 尚未评估 |
simulator: GPT-4o | 0.768 score (0-1) | — | 尚未评估 |
simulator: GPT-4o | 0.873 proportion | — | 尚未评估 |
simulator: GPT-4o | 0.613 proportion-point difference | — | 尚未评估 |
simulator: Qwen2.5-7B-Instruct | 0.783 score (0-1) | — | 尚未评估 |
simulator: Qwen2.5-7B-Instruct | 0.752 score (0-1) | — | 尚未评估 |
simulator: Qwen2.5-7B-Instruct | 0.856 proportion | — | 尚未评估 |
simulator: Qwen2.5-7B-Instruct | 0.636 proportion-point difference | — | 尚未评估 |
simulator: Qwen2.5-3B-Instruct | 0.677 score (0-1) | — | 尚未评估 |
simulator: Qwen2.5-3B-Instruct | 0.536 score (0-1) | — | 尚未评估 |
simulator: Qwen2.5-3B-Instruct | 0.462 proportion | — | 尚未评估 |
simulator: Qwen2.5-3B-Instruct | 0.332 proportion-point difference | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/33 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: Direct · backbone: Qwen2.5-7B | 0.312 score (0-1) | — | 尚未评估 |
method: Direct · backbone: Qwen2.5-7B | 0 proportion | — | 尚未评估 |
method: Direct · backbone: Qwen2.5-7B | 0 proportion | — | 尚未评估 |
method: FATA · backbone: Qwen2.5-7B | 0.177 score (0-1) | — | 尚未评估 |
method: FATA · backbone: Qwen2.5-7B | 0.177 score (0-1) | — | 尚未评估 |
method: FATA · backbone: Qwen2.5-7B | 1 proportion | — | 尚未评估 |
method: FATA · backbone: Qwen2.5-7B | 1 proportion | — | 尚未评估 |
method: ReAct · backbone: Qwen2.5-7B | 0.481 score (0-1) | — | 尚未评估 |
method: ReAct · backbone: Qwen2.5-7B | 0.418 score (0-1) | — | 尚未评估 |
method: ReAct · backbone: Qwen2.5-7B | 0.329 proportion | — | 尚未评估 |
method: ReAct · backbone: Qwen2.5-7B | 0.41 proportion | — | 尚未评估 |
method: SFT · backbone: Qwen2.5-3B | 0.543 score (0-1) | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/16 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.795 score (0-1) | — | 尚未评估 |
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.768 score (0-1) | — | 尚未评估 |
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.873 proportion | — | 尚未评估 |
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.26 proportion | — | 尚未评估 |
CIG: true · variant: -asym · asymmetricBonus: false | 0.736 score (0-1) | — | 尚未评估 |
CIG: true · variant: -asym · asymmetricBonus: false | 0.733 score (0-1) | — | 尚未评估 |
CIG: true · variant: -asym · asymmetricBonus: false | 0.171 proportion | — | 尚未评估 |
CIG: true · variant: -asym · asymmetricBonus: false | 0.036 proportion | — | 尚未评估 |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.744 score (0-1) | — | 尚未评估 |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.7 score (0-1) | — | 尚未评估 |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.763 proportion | — | 尚未评估 |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.167 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
reference: frozen | 0.795 score (0-1) | — | 尚未评估 |
reference: frozen | 0.768 score (0-1) | — | 尚未评估 |
reference: frozen | 0.873 proportion | — | 尚未评估 |
reference: frozen | 0.613 proportion-point difference | — | 尚未评估 |
steps: 1-100 · reference: frozen | 0.62 reward units | — | 尚未评估 |
steps: 200-300 · reference: frozen | 0.78 reward units | — | 尚未评估 |
reference: current actor | 0.607 score (0-1) | — | 尚未评估 |
reference: current actor | 0.359 score (0-1) | — | 尚未评估 |
reference: current actor | 0.233 proportion | — | 尚未评估 |
reference: current actor | 0.135 proportion-point difference | — | 尚未评估 |
steps: 1-100 · reference: current actor | 0.47 reward units | — | 尚未评估 |
steps: 200-300 · reference: current actor | 0.53 reward units | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/7 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: Direct · backbone: Qwen2.5-7B | 0.32 score (0-1) | — | 尚未评估 |
method: FATA · backbone: Qwen2.5-7B | 0.228 score (0-1) | — | 尚未评估 |
method: ReAct · backbone: Qwen2.5-7B | 0.554 score (0-1) | — | 尚未评估 |
method: SFT · backbone: Qwen2.5-3B | 0.589 score (0-1) | — | 尚未评估 |
method: SFT · backbone: Qwen2.5-7B | 0.638 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Qwen2.5-3B | 0.635 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Qwen2.5-7B | 0.724 score (0-1) | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: ACT · status: externally reported · backbone: Zephyr-7B | 0.681 score (0-1) | — | 尚未评估 |
method: ACT · status: externally reported · backbone: Zephyr-7B | 0.62 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Zephyr-7B | 0.709 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Zephyr-7B | 0.686 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Zephyr-7B | 0.712 proportion | — | 尚未评估 |
method: CIGAsk · backbone: Zephyr-7B | 0.044 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/11 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
split: train · dataset: PACIFIC | 15,087 examples | — | 尚未评估 |
split: fullval · dataset: PACIFIC | 1,952 examples | — | 尚未评估 |
dataset: PACIFIC | 16% | — | 尚未评估 |
split: fullval · dataset: PACIFIC | 316 examples | — | 尚未评估 |
split: fullval · dataset: PACIFIC | 1,636 examples | — | 尚未评估 |
split: train · dataset: AbgCoQA | 7,269 examples | — | 尚未评估 |
split: fullval · dataset: AbgCoQA | 1,184 examples | — | 尚未评估 |
dataset: AbgCoQA | 22% | — | 尚未评估 |
split: train · dataset: AmbigNQ | 19,244 examples | — | 尚未评估 |
split: test · dataset: AmbigNQ | 4,377 examples | — | 尚未评估 |
dataset: AmbigNQ | 79% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
simulator: Qwen2.5-7B-Instruct · auditedCases: 8 | 8 cases | — | 尚未评估 |
simulator: Qwen2.5-3B-Instruct · auditedCases: 8 | 1 cases | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
series: IG on clarify turns · trainingStep: 5 | 0.17 CIG signal (nats) | — | 尚未评估 |
series: IG on clarify turns · trainingStep: 250 | 0.47 CIG signal (nats) | — | 尚未评估 |
series: IG batch-weighted · trainingStep: 5 | 0.02 CIG signal (nats) | — | 尚未评估 |
series: IG batch-weighted · trainingStep: 250 | 0.23 CIG signal (nats) | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/10 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: SGP · status: externally reported · backbone: Gemma-2-9B | 0.726 score (0-1) | — | 尚未评估 |
method: SGP · status: externally reported · backbone: Gemma-2-9B | 0.429 proportion | — | 尚未评估 |
method: SGP · status: externally reported · backbone: Gemma-2-9B | 0.206 proportion | — | 尚未评估 |
method: SGP-Oracle · status: externally reported · backbone: Gemma-2-9B · usesGoldAmbiguityLabels: true | 0.787 score (0-1) | — | 尚未评估 |
method: SGP-Oracle · status: externally reported · backbone: Gemma-2-9B · usesGoldAmbiguityLabels: true | 0.435 proportion | — | 尚未评估 |
method: SGP-Oracle · status: externally reported · backbone: Gemma-2-9B · usesGoldAmbiguityLabels: true | 0.156 proportion | — | 尚未评估 |
method: CIGAsk · backbone: Gemma-2-9B | 0.774 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Gemma-2-9B | 0.749 score (0-1) | — | 尚未评估 |
method: CIGAsk · backbone: Gemma-2-9B | 0.813 proportion | — | 尚未评估 |
method: CIGAsk · backbone: Gemma-2-9B | 0.031 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/9 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: Direct · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.058 score (0-1) | — | 尚未评估 |
method: FATA · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.076 score (0-1) | — | 尚未评估 |
method: ReAct · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.076 score (0-1) | — | 尚未评估 |
method: SFT · dataset: AmbigNQ · backbone: Qwen2.5-3B | 0.056 score (0-1) | — | 尚未评估 |
method: SFT · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.109 score (0-1) | — | 尚未评估 |
method: SGP · status: externally reported anchor · dataset: AmbigQA · backbone: Gemma-2-9B · protocol: 50/50 balanced dev | 0.306 score (0-1) | — | 尚未评估 |
method: SGP-Oracle · status: externally reported anchor · dataset: AmbigQA · backbone: Gemma-2-9B · protocol: 50/50 balanced dev · usesGoldAmbiguityLabels: true | 0.359 score (0-1) | — | 尚未评估 |
method: CIGAsk · dataset: AmbigNQ · backbone: Qwen2.5-3B | 0.271 score (0-1) | — | 尚未评估 |
method: CIGAsk · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.511 score (0-1) | — | 尚未评估 |
The printed transcript is a qualitative case, not an automatically reproducible aggregate. Regenerating it requires the unreleased trained state and an exact stochastic simulator snapshot; the schema has no source-defined rubric for deciding transcript equivalence, so the actual targeted-versus-generic comparison is retained without fabricating a number.
The paper omits the late-checkpoint set, point estimates, bootstrap resampling procedure, and interval endpoints. Those missing comparison definitions prevent a concrete automatic package without fabricating a decision rule.
The appendix prints three sensible examples but supplies no sample-wide coding rubric or human-evaluation decision rule. They are carried into the OOD objective as qualitative context, not converted to a scalar target.
The paper gives a qualitative matched-seed trajectory claim but no exact values, checkpoint range, curve, or registered numeric target. The comparison is retained as source context for the PACIFIC 3B workflow; a deterministic decision rule must be established from additional source evidence before automatic verification.
The qualitative “roughly doubles” finding has no counts, denominator, coding rubric, or exact rate. Self-completion and declarative reformulation remain source context for the component ablation, but no scalar target is invented.
This is method context rather than a separately reported empirical result. Its equations and controls are carried into every training objective; correctness is assessed through the bound empirical packages rather than by inventing a method-only scalar.
The English-only scope is descriptive of the evaluated datasets; multilingual behavior is unreported and is not added as a new paper-bound experiment.
This is an ethical/generalization limitation about annotator-dependent ambiguity labels. The rho sensitivity package preserves the source labels but cannot establish behavior for unobserved user populations.
The stated lack of robustness to noisy, adversarial, diverse, or human users is an open scope limitation. The simulator packages test only the source conditions and cannot establish the untested populations.
This is a source limitation, not an independent empirical result. Every training package preserves seed 42 and does not imply variance estimation.
The paper explicitly marks ACT/SGP values as original-protocol context anchors. The catalog binds reconstructed CIGAsk and in-protocol prompting controls only and does not treat backbone matching as protocol matching.
This is a direct method/input requirement from the equations: every planned reconstruction requires the benchmark ambiguity label during training; no separate result is claimed.
The absence of experiments above 9B is an explicit scope limitation, not a positive empirical result to test by substituting an unreported larger model.