Explore ArkGraph and select the steps to run.
The paper’s claims are available in Research claims.
0 / 26 claims verified
the rest still being verified
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
rho: 0.1 | 0.702 score (0-1) | — | Not assessed |
rho: 0.1 | 0.681 proportion | — | Not assessed |
rho: 0.1 | 0.175 proportion | — | Not assessed |
rho: 0.1 | 0.506 proportion-point difference | — | Not assessed |
rho: 0.3 | 0.795 score (0-1) | — | Not assessed |
rho: 0.3 | 0.873 proportion | — | Not assessed |
rho: 0.3 | 0.26 proportion | — | Not assessed |
rho: 0.3 | 0.613 proportion-point difference | — | Not assessed |
rho: 0.5 | 0.703 score (0-1) | — | Not assessed |
rho: 0.5 | 0.789 proportion | — | Not assessed |
rho: 0.5 | 0.188 proportion | — | Not assessed |
rho: 0.5 | 0.601 proportion-point difference | — | Not assessed |
Experiment plan ready; no runs yet.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: CIGAsk 7B · prompt: forced-direct · dataset: TriviaQA | 56.2 F1 points | — | Not assessed |
model: Qwen2.5-7B base · prompt: forced-direct · dataset: TriviaQA | 54 F1 points | — | Not assessed |
model: CIGAsk 7B · prompt: multi-action · dataset: TriviaQA | 5.6% | — | Not assessed |
model: CIGAsk 7B · prompt: forced-direct · dataset: NQ Open | 25.9 F1 points | — | Not assessed |
model: Qwen2.5-7B base · prompt: forced-direct · dataset: NQ Open | 25.1 F1 points | — | Not assessed |
model: CIGAsk 7B · prompt: multi-action · dataset: NQ Open | 9.4% | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
simulator: GPT-4o | 0.795 score (0-1) | — | Not assessed |
simulator: GPT-4o | 0.768 score (0-1) | — | Not assessed |
simulator: GPT-4o | 0.873 proportion | — | Not assessed |
simulator: GPT-4o | 0.613 proportion-point difference | — | Not assessed |
simulator: Qwen2.5-7B-Instruct | 0.783 score (0-1) | — | Not assessed |
simulator: Qwen2.5-7B-Instruct | 0.752 score (0-1) | — | Not assessed |
simulator: Qwen2.5-7B-Instruct | 0.856 proportion | — | Not assessed |
simulator: Qwen2.5-7B-Instruct | 0.636 proportion-point difference | — | Not assessed |
simulator: Qwen2.5-3B-Instruct | 0.677 score (0-1) | — | Not assessed |
simulator: Qwen2.5-3B-Instruct | 0.536 score (0-1) | — | Not assessed |
simulator: Qwen2.5-3B-Instruct | 0.462 proportion | — | Not assessed |
simulator: Qwen2.5-3B-Instruct | 0.332 proportion-point difference | — | Not assessed |
Experiment plan ready; no runs yet.
0/33 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: Direct · backbone: Qwen2.5-7B | 0.312 score (0-1) | — | Not assessed |
method: Direct · backbone: Qwen2.5-7B | 0 proportion | — | Not assessed |
method: Direct · backbone: Qwen2.5-7B | 0 proportion | — | Not assessed |
method: FATA · backbone: Qwen2.5-7B | 0.177 score (0-1) | — | Not assessed |
method: FATA · backbone: Qwen2.5-7B | 0.177 score (0-1) | — | Not assessed |
method: FATA · backbone: Qwen2.5-7B | 1 proportion | — | Not assessed |
method: FATA · backbone: Qwen2.5-7B | 1 proportion | — | Not assessed |
method: ReAct · backbone: Qwen2.5-7B | 0.481 score (0-1) | — | Not assessed |
method: ReAct · backbone: Qwen2.5-7B | 0.418 score (0-1) | — | Not assessed |
method: ReAct · backbone: Qwen2.5-7B | 0.329 proportion | — | Not assessed |
method: ReAct · backbone: Qwen2.5-7B | 0.41 proportion | — | Not assessed |
method: SFT · backbone: Qwen2.5-3B | 0.543 score (0-1) | — | Not assessed |
Experiment plan ready; no runs yet.
0/16 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.795 score (0-1) | — | Not assessed |
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.768 score (0-1) | — | Not assessed |
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.873 proportion | — | Not assessed |
CIG: true · variant: CIGAsk 7B · asymmetricBonus: true | 0.26 proportion | — | Not assessed |
CIG: true · variant: -asym · asymmetricBonus: false | 0.736 score (0-1) | — | Not assessed |
CIG: true · variant: -asym · asymmetricBonus: false | 0.733 score (0-1) | — | Not assessed |
CIG: true · variant: -asym · asymmetricBonus: false | 0.171 proportion | — | Not assessed |
CIG: true · variant: -asym · asymmetricBonus: false | 0.036 proportion | — | Not assessed |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.744 score (0-1) | — | Not assessed |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.7 score (0-1) | — | Not assessed |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.763 proportion | — | Not assessed |
CIG: false · variant: -CIG · asymmetricBonus: true | 0.167 proportion | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
reference: frozen | 0.795 score (0-1) | — | Not assessed |
reference: frozen | 0.768 score (0-1) | — | Not assessed |
reference: frozen | 0.873 proportion | — | Not assessed |
reference: frozen | 0.613 proportion-point difference | — | Not assessed |
steps: 1-100 · reference: frozen | 0.62 reward units | — | Not assessed |
steps: 200-300 · reference: frozen | 0.78 reward units | — | Not assessed |
reference: current actor | 0.607 score (0-1) | — | Not assessed |
reference: current actor | 0.359 score (0-1) | — | Not assessed |
reference: current actor | 0.233 proportion | — | Not assessed |
reference: current actor | 0.135 proportion-point difference | — | Not assessed |
steps: 1-100 · reference: current actor | 0.47 reward units | — | Not assessed |
steps: 200-300 · reference: current actor | 0.53 reward units | — | Not assessed |
Experiment plan ready; no runs yet.
0/7 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: Direct · backbone: Qwen2.5-7B | 0.32 score (0-1) | — | Not assessed |
method: FATA · backbone: Qwen2.5-7B | 0.228 score (0-1) | — | Not assessed |
method: ReAct · backbone: Qwen2.5-7B | 0.554 score (0-1) | — | Not assessed |
method: SFT · backbone: Qwen2.5-3B | 0.589 score (0-1) | — | Not assessed |
method: SFT · backbone: Qwen2.5-7B | 0.638 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Qwen2.5-3B | 0.635 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Qwen2.5-7B | 0.724 score (0-1) | — | Not assessed |
Experiment plan ready; no runs yet.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: ACT · status: externally reported · backbone: Zephyr-7B | 0.681 score (0-1) | — | Not assessed |
method: ACT · status: externally reported · backbone: Zephyr-7B | 0.62 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Zephyr-7B | 0.709 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Zephyr-7B | 0.686 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Zephyr-7B | 0.712 proportion | — | Not assessed |
method: CIGAsk · backbone: Zephyr-7B | 0.044 proportion | — | Not assessed |
Experiment plan ready; no runs yet.
0/11 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
split: train · dataset: PACIFIC | 15,087 examples | — | Not assessed |
split: fullval · dataset: PACIFIC | 1,952 examples | — | Not assessed |
dataset: PACIFIC | 16% | — | Not assessed |
split: fullval · dataset: PACIFIC | 316 examples | — | Not assessed |
split: fullval · dataset: PACIFIC | 1,636 examples | — | Not assessed |
split: train · dataset: AbgCoQA | 7,269 examples | — | Not assessed |
split: fullval · dataset: AbgCoQA | 1,184 examples | — | Not assessed |
dataset: AbgCoQA | 22% | — | Not assessed |
split: train · dataset: AmbigNQ | 19,244 examples | — | Not assessed |
split: test · dataset: AmbigNQ | 4,377 examples | — | Not assessed |
dataset: AmbigNQ | 79% | — | Not assessed |
Experiment plan ready; no runs yet.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
simulator: Qwen2.5-7B-Instruct · auditedCases: 8 | 8 cases | — | Not assessed |
simulator: Qwen2.5-3B-Instruct · auditedCases: 8 | 1 cases | — | Not assessed |
Experiment plan ready; no runs yet.
0/4 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
series: IG on clarify turns · trainingStep: 5 | 0.17 CIG signal (nats) | — | Not assessed |
series: IG on clarify turns · trainingStep: 250 | 0.47 CIG signal (nats) | — | Not assessed |
series: IG batch-weighted · trainingStep: 5 | 0.02 CIG signal (nats) | — | Not assessed |
series: IG batch-weighted · trainingStep: 250 | 0.23 CIG signal (nats) | — | Not assessed |
Experiment plan ready; no runs yet.
0/10 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: SGP · status: externally reported · backbone: Gemma-2-9B | 0.726 score (0-1) | — | Not assessed |
method: SGP · status: externally reported · backbone: Gemma-2-9B | 0.429 proportion | — | Not assessed |
method: SGP · status: externally reported · backbone: Gemma-2-9B | 0.206 proportion | — | Not assessed |
method: SGP-Oracle · status: externally reported · backbone: Gemma-2-9B · usesGoldAmbiguityLabels: true | 0.787 score (0-1) | — | Not assessed |
method: SGP-Oracle · status: externally reported · backbone: Gemma-2-9B · usesGoldAmbiguityLabels: true | 0.435 proportion | — | Not assessed |
method: SGP-Oracle · status: externally reported · backbone: Gemma-2-9B · usesGoldAmbiguityLabels: true | 0.156 proportion | — | Not assessed |
method: CIGAsk · backbone: Gemma-2-9B | 0.774 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Gemma-2-9B | 0.749 score (0-1) | — | Not assessed |
method: CIGAsk · backbone: Gemma-2-9B | 0.813 proportion | — | Not assessed |
method: CIGAsk · backbone: Gemma-2-9B | 0.031 proportion | — | Not assessed |
Experiment plan ready; no runs yet.
0/9 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: Direct · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.058 score (0-1) | — | Not assessed |
method: FATA · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.076 score (0-1) | — | Not assessed |
method: ReAct · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.076 score (0-1) | — | Not assessed |
method: SFT · dataset: AmbigNQ · backbone: Qwen2.5-3B | 0.056 score (0-1) | — | Not assessed |
method: SFT · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.109 score (0-1) | — | Not assessed |
method: SGP · status: externally reported anchor · dataset: AmbigQA · backbone: Gemma-2-9B · protocol: 50/50 balanced dev | 0.306 score (0-1) | — | Not assessed |
method: SGP-Oracle · status: externally reported anchor · dataset: AmbigQA · backbone: Gemma-2-9B · protocol: 50/50 balanced dev · usesGoldAmbiguityLabels: true | 0.359 score (0-1) | — | Not assessed |
method: CIGAsk · dataset: AmbigNQ · backbone: Qwen2.5-3B | 0.271 score (0-1) | — | Not assessed |
method: CIGAsk · dataset: AmbigNQ · backbone: Qwen2.5-7B | 0.511 score (0-1) | — | Not assessed |
The printed transcript is a qualitative case, not an automatically reproducible aggregate. Regenerating it requires the unreleased trained state and an exact stochastic simulator snapshot; the schema has no source-defined rubric for deciding transcript equivalence, so the actual targeted-versus-generic comparison is retained without fabricating a number.
The paper omits the late-checkpoint set, point estimates, bootstrap resampling procedure, and interval endpoints. Those missing comparison definitions prevent a concrete automatic package without fabricating a decision rule.
The appendix prints three sensible examples but supplies no sample-wide coding rubric or human-evaluation decision rule. They are carried into the OOD objective as qualitative context, not converted to a scalar target.
The paper gives a qualitative matched-seed trajectory claim but no exact values, checkpoint range, curve, or registered numeric target. The comparison is retained as source context for the PACIFIC 3B workflow; a deterministic decision rule must be established from additional source evidence before automatic verification.
The qualitative “roughly doubles” finding has no counts, denominator, coding rubric, or exact rate. Self-completion and declarative reformulation remain source context for the component ablation, but no scalar target is invented.
This is method context rather than a separately reported empirical result. Its equations and controls are carried into every training objective; correctness is assessed through the bound empirical packages rather than by inventing a method-only scalar.
The English-only scope is descriptive of the evaluated datasets; multilingual behavior is unreported and is not added as a new paper-bound experiment.
This is an ethical/generalization limitation about annotator-dependent ambiguity labels. The rho sensitivity package preserves the source labels but cannot establish behavior for unobserved user populations.
The stated lack of robustness to noisy, adversarial, diverse, or human users is an open scope limitation. The simulator packages test only the source conditions and cannot establish the untested populations.
This is a source limitation, not an independent empirical result. Every training package preserves seed 42 and does not imply variance estimation.
The paper explicitly marks ACT/SGP values as original-protocol context anchors. The catalog binds reconstructed CIGAsk and in-protocol prompting controls only and does not treat backbone matching as protocol matching.
This is a direct method/input requirement from the equations: every planned reconstruction requires the benchmark ambiguity label during training; no separate result is claimed.
The absence of experiments above 9B is an explicit scope limitation, not a positive empirical result to test by substituting an unreported larger model.