Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/18
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
18
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 8 claims with no independent reproduction scheduled in this plan.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
Sensitivity to the ambiguity-bonus weight is non-monotonic on PACIFIC: among rho values 0.1, 0.3, and 0.5, rho=0.3 has the highest overall F1, ambiguous-query clarification rate, and reported selectivity. Raising rho to 0.5 reduces F1 and recall and does not improve selectivity over the main setting.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On unseen closed-book TriviaQA and NQ Open subsets, forced-direct CIGAsk 7B slightly exceeds the Qwen2.5-7B base F1, while under a multi-action prompt it clarifies only 5.6% and 9.4% of queries, respectively.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
CIGAsk OOD retention and clarification selectivity
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Replacing GPT-4o with Qwen2.5-7B-Instruct as the fixed latent-intent-aware simulator leaves PACIFIC performance close to the main setting and slightly raises selectivity, whereas Qwen2.5-3B-Instruct substantially lowers post-clarification F1 and ambiguous-query recall.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In the PACIFIC deferred-tax-asset case, a generic prompted clarification elicits only a restatement and produces the incorrect aggregate 131,028. CIGAsk asks which period is intended, receives 'as of December 31, 2019,' and returns the gold answer 73,260 in one clarification turn.
Information insufficientqualitative-case-0010 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On PACIFIC fullval, CIGAsk 7B reports the best overall F1, post-clarification F1, and ambiguous-query clarification rate in Table 2; CIGAsk 3B also exceeds all prompting and SFT baselines. Scaling CIGAsk from 3B to 7B raises F1, post-clarification F1, and ambiguous-query recall while also raising the clear-query false-positive rate from 0.179 to 0.260. Prompting and SFT expose negative behaviors: Direct and SFT never clarify, FATA clarifies every input and has low F1, and ReAct clarifies more often on clear than ambiguous inputs.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
PACIFIC direct-answer baseline
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
PACIFIC prompted clarification controls
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
PACIFIC CIGAsk and SFT 3B reconstruction
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The PACIFIC 7B component ablation separates the reward roles: retaining CIG but removing the asymmetric bonus yields useful post-clarification answers but low clarification recall; retaining the asymmetric bonus but removing CIG largely preserves recall while reducing post-clarification F1. Removing both performs worst overall among the four variants.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Replacing the frozen CIG reference with the current actor while holding other settings fixed substantially reduces overall F1, post-clarification F1, ambiguous-query clarification recall, and selectivity on PACIFIC. The current-actor variant's mean training reward changes only modestly from early to late training.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
When the same recipe is retrained separately on passage-grounded AbgCoQA without per-dataset hyperparameter search, CIGAsk 7B reaches F1 0.724, above the 3B variant and all reported prompting and SFT baselines in Table 2.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
AbgCoQA CIGAsk and SFT 3B reconstruction
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
AbgCoQA CIGAsk and SFT 7B reconstruction
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
AbgCoQA prompted clarification controls
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On the Zephyr-7B backbone, CIGAsk reports higher overall and post-clarification F1 than ACT and additionally reports a 0.712 ambiguous-query clarification rate with a 0.044 clear-query clarification rate.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Across late-training checkpoints, PACIFIC items that the policy chose to clarify had higher F1 than items answered directly, with bootstrap 95% confidence intervals for the difference above zero throughout.
Information insufficientlate-checkpoint-productivity-0010 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On the small OOD subsets where CIGAsk chooses to clarify, the appendix describes targeted questions about lexical ambiguity, year, temporal scope, or episode identity rather than generic requests for context, indicating qualitative transfer from table-grounded to text-grounded QA without retraining.
Information insufficientood-clarification-quality-0011 plan0 runs
Scientific conclusionNot assessed
CIGAsk OOD retention and clarification selectivity
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In a matched-seed 3B comparison, CIGAsk and outcome-only GRPO learned to clarify at similar rates, but only the CIG-equipped model consistently converted those turns into improved post-clarification F1; the outcome-only variant fell to near zero.
Information insufficientcig-effect-3b-0010 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The three clarification benchmarks differ substantially in size and ambiguity prevalence: PACIFIC is table-grounded financial QA, AbgCoQA is passage-grounded conversational QA, and AmbigNQ is ambiguity-heavy open-domain QA.
Benchmark identity, split size, and ambiguity profile
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In a manual audit of eight ambiguous PACIFIC cases, Qwen2.5-7B-Instruct returned an intent-consistent response in every case, whereas Qwen2.5-3B-Instruct did so in only one case; all three examples printed in Table 11 show the 7B reply matching the latent intent and the 3B reply contradicting it.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For CIGAsk 7B, both the mean CIG on clarification turns and the batch-weighted CIG rise over training; the paper reports that this rise coincides with the emergence of the ambiguous-recall versus clear-query false-positive-rate gap.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On the Gemma-2-9B backbone, CIGAsk reports higher F1 than deployable SGP and much higher ambiguous-query clarification with lower clear-query over-clarification, while its F1 remains below the gold-label SGP-Oracle upper bound.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On the ambiguity-heavy AmbigNQ subset, CIGAsk 7B reports EM 0.511, substantially above the prompting and SFT baselines and the CIGAsk 3B variant; it also exceeds the contextual SGP and SGP-Oracle AmbigQA values placed in this column by the paper.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
AmbigNQ direct-answer baseline
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
AmbigNQ prompted clarification controls
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
AmbigNQ CIGAsk and SFT 3B reconstruction
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Removing the asymmetric bonus roughly doubles the failure rate among generated clarification turns. The dominant failures are self-completion, where the policy supplies a question and plausible answer without yielding to the simulator, and declarative reformulation, which does not ask an actionable question.
Information insufficientclarification-failures-0010 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Community reproductions
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.