Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/4
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
4
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
AGENTIC R AG-R1 reports higher or near-highest F1 than listed baselines across seven multi-hop, open-domain, and agentic QA benchmarks.
Information insufficientclaim-main-benchmark1 plan0 runs
Scientific conclusionNot assessed
Optional independent reconstruction of the main QA benchmark
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The paper reports increasing AGENTIC R AG-R1 F1 as the maximum reasoning-step budget grows on selected 2Wiki, Bamboogle, and TriviaQA evaluations.
Information insufficientclaim-long-horizon1 plan0 runs
Scientific conclusionNot assessed
Optional independent reconstruction of step-budget scaling
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The paper reports lower average inference time than TC-RAG and downstream results on MedQA, DeepResearch Bench, ALFWorld, and WebShop.
Information insufficientclaim-efficiency-generalization0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The paper reports that removing memory actions, rollout rejection, or both lowers average F1, and that removing both retrieval and memory rewards gives the lowest ablation average.
Information insufficientclaim-component-ablation1 plan0 runs
Scientific conclusionNot assessed
Optional independent reconstruction of component ablations
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Community reproductions
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.