The paper’s claims are available in Research claims.
0 / 4 claims verified
the rest still being verified
No verified repository, trained checkpoint, complete curated training data, retrieval implementation, or strict evaluator is available in the fixed inputs.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
datasets: ["2WikiMultiHopQA","HotpotQA","Bamboogle","FRAMES","MusiQue","NQ","TriviaQA"] · aggregation: arithmetic average · model_variant: Qwen2.5-3B | 33.55 percentage_points | — | Not assessed |
split: test · dataset: HotpotQA · aggregation: single reported evaluation · model_variant: Qwen2.5-3B | 44 percentage_points | — | Not assessed |
split: test · dataset: Bamboogle · aggregation: single reported evaluation · model_variant: Qwen2.5-7B | 49.21 percentage_points | — | Not assessed |
The missing checkpoint, retrieval service, and complete evaluation protocol prevent strict comparable evidence; execution resources are not the blocker.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
dataset: 2WikiMultiHopQA · max_steps: 10 · aggregation: single reported evaluation · model_variant: Qwen2.5-3B | 23.91 percentage_points | — | Not assessed |
dataset: 2WikiMultiHopQA · max_steps: 30 · aggregation: single reported evaluation · model_variant: Qwen2.5-3B | 47.45 percentage_points | — | Not assessed |
The original hardware-sensitive benchmark conditions are underspecified by the fixed paper and repository, so a strict comparable verdict is not available.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
hardware: NVIDIA A800 server · aggregation: average · model_variant: Qwen2.5-3B | 10.15 seconds | — | Not assessed |
hardware: NVIDIA A800 server · aggregation: average · model_variant: Qwen2.5-3B | 11.86 seconds | — | Not assessed |
judge: Qwen2.5-72B · dataset: MedQA test · model_variant: Qwen2.5-7B-Instruct | 87 percentage_points | — | Not assessed |
The trained variants, reward model prompts, curated training set, optimizer implementation, and exact seeds are unavailable.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
condition: full model · aggregation: average across seven benchmarks · model_variant: Qwen2.5-3B | 33.55 percentage_points | — | Not assessed |
condition: without retrieval and memory rewards · aggregation: average across seven benchmarks · model_variant: Qwen2.5-3B | 19.94 percentage_points | — | Not assessed |