论文中的结论可在「研究结论」中查看。
0 / 4 条结论已通过验证
其余仍在验证中
No verified repository, trained checkpoint, complete curated training data, retrieval implementation, or strict evaluator is available in the fixed inputs.
已评估 0/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
datasets: ["2WikiMultiHopQA","HotpotQA","Bamboogle","FRAMES","MusiQue","NQ","TriviaQA"] · aggregation: arithmetic average · model_variant: Qwen2.5-3B | 33.55 percentage_points | — | 尚未评估 |
split: test · dataset: HotpotQA · aggregation: single reported evaluation · model_variant: Qwen2.5-3B | 44 percentage_points | — | 尚未评估 |
split: test · dataset: Bamboogle · aggregation: single reported evaluation · model_variant: Qwen2.5-7B | 49.21 percentage_points | — | 尚未评估 |
The missing checkpoint, retrieval service, and complete evaluation protocol prevent strict comparable evidence; execution resources are not the blocker.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
dataset: 2WikiMultiHopQA · max_steps: 10 · aggregation: single reported evaluation · model_variant: Qwen2.5-3B | 23.91 percentage_points | — | 尚未评估 |
dataset: 2WikiMultiHopQA · max_steps: 30 · aggregation: single reported evaluation · model_variant: Qwen2.5-3B | 47.45 percentage_points | — | 尚未评估 |
The original hardware-sensitive benchmark conditions are underspecified by the fixed paper and repository, so a strict comparable verdict is not available.
已评估 0/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
hardware: NVIDIA A800 server · aggregation: average · model_variant: Qwen2.5-3B | 10.15 seconds | — | 尚未评估 |
hardware: NVIDIA A800 server · aggregation: average · model_variant: Qwen2.5-3B | 11.86 seconds | — | 尚未评估 |
judge: Qwen2.5-72B · dataset: MedQA test · model_variant: Qwen2.5-7B-Instruct | 87 percentage_points | — | 尚未评估 |
The trained variants, reward model prompts, curated training set, optimizer implementation, and exact seeds are unavailable.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
condition: full model · aggregation: average across seven benchmarks · model_variant: Qwen2.5-3B | 33.55 percentage_points | — | 尚未评估 |
condition: without retrieval and memory rewards · aggregation: average across seven benchmarks · model_variant: Qwen2.5-3B | 19.94 percentage_points | — | 尚未评估 |