复现资料准备
已准备 0/1 项资料
科研资料尚未闭合:configuration。这些问题只影响预下载缓存,不会阻止最终复现 Agent。
准备方式将在链接检查和资料清单确认后确定。
本次复现积分预估
预计消耗
127
最大预留
508
预计算力时长
约 5 分钟
实际消耗 475 积分,扣除 0 积分;算力运行约 22 分钟,模型 Token 18963722。
论文中的结论可在「研究结论」中查看。
0 / 61 条结论已通过验证
其余仍在验证中
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/27 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 07 · split: Train | 65% | — | 尚未评估 |
round: Round 07 · split: Train | 99.7 months | — | 尚未评估 |
round: Round 07 · split: Train | 40.722 million_usd | — | 尚未评估 |
round: Round 07 · split: Train | 100% | — | 尚未评估 |
round: Round 07 · split: Train | 75% | — | 尚未评估 |
round: Round 07 · split: Train | 65% | — | 尚未评估 |
round: Round 07 · split: Train | 3.19 calls_per_month | — | 尚未评估 |
round: Round 07 · split: Train | 100.7 actions | — | 尚未评估 |
round: Round 07 · split: Train | 40.46 million_usd | — | 尚未评估 |
round: Round 07 · split: Validate | 80% | — | 尚未评估 |
round: Round 07 · split: Validate | 114.15 months | — | 尚未评估 |
round: Round 07 · split: Validate | 47.607 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 31.3% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 34.78% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 37.39% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 32.17% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 26.09% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 33.91% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 38.26% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: τ-bench | 44.35% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 83.1% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 82.5% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 82.3% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 83.4% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 83.1% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 82.5% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 82.7% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: HotpotQA | 84.5% | — | 尚未评估 |
报告
81.8%
观测
—
Percentage reduction in validation tool invocations per month ((17.23 - 3.13) / 17.23 = 81.8%) derived directly from the EnterpriseArena self-evolution metrics in Table 11.
已评估 0/1 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 81.8% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/32 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline | 44% | — | 尚未评估 |
configuration: Baseline | 89.8 months | — | 尚未评估 |
configuration: Baseline | 78.86 million_usd | — | 尚未评估 |
configuration: Baseline | 100% | — | 尚未评估 |
configuration: Baseline | 78% | — | 尚未评估 |
configuration: Baseline | 52% | — | 尚未评估 |
configuration: Baseline | 0.13 calls_per_month | — | 尚未评估 |
configuration: Baseline | 152.12 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 50% | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 93.24 months | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 76.76 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 100% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 10 · split: Train | 90% | — | 尚未评估 |
round: Round 10 · split: Train | 121.15 months | — | 尚未评估 |
round: Round 10 · split: Train | 58.418 million_usd | — | 尚未评估 |
round: Round 10 · split: Train | 100% | — | 尚未评估 |
round: Round 10 · split: Train | 90% | — | 尚未评估 |
round: Round 10 · split: Train | 90% | — | 尚未评估 |
round: Round 10 · split: Train | 3.07 calls_per_month | — | 尚未评估 |
round: Round 10 · split: Train | 122.2 actions | — | 尚未评估 |
round: Round 10 · split: Train | 88.32 million_usd | — | 尚未评估 |
round: Round 10 · split: Validate | 85% | — | 尚未评估 |
round: Round 10 · split: Validate | 121.2 months | — | 尚未评估 |
round: Round 10 · split: Validate | 56.386 million_usd | — | 尚未评估 |
Comparative statistical test (Fisher exact test p-value and peak round comparison) derived from the EnterpriseArena self-evolution simulation test results in Table 11.
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 85% | — | 尚未评估 | |
| 0% | — | 尚未评估 | |
| 0 probability | — | 尚未评估 | |
| 95% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 33.4% | — | 尚未评估 | |
| 55.4% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 72.17% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 67.83% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 73.04% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 64.35% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 65.22% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 67.83% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 64.35% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: τ-bench | 80% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 52% | — | 尚未评估 |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 56% | — | 尚未评估 |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 55% | — | 尚未评估 |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 52% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 56% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 52% | — | 尚未评估 |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 54% | — | 尚未评估 |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 60% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Baseline · split: Validate | 0% | — | 尚未评估 |
round: Baseline · split: Validate | 34.8 months | — | 尚未评估 |
round: Baseline · split: Validate | 28.377 million_usd | — | 尚未评估 |
round: Baseline · split: Validate | 100% | — | 尚未评估 |
round: Baseline · split: Validate | 0% | — | 尚未评估 |
round: Baseline · split: Validate | 0% | — | 尚未评估 |
round: Baseline · split: Validate | 17.23 calls_per_month | — | 尚未评估 |
round: Baseline · split: Validate | 35.7 actions | — | 尚未评估 |
round: Baseline · split: Validate | 0.47 million_usd | — | 尚未评估 |
round: Baseline · split: Test | 0% | — | 尚未评估 |
round: Baseline · split: Test | 33.3 months | — | 尚未评估 |
round: Baseline · split: Test | 28.754 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 57.23 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 55.19 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 66.79 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 65.66 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 53.58 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 61.55 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 68.15 points | — | 尚未评估 |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: GDPval | 71.19 points | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 26.12% | — | 尚未评估 |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 23.88% | — | 尚未评估 |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 14.93% | — | 尚未评估 |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 42.54% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 30.6% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 17.16% | — | 尚未评估 |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 18.66% | — | 尚未评估 |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: ALFWorld | 39.55% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline (no graph) | 54.8 points | — | 尚未评估 |
configuration: Baseline (no graph) | 275,638 tokens | — | 尚未评估 |
configuration: Baseline (no graph) | 28.2 steps | — | 尚未评估 |
configuration: Full graph, raw injection | 57.17 points | — | 尚未评估 |
configuration: Full graph, raw injection | 264,680 tokens | — | 尚未评估 |
configuration: Full graph, raw injection | 33.55 steps | — | 尚未评估 |
configuration: Full graph, generative | 56.75 points | — | 尚未评估 |
configuration: Full graph, generative | 448,972 tokens | — | 尚未评估 |
configuration: Full graph, generative | 22.07 steps | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 63.99 points | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 367,738 tokens | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 18.57 steps | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/14 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
benchmark: HotpotQA | 9 nodes | — | 尚未评估 |
benchmark: HotpotQA | 9 triplets | — | 尚未评估 |
benchmark: MultiChallenge | 7 nodes | — | 尚未评估 |
benchmark: MultiChallenge | 7 triplets | — | 尚未评估 |
benchmark: GDPval | 15 nodes | — | 尚未评估 |
benchmark: GDPval | 22 triplets | — | 尚未评估 |
benchmark: ALFWorld | 11 nodes | — | 尚未评估 |
benchmark: ALFWorld | 27 triplets | — | 尚未评估 |
benchmark: τ-bench | 17 nodes | — | 尚未评估 |
benchmark: τ-bench | 18 triplets | — | 尚未评估 |
benchmark: BFCL v3 | 131 nodes | — | 尚未评估 |
benchmark: BFCL v3 | 265 triplets | — | 尚未评估 |
实验执行出错 · 不会自动重试,需要人工决定后续处理。
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 94.78% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 85.07% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 80.6% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 97.76% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 91.04% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 99.25% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 95.52% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: ALFWorld | 100% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 06 · split: Train | 80% | — | 尚未评估 |
round: Round 06 · split: Train | 114.9 months | — | 尚未评估 |
round: Round 06 · split: Train | 46.345 million_usd | — | 尚未评估 |
round: Round 06 · split: Train | 100% | — | 尚未评估 |
round: Round 06 · split: Train | 90% | — | 尚未评估 |
round: Round 06 · split: Train | 80% | — | 尚未评估 |
round: Round 06 · split: Train | 3.09 calls_per_month | — | 尚未评估 |
round: Round 06 · split: Train | 115.9 actions | — | 尚未评估 |
round: Round 06 · split: Train | 56.52 million_usd | — | 尚未评估 |
round: Round 06 · split: Validate | 65% | — | 尚未评估 |
round: Round 06 · split: Validate | 108.75 months | — | 尚未评估 |
round: Round 06 · split: Validate | 43.831 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 56.39 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 61.97 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 71.37 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 64.69 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 56.18 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 63.85 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 69.1 points | — | 尚未评估 |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: GDPval | 78.78 points | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 86.57% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 86.57% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 79.1% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 91.79% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 84.33% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 67.16% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 87.31% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: ALFWorld | 93.28% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 56% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 54% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 54% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 58% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 56% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 52% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 58% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 67% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 59% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 63% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 57% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 63% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 63% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 64% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 61% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 66% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 72.7% | — | 尚未评估 |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 70.7% | — | 尚未评估 |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 70.2% | — | 尚未评估 |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 73.5% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 70% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 74.3% | — | 尚未评估 |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 73.1% | — | 尚未评估 |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: HotpotQA | 74.4% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/32 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline | 0% | — | 尚未评估 |
configuration: Baseline | 33.58 months | — | 尚未评估 |
configuration: Baseline | 28.59 million_usd | — | 尚未评估 |
configuration: Baseline | 100% | — | 尚未评估 |
configuration: Baseline | 0% | — | 尚未评估 |
configuration: Baseline | 0% | — | 尚未评估 |
configuration: Baseline | 18.94 calls_per_month | — | 尚未评估 |
configuration: Baseline | 0 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 0% | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 34.04 months | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 28.75 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 100% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/30 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 86.96% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 86.67% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 71.43% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 100% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 87.5% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 52.17% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 66.67% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 42.86% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 72.73% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 58.93% | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 47.83% | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 53.33% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 81.33% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 89.16% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 89.16% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 89.16% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 89.16% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 91.57% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 91.57% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 91.57% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| -15% | — | 尚未评估 | |
| -10.45 months | — | 尚未评估 | |
| -5% | — | 尚未评估 | |
| -0.184 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/27 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 09 · split: Train | 95% | — | 尚未评估 |
round: Round 09 · split: Train | 126.05 months | — | 尚未评估 |
round: Round 09 · split: Train | 58.188 million_usd | — | 尚未评估 |
round: Round 09 · split: Train | 100% | — | 尚未评估 |
round: Round 09 · split: Train | 95% | — | 尚未评估 |
round: Round 09 · split: Train | 95% | — | 尚未评估 |
round: Round 09 · split: Train | 3.08 calls_per_month | — | 尚未评估 |
round: Round 09 · split: Train | 127 actions | — | 尚未评估 |
round: Round 09 · split: Train | 97.84 million_usd | — | 尚未评估 |
round: Round 09 · split: Validate | 90% | — | 尚未评估 |
round: Round 09 · split: Validate | 121.2 months | — | 尚未评估 |
round: Round 09 · split: Validate | 56.57 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/32 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline | 6% | — | 尚未评估 |
configuration: Baseline | 50.28 months | — | 尚未评估 |
configuration: Baseline | 31.85 million_usd | — | 尚未评估 |
configuration: Baseline | 100% | — | 尚未评估 |
configuration: Baseline | 38% | — | 尚未评估 |
configuration: Baseline | 8% | — | 尚未评估 |
configuration: Baseline | 0.89 calls_per_month | — | 尚未评估 |
configuration: Baseline | 21.82 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 6% | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 51.34 months | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 32.27 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 100% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 04 · split: Train | 80% | — | 尚未评估 |
round: Round 04 · split: Train | 115.05 months | — | 尚未评估 |
round: Round 04 · split: Train | 46.631 million_usd | — | 尚未评估 |
round: Round 04 · split: Train | 100% | — | 尚未评估 |
round: Round 04 · split: Train | 90% | — | 尚未评估 |
round: Round 04 · split: Train | 80% | — | 尚未评估 |
round: Round 04 · split: Train | 3.06 calls_per_month | — | 尚未评估 |
round: Round 04 · split: Train | 116 actions | — | 尚未评估 |
round: Round 04 · split: Train | 53.04 million_usd | — | 尚未评估 |
round: Round 04 · split: Validate | 75% | — | 尚未评估 |
round: Round 04 · split: Validate | 109.2 months | — | 尚未评估 |
round: Round 04 · split: Validate | 49.831 million_usd | — | 尚未评估 |
Meta-statistical summary aggregation (win/tie/loss counts and binomial sign test) derived deterministically across the 24 model-benchmark evaluations in Table 1.
已评估 0/5 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
totalSettings: 24 | 21 settings | — | 尚未评估 |
| 19 settings | — | 尚未评估 | |
| 2 settings | — | 尚未评估 | |
| 3 settings | — | 尚未评估 | |
| 0.0004 probability | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 64.35% | — | 尚未评估 |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 58.26% | — | 尚未评估 |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 64.35% | — | 尚未评估 |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 63.48% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 61.74% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 63.48% | — | 尚未评估 |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 68.7% | — | 尚未评估 |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: τ-bench | 67.83% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/27 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 02 · split: Train | 85% | — | 尚未评估 |
round: Round 02 · split: Train | 117.75 months | — | 尚未评估 |
round: Round 02 · split: Train | 46.524 million_usd | — | 尚未评估 |
round: Round 02 · split: Train | 100% | — | 尚未评估 |
round: Round 02 · split: Train | 90% | — | 尚未评估 |
round: Round 02 · split: Train | 85% | — | 尚未评估 |
round: Round 02 · split: Train | 3.39 calls_per_month | — | 尚未评估 |
round: Round 02 · split: Train | 118.8 actions | — | 尚未评估 |
round: Round 02 · split: Train | 57.27 million_usd | — | 尚未评估 |
round: Round 02 · split: Validate | 80% | — | 尚未评估 |
round: Round 02 · split: Validate | 112.6 months | — | 尚未评估 |
round: Round 02 · split: Validate | 46.038 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/32 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline | 26% | — | 尚未评估 |
configuration: Baseline | 63.76 months | — | 尚未评估 |
configuration: Baseline | 39.42 million_usd | — | 尚未评估 |
configuration: Baseline | 100% | — | 尚未评估 |
configuration: Baseline | 42% | — | 尚未评估 |
configuration: Baseline | 28% | — | 尚未评估 |
configuration: Baseline | 0.47 calls_per_month | — | 尚未评估 |
configuration: Baseline | 27.24 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 28% | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 64.08 months | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 35.47 million_usd | — | 尚未评估 |
configuration: RAP (Kagaya et al., 2024) | 100% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 68.07% | — | 尚未评估 |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 84.94% | — | 尚未评估 |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 80.12% | — | 尚未评估 |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 84.94% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 80.72% | — | 尚未评估 |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 83.73% | — | 尚未评估 |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 83.13% | — | 尚未评估 |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 86.75% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 82.84% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 76.87% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 76.12% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 90.3% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 80.6% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 82.09% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 81.34% | — | 尚未评估 |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: ALFWorld | 94.03% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 87.95% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 95.18% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 94.58% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 92.77% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 93.37% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 93.98% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 95.18% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 95.78% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
mode: Mode 5 · generation: 1 | 77.59% | — | 尚未评估 |
mode: Mode 5 · generation: 10 | 83.31% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline (no graph) | 72.58% | — | 尚未评估 |
configuration: Baseline (no graph) | 18,055 tokens | — | 尚未评估 |
configuration: Baseline (no graph) | 21.84 steps | — | 尚未评估 |
configuration: Full graph, raw injection | 70.34% | — | 尚未评估 |
configuration: Full graph, raw injection | 21,062 tokens | — | 尚未评估 |
configuration: Full graph, raw injection | 25 steps | — | 尚未评估 |
configuration: Full graph, generative | 54.48% | — | 尚未评估 |
configuration: Full graph, generative | 96,360 tokens | — | 尚未评估 |
configuration: Full graph, generative | 30.05 steps | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 81.53% | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 28,064 tokens | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 18.8 steps | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 59.33 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 59.27 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 60.23 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 62.45 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 50.29 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 50.25 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 54 points | — | 尚未评估 |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: GDPval | 64.42 points | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
mode: Mode 3 · phase: initial | 54% | — | 尚未评估 |
mode: Mode 3 · generation: 5 · initialEdgeCount: 13 | 10 edges | — | 尚未评估 |
mode: Mode 3 · phase: restructured | 93.9% | — | 尚未评估 |
mode: Mode 5 (scratch) | 94.9% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 61% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 62% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 65% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 60% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 59% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 62% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 60% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 67% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 74.6% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 74.2% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 71.8% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 75.2% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 75.4% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 73.9% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 74.3% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: HotpotQA | 74.5% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 66.09% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 63.48% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 60.87% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 71.3% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 65.22% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 66.96% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 61.74% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: τ-bench | 73.91% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 58.8% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 71.21% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 62.8% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 76.61% | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 63.8% | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 77.16% | — | 尚未评估 |
constructionMode: Mode 3: Expert + Online Evolution | 63.1% | — | 尚未评估 |
constructionMode: Mode 3: Expert + Online Evolution | 76.34% | — | 尚未评估 |
constructionMode: Mode 4: Scratch + Static Build | 55.4% | — | 尚未评估 |
constructionMode: Mode 4: Scratch + Static Build | 69.49% | — | 尚未评估 |
constructionMode: Mode 5: Scratch + Online Evolution | 66.3% | — | 尚未评估 |
constructionMode: Mode 5: Scratch + Online Evolution | 78.79% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/30 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 87.5% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 7,403.98 tokens | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 7.05 steps | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 1.05 failures_per_sample | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 127.59 seconds | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 58.93% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 11,039.82 tokens | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 6.6 steps | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 0.88 failures_per_sample | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 117.7 seconds | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 53.57% | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 5,990.79 tokens | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/27 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 01 · split: Train | 0% | — | 尚未评估 |
round: Round 01 · split: Train | 33.4 months | — | 尚未评估 |
round: Round 01 · split: Train | 28.777 million_usd | — | 尚未评估 |
round: Round 01 · split: Train | 100% | — | 尚未评估 |
round: Round 01 · split: Train | 0% | — | 尚未评估 |
round: Round 01 · split: Train | 0% | — | 尚未评估 |
round: Round 01 · split: Train | 16.25 calls_per_month | — | 尚未评估 |
round: Round 01 · split: Train | 34.4 actions | — | 尚未评估 |
round: Round 01 · split: Train | 0 million_usd | — | 尚未评估 |
round: Round 01 · split: Validate | 45% | — | 尚未评估 |
round: Round 01 · split: Validate | 88.9 months | — | 尚未评估 |
round: Round 01 · split: Validate | 38.594 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/27 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 08 · split: Train | 85% | — | 尚未评估 |
round: Round 08 · split: Train | 120.35 months | — | 尚未评估 |
round: Round 08 · split: Train | 49.781 million_usd | — | 尚未评估 |
round: Round 08 · split: Train | 100% | — | 尚未评估 |
round: Round 08 · split: Train | 90% | — | 尚未评估 |
round: Round 08 · split: Train | 90% | — | 尚未评估 |
round: Round 08 · split: Train | 3.12 calls_per_month | — | 尚未评估 |
round: Round 08 · split: Train | 121.3 actions | — | 尚未评估 |
round: Round 08 · split: Train | 77.33 million_usd | — | 尚未评估 |
round: Round 08 · split: Validate | 90% | — | 尚未评估 |
round: Round 08 · split: Validate | 121.2 months | — | 尚未评估 |
round: Round 08 · split: Validate | 57.371 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 83.73% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 89.16% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 86.75% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 89.16% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 88.55% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 89.76% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 87.35% | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 89.76% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 05 · split: Train | 80% | — | 尚未评估 |
round: Round 05 · split: Train | 114.05 months | — | 尚未评估 |
round: Round 05 · split: Train | 44.565 million_usd | — | 尚未评估 |
round: Round 05 · split: Train | 100% | — | 尚未评估 |
round: Round 05 · split: Train | 90% | — | 尚未评估 |
round: Round 05 · split: Train | 80% | — | 尚未评估 |
round: Round 05 · split: Train | 3.12 calls_per_month | — | 尚未评估 |
round: Round 05 · split: Train | 115 actions | — | 尚未评估 |
round: Round 05 · split: Train | 52.85 million_usd | — | 尚未评估 |
round: Round 05 · split: Validate (carried forward R04) | 75% | — | 尚未评估 |
round: Round 05 · split: Validate (carried forward R04) | 109.2 months | — | 尚未评估 |
round: Round 05 · split: Validate (carried forward R04) | 49.831 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 46.23 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 50.13 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 40.97 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 48.22 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 41.39 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 40.61 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 42.14 points | — | 尚未评估 |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: GDPval | 51.49 points | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: Baseline (no graph) | 80.27% | — | 尚未评估 |
configuration: Baseline (no graph) | 6,629 tokens | — | 尚未评估 |
configuration: Baseline (no graph) | 3.87 steps | — | 尚未评估 |
configuration: Full graph, raw injection | 86.6% | — | 尚未评估 |
configuration: Full graph, raw injection | 10,164 tokens | — | 尚未评估 |
configuration: Full graph, raw injection | 4.54 steps | — | 尚未评估 |
configuration: Full graph, generative | 87.35% | — | 尚未评估 |
configuration: Full graph, generative | 14,434 tokens | — | 尚未评估 |
configuration: Full graph, generative | 3.08 steps | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 89.31% | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 12,295 tokens | — | 尚未评估 |
configuration: Subgraph, generative (Ours) | 4.22 steps | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 85.9% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 85.1% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 83.6% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 86% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 85.8% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 85.1% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 85% | — | 尚未评估 |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: HotpotQA | 87.3% | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
round: Round 03 · split: Train | 80% | — | 尚未评估 |
round: Round 03 · split: Train | 113.35 months | — | 尚未评估 |
round: Round 03 · split: Train | 46.857 million_usd | — | 尚未评估 |
round: Round 03 · split: Train | 100% | — | 尚未评估 |
round: Round 03 · split: Train | 85% | — | 尚未评估 |
round: Round 03 · split: Train | 80% | — | 尚未评估 |
round: Round 03 · split: Train | 3.08 calls_per_month | — | 尚未评估 |
round: Round 03 · split: Train | 114.3 actions | — | 尚未评估 |
round: Round 03 · split: Train | 51.88 million_usd | — | 尚未评估 |
round: Round 03 · split: Validate | 65% | — | 尚未评估 |
round: Round 03 · split: Validate | 102.15 months | — | 尚未评估 |
round: Round 03 · split: Validate | 43.965 million_usd | — | 尚未评估 |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
已评估 0/30 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 71.21% | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 4,003.24 tokens | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 4.88 steps | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 0.016 failures_per_sample | — | 尚未评估 |
constructionMode: Unguided Baseline (w/o PG) | 18.06 seconds | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 76.61% | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 9,045.69 tokens | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 4.12 steps | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 0.007 failures_per_sample | — | 尚未评估 |
constructionMode: Mode 1: Hand-crafted Expert | 39.34 seconds | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 77.16% | — | 尚未评估 |
constructionMode: Mode 2: Expert + Static Update | 8,942.79 tokens | — | 尚未评估 |
Qualitative behavioral analysis and case study from execution traces illustrating agent decision dynamics rather than an independently evaluated scalar metric target.
Qualitative behavioral analysis and case study from execution traces illustrating agent decision dynamics rather than an independently evaluated scalar metric target.
Qualitative behavioral analysis and case study from execution traces illustrating agent decision dynamics rather than an independently evaluated scalar metric target.
Algorithmic specification and architectural definition provided as methodological framework rather than an independently testable empirical evaluation hypothesis.
Algorithmic specification and architectural definition provided as methodological framework rather than an independently testable empirical evaluation hypothesis.
Expository discussion and analytical limitation characterizing experimental design and architectural behavior without standalone quantitative test targets.
Expository discussion and analytical limitation characterizing experimental design and architectural behavior without standalone quantitative test targets.