Reproduction materials
Prepared 0/1 materials
科研资料尚未闭合:configuration。这些问题只影响预下载缓存,不会阻止最终复现 Agent。
The preparation method is selected after link checks and inventory review.
Credit estimate for this reproduction
Expected usage
127
Maximum reservation
508
Estimated compute time
About 5 min
Actual usage was 475 credits; 0 credits were charged. Compute ran for about 22 minutes and the model used 18963722 tokens.
The paper’s claims are available in Research claims.
0 / 61 claims verified
the rest still being verified
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/27 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 07 · split: Train | 65% | — | Not assessed |
round: Round 07 · split: Train | 99.7 months | — | Not assessed |
round: Round 07 · split: Train | 40.722 million_usd | — | Not assessed |
round: Round 07 · split: Train | 100% | — | Not assessed |
round: Round 07 · split: Train | 75% | — | Not assessed |
round: Round 07 · split: Train | 65% | — | Not assessed |
round: Round 07 · split: Train | 3.19 calls_per_month | — | Not assessed |
round: Round 07 · split: Train | 100.7 actions | — | Not assessed |
round: Round 07 · split: Train | 40.46 million_usd | — | Not assessed |
round: Round 07 · split: Validate | 80% | — | Not assessed |
round: Round 07 · split: Validate | 114.15 months | — | Not assessed |
round: Round 07 · split: Validate | 47.607 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 31.3% | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 34.78% | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 37.39% | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 32.17% | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 26.09% | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 33.91% | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 38.26% | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: τ-bench | 44.35% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 83.1% | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 82.5% | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 82.3% | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 83.4% | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 83.1% | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 82.5% | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 82.7% | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: HotpotQA | 84.5% | — | Not assessed |
Reported
81.8%
Observed
—
Percentage reduction in validation tool invocations per month ((17.23 - 3.13) / 17.23 = 81.8%) derived directly from the EnterpriseArena self-evolution metrics in Table 11.
0/1 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 81.8% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/32 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline | 44% | — | Not assessed |
configuration: Baseline | 89.8 months | — | Not assessed |
configuration: Baseline | 78.86 million_usd | — | Not assessed |
configuration: Baseline | 100% | — | Not assessed |
configuration: Baseline | 78% | — | Not assessed |
configuration: Baseline | 52% | — | Not assessed |
configuration: Baseline | 0.13 calls_per_month | — | Not assessed |
configuration: Baseline | 152.12 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 50% | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 93.24 months | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 76.76 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 100% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 10 · split: Train | 90% | — | Not assessed |
round: Round 10 · split: Train | 121.15 months | — | Not assessed |
round: Round 10 · split: Train | 58.418 million_usd | — | Not assessed |
round: Round 10 · split: Train | 100% | — | Not assessed |
round: Round 10 · split: Train | 90% | — | Not assessed |
round: Round 10 · split: Train | 90% | — | Not assessed |
round: Round 10 · split: Train | 3.07 calls_per_month | — | Not assessed |
round: Round 10 · split: Train | 122.2 actions | — | Not assessed |
round: Round 10 · split: Train | 88.32 million_usd | — | Not assessed |
round: Round 10 · split: Validate | 85% | — | Not assessed |
round: Round 10 · split: Validate | 121.2 months | — | Not assessed |
round: Round 10 · split: Validate | 56.386 million_usd | — | Not assessed |
Comparative statistical test (Fisher exact test p-value and peak round comparison) derived from the EnterpriseArena self-evolution simulation test results in Table 11.
0/4 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 85% | — | Not assessed | |
| 0% | — | Not assessed | |
| 0 probability | — | Not assessed | |
| 95% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 33.4% | — | Not assessed | |
| 55.4% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 72.17% | — | Not assessed |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 67.83% | — | Not assessed |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 73.04% | — | Not assessed |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 64.35% | — | Not assessed |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 65.22% | — | Not assessed |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 67.83% | — | Not assessed |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 64.35% | — | Not assessed |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: τ-bench | 80% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 52% | — | Not assessed |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 56% | — | Not assessed |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 55% | — | Not assessed |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 52% | — | Not assessed |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 56% | — | Not assessed |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 52% | — | Not assessed |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 54% | — | Not assessed |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 60% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Baseline · split: Validate | 0% | — | Not assessed |
round: Baseline · split: Validate | 34.8 months | — | Not assessed |
round: Baseline · split: Validate | 28.377 million_usd | — | Not assessed |
round: Baseline · split: Validate | 100% | — | Not assessed |
round: Baseline · split: Validate | 0% | — | Not assessed |
round: Baseline · split: Validate | 0% | — | Not assessed |
round: Baseline · split: Validate | 17.23 calls_per_month | — | Not assessed |
round: Baseline · split: Validate | 35.7 actions | — | Not assessed |
round: Baseline · split: Validate | 0.47 million_usd | — | Not assessed |
round: Baseline · split: Test | 0% | — | Not assessed |
round: Baseline · split: Test | 33.3 months | — | Not assessed |
round: Baseline · split: Test | 28.754 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 57.23 points | — | Not assessed |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 55.19 points | — | Not assessed |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 66.79 points | — | Not assessed |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 65.66 points | — | Not assessed |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 53.58 points | — | Not assessed |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 61.55 points | — | Not assessed |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 68.15 points | — | Not assessed |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: GDPval | 71.19 points | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 26.12% | — | Not assessed |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 23.88% | — | Not assessed |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 14.93% | — | Not assessed |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 42.54% | — | Not assessed |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 30.6% | — | Not assessed |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 17.16% | — | Not assessed |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 18.66% | — | Not assessed |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: ALFWorld | 39.55% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline (no graph) | 54.8 points | — | Not assessed |
configuration: Baseline (no graph) | 275,638 tokens | — | Not assessed |
configuration: Baseline (no graph) | 28.2 steps | — | Not assessed |
configuration: Full graph, raw injection | 57.17 points | — | Not assessed |
configuration: Full graph, raw injection | 264,680 tokens | — | Not assessed |
configuration: Full graph, raw injection | 33.55 steps | — | Not assessed |
configuration: Full graph, generative | 56.75 points | — | Not assessed |
configuration: Full graph, generative | 448,972 tokens | — | Not assessed |
configuration: Full graph, generative | 22.07 steps | — | Not assessed |
configuration: Subgraph, generative (Ours) | 63.99 points | — | Not assessed |
configuration: Subgraph, generative (Ours) | 367,738 tokens | — | Not assessed |
configuration: Subgraph, generative (Ours) | 18.57 steps | — | Not assessed |
Experiment plan ready; no runs yet.
0/14 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
benchmark: HotpotQA | 9 nodes | — | Not assessed |
benchmark: HotpotQA | 9 triplets | — | Not assessed |
benchmark: MultiChallenge | 7 nodes | — | Not assessed |
benchmark: MultiChallenge | 7 triplets | — | Not assessed |
benchmark: GDPval | 15 nodes | — | Not assessed |
benchmark: GDPval | 22 triplets | — | Not assessed |
benchmark: ALFWorld | 11 nodes | — | Not assessed |
benchmark: ALFWorld | 27 triplets | — | Not assessed |
benchmark: τ-bench | 17 nodes | — | Not assessed |
benchmark: τ-bench | 18 triplets | — | Not assessed |
benchmark: BFCL v3 | 131 nodes | — | Not assessed |
benchmark: BFCL v3 | 265 triplets | — | Not assessed |
Experiment execution error · No automatic retry; a person must decide what to do next.
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 94.78% | — | Not assessed |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 85.07% | — | Not assessed |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 80.6% | — | Not assessed |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 97.76% | — | Not assessed |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 91.04% | — | Not assessed |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 99.25% | — | Not assessed |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 95.52% | — | Not assessed |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: ALFWorld | 100% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 06 · split: Train | 80% | — | Not assessed |
round: Round 06 · split: Train | 114.9 months | — | Not assessed |
round: Round 06 · split: Train | 46.345 million_usd | — | Not assessed |
round: Round 06 · split: Train | 100% | — | Not assessed |
round: Round 06 · split: Train | 90% | — | Not assessed |
round: Round 06 · split: Train | 80% | — | Not assessed |
round: Round 06 · split: Train | 3.09 calls_per_month | — | Not assessed |
round: Round 06 · split: Train | 115.9 actions | — | Not assessed |
round: Round 06 · split: Train | 56.52 million_usd | — | Not assessed |
round: Round 06 · split: Validate | 65% | — | Not assessed |
round: Round 06 · split: Validate | 108.75 months | — | Not assessed |
round: Round 06 · split: Validate | 43.831 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 56.39 points | — | Not assessed |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 61.97 points | — | Not assessed |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 71.37 points | — | Not assessed |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 64.69 points | — | Not assessed |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 56.18 points | — | Not assessed |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 63.85 points | — | Not assessed |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 69.1 points | — | Not assessed |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: GDPval | 78.78 points | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 86.57% | — | Not assessed |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 86.57% | — | Not assessed |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 79.1% | — | Not assessed |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 91.79% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 84.33% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 67.16% | — | Not assessed |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 87.31% | — | Not assessed |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: ALFWorld | 93.28% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 56% | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 54% | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 54% | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 58% | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 56% | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 52% | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 58% | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 67% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 59% | — | Not assessed |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 63% | — | Not assessed |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 57% | — | Not assessed |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 63% | — | Not assessed |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 63% | — | Not assessed |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 64% | — | Not assessed |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 61% | — | Not assessed |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 66% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 72.7% | — | Not assessed |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 70.7% | — | Not assessed |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 70.2% | — | Not assessed |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 73.5% | — | Not assessed |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 70% | — | Not assessed |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 74.3% | — | Not assessed |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 73.1% | — | Not assessed |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: HotpotQA | 74.4% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/32 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline | 0% | — | Not assessed |
configuration: Baseline | 33.58 months | — | Not assessed |
configuration: Baseline | 28.59 million_usd | — | Not assessed |
configuration: Baseline | 100% | — | Not assessed |
configuration: Baseline | 0% | — | Not assessed |
configuration: Baseline | 0% | — | Not assessed |
configuration: Baseline | 18.94 calls_per_month | — | Not assessed |
configuration: Baseline | 0 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 0% | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 34.04 months | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 28.75 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 100% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 86.96% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 86.67% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 71.43% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 100% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 87.5% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 52.17% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 66.67% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 42.86% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 72.73% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 58.93% | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 47.83% | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 53.33% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 81.33% | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 89.16% | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 89.16% | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 89.16% | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 89.16% | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 91.57% | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 91.57% | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 91.57% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/4 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| -15% | — | Not assessed | |
| -10.45 months | — | Not assessed | |
| -5% | — | Not assessed | |
| -0.184 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/27 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 09 · split: Train | 95% | — | Not assessed |
round: Round 09 · split: Train | 126.05 months | — | Not assessed |
round: Round 09 · split: Train | 58.188 million_usd | — | Not assessed |
round: Round 09 · split: Train | 100% | — | Not assessed |
round: Round 09 · split: Train | 95% | — | Not assessed |
round: Round 09 · split: Train | 95% | — | Not assessed |
round: Round 09 · split: Train | 3.08 calls_per_month | — | Not assessed |
round: Round 09 · split: Train | 127 actions | — | Not assessed |
round: Round 09 · split: Train | 97.84 million_usd | — | Not assessed |
round: Round 09 · split: Validate | 90% | — | Not assessed |
round: Round 09 · split: Validate | 121.2 months | — | Not assessed |
round: Round 09 · split: Validate | 56.57 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/32 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline | 6% | — | Not assessed |
configuration: Baseline | 50.28 months | — | Not assessed |
configuration: Baseline | 31.85 million_usd | — | Not assessed |
configuration: Baseline | 100% | — | Not assessed |
configuration: Baseline | 38% | — | Not assessed |
configuration: Baseline | 8% | — | Not assessed |
configuration: Baseline | 0.89 calls_per_month | — | Not assessed |
configuration: Baseline | 21.82 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 6% | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 51.34 months | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 32.27 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 100% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 04 · split: Train | 80% | — | Not assessed |
round: Round 04 · split: Train | 115.05 months | — | Not assessed |
round: Round 04 · split: Train | 46.631 million_usd | — | Not assessed |
round: Round 04 · split: Train | 100% | — | Not assessed |
round: Round 04 · split: Train | 90% | — | Not assessed |
round: Round 04 · split: Train | 80% | — | Not assessed |
round: Round 04 · split: Train | 3.06 calls_per_month | — | Not assessed |
round: Round 04 · split: Train | 116 actions | — | Not assessed |
round: Round 04 · split: Train | 53.04 million_usd | — | Not assessed |
round: Round 04 · split: Validate | 75% | — | Not assessed |
round: Round 04 · split: Validate | 109.2 months | — | Not assessed |
round: Round 04 · split: Validate | 49.831 million_usd | — | Not assessed |
Meta-statistical summary aggregation (win/tie/loss counts and binomial sign test) derived deterministically across the 24 model-benchmark evaluations in Table 1.
0/5 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
totalSettings: 24 | 21 settings | — | Not assessed |
| 19 settings | — | Not assessed | |
| 2 settings | — | Not assessed | |
| 3 settings | — | Not assessed | |
| 0.0004 probability | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 64.35% | — | Not assessed |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 58.26% | — | Not assessed |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 64.35% | — | Not assessed |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 63.48% | — | Not assessed |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 61.74% | — | Not assessed |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 63.48% | — | Not assessed |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 68.7% | — | Not assessed |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: τ-bench | 67.83% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/27 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 02 · split: Train | 85% | — | Not assessed |
round: Round 02 · split: Train | 117.75 months | — | Not assessed |
round: Round 02 · split: Train | 46.524 million_usd | — | Not assessed |
round: Round 02 · split: Train | 100% | — | Not assessed |
round: Round 02 · split: Train | 90% | — | Not assessed |
round: Round 02 · split: Train | 85% | — | Not assessed |
round: Round 02 · split: Train | 3.39 calls_per_month | — | Not assessed |
round: Round 02 · split: Train | 118.8 actions | — | Not assessed |
round: Round 02 · split: Train | 57.27 million_usd | — | Not assessed |
round: Round 02 · split: Validate | 80% | — | Not assessed |
round: Round 02 · split: Validate | 112.6 months | — | Not assessed |
round: Round 02 · split: Validate | 46.038 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/32 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline | 26% | — | Not assessed |
configuration: Baseline | 63.76 months | — | Not assessed |
configuration: Baseline | 39.42 million_usd | — | Not assessed |
configuration: Baseline | 100% | — | Not assessed |
configuration: Baseline | 42% | — | Not assessed |
configuration: Baseline | 28% | — | Not assessed |
configuration: Baseline | 0.47 calls_per_month | — | Not assessed |
configuration: Baseline | 27.24 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 28% | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 64.08 months | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 35.47 million_usd | — | Not assessed |
configuration: RAP (Kagaya et al., 2024) | 100% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Grok 4.1 Fast · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 68.07% | — | Not assessed |
model: Grok 4.1 Fast · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 84.94% | — | Not assessed |
model: Grok 4.1 Fast · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 80.12% | — | Not assessed |
model: Grok 4.1 Fast · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 84.94% | — | Not assessed |
model: Grok 4.1 Fast · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 80.72% | — | Not assessed |
model: Grok 4.1 Fast · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 83.73% | — | Not assessed |
model: Grok 4.1 Fast · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 83.13% | — | Not assessed |
model: Grok 4.1 Fast · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 86.75% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: ALFWorld | 82.84% | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: ALFWorld | 76.87% | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: ALFWorld | 76.12% | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: ALFWorld | 90.3% | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: ALFWorld | 80.6% | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: ALFWorld | 82.09% | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: ALFWorld | 81.34% | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: ALFWorld | 94.03% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 87.95% | — | Not assessed |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 95.18% | — | Not assessed |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 94.58% | — | Not assessed |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 92.77% | — | Not assessed |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 93.37% | — | Not assessed |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 93.98% | — | Not assessed |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 95.18% | — | Not assessed |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 95.78% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
mode: Mode 5 · generation: 1 | 77.59% | — | Not assessed |
mode: Mode 5 · generation: 10 | 83.31% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline (no graph) | 72.58% | — | Not assessed |
configuration: Baseline (no graph) | 18,055 tokens | — | Not assessed |
configuration: Baseline (no graph) | 21.84 steps | — | Not assessed |
configuration: Full graph, raw injection | 70.34% | — | Not assessed |
configuration: Full graph, raw injection | 21,062 tokens | — | Not assessed |
configuration: Full graph, raw injection | 25 steps | — | Not assessed |
configuration: Full graph, generative | 54.48% | — | Not assessed |
configuration: Full graph, generative | 96,360 tokens | — | Not assessed |
configuration: Full graph, generative | 30.05 steps | — | Not assessed |
configuration: Subgraph, generative (Ours) | 81.53% | — | Not assessed |
configuration: Subgraph, generative (Ours) | 28,064 tokens | — | Not assessed |
configuration: Subgraph, generative (Ours) | 18.8 steps | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 59.33 points | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 59.27 points | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 60.23 points | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 62.45 points | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 50.29 points | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 50.25 points | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 54 points | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: GDPval | 64.42 points | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/4 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
mode: Mode 3 · phase: initial | 54% | — | Not assessed |
mode: Mode 3 · generation: 5 · initialEdgeCount: 13 | 10 edges | — | Not assessed |
mode: Mode 3 · phase: restructured | 93.9% | — | Not assessed |
mode: Mode 5 (scratch) | 94.9% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: BFCL v3 | 61% | — | Not assessed |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: BFCL v3 | 62% | — | Not assessed |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: BFCL v3 | 65% | — | Not assessed |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: BFCL v3 | 60% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: BFCL v3 | 59% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: BFCL v3 | 62% | — | Not assessed |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: BFCL v3 | 60% | — | Not assessed |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: BFCL v3 | 67% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 74.6% | — | Not assessed |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 74.2% | — | Not assessed |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 71.8% | — | Not assessed |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 75.2% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 75.4% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 73.9% | — | Not assessed |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 74.3% | — | Not assessed |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: HotpotQA | 74.5% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: τ-bench | 66.09% | — | Not assessed |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: τ-bench | 63.48% | — | Not assessed |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: τ-bench | 60.87% | — | Not assessed |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: τ-bench | 71.3% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: τ-bench | 65.22% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: τ-bench | 66.96% | — | Not assessed |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: τ-bench | 61.74% | — | Not assessed |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: τ-bench | 73.91% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 58.8% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 71.21% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 62.8% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 76.61% | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 63.8% | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 77.16% | — | Not assessed |
constructionMode: Mode 3: Expert + Online Evolution | 63.1% | — | Not assessed |
constructionMode: Mode 3: Expert + Online Evolution | 76.34% | — | Not assessed |
constructionMode: Mode 4: Scratch + Static Build | 55.4% | — | Not assessed |
constructionMode: Mode 4: Scratch + Static Build | 69.49% | — | Not assessed |
constructionMode: Mode 5: Scratch + Online Evolution | 66.3% | — | Not assessed |
constructionMode: Mode 5: Scratch + Online Evolution | 78.79% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 87.5% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 7,403.98 tokens | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 7.05 steps | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 1.05 failures_per_sample | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 127.59 seconds | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 58.93% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 11,039.82 tokens | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 6.6 steps | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 0.88 failures_per_sample | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 117.7 seconds | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 53.57% | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 5,990.79 tokens | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/27 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 01 · split: Train | 0% | — | Not assessed |
round: Round 01 · split: Train | 33.4 months | — | Not assessed |
round: Round 01 · split: Train | 28.777 million_usd | — | Not assessed |
round: Round 01 · split: Train | 100% | — | Not assessed |
round: Round 01 · split: Train | 0% | — | Not assessed |
round: Round 01 · split: Train | 0% | — | Not assessed |
round: Round 01 · split: Train | 16.25 calls_per_month | — | Not assessed |
round: Round 01 · split: Train | 34.4 actions | — | Not assessed |
round: Round 01 · split: Train | 0 million_usd | — | Not assessed |
round: Round 01 · split: Validate | 45% | — | Not assessed |
round: Round 01 · split: Validate | 88.9 months | — | Not assessed |
round: Round 01 · split: Validate | 38.594 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/27 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 08 · split: Train | 85% | — | Not assessed |
round: Round 08 · split: Train | 120.35 months | — | Not assessed |
round: Round 08 · split: Train | 49.781 million_usd | — | Not assessed |
round: Round 08 · split: Train | 100% | — | Not assessed |
round: Round 08 · split: Train | 90% | — | Not assessed |
round: Round 08 · split: Train | 90% | — | Not assessed |
round: Round 08 · split: Train | 3.12 calls_per_month | — | Not assessed |
round: Round 08 · split: Train | 121.3 actions | — | Not assessed |
round: Round 08 · split: Train | 77.33 million_usd | — | Not assessed |
round: Round 08 · split: Validate | 90% | — | Not assessed |
round: Round 08 · split: Validate | 121.2 months | — | Not assessed |
round: Round 08 · split: Validate | 57.371 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: MultiChallenge | 83.73% | — | Not assessed |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: MultiChallenge | 89.16% | — | Not assessed |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: MultiChallenge | 86.75% | — | Not assessed |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: MultiChallenge | 89.16% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: MultiChallenge | 88.55% | — | Not assessed |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: MultiChallenge | 89.76% | — | Not assessed |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: MultiChallenge | 87.35% | — | Not assessed |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: MultiChallenge | 89.76% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 05 · split: Train | 80% | — | Not assessed |
round: Round 05 · split: Train | 114.05 months | — | Not assessed |
round: Round 05 · split: Train | 44.565 million_usd | — | Not assessed |
round: Round 05 · split: Train | 100% | — | Not assessed |
round: Round 05 · split: Train | 90% | — | Not assessed |
round: Round 05 · split: Train | 80% | — | Not assessed |
round: Round 05 · split: Train | 3.12 calls_per_month | — | Not assessed |
round: Round 05 · split: Train | 115 actions | — | Not assessed |
round: Round 05 · split: Train | 52.85 million_usd | — | Not assessed |
round: Round 05 · split: Validate (carried forward R04) | 75% | — | Not assessed |
round: Round 05 · split: Validate (carried forward R04) | 109.2 months | — | Not assessed |
round: Round 05 · split: Validate (carried forward R04) | 49.831 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Claude Sonnet 4.6 · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: GDPval | 46.23 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: MemoryBank (Zhong et al., 2024) · benchmark: GDPval | 50.13 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: RAP (Kagaya et al., 2024) · benchmark: GDPval | 40.97 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: ExpeL (Zhao et al., 2024) · benchmark: GDPval | 48.22 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: AutoGuide (Fu et al., 2024) · benchmark: GDPval | 41.39 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: AWM (Wang et al., 2025b) · benchmark: GDPval | 40.61 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: KnowAgent (Zhu et al., 2025) · benchmark: GDPval | 42.14 points | — | Not assessed |
model: Claude Sonnet 4.6 · method: Procedural Graph (Ours) · benchmark: GDPval | 51.49 points | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: Baseline (no graph) | 80.27% | — | Not assessed |
configuration: Baseline (no graph) | 6,629 tokens | — | Not assessed |
configuration: Baseline (no graph) | 3.87 steps | — | Not assessed |
configuration: Full graph, raw injection | 86.6% | — | Not assessed |
configuration: Full graph, raw injection | 10,164 tokens | — | Not assessed |
configuration: Full graph, raw injection | 4.54 steps | — | Not assessed |
configuration: Full graph, generative | 87.35% | — | Not assessed |
configuration: Full graph, generative | 14,434 tokens | — | Not assessed |
configuration: Full graph, generative | 3.08 steps | — | Not assessed |
configuration: Subgraph, generative (Ours) | 89.31% | — | Not assessed |
configuration: Subgraph, generative (Ours) | 12,295 tokens | — | Not assessed |
configuration: Subgraph, generative (Ours) | 4.22 steps | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.1 Pro · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 85.9% | — | Not assessed |
model: Gemini 3.1 Pro · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 85.1% | — | Not assessed |
model: Gemini 3.1 Pro · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 83.6% | — | Not assessed |
model: Gemini 3.1 Pro · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 86% | — | Not assessed |
model: Gemini 3.1 Pro · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 85.8% | — | Not assessed |
model: Gemini 3.1 Pro · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 85.1% | — | Not assessed |
model: Gemini 3.1 Pro · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 85% | — | Not assessed |
model: Gemini 3.1 Pro · method: Procedural Graph (Ours) · benchmark: HotpotQA | 87.3% | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
round: Round 03 · split: Train | 80% | — | Not assessed |
round: Round 03 · split: Train | 113.35 months | — | Not assessed |
round: Round 03 · split: Train | 46.857 million_usd | — | Not assessed |
round: Round 03 · split: Train | 100% | — | Not assessed |
round: Round 03 · split: Train | 85% | — | Not assessed |
round: Round 03 · split: Train | 80% | — | Not assessed |
round: Round 03 · split: Train | 3.08 calls_per_month | — | Not assessed |
round: Round 03 · split: Train | 114.3 actions | — | Not assessed |
round: Round 03 · split: Train | 51.88 million_usd | — | Not assessed |
round: Round 03 · split: Validate | 65% | — | Not assessed |
round: Round 03 · split: Validate | 102.15 months | — | Not assessed |
round: Round 03 · split: Validate | 43.965 million_usd | — | Not assessed |
Reproduction blocked by unavailability of proprietary foundation model API access required for agent action generation and trajectory evaluation.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
constructionMode: Unguided Baseline (w/o PG) | 71.21% | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 4,003.24 tokens | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 4.88 steps | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 0.016 failures_per_sample | — | Not assessed |
constructionMode: Unguided Baseline (w/o PG) | 18.06 seconds | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 76.61% | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 9,045.69 tokens | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 4.12 steps | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 0.007 failures_per_sample | — | Not assessed |
constructionMode: Mode 1: Hand-crafted Expert | 39.34 seconds | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 77.16% | — | Not assessed |
constructionMode: Mode 2: Expert + Static Update | 8,942.79 tokens | — | Not assessed |
Qualitative behavioral analysis and case study from execution traces illustrating agent decision dynamics rather than an independently evaluated scalar metric target.
Qualitative behavioral analysis and case study from execution traces illustrating agent decision dynamics rather than an independently evaluated scalar metric target.
Qualitative behavioral analysis and case study from execution traces illustrating agent decision dynamics rather than an independently evaluated scalar metric target.
Algorithmic specification and architectural definition provided as methodological framework rather than an independently testable empirical evaluation hypothesis.
Algorithmic specification and architectural definition provided as methodological framework rather than an independently testable empirical evaluation hypothesis.
Expository discussion and analytical limitation characterizing experimental design and architectural behavior without standalone quantitative test targets.
Expository discussion and analytical limitation characterizing experimental design and architectural behavior without standalone quantitative test targets.