Source: paper:p7-b74
No immutable Assessment has been published for this Claim yet.
No execution has been linked to this Claim yet.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Gemini 3.5 Flash · method: Vanilla ReAct (Yao et al., 2023b) · benchmark: HotpotQA | 83.1% | — | Not assessed |
model: Gemini 3.5 Flash · method: MemoryBank (Zhong et al., 2024) · benchmark: HotpotQA | 82.5% | — | Not assessed |
model: Gemini 3.5 Flash · method: RAP (Kagaya et al., 2024) · benchmark: HotpotQA | 82.3% | — | Not assessed |
model: Gemini 3.5 Flash · method: ExpeL (Zhao et al., 2024) · benchmark: HotpotQA | 83.4% | — | Not assessed |
model: Gemini 3.5 Flash · method: AutoGuide (Fu et al., 2024) · benchmark: HotpotQA | 83.1% | — | Not assessed |
model: Gemini 3.5 Flash · method: AWM (Wang et al., 2025b) · benchmark: HotpotQA | 82.5% | — | Not assessed |
model: Gemini 3.5 Flash · method: KnowAgent (Zhu et al., 2025) · benchmark: HotpotQA | 82.7% | — | Not assessed |
model: Gemini 3.5 Flash · method: Procedural Graph (Ours) · benchmark: HotpotQA | 84.5% | — | Not assessed |