Explore ArkGraph and select the steps to run.
The paper’s claims are available in Research claims.
0 / 20 claims verified
the rest still being verified
Experiment plan ready; no runs yet.
0/20 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
outcome: Evidence-grounding failure | 33 candidates | — | Not assessed |
outcome: Evidence-grounding failure | 33% | — | Not assessed |
outcome: Ambiguity/non-uniqueness | 22 candidates | — | Not assessed |
outcome: Ambiguity/non-uniqueness | 22% | — | Not assessed |
outcome: Query leakage or construction issue | 16 candidates | — | Not assessed |
outcome: Query leakage or construction issue | 16% | — | Not assessed |
outcome: Task-consistency failure | 14 candidates | — | Not assessed |
outcome: Task-consistency failure | 14% | — | Not assessed |
outcome: Coverage/completeness failure | 11 candidates | — | Not assessed |
outcome: Coverage/completeness failure | 11% | — | Not assessed |
outcome: Correctly rejected (subtotal) | 96 candidates | — | Not assessed |
outcome: Correctly rejected (subtotal) | 96% | — | Not assessed |
Experiment plan ready; no runs yet.
0/21 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
comparison: two primary annotators before adjudication | 96% | — | Not assessed |
comparison: two primary annotators before adjudication | 0.916 dimensionless | — | Not assessed |
evaluator: Human reference | 62% | — | Not assessed |
evaluator: MiniMax-2.5 · reference: adjudicated human labels | 58.5% | — | Not assessed |
evaluator: MiniMax-2.5 · reference: adjudicated human labels | 91.5% | — | Not assessed |
evaluator: MiniMax-2.5 · reference: adjudicated human labels | 0.823 dimensionless | — | Not assessed |
evaluator: MiniMax-2.5 · reference: adjudicated human labels | 95.7% | — | Not assessed |
evaluator: MiniMax-2.5 · reference: adjudicated human labels | 90.3% | — | Not assessed |
evaluator: MiniMax-2.5 · reference: adjudicated human labels | 92.9% | — | Not assessed |
evaluator: GPT-5.5 · reference: adjudicated human labels | 60% | — | Not assessed |
evaluator: GPT-5.5 · reference: adjudicated human labels | 92% | — | Not assessed |
evaluator: GPT-5.5 · reference: adjudicated human labels | 0.832 dimensionless | — | Not assessed |
Experiment plan ready; no runs yet.
0/20 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
variant: w/o Memorizer · category: overall · benchmark: locomo | 49.06 F1 points (0–100) | — | Not assessed |
variant: w/o Memorizer · category: overall · benchmark: longmemeval | 64% | — | Not assessed |
variant: w/o Memorizer · category: overall · benchmark: narrativeqa | 38.87 F1 points (0–100) | — | Not assessed |
variant: w/o Memorizer · category: overall · benchmark: hotpotqa | 57.72 F1 points (0–100) | — | Not assessed |
variant: w/o Researcher · category: overall · benchmark: locomo | 43.3 F1 points (0–100) | — | Not assessed |
variant: w/o Researcher · category: overall · benchmark: longmemeval | 58.4% | — | Not assessed |
variant: w/o Researcher · category: overall · benchmark: narrativeqa | 32.51 F1 points (0–100) | — | Not assessed |
variant: w/o Researcher · category: overall · benchmark: hotpotqa | 39.1 F1 points (0–100) | — | Not assessed |
setting: Flat Raw Store · benchmark: LoCoMo | 0 tokens/workspace | — | Not assessed |
setting: Flat Raw Store · benchmark: LoCoMo | 0 s/workspace | — | Not assessed |
setting: Flat Raw Store · benchmark: LoCoMo | 8,174.29 tokens/query | — | Not assessed |
setting: Flat Raw Store · benchmark: LoCoMo | 17.11 s/query | — | Not assessed |
Experiment plan ready; no runs yet.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
variant: w/o BM25 · category: overall · benchmark: locomo | 45.9 F1 points (0–100) | — | Not assessed |
variant: w/o BM25 · category: overall · benchmark: longmemeval | 62.4% | — | Not assessed |
variant: w/o BM25 · category: overall · benchmark: narrativeqa | 38.88 F1 points (0–100) | — | Not assessed |
variant: w/o BM25 · category: overall · benchmark: hotpotqa | 52.14 F1 points (0–100) | — | Not assessed |
variant: w/o Embedding · category: overall · benchmark: locomo | 50.86 F1 points (0–100) | — | Not assessed |
variant: w/o Embedding · category: overall · benchmark: longmemeval | 61.8% | — | Not assessed |
variant: w/o Embedding · category: overall · benchmark: narrativeqa | 39.45 F1 points (0–100) | — | Not assessed |
variant: w/o Embedding · category: overall · benchmark: hotpotqa | 58.13 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/28 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
variant: w/o SFT · category: overall · benchmark: locomo | 48.59 F1 points (0–100) | — | Not assessed |
variant: w/o SFT · category: overall · benchmark: longmemeval | 48.8% | — | Not assessed |
variant: w/o SFT · category: overall · benchmark: narrativeqa | 34.98 F1 points (0–100) | — | Not assessed |
variant: w/o SFT · category: overall · benchmark: hotpotqa | 45.44 F1 points (0–100) | — | Not assessed |
variant: MemAgent data · category: overall · benchmark: locomo | 48.6 F1 points (0–100) | — | Not assessed |
variant: MemAgent data · category: overall · benchmark: longmemeval | 55.2% | — | Not assessed |
variant: MemAgent data · category: overall · benchmark: narrativeqa | 31.98 F1 points (0–100) | — | Not assessed |
variant: MemAgent data · category: overall · benchmark: hotpotqa | 54.53 F1 points (0–100) | — | Not assessed |
variant: Memory-R1 data · category: overall · benchmark: locomo | 48.9 F1 points (0–100) | — | Not assessed |
variant: Memory-R1 data · category: overall · benchmark: longmemeval | 57.2% | — | Not assessed |
variant: Memory-R1 data · category: overall · benchmark: narrativeqa | 32.42 F1 points (0–100) | — | Not assessed |
variant: Memory-R1 data · category: overall · benchmark: hotpotqa | 49.81 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/20 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
benchmark: LoCoMo · maximumActionRounds: 20 | 95.58% | — | Not assessed |
benchmark: LoCoMo · maximumActionRounds: 20 | 3.16 rounds/query | — | Not assessed |
benchmark: LoCoMo · maximumActionRounds: 20 | 3 rounds/query | — | Not assessed |
benchmark: LoCoMo · maximumActionRounds: 20 | 4 rounds/query | — | Not assessed |
benchmark: LoCoMo · maximumActionRounds: 20 | 0% | — | Not assessed |
benchmark: LongMemEval · maximumActionRounds: 20 | 82.8% | — | Not assessed |
benchmark: LongMemEval · maximumActionRounds: 20 | 3.82 rounds/query | — | Not assessed |
benchmark: LongMemEval · maximumActionRounds: 20 | 3 rounds/query | — | Not assessed |
benchmark: LongMemEval · maximumActionRounds: 20 | 6 rounds/query | — | Not assessed |
benchmark: LongMemEval · maximumActionRounds: 20 | 0.4% | — | Not assessed |
benchmark: NarrativeQA · maximumActionRounds: 20 | 84% | — | Not assessed |
benchmark: NarrativeQA · maximumActionRounds: 20 | 3.92 rounds/query | — | Not assessed |
Experiment plan ready; no runs yet.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: A-MEM · benchmark: LoCoMo | 200.42 s/workspace | — | Not assessed |
method: A-MEM · benchmark: LoCoMo | 0.48 s/query | — | Not assessed |
method: A-MEM · benchmark: LoCoMo | 41.17 F1 points (0–100) | — | Not assessed |
method: Mem0 · benchmark: LoCoMo | 90.67 s/workspace | — | Not assessed |
method: Mem0 · benchmark: LoCoMo | 0.2 s/query | — | Not assessed |
method: Mem0 · benchmark: LoCoMo | 32.1 F1 points (0–100) | — | Not assessed |
method: MemoryOS · benchmark: LoCoMo | 41.36 s/workspace | — | Not assessed |
method: MemoryOS · benchmark: LoCoMo | 0.52 s/query | — | Not assessed |
method: MemoryOS · benchmark: LoCoMo | 31.67 F1 points (0–100) | — | Not assessed |
method: LightMem · benchmark: LoCoMo | 6.82 s/workspace | — | Not assessed |
method: LightMem · benchmark: LoCoMo | 0.22 s/query | — | Not assessed |
method: LightMem · benchmark: LoCoMo | 32.97 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/100 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: VANILLA · baseLLM: Qwen3.5-4B · category: SH · benchmark: locomo | 48.15 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: MH · benchmark: locomo | 32.98 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: TE · benchmark: locomo | 42.66 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: OD · benchmark: locomo | 19.63 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: OA · benchmark: locomo | 42.45 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: LME · benchmark: longmemeval | 54.8% | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: NAQA · benchmark: narrativeqa | 31.23 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: 56k · benchmark: hotpotqa-56k | 63.56 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: 112k · benchmark: hotpotqa-112k | 53.04 F1 points (0–100) | — | Not assessed |
method: VANILLA · baseLLM: Qwen3.5-4B · category: 224k · benchmark: hotpotqa-224k | 36.97 F1 points (0–100) | — | Not assessed |
method: RAG · baseLLM: Qwen3.5-4B · category: SH · benchmark: locomo | 49.04 F1 points (0–100) | — | Not assessed |
method: RAG · baseLLM: Qwen3.5-4B · category: MH · benchmark: locomo | 34.39 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/20 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
benchmark: LoCoMo F1 · statisticOrRun: Run 1 | 52.09 F1 points (0–100) | — | Not assessed |
benchmark: LoCoMo F1 · statisticOrRun: Run 2 | 52.68 F1 points (0–100) | — | Not assessed |
benchmark: LoCoMo F1 · statisticOrRun: Run 3 | 51.1 F1 points (0–100) | — | Not assessed |
benchmark: LoCoMo F1 · statisticOrRun: Mean | 51.96 F1 points (0–100) | — | Not assessed |
benchmark: LoCoMo F1 · statisticOrRun: Std. | 0.8 F1 points (0–100) | — | Not assessed |
benchmark: LongMemEval Acc. · statisticOrRun: Run 1 | 65.2% | — | Not assessed |
benchmark: LongMemEval Acc. · statisticOrRun: Run 2 | 64.6% | — | Not assessed |
benchmark: LongMemEval Acc. · statisticOrRun: Run 3 | 65% | — | Not assessed |
benchmark: LongMemEval Acc. · statisticOrRun: Mean | 64.93% | — | Not assessed |
benchmark: LongMemEval Acc. · statisticOrRun: Std. | 0.31 percentage points | — | Not assessed |
benchmark: NarrativeQA F1 · statisticOrRun: Run 1 | 45.72 F1 points (0–100) | — | Not assessed |
benchmark: NarrativeQA F1 · statisticOrRun: Run 2 | 43.37 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-4B (default) | 79.53 s/workspace | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-4B (default) | 13.81 s/query | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-4B (default) | 3.16 rounds/query | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-4B (default) | 52.09 F1 points (0–100) | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-122B-A10B | 132.43 s/workspace | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-122B-A10B | 14.36 s/query | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-122B-A10B | 3.25 rounds/query | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: Qwen3.5-122B-A10B | 51.3 F1 points (0–100) | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: GPT-5.5 | 117.65 s/workspace | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: GPT-5.5 | 13.19 s/query | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: GPT-5.5 | 3.03 rounds/query | — | Not assessed |
researcher: trained Qwen3.5-4B · memorizerBackbone: GPT-5.5 | 52.62 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/42 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
criterion: Evidence support | 200 instances | — | Not assessed |
criterion: Evidence support | 196 instances | — | Not assessed |
criterion: Evidence support | 98% | — | Not assessed |
criterion: Evidence support | 99.5% | — | Not assessed |
criterion: Evidence support | 0.886 dimensionless | — | Not assessed |
criterion: Answer uniqueness | 134 instances | — | Not assessed |
criterion: Answer uniqueness | 131 instances | — | Not assessed |
criterion: Answer uniqueness | 97.8% | — | Not assessed |
criterion: Answer uniqueness | 99.3% | — | Not assessed |
criterion: Answer uniqueness | 0.853 dimensionless | — | Not assessed |
criterion: Leakage-free | 200 instances | — | Not assessed |
criterion: Leakage-free | 199 instances | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
contextSubset: 64K · trainingStage: Untrained | 56.58% | — | Not assessed |
contextSubset: 128K · trainingStage: Untrained | 55.43% | — | Not assessed |
contextSubset: 256K · trainingStage: Untrained | 49.23% | — | Not assessed |
contextSubset: Overall · trainingStage: Untrained | 54.08% | — | Not assessed |
contextSubset: 64K · trainingStage: SFT only | 68.42% | — | Not assessed |
contextSubset: 128K · trainingStage: SFT only | 64.13% | — | Not assessed |
contextSubset: 256K · trainingStage: SFT only | 64.62% | — | Not assessed |
contextSubset: Overall · trainingStage: SFT only | 65.67% | — | Not assessed |
contextSubset: 64K · trainingStage: SFT+RL | 76.32% | — | Not assessed |
contextSubset: 128K · trainingStage: SFT+RL | 76.09% | — | Not assessed |
contextSubset: 256K · trainingStage: SFT+RL | 73.85% | — | Not assessed |
contextSubset: Overall · trainingStage: SFT+RL | 75.54% | — | Not assessed |
Experiment plan ready; no runs yet.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 59.4% | — | Not assessed | |
representation: Figure 2 printed rate | 0.59 proportion | — | Not assessed |
Experiment plan ready; no runs yet.
0/16 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
variant: w/o Training · category: overall · benchmark: locomo | 48.59 F1 points (0–100) | — | Not assessed |
variant: w/o Training · category: overall · benchmark: longmemeval | 48.8% | — | Not assessed |
variant: w/o Training · category: overall · benchmark: narrativeqa | 34.98 F1 points (0–100) | — | Not assessed |
variant: w/o Training · category: overall · benchmark: hotpotqa | 45.44 F1 points (0–100) | — | Not assessed |
variant: SFT Only · category: overall · benchmark: locomo | 50.7 F1 points (0–100) | — | Not assessed |
variant: SFT Only · category: overall · benchmark: longmemeval | 60.2% | — | Not assessed |
variant: SFT Only · category: overall · benchmark: narrativeqa | 41.59 F1 points (0–100) | — | Not assessed |
variant: SFT Only · category: overall · benchmark: hotpotqa | 52.08 F1 points (0–100) | — | Not assessed |
variant: GRPO w/o Hint · category: overall · benchmark: locomo | 51.42 F1 points (0–100) | — | Not assessed |
variant: GRPO w/o Hint · category: overall · benchmark: longmemeval | 62.4% | — | Not assessed |
variant: GRPO w/o Hint · category: overall · benchmark: narrativeqa | 43.62 F1 points (0–100) | — | Not assessed |
variant: GRPO w/o Hint · category: overall · benchmark: hotpotqa | 55.64 F1 points (0–100) | — | Not assessed |
Experiment plan ready; no runs yet.
0/29 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
method: RAG · benchmark: LoCoMo · sourceBlock: Evaluated with our judge | 58.25% | — | Not assessed |
method: RAG · benchmark: NarrativeQA · sourceBlock: Evaluated with our judge | 48% | — | Not assessed |
method: RAG · benchmark: HotpotQA · sourceBlock: Evaluated with our judge | 53.13% | — | Not assessed |
method: A-MEM · benchmark: LoCoMo · sourceBlock: Evaluated with our judge | 54.74% | — | Not assessed |
method: A-MEM · benchmark: NarrativeQA · sourceBlock: Evaluated with our judge | 45% | — | Not assessed |
method: A-MEM · benchmark: HotpotQA · sourceBlock: Evaluated with our judge | 35.68% | — | Not assessed |
method: Mem0 · benchmark: LoCoMo · sourceBlock: Evaluated with our judge | 46.23% | — | Not assessed |
method: Mem0 · benchmark: NarrativeQA · sourceBlock: Evaluated with our judge | 43% | — | Not assessed |
method: Mem0 · benchmark: HotpotQA · sourceBlock: Evaluated with our judge | 38.02% | — | Not assessed |
method: MemoryOS · benchmark: LoCoMo · sourceBlock: Evaluated with our judge | 47.66% | — | Not assessed |
method: MemoryOS · benchmark: NarrativeQA · sourceBlock: Evaluated with our judge | 40% | — | Not assessed |
method: MemoryOS · benchmark: HotpotQA · sourceBlock: Evaluated with our judge | 30.21% | — | Not assessed |
The proposed observation comparison is retained, but a deterministic decision rule is unresolved.
The proposed observation comparison is retained, but a deterministic decision rule is unresolved.
The proposed observation comparison is retained, but a deterministic decision rule is unresolved.
Source-stated local limitation/interpretation, not an independently reported empirical target. Preserve as protocol context: The paper keeps the Memorizer fixed during Researcher training and does not study Memorizer fine-tuning or joint Memorizer–Researcher optimization. The backbone comparison isolates model choice with a fixed trained Researcher.
Source-stated local limitation/interpretation, not an independently reported empirical target. Preserve as protocol context: Reaching the action cap does not imply an incorrect answer, and a positive sufficiency assessment does not guarantee complete evidence recovery. Budget-exhausted queries receive the same finalization step over collected evidence, without further exploration.