Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/17
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
17
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 3 claims with no independent reproduction scheduled in this plan.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
Removing the current task from the curator input lowers untrained JITMEM-base by up to 3.1 ALFWorld and 4.6 WebShop success-rate points; after RL training, the maximum drops widen to 11.4 and 10.4 points, which the paper interprets as evidence that RL learns to exploit the task signal.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
RL-trained curator task/retrieval ablations on WebShop
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Training-free JITMEM design ablations on ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
RL-trained curator task/retrieval ablations on ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Appendix examples show task-specific payload synthesis on all three benchmarks: combining partially relevant episodes into an ALFWorld clean-and-place procedure, converting unrelated product searches into WebShop search and attribute-selection advice, and distilling MMS cases into an ordered tau2-bench diagnostic that retains the policy requirement for explicit approval before a mutating call.
Information insufficientclaim-example-payloads-0160 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On tau2-bench with GPT-5.4 executor, prompted JITMEM-gpt reports 75.6 micro-average success rate, 3.9 points above the strongest baseline, driven by 72.6 on Telecom; on Airline and Retail, the paper says no memory method improves over no memory beyond variance, and JITMEM remains on par with baselines.
CiteArk reconstructionclaim-tau2-0051 plan0 runs
Scientific conclusionNot assessed
Training-free JITMEM on tau2-bench
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
With Gemini-2.5-Pro as executor, transferred RL-trained JITMEM reports 86.2 ALFWorld success rate, 61.0 WebShop score, and 50.5 WebShop success rate; prompted JITMEM-gemini reports the highest WebShop values in the block, 72.1 score and 61.0 success rate.
Main WebShop comparison with Gemini-2.5-Pro executor
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Main ALFWorld comparison with Gemini-2.5-Pro executor
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On ALFWorld, a curator trained with Qwen3-8B and transferred to GPT-5.4 reaches 86.7 success rate, 1.4 points below a curator trained directly with GPT-5.4; with Qwen3-8B test executor the same training raises success from 60.5 to 77.4.
ALFWorld curator transfer across Qwen3-8B and GPT-5.4
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Changing the number of retrieved WebShop trajectories from k=3 to k=5 changes success rate by less than two points for every executor; the differences are within standard deviation for Qwen3-8B and GPT-5.4, and the main experiments use k=3.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For untrained JITMEM-base, storing all trajectories with correctness labels rather than filtering to successful trajectories lowers success by 1.5-2.9 points on ALFWorld and 2.3-3.4 on WebShop across executors; replacing raw traces with write-time ReasoningBank-style distillations lowers success by 1.7-2.9 on ALFWorld and 6.8-8.2 on WebShop.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Training-free JITMEM design ablations on ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
With Qwen3-8B as executor, RL-trained JITMEM reports 77.4 ALFWorld success rate, 61.1 WebShop score, and 32.8 WebShop success rate, exceeding the strongest listed write-time baseline by 16.2, 20.5, and 16.3 points respectively; the untrained JITMEM-base is competitive on ALFWorld but has lower WebShop score and success than SkillOS-base.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Stronger-curator control on Qwen3-8B ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
RL-trained JITMEM versus SkillOS on ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
RL-trained JITMEM versus SkillOS on WebShop
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Training-free read-time curation on ALFWorld with Qwen3-8B
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Training-free read-time curation on WebShop with Qwen3-8B
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On ALFWorld with GPT-5.4 executor, JITMEM uses 9.8K input tokens, 0.87K output tokens, and 11.6 steps per task; relative to JITMEM-base, the paper reports reductions of 10.1%, 13.0%, and 12.1%, while all memory methods add input context but generally reduce output tokens and steps versus no memory.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On WebShop, rebuilding the training bank after 100 GRPO steps and training 50 more raises success by 2.8 points for Qwen3-8B and 0.9 for GPT-5.4 but leaves Gemini-2.5-Pro unchanged; warm-starting the test bank with 100 training trajectories changes success by at most 1.3 points and remains within standard deviation.
CiteArk reconstructionclaim-bank-0111 plan0 runs
Scientific conclusionNot assessed
WebShop training-bank refresh and test-bank warm start
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Forcing the RL-trained curator to receive no retrieved trajectories degrades it to or below untrained JITMEM-base across all executors, with maximum reported success-rate drops of 14.8 points on ALFWorld and 15.2 on WebShop; the paper takes this as evidence that RL learns to distill retrieved experience rather than only emit hints from parametric knowledge.
RL-trained curator task/retrieval ablations on WebShop
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
RL-trained curator task/retrieval ablations on ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The authors' reproduced no-memory baselines are at or below the SkillOS-reported values. For Qwen3-8B, thinking mode is closer on ALFWorld, non-thinking is closer on WebShop, and WebShop additionally uses calibrated search guidance because the reported no-memory result could not otherwise be fully reproduced.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-8B WebShop no-memory calibration
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Gemini-2.5-Pro ALFWorld no-memory calibration
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Gemini-2.5-Pro WebShop no-memory calibration
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For an identical input, the untrained curator emits a generic action sequence, while the RL-trained curator supplies an ALFWorld-specific workflow: move to the desklamp and then examine the bowl with it; the paper characterizes this as emergence of environment-specific procedural semantics.
Information insufficientclaim-qualitative-rl-0150 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Two ALFWorld tasks retrieving the same past experience receive different payloads: the hot-potato task foregrounds heat/cool state-transition guidance, whereas the newspaper task foregrounds placement and target-location verification.
Information insufficientclaim-qualitative-adaptation-0140 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
With GPT-5.4 as executor, transferred RL-trained JITMEM reports 86.7 ALFWorld success rate, 53.8 WebShop score, and 45.4 WebShop success rate, while prompted JITMEM-gpt reports 83.3, 57.0, and 47.5 respectively; JITMEM-base with the weaker Qwen3-8B curator exceeds the listed stronger-curator write-time baselines on ALFWorld.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Main ALFWorld comparison with GPT-5.4 executor
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Main evaluation uses 140 ALFWorld test tasks, 500 WebShop test instances, and tau2-bench airline, retail, and telecom domains; banks start empty, tasks are streamed in batches, ground-truth verifiers score tasks, and the executor-as-judge controls bank admission.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Reference-hardware WebShop GRPO training time
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Reference-hardware ALFWorld GRPO training time
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
WebShop evaluation split and verifier identity
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Across 100 GRPO steps on ALFWorld and WebShop, the plotted training and validation success rates generally rise while executor turns fall; the paper describes training as stable under task reward alone, without auxiliary content-quality reward, task grouping, or return shaping.
Information insufficientclaim-training-curves-0172 plans0 runs
Scientific conclusionNot assessed
RL-trained JITMEM versus SkillOS on ALFWorld
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
RL-trained JITMEM versus SkillOS on WebShop
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Community reproductions
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.