Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/4
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
1
Inconclusive
3
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
ClaimResultFinishedStatus
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 1 claims with no independent reproduction scheduled in this plan.
1 targets·1 with evidence·0 active·0 need attention
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
Prevention-based defenses that perform well on short-context benchmarks retain substantial attack success on the paper's long-context synthetic benchmark, although PromptLocate and MetaSecAlign 8B are lower than several simpler defenses on some tasks.
Information insufficientclaim-prevention-degradation0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The paper reports that detection-based defenses exhibit an extreme false-positive/false-negative trade-off on long-context inputs, with some methods flagging benign inputs and others missing attacks.
Information insufficientclaim-detection-tradeoff0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The paper reports that GCG and its universal variant achieve high attack success across all four task suites and outperform heuristic attacks on several tasks.
Information insufficientclaim-optimization-attacks0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On long-context document tasks, heuristic prompt-injection attacks substantially increase attack success rate over no-attack baselines, and Authority spoof is generally the strongest listed heuristic.
CiteArk reconstructionclaim-heuristic-attacks-effective1 plan1 run
Scientific conclusionInconclusive
Independent approximate paper-review Authority spoof ASR on controlled long contexts
exp-reconstruct-paper-review-authority2 attempts
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
Community reproductions
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.