Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/15
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
1
Inconclusive
14
Not assessed
Supporting Assessment
Measurement confirmed
The MCP census used a full official-registry snapshot from 2026-07-27, retained the latest version of each distinct server, anonymously queried tools/list on all remote targets, made no tool calls, and obtained 98,291 tools from the targets that returned at least one tool.
Reported
59625 entries
Observed
59625 entries
Difference 0
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
ClaimResultFinishedStatus
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 7 claims with no independent reproduction scheduled in this plan.
5 targets·5 with evidence·0 active·0 need attention
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
Atomix is classified as partially covering A1 and A3, covering compensation-safe A4 and speculation-safe A5-A6, and partially covering externally mediated A7; its known-outcome path assumes idempotence, it lacks crash-safe exactly-once behavior, and it has an A2 gap for multi-effect release.
Information insufficientclaim-runtime-atomix0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
None of the eight boundary capabilities required by the anomaly catalog is fully expressible in current standard MCP tool annotations: A1 has only a limited idempotence hint, A7 exposes externality only, and A2-A6 and A8 have no corresponding expression.
Information insufficientclaim-mcp-capability-expressibility0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Shepherd supplies per-effect reversibility tiers and observed A7 in a supervisor use case, but is described as a meta-agent substrate rather than a runtime guarantee; irreversible effects are logged.
Information insufficientclaim-runtime-shepherd0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
CoAgent is classified as partially covering externally mediated A7: reordering prevents some conflicts, registered inverses enable repair, and irreversible effects are gated, but this is an achievability condition rather than an unconditional guarantee.
Information insufficientclaim-runtime-coagent0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Tool-level signatures over true, false, or omitted values for the four annotation fields were highly concentrated: 66 of 81 possible signatures occurred, the leading signature covered 39.9% of tools, no annotations was second at 26.0%, and the top three covered 75.8%.
Official implementationclaim-tool-signature-concentration1 plan0 runs
Scientific conclusionNot assessed
Offline evaluation of the released 2026-07-27 MCP census artifact
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Annotation differentiation within servers was common but coarse: among multi-tool servers emitting at least one field, 76.7% used more than one signature, while the median dominant signature accounted for 79.4% of an emitting server's tools.
Official implementationclaim-server-signature-differentiation1 plan0 runs
Scientific conclusionNot assessed
Offline evaluation of the released 2026-07-27 MCP census artifact
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
RAC is classified as partially compensation-safe for A4, but it has no gating or idempotency and therefore does not address A1-A3.
Information insufficientclaim-runtime-rac0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Cordon is classified as partially covering A1 and A3, covering compensation-safe A4, and partially covering speculation-safe A5, conditional on idempotent tools and task-scoped staged effects.
Information insufficientclaim-runtime-cordon0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The paper identifies four boundaries at which black-box tool invocation alone is argued to be insufficient for a general guarantee: ambiguous outcomes, irreversible non-commuting effects without mediation, open-world reactions, and atomic release of multiple irreversible effects across tools.
Information insufficientclaim-black-box-guarantee-boundaries0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The catalog defines eight external-effect anomalies in three structurally motivated families: uncertainty (A1 duplicated, A2 missing, A3 orphaned), workflow (A4 residue, A5 premature, A6 contaminated), and interaction (A7 conflicting, A8 phantom).
Information insufficientclaim-anomaly-catalog0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
ACRFence is classified as partially targeting unknown-safe execution for A1; the proposed attack was validated, but the system was not implemented.
Information insufficientclaim-runtime-acrfence0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Although destructiveHint was serialized on 65.8% of tools, only 12.9% of all tools carried an applicable destructive classification and only 3.1% asserted that the operation was destructive.
Official implementationclaim-destructive-hint-applicability1 plan0 runs
Scientific conclusionNot assessed
Offline evaluation of the released 2026-07-27 MCP census artifact
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Across the six reviewed systems, coverage depends on runtime-owned declarations rather than a reusable shared contract; none controls exogenous reactions (A8), and no system covers the full anomaly catalog.
Information insufficientclaim-runtime-aggregate-gaps0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Standard MCP annotation fields were widely serialized, but a majority pairing made destructiveHint inapplicable because readOnlyHint was true, and the paper cautions that emitted values may come from defaults or templates rather than deliberate declarations.
Official implementationclaim-annotation-field-emission1 plan0 runs
Scientific conclusionNot assessed
Offline evaluation of the released 2026-07-27 MCP census artifact
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
The MCP census used a full official-registry snapshot from 2026-07-27, retained the latest version of each distinct server, anonymously queried tools/list on all remote targets, made no tool calls, and obtained 98,291 tools from the targets that returned at least one tool.
Official implementationclaim-census-sampling1 plan1 run
Scientific conclusionInconclusive
Offline evaluation of the released 2026-07-27 MCP census artifact