浏览 ArkGraph,选择本次要执行的步骤。
论文中的结论可在「研究结论」中查看。
0 / 15 条结论已通过验证
其余仍在验证中
实验方案已生成,还没有运行记录
已评估 0/26 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 200 items | — | 尚未评估 | |
defense: None · question: Copy | 6.8% | — | 尚未评估 |
defense: None · question: First | 20% | — | 尚未评估 |
defense: None · question: Count | 6.4% | — | 尚未评估 |
defense: None · question: Redact | 0.8% | — | 尚未评估 |
defense: None · question: All | 8.5% | — | 尚未评估 |
defense: Strip · question: Copy | 0% | — | 尚未评估 |
defense: Strip · question: First | 8.8% | — | 尚未评估 |
defense: Strip · question: Count | 0% | — | 尚未评估 |
defense: Strip · question: Redact | 0% | — | 尚未评估 |
defense: Strip · question: All | 2.2% | — | 尚未评估 |
defense: Mask · question: Copy | 0% | — | 尚未评估 |
The text-interface boundary is methodological context, and the duplicated 46.9%/68.0% forged-system observation is derived from the represented task-only attack claim rather than a separate experiment.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
attack: Forged system · tokenization: standard · systemMessage: task-only | 46.9% | — | 尚未评估 |
attack: Forged system · tokenization: nameless · systemMessage: task-only | 68% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/31 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
row: Clean accuracy · tokenization: standard · trainingSeparation: None | 86.1% | — | 尚未评估 |
row: Naive · tokenization: standard · trainingSeparation: None | 7.4% | — | 尚未评估 |
row: Lookalike · tokenization: standard · trainingSeparation: None | 74% | — | 尚未评估 |
row: Forged turn · tokenization: standard · trainingSeparation: None | 61.1% | — | 尚未评估 |
row: Forged system · tokenization: standard · trainingSeparation: None | 98% | — | 尚未评估 |
row: Clean accuracy · tokenization: nameless · trainingSeparation: None | 86.1% | — | 尚未评估 |
row: Naive · tokenization: nameless · trainingSeparation: None | 7.4% | — | 尚未评估 |
row: Lookalike · tokenization: nameless · trainingSeparation: None | 74% | — | 尚未评估 |
row: Forged turn · tokenization: nameless · trainingSeparation: None | 60.8% | — | 尚未评估 |
row: Forged system · tokenization: nameless · trainingSeparation: None | 66.6% | — | 尚未评估 |
row: Clean accuracy · tokenization: standard · trainingSeparation: ISE | 85.1% | — | 尚未评估 |
row: Naive · tokenization: standard · trainingSeparation: ISE | 15.5% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/5 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 16,415 perturbations | — | 尚未评估 | |
defense: none | 27.3% | — | 尚未评估 |
scope: all tested models · defense: split_special_tokens | 5.1% | — | 尚未评估 |
model: Qwen3.8-27B · defense: split_special_tokens | 18.3% | — | 尚未评估 |
defense: nameless tokenization | 0% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/60 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemma-4-31B · component: Absent · objective: Objective A | 57.8% | — | 尚未评估 |
model: Gemma-4-31B · component: Text · objective: Objective A | 99% | — | 尚未评估 |
model: Gemma-4-31B · component: Reserved · objective: Objective A | 99.7% | — | 尚未评估 |
model: Gemma-4-31B · component: Surface · objective: Objective A | 41.2 percentage points | — | 尚未评估 |
model: Gemma-4-31B · component: Identifier · objective: Objective A | 0.7 percentage points | — | 尚未评估 |
model: Llama-3.1-8B · component: Absent · objective: Objective A | 63.2% | — | 尚未评估 |
model: Llama-3.1-8B · component: Text · objective: Objective A | 99.7% | — | 尚未评估 |
model: Llama-3.1-8B · component: Reserved · objective: Objective A | 100% | — | 尚未评估 |
model: Llama-3.1-8B · component: Surface · objective: Objective A | 36.5 percentage points | — | 尚未评估 |
model: Llama-3.1-8B · component: Identifier · objective: Objective A | 0.3 percentage points | — | 尚未评估 |
model: Ministral-3-8B · component: Absent · objective: Objective A | 3% | — | 尚未评估 |
model: Ministral-3-8B · component: Text · objective: Objective A | 67.2% | — | 尚未评估 |
The visible Figure 1 document supports a real paired probe. The other 23 documents are absent, so the catalog retains a clearly labeled proxy route and permits exact rate comparison only if the original manifest is found.
已评估 0/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 24 documents | — | 尚未评估 | |
tokenization: standard | 100% | — | 尚未评估 |
tokenization: nameless | 4.2% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/38 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 400 repositories | — | 尚未评估 | |
| 120 repositories | — | 尚未评估 | |
| 24 repositories | — | 尚未评估 | |
| 256 tokenizers | — | 尚未评估 | |
| 99.6 percent of analyzed tokenizers | — | 尚未评估 | |
| 56.2 percent of analyzed tokenizers | — | 尚未评估 | |
family: Qwen | 59 tokenizers | — | 尚未评估 |
family: Qwen · condition: default | 100% | — | 尚未评估 |
family: Qwen · condition: split_special_tokens enabled | 84.7% | — | 尚未评估 |
family: Nvidia | 11 tokenizers | — | 尚未评估 |
family: Nvidia · condition: default | 100% | — | 尚未评估 |
family: Nvidia · condition: split_special_tokens enabled | 81.8% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/16 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 296 items | — | 尚未评估 | |
task: sentiment classification | 100 items | — | 尚未评估 |
task: natural-language inference | 100 items | — | 尚未评估 |
task: extractive question answering | 96 items | — | 尚未评估 |
| 0.42% | — | 尚未评估 | |
| 4 models | — | 尚未评估 | |
defense: None | 100% | — | 尚未评估 |
defense: Strip | 100% | — | 尚未评估 |
defense: Mask | 100% | — | 尚未评估 |
defense: Escape | 100% | — | 尚未评估 |
defense: Nameless | 100% | — | 尚未评估 |
defense: None | 89.1% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/110 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemma-4-31B · attack: Naive · defense: None | 58.1% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Strip | 57.8% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Mask | 57.8% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Escape | 57.8% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Nameless | 57.8% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: None | 94.3% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Strip | 94.3% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Mask | 94.3% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Escape | 94.3% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Nameless | 94.3% | — | 尚未评估 |
model: Gemma-4-31B · attack: Forged turn · defense: None | 99.7% | — | 尚未评估 |
model: Gemma-4-31B · attack: Forged turn · defense: Strip | 76% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/85 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Gemma-4-31B · attack: Naive · defense: None | 0% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Strip | 0% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Mask | 0% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Escape | 0% | — | 尚未评估 |
model: Gemma-4-31B · attack: Naive · defense: Nameless | 0% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: None | 2% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Strip | 2% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Mask | 2% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Escape | 2% | — | 尚未评估 |
model: Gemma-4-31B · attack: Lookalike · defense: Nameless | 2% | — | 尚未评估 |
model: Gemma-4-31B · attack: Forged turn · defense: None | 8.1% | — | 尚未评估 |
model: Gemma-4-31B · attack: Forged turn · defense: Strip | 0.7% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/10 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: gemma-4-31B-it · owner: google | 13 identifiers | — | 尚未评估 |
model: gemma-4-31B-it · toolRole: no | 0 identifiers | — | 尚未评估 |
model: Llama-3.1-8B-Instruct · owner: meta-llama | 4 identifiers | — | 尚未评估 |
model: Llama-3.1-8B-Instruct · toolRole: no | 0 identifiers | — | 尚未评估 |
model: Ministral-3-8B-Instruct-2512 · owner: mistralai | 12 identifiers | — | 尚未评估 |
model: Ministral-3-8B-Instruct-2512 · toolRole: yes | 0 identifiers | — | 尚未评估 |
model: Qwen3.8-27B · owner: Qwen | 8 identifiers | — | 尚未评估 |
model: Qwen3.8-27B · toolRole: yes | 6 identifiers | — | 尚未评估 |
model: gpt-oss-20b · owner: openai | 7 identifiers | — | 尚未评估 |
model: gpt-oss-20b · toolRole: no | 0 identifiers | — | 尚未评估 |
This is source context limiting interpretation to five models, three host tasks, two objectives, and one defensive sentence, not an additional independent empirical result. The represented experiment packages preserve those boundaries.
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 5 models | — | 尚未评估 | |
| 3 tasks | — | 尚未评估 | |
| 2 objectives | — | 尚未评估 | |
| 1 sentences | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/29 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
attack: Naive · defense: None | 62.5% | — | 尚未评估 |
attack: Naive · defense: Strip | 62.2% | — | 尚未评估 |
attack: Naive · defense: Mask | 62.2% | — | 尚未评估 |
attack: Naive · defense: Escape | 62.3% | — | 尚未评估 |
attack: Naive · defense: Nameless | 62.4% | — | 尚未评估 |
attack: Lookalike · defense: None | 93% | — | 尚未评估 |
attack: Lookalike · defense: Strip | 93% | — | 尚未评估 |
attack: Lookalike · defense: Mask | 93% | — | 尚未评估 |
attack: Lookalike · defense: Escape | 93% | — | 尚未评估 |
attack: Lookalike · defense: Nameless | 93% | — | 尚未评估 |
attack: Forged turn · defense: None | 99.9% | — | 尚未评估 |
attack: Forged turn · defense: None | 99.8% | — | 尚未评估 |
The independent renderer is inspected and challenged by the finite perturbation property test. Execution must state that passing finite tests can falsify defects but cannot prove the universal text-to-identifier guarantee.
This is an actual limitation on external validity—synthetic content frequency is not estimated and custom-code tokenizers are excluded—not an independently testable result. Both constraints remain explicit in the corresponding objectives.