浏览 ArkGraph,选择本次要执行的步骤。
论文中的结论可在「研究结论」中查看。
0 / 28 条结论已通过验证
其余仍在验证中
实验方案已生成,还没有运行记录
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3.6-35B · notice: next-turn | 0.98 proportion | — | 尚未评估 |
model: Qwen3.6-35B · notice: same-turn | 0.97 proportion | — | 尚未评估 |
model: Qwen3-235B · notice: next-turn | 0.98 proportion | — | 尚未评估 |
model: Qwen3-235B · notice: same-turn | 0.98 proportion | — | 尚未评估 |
model: Qwen3.6-35B · notice: next-turn · denominator: 86 | 70 replies | — | 尚未评估 |
model: Qwen3.6-35B · notice: same-turn · denominator: 89 | 0 replies | — | 尚未评估 |
model: Qwen3-235B · notice: next-turn · denominator: 27 | 9 replies | — | 尚未评估 |
model: Qwen3-235B · notice: same-turn · denominator: 17 | 0 replies | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/72 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.38 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.5 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 35 events per 128 episodes | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.5 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.66 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 24 events per 128 episodes | — | 尚未评估 |
arm: top-k retrieval, replace · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0 proportion | — | 尚未评估 |
arm: top-k retrieval, replace · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0 proportion | — | 尚未评估 |
arm: top-k retrieval, replace · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 157 events per 128 episodes | — | 尚未评估 |
arm: top-k retrieval, replace · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0 p-value | — | 尚未评估 |
arm: top-k retrieval, replace · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.02 proportion | — | 尚未评估 |
arm: top-k retrieval, replace · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.02 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/22 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.5 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.79 proportion | — | 尚未评估 |
arm: directive gate · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.9 proportion | — | 尚未评估 |
arm: enforcement gate · brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.98 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.66 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.77 proportion | — | 尚未评估 |
arm: directive gate · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.92 proportion | — | 尚未评估 |
arm: enforcement gate · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 0.98 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 15 · density: 0.3 · reasoning: no-think | 0.34 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 15 · density: 0.3 · reasoning: no-think | 0.45 proportion | — | 尚未评估 |
arm: directive gate · brief: amended · model: Qwen3.6-35B · steps: 15 · density: 0.3 · reasoning: no-think | 0.77 proportion | — | 尚未评估 |
arm: enforcement gate · brief: amended · model: Qwen3.6-35B · steps: 15 · density: 0.3 · reasoning: no-think | 0.97 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/203 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.26 proportion | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.49 proportion of labeled turns | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.28 proportion of labeled turns | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.06 proportion of labeled turns | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.04 proportion of labeled turns | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0 proportion of labeled turns | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 4 executions per 128 episodes | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.39 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.15 proportion of labeled turns | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.35 proportion of labeled turns | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0.02 proportion of labeled turns | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · reasoning: no-think · configuration: S=5 | 0 proportion of labeled turns | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/44 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-235B · variant: fresh work order for finished step · protocol: text | 0.21 proportion | — | 尚未评估 |
model: Qwen3-235B · variant: fresh work order for finished step · protocol: tools | 0.09 proportion | — | 尚未评估 |
model: Qwen3.6-35B · variant: fresh work order for finished step · protocol: text | 0.47 proportion | — | 尚未评估 |
model: Qwen3.6-35B · variant: fresh work order for finished step · protocol: tools | 0.02 proportion | — | 尚未评估 |
model: Qwen3-235B · variant: confirmation wording, same work order · protocol: text | 0.35 proportion | — | 尚未评估 |
model: Qwen3-235B · variant: confirmation wording, same work order · protocol: tools | 0.08 proportion | — | 尚未评估 |
model: Qwen3.6-35B · variant: confirmation wording, same work order · protocol: text | 0.44 proportion | — | 尚未评估 |
model: Qwen3.6-35B · variant: confirmation wording, same work order · protocol: tools | 0 proportion | — | 尚未评估 |
model: Qwen3-235B · variant: status-question wording, same work order · protocol: text | 0.16 proportion | — | 尚未评估 |
model: Qwen3-235B · variant: status-question wording, same work order · protocol: tools | 0 proportion | — | 尚未评估 |
model: Qwen3.6-35B · variant: status-question wording, same work order · protocol: text | 0.38 proportion | — | 尚未评估 |
model: Qwen3.6-35B · variant: status-question wording, same work order · protocol: tools | 0 proportion | — | 尚未评估 |
At least one source-grounded, automatically runnable reconstruction package binds every reported measurement. Exact commands and asset availability remain preflight questions; observed evidence, not the plan, determines support.
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: prompt manifest | 0.87 proportion | — | 尚未评估 |
configuration: guarded trigger with transcript sanitization | 0 proportion | — | 尚未评估 |
configuration: guarded trigger | 0.88 proportion | — | 尚未评估 |
configuration: guarded trigger without transcript sanitization | 0.64 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/5 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 75 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 70 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 15 · density: 0.15 · reasoning: no-think | 82 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 64 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 18 · density: 0.15 · reasoning: no-think | 53 events per 128 episodes | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/11 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript | 0.48 proportion | — | 尚未评估 |
arm: raw transcript | 0.63 proportion | — | 尚未评估 |
arm: text checklist | 0.44 proportion | — | 尚未评估 |
arm: text checklist | 0.61 proportion | — | 尚未评估 |
arm: directive gate | 0.78 proportion | — | 尚未评估 |
arm: directive gate | 0.82 proportion | — | 尚未评估 |
arm: enforcement gate | 0.93 proportion | — | 尚未评估 |
arm: enforcement gate | 0.93 proportion | — | 尚未评估 |
contrast: raw transcript to text checklist | 0.6 p-value | — | 尚未评估 |
contrast: text checklist to directive gate | 0 p-value | — | 尚未评估 |
contrast: directive gate to enforcement gate | 0.0006 p-value | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/104 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
rung: raw transcript · runs: 1 · model: Qwen3-235B | 0.67 F1 | — | 尚未评估 |
sd: 0.03 · rung: raw transcript · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 0.67 F1 | — | 尚未评估 |
sd: 2 · rung: raw transcript · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 56% | — | 尚未评估 |
sd: 2 · rung: raw transcript · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 42% | — | 尚未评估 |
sd: 3 · rung: raw transcript · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 8 events per run | — | 尚未评估 |
sd: 2 · rung: raw transcript · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 4 events per run | — | 尚未评估 |
sd: 2 · rung: raw transcript · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 28 queries per run | — | 尚未评估 |
rung: self-written ledger · runs: 1 · model: Qwen3-235B | 0.73 F1 | — | 尚未评估 |
sd: 0.04 · rung: self-written ledger · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 0.71 F1 | — | 尚未评估 |
sd: 5 · rung: self-written ledger · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 60% | — | 尚未评估 |
sd: 4 · rung: self-written ledger · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 37% | — | 尚未评估 |
sd: 1 · rung: self-written ledger · runs: 4 · week: v9 · model: Qwen3-235B · temperature: 0.7 | 4 events per run | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 200 episode-arms | — | 尚未评估 | |
| 0 episode-arms | — | 尚未评估 | |
graph: generator · policy: memoryless always-book | 1 proportion | — | 尚未评估 |
graph: compiled · policy: memoryless always-book | 0.99 proportion | — | 尚未评估 |
| 90 items | — | 尚未评估 | |
| 88 items | — | 尚未评估 | |
| 0 items | — | 尚未评估 | |
| 4 items | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/44 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.27,0.43] | 0.34 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.36,0.53] | 0.45 proportion | — | 尚未评估 |
arm: directive gate · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.77,0.9] | 0.84 proportion | — | 尚未评估 |
arm: enforcement gate · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.84,0.95] | 0.91 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: thinking · wilson95Ci: [0.69,0.83] | 0.77 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: thinking · wilson95Ci: [0.59,0.75] | 0.68 proportion | — | 尚未评估 |
arm: directive gate · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: thinking · wilson95Ci: [0.87,0.96] | 0.93 proportion | — | 尚未评估 |
arm: enforcement gate · brief: amended · model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: thinking · wilson95Ci: [0.87,0.96] | 0.93 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.27,0.44] | 0.35 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.52,0.69] | 0.61 proportion | — | 尚未评估 |
arm: directive gate · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.81,0.92] | 0.88 proportion | — | 尚未评估 |
arm: enforcement gate · brief: amended · model: Qwen3-235B · steps: 15 · density: 0.15 · reasoning: no-think · wilson95Ci: [0.9,0.98] | 0.95 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/56 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-235B · steps: 15 · issued: 175 · obeyed: 158 · density: 0.3 · directive: CANCELLED · reasoning: no-think · wilson95Ci: [0.85,0.938] | 0.903 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 276 · obeyed: 275 · density: 0.3 · directive: already DONE · reasoning: no-think · wilson95Ci: [0.98,0.999] | 0.996 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 259 · obeyed: 240 · density: 0.3 · directive: BLOCKED · reasoning: no-think · wilson95Ci: [0.888,0.953] | 0.927 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 710 · obeyed: 673 · density: 0.3 · directive: all traps · reasoning: no-think · wilson95Ci: [0.929,0.962] | 0.948 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · numerator: 106 · reasoning: no-think · denominator: 128 | 0.828 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · numerator: 2688 · reasoning: no-think · denominator: 2688 | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 175 · obeyed: 171 · density: 0.3 · directive: CANCELLED · reasoning: thinking · wilson95Ci: [0.943,0.991] | 0.977 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 261 · obeyed: 253 · density: 0.3 · directive: already DONE · reasoning: thinking · wilson95Ci: [0.941,0.984] | 0.969 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 272 · obeyed: 263 · density: 0.3 · directive: BLOCKED · reasoning: thinking · wilson95Ci: [0.938,0.982] | 0.967 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · issued: 708 · obeyed: 687 · density: 0.3 · directive: all traps · reasoning: thinking · wilson95Ci: [0.955,0.981] | 0.97 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/38 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-235B · steps: 5 · density: 0.15 · compiledEpisodes: 128 | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · compiledEpisodes: 128 | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · compiledEpisodes: 128 | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · compiledEpisodes: 128 | 0.891 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · compiledEpisodes: 128 | 0.969 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · compiledEpisodes: 128 | 0.996 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · compiledEpisodes: 128 | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · compiledEpisodes: 128 | 1 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · compiledEpisodes: 128 | 0.867 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · compiledEpisodes: 128 | 0.969 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.15 · compiledEpisodes: 128 | 0.998 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.15 · compiledEpisodes: 128 | 0.999 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/24 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · model: Qwen3-235B · steps: 10 · reasoning: no-think · wilson95Ci: [0.3,0.47] | 0.38 proportion | — | 尚未评估 |
arm: text checklist · model: Qwen3-235B · steps: 10 · reasoning: no-think · wilson95Ci: [0.47,0.64] | 0.55 proportion | — | 尚未评估 |
arm: directive gate · model: Qwen3-235B · steps: 10 · reasoning: no-think · wilson95Ci: [0.76,0.89] | 0.84 proportion | — | 尚未评估 |
arm: enforcement gate · model: Qwen3-235B · steps: 10 · reasoning: no-think · wilson95Ci: [0.94,1] | 0.98 proportion | — | 尚未评估 |
arm: raw transcript · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think · wilson95Ci: [0.41,0.58] | 0.5 proportion | — | 尚未评估 |
arm: text checklist · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think · wilson95Ci: [0.51,0.68] | 0.59 proportion | — | 尚未评估 |
arm: directive gate · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think · wilson95Ci: [0.8,0.92] | 0.87 proportion | — | 尚未评估 |
arm: enforcement gate · model: Qwen3-235B · steps: 15 · density: 0.3 · reasoning: no-think · wilson95Ci: [0.94,1] | 0.98 proportion | — | 尚未评估 |
arm: raw transcript · model: Qwen3-235B · steps: 10 · reasoning: thinking · wilson95Ci: [0.72,0.86] | 0.8 proportion | — | 尚未评估 |
arm: text checklist · model: Qwen3-235B · steps: 10 · reasoning: thinking · wilson95Ci: [0.73,0.86] | 0.8 proportion | — | 尚未评估 |
arm: directive gate · model: Qwen3-235B · steps: 10 · reasoning: thinking · wilson95Ci: [0.81,0.92] | 0.88 proportion | — | 尚未评估 |
arm: enforcement gate · model: Qwen3-235B · steps: 10 · reasoning: thinking · wilson95Ci: [0.9,0.98] | 0.95 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/125 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 4 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 0 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 0 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 4 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 2 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 0 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 2 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 2 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think | 0 events | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think · denominator: 128 | 7 episodes | — | 尚未评估 |
model: Qwen3-235B · steps: 5 · density: 0.15 · reasoning: no-think · denominator: 128 | 15 episodes | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0 events | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/34 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · scoring: strict · phrasing: canonical · wilson95Ci: [0.29,0.53] | 0.41 proportion | — | 尚未评估 |
arm: raw transcript · scoring: strict · phrasing: held-out · wilson95Ci: [0.38,0.62] | 0.5 proportion | — | 尚未评估 |
arm: raw transcript · scoring: strict · phrasing: held-out plus ambiguous · wilson95Ci: [0.21,0.43] | 0.31 proportion | — | 尚未评估 |
arm: raw transcript · scoring: decline-aware · phrasing: canonical · wilson95Ci: [0.46,0.69] | 0.58 proportion | — | 尚未评估 |
arm: raw transcript · scoring: decline-aware · phrasing: held-out · wilson95Ci: [0.52,0.75] | 0.64 proportion | — | 尚未评估 |
arm: raw transcript · scoring: decline-aware · phrasing: held-out plus ambiguous · wilson95Ci: [0.28,0.51] | 0.39 proportion | — | 尚未评估 |
arm: text checklist · scoring: strict · phrasing: canonical · wilson95Ci: [0.5,0.73] | 0.62 proportion | — | 尚未评估 |
arm: text checklist · scoring: strict · phrasing: held-out · wilson95Ci: [0.49,0.72] | 0.61 proportion | — | 尚未评估 |
arm: text checklist · scoring: strict · phrasing: held-out plus ambiguous · wilson95Ci: [0.21,0.43] | 0.31 proportion | — | 尚未评估 |
arm: text checklist · scoring: decline-aware · phrasing: canonical · wilson95Ci: [0.67,0.86] | 0.78 proportion | — | 尚未评估 |
arm: text checklist · scoring: decline-aware · phrasing: held-out · wilson95Ci: [0.61,0.83] | 0.73 proportion | — | 尚未评估 |
arm: text checklist · scoring: decline-aware · phrasing: held-out plus ambiguous · wilson95Ci: [0.31,0.54] | 0.42 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/43 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think | 0.72 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think | 0.84 events | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think | 15 events | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think | 12 events | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think | 27 events | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think | 0.56 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think · denominator: 36 | 17 episodes | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.001 · episodes: 64 · reasoning: no-think · denominator: 28 | 1 episodes | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.003 · episodes: 64 · reasoning: no-think | 0.41 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.003 · episodes: 64 · reasoning: no-think | 1.95 events | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.003 · episodes: 64 · reasoning: no-think | 38 events | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · epsilon: 0.003 · episodes: 64 · reasoning: no-think | 30 events | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/45 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-235B · steps: 10 · density: 0.15 · variant: text checklist append · wilson95Ci: [0.47,0.64] | 0.55 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · variant: text checklist append · wilson95Ci: [0.51,0.68] | 0.59 proportion | — | 尚未评估 |
model: Qwen3.6-35B · steps: 15 · density: 0.3 · variant: text checklist append · wilson95Ci: [0.28,0.45] | 0.36 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · variant: text checklist append | 215 thousand tokens | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · variant: checklist replace mode · wilson95Ci: [0.44,0.61] | 0.52 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · variant: checklist replace mode · comparison: paper arm | 0.704 p-value | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · variant: checklist replace mode · wilson95Ci: [0.38,0.55] | 0.46 proportion | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · variant: checklist replace mode · comparison: paper arm | 0.119 p-value | — | 尚未评估 |
model: Qwen3.6-35B · steps: 15 · density: 0.3 · variant: checklist replace mode · wilson95Ci: [0.2,0.35] | 0.27 proportion | — | 尚未评估 |
model: Qwen3.6-35B · steps: 15 · density: 0.3 · variant: checklist replace mode · comparison: paper arm | 0.126 p-value | — | 尚未评估 |
model: Qwen3-235B · steps: 15 · density: 0.3 · variant: checklist replace mode | 127 thousand tokens | — | 尚未评估 |
model: Qwen3-235B · steps: 10 · density: 0.15 · variant: checklist plus matched step · wilson95Ci: [0.26,0.42] | 0.34 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/76 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.39,0.65] | 0.52 proportion | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.39,0.65] | 49 events | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.39,0.65] | 13 events | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.39,0.65] | 7 events | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.39,0.65] | 5 events | — | 尚未评估 |
arm: raw transcript · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.39,0.65] | 0 events | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.44,0.71] | 0.58 proportion | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.44,0.71] | 0.61 p-value | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.44,0.71] | 51 events | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.44,0.71] | 14 events | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.44,0.71] | 14 events | — | 尚未评估 |
arm: text checklist · brief: amended · model: Qwen3.6-35B · steps: 10 · episodes: 50 · wilson95Ci: [0.44,0.71] | 0 events | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/24 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · episodes: 128 · reasoning: no-think · wilson95Ci: [0.2,0.35] | 0.27 proportion | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 15 · density: 0.3 · episodes: 128 · reasoning: no-think · wilson95Ci: [0.27,0.43] | 0.34 proportion | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · episodes: 128 · reasoning: thinking · wilson95Ci: [0.22,0.37] | 0.29 proportion | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 15 · density: 0.3 · episodes: 128 · reasoning: thinking · wilson95Ci: [0.21,0.37] | 0.28 proportion | — | 尚未评估 |
arm: raw transcript · brief: original · model: DeepSeek V4 Pro · steps: 15 · density: 0.3 · episodes: 32 · reasoning: no-think · wilson95Ci: [0.51,0.82] | 0.69 proportion | — | 尚未评估 |
arm: raw transcript · brief: original · model: DeepSeek V4 Pro · steps: 15 · density: 0.3 · episodes: 32 · reasoning: thinking · wilson95Ci: [0.23,0.55] | 0.38 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · episodes: 128 · reasoning: no-think · wilson95Ci: [0.48,0.65] | 0.56 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 15 · density: 0.3 · episodes: 128 · reasoning: no-think · wilson95Ci: [0.45,0.62] | 0.54 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · episodes: 128 · reasoning: thinking · wilson95Ci: [0.55,0.71] | 0.63 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 15 · density: 0.3 · episodes: 128 · reasoning: thinking · wilson95Ci: [0.61,0.77] | 0.7 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: DeepSeek V4 Pro · steps: 15 · density: 0.3 · episodes: 32 · reasoning: no-think · wilson95Ci: [0.68,0.93] | 0.84 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: DeepSeek V4 Pro · steps: 15 · density: 0.3 · episodes: 32 · reasoning: thinking · wilson95Ci: [0.8,0.98] | 0.94 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/78 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · scoring: primary · reasoning: no-think | 121 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · scoring: primary · reasoning: no-think | 70 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · scoring: decline-aware · reasoning: no-think | 64 events per 128 episodes | — | 尚未评估 |
brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · scoring: primary · reasoning: thinking | 140 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 10 · density: 0.15 · scoring: primary · reasoning: thinking | 17 events per 128 episodes | — | 尚未评估 |
brief: original · model: Qwen3-235B · steps: 15 · density: 0.3 · scoring: primary · reasoning: no-think | 97 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · scoring: primary · reasoning: no-think | 64 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · scoring: decline-aware · reasoning: no-think | 45 events per 128 episodes | — | 尚未评估 |
brief: original · model: Qwen3-235B · steps: 15 · density: 0.3 · scoring: primary · reasoning: thinking | 136 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3-235B · steps: 15 · density: 0.3 · scoring: primary · reasoning: thinking | 29 events per 128 episodes | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · reasoning: no-think | 133 events per 128 episodes | — | 尚未评估 |
brief: amended · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · reasoning: no-think | 120 events per 128 episodes | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/230 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · numerator: 16 · reasoning: no-think · denominator: 128 | 0.12 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · numerator: 45 · reasoning: no-think · denominator: 128 | 0.35 proportion | — | 尚未评估 |
arm: directive gate · brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · numerator: 90 · reasoning: no-think · denominator: 128 | 0.7 proportion | — | 尚未评估 |
arm: enforcement gate · brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · numerator: 127 · reasoning: no-think · denominator: 128 | 0.99 proportion | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · contrast: raw transcript to text checklist · reasoning: no-think | 0 p-value | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · contrast: raw transcript to text checklist · reasoning: no-think | 0.0003 p-value | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · contrast: text checklist to directive gate · reasoning: no-think | 0 p-value | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · contrast: text checklist to directive gate · reasoning: no-think | 0 p-value | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · contrast: directive gate to enforcement gate · reasoning: no-think | 0 p-value | — | 尚未评估 |
brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · contrast: directive gate to enforcement gate · reasoning: no-think | 0 p-value | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · numerator: 2 · reasoning: thinking · denominator: 68 | 0.03 proportion | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3.6-35B · steps: 15 · density: 0.3 · scoring: primary · numerator: 4 · reasoning: thinking · denominator: 19 | 0.21 proportion | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/46 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 103.6 thousand tokens | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 3 thousand tokens | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 106.6 thousand tokens | — | 尚未评估 |
arm: raw transcript · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 1 ratio | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 140.1 thousand tokens | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 2.7 thousand tokens | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 142.8 thousand tokens | — | 尚未评估 |
arm: text checklist · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 1.34 ratio | — | 尚未评估 |
arm: directive gate · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 95.7 thousand tokens | — | 尚未评估 |
arm: directive gate · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 2.2 thousand tokens | — | 尚未评估 |
arm: directive gate · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 97.9 thousand tokens | — | 尚未评估 |
arm: directive gate · brief: original · model: Qwen3-235B · steps: 10 · density: 0.15 · reasoning: no-think | 0.92 ratio | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/87 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
arm: ungated · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 0.39 proportion | — | 尚未评估 |
arm: ungated · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 0.54 proportion | — | 尚未评估 |
arm: ungated · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 9 episodes | — | 尚未评估 |
arm: ungated · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 35 episodes | — | 尚未评估 |
arm: ungated · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 10 episodes | — | 尚未评估 |
arm: ungated · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 26 writes | — | 尚未评估 |
arm: ledger (state shown) · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 0.52 proportion | — | 尚未评估 |
arm: ledger (state shown) · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 0.38 proportion | — | 尚未评估 |
arm: ledger (state shown) · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 14 episodes | — | 尚未评估 |
arm: ledger (state shown) · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 14 episodes | — | 尚未评估 |
arm: ledger (state shown) · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 10 episodes | — | 尚未评估 |
arm: ledger (state shown) · model: Qwen3-235B · domain: tau2-bench airline · episodes: 100 | 4 writes | — | 尚未评估 |
This is source context delimiting what the paper did not test, not an independently reported empirical result. It is preserved as a limitation and informs every experiment objective.
This is source context delimiting what the paper did not test, not an independently reported empirical result. It is preserved as a limitation and informs every experiment objective.
This is source context delimiting what the paper did not test, not an independently reported empirical result. It is preserved as a limitation and informs every experiment objective.
This is source context delimiting what the paper did not test, not an independently reported empirical result. It is preserved as a limitation and informs every experiment objective.