论文中的结论可在「研究结论」中查看。
0 / 14 条结论已通过验证
其余仍在验证中
The offline episode traces, prompt sharding scripts for Math and Code from Lost in Conversation, and exact hidden state trajectory extraction codebase are not provided in the paper snapshot or any public repository.
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
task: Math · model: Llama-3.1-8B | 0.047 scalar | — | 尚未评估 |
task: Code · model: Qwen3-14B | 0.148 scalar | — | 尚未评估 |
task: Retail · model: Qwen3-14B | 0.03 scalar | — | 尚未评估 |
task: Airline · model: Qwen3-14B | 0.049 scalar | — | 尚未评估 |
task: Retail · model: Qwen3-32B | 0.035 scalar | — | 尚未评估 |
task: Airline · model: Qwen3-32B | 0.088 scalar | — | 尚未评估 |
The paper does not provide the classifier model, training code, loss function, evaluation split, or feature representation.
已评估 0/4 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
action: Read | 0.826 fraction | — | 尚未评估 |
action: Respond | 0.711 fraction | — | 尚未评估 |
action: Write | 0.7 fraction | — | 尚未评估 |
action: Transfer | 0.531 fraction | — | 尚未评估 |
The learned action classifier (predicting next action a_hat_t) is not released, nor are the training data, features, architecture, or tau-Bench integration harness.
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-14B · domain: Retail · trigger_action: Read | 0.393 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · trigger_action: Write | 0.381 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · trigger_action: Respond | 0.29 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · trigger_action: Respond | 0.372 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · trigger_action: Read | 0.352 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · trigger_action: Write | 0.34 fraction | — | 尚未评估 |
Complete sliding window logs across episodes and domain trajectories are not available.
已评估 0/8 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
group: correct · domain: Retail · pattern: Read->Read->Read · chain_id: 1 | 19.6 percentage_points | — | 尚未评估 |
group: incorrect · domain: Retail · pattern: Read->Read->Read · chain_id: 1 | 22.2 percentage_points | — | 尚未评估 |
group: correct · domain: Airline · pattern: Read->Read->Read · chain_id: 1 | 17.1 percentage_points | — | 尚未评估 |
group: incorrect · domain: Airline · pattern: Read->Read->Read · chain_id: 1 | 14.5 percentage_points | — | 尚未评估 |
group: correct · domain: Retail · pattern: Resp->Resp->Resp · chain_id: 2 | 3.6 percentage_points | — | 尚未评估 |
group: incorrect · domain: Retail · pattern: Resp->Resp->Resp · chain_id: 2 | 8.2 percentage_points | — | 尚未评估 |
group: correct · domain: Airline · pattern: Resp->Resp->Resp · chain_id: 2 | 9.1 percentage_points | — | 尚未评估 |
group: incorrect · domain: Airline · pattern: Resp->Resp->Resp · chain_id: 2 | 23.6 percentage_points | — | 尚未评估 |
Raw episode traces, prompt sharding scripts, and hidden state trajectory extraction codebase are unreleased.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
task: Airline · model: Qwen3-14B | 0.3 scalar | — | 尚未评估 |
task: Airline · model: Qwen3-32B | 0.057 scalar | — | 尚未评估 |
The paper does not provide an official repository or executable code release (repository-identity.json indicates available: false). Furthermore, key protocol components including prompt templates, user simulator conversation driver scripts, probe extraction and logging pipelines, the exact classifier training split for Learned-HST, and detailed per-task configurations are deferred to an unreleased Appendix.
已评估 0/5 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
policy: never-think · benchmark: tau-bench | 0.241 fraction | — | 尚未评估 |
policy: geometry-conditioned · benchmark: tau-bench | 0.396 fraction | — | 尚未评估 |
policy: never-think · benchmark: tau-bench | 104.8 thousands_tokens | — | 尚未评估 |
policy: geometry-conditioned · benchmark: tau-bench | 93 thousands_tokens | — | 尚未评估 |
benchmark: tau-bench | 11.2 percentage_points | — | 尚未评估 |
Official execution artifacts, dialogue simulators, user simulator policy prompts, and trajectory monitoring hooks are not released.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-32B · domain: Airline · policy: Never-think | 0.136 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Warm-up | 0.28 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Always-think | 0.364 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Variance-slope | 0.328 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Temporal-curvature · threshold: -0.1 | 0.352 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Learned-HST | 0.364 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Never-think | 98.2 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Warm-up | 92.6 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Always-think | 100.6 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Variance-slope | 86.3 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Temporal-curvature · threshold: -0.1 | 83 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Airline · policy: Learned-HST | 87.5 thousands_tokens | — | 尚未评估 |
Official execution artifacts, dialogue simulators, user simulator policy prompts, and trajectory monitoring hooks are not released.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-14B · domain: Airline · policy: Never-think | 0.196 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Warm-up | 0.348 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Always-think | 0.356 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Variance-slope | 0.392 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Temporal-curvature · threshold: -0.15 | 0.384 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Learned-HST | 0.42 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Never-think | 82.5 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Warm-up | 87.1 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Always-think | 83.5 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Variance-slope | 83.3 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Temporal-curvature · threshold: -0.15 | 78.9 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Airline · policy: Learned-HST | 85 thousands_tokens | — | 尚未评估 |
Official execution artifacts, dialogue simulators, user simulator policy prompts, and trajectory monitoring hooks are not released.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-32B · domain: Retail · policy: Never-think | 0.347 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Warm-up | 0.405 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Always-think | 0.384 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Variance-slope | 0.412 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Temporal-curvature · threshold: -0.15 | 0.442 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Learned-HST | 0.405 fraction | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Never-think | 125.7 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Warm-up | 107.1 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Always-think | 119.3 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Variance-slope | 105.1 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Temporal-curvature · threshold: -0.15 | 104.6 thousands_tokens | — | 尚未评估 |
model: Qwen3-32B · domain: Retail · policy: Learned-HST | 108.5 thousands_tokens | — | 尚未评估 |
Evaluation scripts and environment setup for sweeping tau_kappa on tau-Bench Retail are not available.
已评估 0/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-14B · domain: Retail · threshold: -0.15 | 0.388 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · threshold: -0.2 | 0.397 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · threshold: -0.25 | 0.363 fraction | — | 尚未评估 |
The Code task generation, interaction logs, and hidden-state extraction code are not available.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
task: Code · group: correct · model: Qwen3-14B | 0.127 scalar | — | 尚未评估 |
task: Code · group: correct · model: Qwen3-14B | 0.096 scalar | — | 尚未评估 |
The exact Math problem subset, seeds, sharding prompt splits, and model generation outputs from Lost in Conversation are not included in the paper snapshot.
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
task: Math · group: correct | 259 count | — | 尚未评估 |
task: Math · group: incorrect | 227 count | — | 尚未评估 |
Execution requires official code and prompt/dialogue harnesses that are not published with the paper. Missing items include the tau-Bench agent loop modifications, the thinking mode activation protocol for Qwen3-14B, the layer 22 extraction hook, and the specific Learned-HST classifier model and weights.
已评估 0/18 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
model: Qwen3-14B · domain: Retail · policy: Never-think | 0.286 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Warm-up | 0.361 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Always-think | 0.307 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Variance-slope | 0.393 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Temporal-curvature · threshold: -0.2 | 0.397 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Learned-HST | 0.398 fraction | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Never-think | 112.6 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Warm-up | 111.1 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Always-think | 113.1 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Variance-slope | 104.3 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Temporal-curvature · threshold: -0.2 | 101 thousands_tokens | — | 尚未评估 |
model: Qwen3-14B · domain: Retail · policy: Learned-HST | 104.4 thousands_tokens | — | 尚未评估 |