The paper’s claims are available in Research claims.
0 / 14 claims verified
the rest still being verified
The offline episode traces, prompt sharding scripts for Math and Code from Lost in Conversation, and exact hidden state trajectory extraction codebase are not provided in the paper snapshot or any public repository.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
task: Math · model: Llama-3.1-8B | 0.047 scalar | — | Not assessed |
task: Code · model: Qwen3-14B | 0.148 scalar | — | Not assessed |
task: Retail · model: Qwen3-14B | 0.03 scalar | — | Not assessed |
task: Airline · model: Qwen3-14B | 0.049 scalar | — | Not assessed |
task: Retail · model: Qwen3-32B | 0.035 scalar | — | Not assessed |
task: Airline · model: Qwen3-32B | 0.088 scalar | — | Not assessed |
The paper does not provide the classifier model, training code, loss function, evaluation split, or feature representation.
0/4 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
action: Read | 0.826 fraction | — | Not assessed |
action: Respond | 0.711 fraction | — | Not assessed |
action: Write | 0.7 fraction | — | Not assessed |
action: Transfer | 0.531 fraction | — | Not assessed |
The learned action classifier (predicting next action a_hat_t) is not released, nor are the training data, features, architecture, or tau-Bench integration harness.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Qwen3-14B · domain: Retail · trigger_action: Read | 0.393 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · trigger_action: Write | 0.381 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · trigger_action: Respond | 0.29 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · trigger_action: Respond | 0.372 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · trigger_action: Read | 0.352 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · trigger_action: Write | 0.34 fraction | — | Not assessed |
Complete sliding window logs across episodes and domain trajectories are not available.
0/8 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
group: correct · domain: Retail · pattern: Read->Read->Read · chain_id: 1 | 19.6 percentage_points | — | Not assessed |
group: incorrect · domain: Retail · pattern: Read->Read->Read · chain_id: 1 | 22.2 percentage_points | — | Not assessed |
group: correct · domain: Airline · pattern: Read->Read->Read · chain_id: 1 | 17.1 percentage_points | — | Not assessed |
group: incorrect · domain: Airline · pattern: Read->Read->Read · chain_id: 1 | 14.5 percentage_points | — | Not assessed |
group: correct · domain: Retail · pattern: Resp->Resp->Resp · chain_id: 2 | 3.6 percentage_points | — | Not assessed |
group: incorrect · domain: Retail · pattern: Resp->Resp->Resp · chain_id: 2 | 8.2 percentage_points | — | Not assessed |
group: correct · domain: Airline · pattern: Resp->Resp->Resp · chain_id: 2 | 9.1 percentage_points | — | Not assessed |
group: incorrect · domain: Airline · pattern: Resp->Resp->Resp · chain_id: 2 | 23.6 percentage_points | — | Not assessed |
Raw episode traces, prompt sharding scripts, and hidden state trajectory extraction codebase are unreleased.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
task: Airline · model: Qwen3-14B | 0.3 scalar | — | Not assessed |
task: Airline · model: Qwen3-32B | 0.057 scalar | — | Not assessed |
The paper does not provide an official repository or executable code release (repository-identity.json indicates available: false). Furthermore, key protocol components including prompt templates, user simulator conversation driver scripts, probe extraction and logging pipelines, the exact classifier training split for Learned-HST, and detailed per-task configurations are deferred to an unreleased Appendix.
0/5 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
policy: never-think · benchmark: tau-bench | 0.241 fraction | — | Not assessed |
policy: geometry-conditioned · benchmark: tau-bench | 0.396 fraction | — | Not assessed |
policy: never-think · benchmark: tau-bench | 104.8 thousands_tokens | — | Not assessed |
policy: geometry-conditioned · benchmark: tau-bench | 93 thousands_tokens | — | Not assessed |
benchmark: tau-bench | 11.2 percentage_points | — | Not assessed |
Official execution artifacts, dialogue simulators, user simulator policy prompts, and trajectory monitoring hooks are not released.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Qwen3-32B · domain: Airline · policy: Never-think | 0.136 fraction | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Warm-up | 0.28 fraction | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Always-think | 0.364 fraction | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Variance-slope | 0.328 fraction | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Temporal-curvature · threshold: -0.1 | 0.352 fraction | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Learned-HST | 0.364 fraction | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Never-think | 98.2 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Warm-up | 92.6 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Always-think | 100.6 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Variance-slope | 86.3 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Temporal-curvature · threshold: -0.1 | 83 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Airline · policy: Learned-HST | 87.5 thousands_tokens | — | Not assessed |
Official execution artifacts, dialogue simulators, user simulator policy prompts, and trajectory monitoring hooks are not released.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Qwen3-14B · domain: Airline · policy: Never-think | 0.196 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Warm-up | 0.348 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Always-think | 0.356 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Variance-slope | 0.392 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Temporal-curvature · threshold: -0.15 | 0.384 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Learned-HST | 0.42 fraction | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Never-think | 82.5 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Warm-up | 87.1 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Always-think | 83.5 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Variance-slope | 83.3 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Temporal-curvature · threshold: -0.15 | 78.9 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Airline · policy: Learned-HST | 85 thousands_tokens | — | Not assessed |
Official execution artifacts, dialogue simulators, user simulator policy prompts, and trajectory monitoring hooks are not released.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Qwen3-32B · domain: Retail · policy: Never-think | 0.347 fraction | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Warm-up | 0.405 fraction | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Always-think | 0.384 fraction | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Variance-slope | 0.412 fraction | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Temporal-curvature · threshold: -0.15 | 0.442 fraction | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Learned-HST | 0.405 fraction | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Never-think | 125.7 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Warm-up | 107.1 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Always-think | 119.3 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Variance-slope | 105.1 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Temporal-curvature · threshold: -0.15 | 104.6 thousands_tokens | — | Not assessed |
model: Qwen3-32B · domain: Retail · policy: Learned-HST | 108.5 thousands_tokens | — | Not assessed |
Evaluation scripts and environment setup for sweeping tau_kappa on tau-Bench Retail are not available.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Qwen3-14B · domain: Retail · threshold: -0.15 | 0.388 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · threshold: -0.2 | 0.397 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · threshold: -0.25 | 0.363 fraction | — | Not assessed |
The Code task generation, interaction logs, and hidden-state extraction code are not available.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
task: Code · group: correct · model: Qwen3-14B | 0.127 scalar | — | Not assessed |
task: Code · group: correct · model: Qwen3-14B | 0.096 scalar | — | Not assessed |
The exact Math problem subset, seeds, sharding prompt splits, and model generation outputs from Lost in Conversation are not included in the paper snapshot.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
task: Math · group: correct | 259 count | — | Not assessed |
task: Math · group: incorrect | 227 count | — | Not assessed |
Execution requires official code and prompt/dialogue harnesses that are not published with the paper. Missing items include the tau-Bench agent loop modifications, the thinking mode activation protocol for Qwen3-14B, the layer 22 extraction hook, and the specific Learned-HST classifier model and weights.
0/18 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
model: Qwen3-14B · domain: Retail · policy: Never-think | 0.286 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Warm-up | 0.361 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Always-think | 0.307 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Variance-slope | 0.393 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Temporal-curvature · threshold: -0.2 | 0.397 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Learned-HST | 0.398 fraction | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Never-think | 112.6 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Warm-up | 111.1 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Always-think | 113.1 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Variance-slope | 104.3 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Temporal-curvature · threshold: -0.2 | 101 thousands_tokens | — | Not assessed |
model: Qwen3-14B · domain: Retail · policy: Learned-HST | 104.4 thousands_tokens | — | Not assessed |