正在加载页面…
On tau-Bench Airline with Qwen3-32B, trigger policies yield task rewards of 0.136 (Never-think), 0.280 (Warm-up), 0.364 (Always-think), 0.328 (Variance slope beta_t > tau_beta), 0.352 (Temporal curvature kappa_t < -0.10), and 0.364 (Learned-HST), with mean token costs in thousands of 98.2, 92.6, 100.6, 86.3, 83.0, and 87.5 respectively. · CiteArk