Loading page…
On tau-Bench Airline with Qwen3-14B, trigger policies yield task rewards of 0.196 (Never-think), 0.348 (Warm-up), 0.356 (Always-think), 0.392 (Variance slope beta_t > tau_beta), 0.384 (Temporal curvature kappa_t < -0.15), and 0.420 (Learned-HST), with mean token costs in thousands of 82.5, 87.1, 83.5, 83.3, 78.9, and 85.0 respectively. · CiteArk