Loading page…
On tau-Bench Retail with Qwen3-14B, trigger policies yield task rewards of 0.286 (Never-think), 0.361 (Warm-up), 0.307 (Always-think), 0.393 (Variance slope beta_t > tau_beta), 0.397 (Temporal curvature kappa_t < -0.20), and 0.398 (Learned-HST), with mean token costs in thousands of 112.6, 111.1, 113.1, 104.3, 101.0, and 104.4 respectively. · CiteArk