Loading page…
On tau-Bench Retail with Qwen3-32B, trigger policies yield task rewards of 0.347 (Never-think), 0.405 (Warm-up), 0.384 (Always-think), 0.412 (Variance slope beta_t > tau_beta), 0.442 (Temporal curvature kappa_t < -0.15), and 0.405 (Learned-HST), with mean token costs in thousands of 125.7, 107.1, 119.3, 105.1, 104.6, and 108.5 respectively. · CiteArk