Loading page…
Failed episodes exhibit systematically more negative temporal curvature kappa across all six task-LLM pairings, with correct minus incorrect mean differences of: Math (Llama-3.1-8B) +0.047, Code (Qwen3-14B) +0.148, Retail (Qwen3-14B) +0.030, Airline (Qwen3-14B) +0.049, Retail (Qwen3-32B) +0.035, and Airline (Qwen3-32B) +0.088. · CiteArk