正在加载页面…
With Qwen3-8B greedy decoding on eight LongBench-v1 tasks, four tasks agreed within 0.05 points, the largest task deviation was 0.50 points, and the task-macro average was 51.92 for AutoRegressive versus 51.72 for ASPIRE (−0.20 points) across 1,460 examples. The authors attribute the small differences to floating-point nondeterminism in batched-kernel tie-breaking. · CiteArk