Loading page…
For Qwen3-8B using the largest fitting batch, calibrated TP=2 reached 2.16× on AIME25 and 2.07× on LB[16k-18k], above the table’s TP=1 runs. The timing model achieved 5–11% MAPE when refit at either TP degree; reusing TP=1 coefficients at TP=2 raised MAPE to 35–50%, left AIME speedup benign at 2.21×, but reduced LB speedup to 1.51×. · CiteArk