Loading page…
With a fixed 512-token sparse budget on Qwen3-8B/LB[16k-18k], sweeping page size from 1 to 64 produced 1,334–1,451 output tok/s and 81.2–88.3% acceptance. The paper reports throughput varied by under 9% over the 64× granularity range and page size 16 was near the peak, not at the maximum. · CiteArk