Loading page…
Across the three models and five workloads, the paper summarizes ASPIRE as delivering 1.70–4.58× AutoRegressive decoding throughput and roughly 27% higher average speedup than the strongest prior self-speculative baselines. · CiteArk