浏览 ArkGraph,选择本次要执行的步骤。
ASPIRE is a batched self-speculative decoding system for long-context language-model inference. It lets requests independently draft with sparse attention or verify with full attention inside one target-model forward pass, schedules each request using online acceptance and cost estimates, and refreshes sparse context during drafting through one full-attention layer. Experiments span three open models and five reasoning or long-context workloads. The paper reports 1.70–4.58× decode-throughput speedups over autoregressive decoding, generally stronger gains at longer contexts, small greedy-quality differences, and robustness studies covering oracle scheduling, tensor parallelism, refresh position, page size, estimator constants, and draft length.
In plain language, ASPIRE aims to stop every request in a batch from being forced to speculate for the same amount of time. That matters most when long contexts make attention expensive: requests can verify when useful while their batch-mates keep drafting. The reported gains cover decode passes rather than prefill, depend on deployment-specific timing calibration, and have not yet been integrated with a general continuous-batching engine with chunked prefill.
论文中的结论可在「研究结论」中查看。
历史研究报告
此前保存的报告仍可阅读,新报告生成已暂停。