Loading page…
The main throughput evaluation covers Qwen3-1.7B, Qwen3-8B, and DeepSeek-R1-Distill-LLaMA-8B on AIME 2025, CodeElo, LongBench inputs of 16k–18k tokens, and LongBench-v2 inputs of 30k–40k and 80k–100k tokens. Unless otherwise stated it measures decode-only throughput on one H100 NVL 96GB GPU. · CiteArk