The main throughput evaluation covers Qwen3-1.7B, Qwen3-8B, and DeepSeek-R1-Distill-LLaMA-8B on AIME 2025, CodeElo, LongBench inputs of 16k–18k tokens, and LongBench-v2 inputs of 30k–40k and 80k–100k tokens. Unless otherwise stated it measures decode-only throughput on one H100 NVL 96GB GPU. · CiteArk
Not assessedNo independent reproduction scheduledMethodc01-experimental-protocol
The main throughput evaluation covers Qwen3-1.7B, Qwen3-8B, and DeepSeek-R1-Distill-LLaMA-8B on AIME 2025, CodeElo, LongBench inputs of 16k–18k tokens, and LongBench-v2 inputs of 30k–40k and 80k–100k tokens. Unless otherwise stated it measures decode-only throughput on one H100 NVL 96GB GPU.
Source: paper:PDF p.8, §5.1 Experimental Setup
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.