Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/31
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
31
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 5 claims with no independent reproduction scheduled in this plan.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
All per-dataset FIXED-K scores and unweighted row means for five additional architectures across every tested integer budget are reported in Table 6.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Some appendix rule configurations do not land exactly at their target tier: overspending rows are daggered and excluded from Table 3, whereas underspending rows remain visible but a deficit on them is not treated as evidence against the rule.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-30B matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Four-model dynamic-allocation synthesis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
At aggressive target budgets retaining one third to one half of native K, the best eligible dynamic rule improves the eleven-dataset mean on every core model, by 0.90–2.98 points, averaging 2.26 points.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
GPT-OSS-20B matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-30B matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Ling-lite-1.5 matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Four-model dynamic-allocation synthesis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The ordering of dynamic rules changes between conservative and aggressive tiers on the same model, and the spread among rules at a matched tier is wider than the gap between the best rule and FIXED-K.
Information insufficientclaim-rule-ordering-and-spread1 plan0 runs
Scientific conclusionNot assessed
Four-model dynamic-allocation synthesis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On Qwen3-30B-A3B at the conservative tier, damping DIEP’s similarity term raises the suite mean from 54.32 to 61.60, while removing it entirely (NAEE) yields 63.92; conservative DIEP deficits versus FIXED-K do not vary monotonically with native K.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Four-model dynamic-allocation synthesis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Knowledge QA declines gradually, whereas generative reasoning initially holds steady and then drops sharply; mathematical and code reasoning is the least robust family at the aggressive end on eight of nine models.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Nine-model two-thirds retention synthesis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Table 3 reports, per dataset, the unpruned score and the matched-tier FIXED-K and best eligible dynamic-rule scores for all four core models at conservative and aggressive budgets.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
GPT-OSS-20B matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-30B matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Ling-lite-1.5 matched-budget dynamic routing
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Four-model dynamic-allocation synthesis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Table 8 reports every evaluated rule at both conservative and aggressive budget tiers for Qwen3-Next-80B-A3B-Instruct, including per-dataset scores, measured budgets, and suite means.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Table 9 reports every evaluated rule at both conservative and aggressive budget tiers for Ling-lite-1.5, including per-dataset scores, measured budgets, and suite means.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Table 7 reports every evaluated rule at both conservative and aggressive budget tiers for Qwen3-30B-A3B-Instruct, including per-dataset scores, measured budgets, and suite means.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On prompt tokens, Qwen3-VL’s selected routing weights are substantially flatter than the text model’s: its top three hold less of the selected-eight mass, and the selected eight hold less of the full 128-way softmax mass.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Uniformly retaining the paper's stated ceil(2K/3) operating point preserves 98.8% of unpruned performance on average across nine models and three task families; six of 27 model-family combinations match or exceed their unpruned scores.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-VL routing becomes somewhat more concentrated during decoding than during prompting but remains flatter than the text model; bootstrap intervals over layers support stable differences in top-eight mass and entropy.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Table 10 reports every evaluated rule at both conservative and aggressive budget tiers for GPT-OSS-20B, including per-dataset scores, measured budgets, and suite means.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Thinking and Instruct have similar normalized retention from eight to five experts, but Thinking leads on the four hardest generative datasets by 2.7 points at k=4 and 5.1 points at k=3; the reported task breakdown shows little Knowledge QA difference.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
FIXED-K, NAEE, DYNROUTE, DIEP/DIEP-d, and BAN are implemented at the router-selection point; retained gate weights are renormalized, all experts remain resident, and only routed experts may be dropped.
Information insufficientclaim-rule-implementation0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Repeated FIXED-K measurements differ by up to about one point; the paper keeps the older, lower reading in every listed pair so dynamic rules are not credited against a high reference fluctuation.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Uniform truncation remains near-flat at high retained fractions but breaks down sharply at aggressive budgets, and the same absolute k has very different meaning across native router widths.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
GPT-OSS-20B complete uniform expert sweep
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Ling-lite-1.5 complete uniform expert sweep
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Within the multimodal suite, ChartQA falls furthest at every reduced budget; at k=2, the three image-text transcription datasets are at or near zero, while multiple-choice datasets retain roughly one third to one half of their own unpruned scores.
Information insufficientclaim-multimodal-dataset-failure-pattern1 plan0 runs
Scientific conclusionNot assessed
Multimodal versus text pruning sensitivity
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-235B has no normalized-retention advantage over Qwen3-30B from k=7 through k=5 on the four hardest generative datasets, but leads by 2.5 points at k=4 and 14.4 points at k=3.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For both prompt-token concentration measures, the text and multimodal models’ per-dataset values do not overlap, and the text model remains more concentrated in the same direction across all 48 MoE layers.
Information insufficientclaim-router-concentration-consistency1 plan0 runs
Scientific conclusionNot assessed
Text and multimodal router concentration
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The gains from dynamic allocation under aggressive pruning are concentrated in generative tasks, including large recoveries on MATH-500 and LiveCodeBench.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Qwen3-Next-80B complete uniform expert sweep
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
GPT-OSS-20B complete uniform expert sweep
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Ling-lite-1.5 complete uniform expert sweep
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Reported k-bar counts routed experts only; shared experts always execute, so equal routed-expert budgets are not equal total expert evaluations across architectures.
Information insufficientclaim-shared-expert-budget-caveat0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
At conservative budgets, the best eligible dynamic rule differs from FIXED-K by −0.53 to +0.67 percentage points across the four core models, averaging +0.06 points.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
At the tested reduced FIXED-K budgets, prefill and decode throughput improve on both vLLM and HF Transformers, with all per-model ratios reported in Table 2.
Information insufficientclaim-uniform-throughput1 plan0 runs
Scientific conclusionNot assessed
FIXED-K prefill and decode throughput on two backends
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Table 11 reports the full per-dataset FIXED-K sweeps for the Qwen3-30B Instruct reference, its Thinking counterpart, and the Qwen3-235B Instruct scale comparison.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
30B versus 235B pruning resilience
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The twelve reported checkpoints span routed-expert pools of 32–512 and native selections of 4–10 routed experts per token, with model roles and parameter counts reported in Table 1.
Information insufficientclaim-model-architecture-inventory0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The multimodal and text short-answer suites have normalized means within one point from k=8 through k=4, but the multimodal model falls 11 points behind at k=3 and 42 points behind at k=2.