Source: paper:PDF pp. 7–8, Figure 4, Table 3, and main-comparison paragraph
No immutable Assessment has been published for this Claim yet.
No execution has been linked to this Claim yet.
0/45 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
bpd: 0.5 · side: K-only · model: LLaMA-3.1-8B · objective: MSE · datasetOrTask: mean(Wikitext-2, PTB, C4) | 31.17 perplexity | — | Not assessed |
bpd: 0.5 · side: K-only · model: LLaMA-3.1-8B · objective: KL · datasetOrTask: mean(Wikitext-2, PTB, C4) | 37.31 perplexity | — | Not assessed |
bpd: 0.5 · side: K-only · model: Mistral-7B-v0.3 · objective: MSE · datasetOrTask: mean(Wikitext-2, PTB, C4) | 36.5 perplexity | — | Not assessed |
bpd: 0.5 · side: K-only · model: Mistral-7B-v0.3 · objective: KL · datasetOrTask: mean(Wikitext-2, PTB, C4) | 36.24 perplexity | — | Not assessed |
bpd: 0.5 · side: K-only · model: Qwen2.5-7B-Instruct · objective: MSE · datasetOrTask: mean(Wikitext-2, PTB, C4) | 16.33 perplexity | — | Not assessed |
bpd: 0.5 · side: K-only · model: Qwen2.5-7B-Instruct · objective: KL · datasetOrTask: mean(Wikitext-2, PTB, C4) | 17.17 perplexity | — | Not assessed |
bpd: 0.5 · side: K-only · model: LLaMA-3.1-8B · objective: MSE · datasetOrTask: mean(five tasks) | 48.8% | — | Not assessed |
bpd: 0.5 · side: K-only · model: LLaMA-3.1-8B · objective: KL · datasetOrTask: mean(five tasks) | 48.5% | — | Not assessed |
bpd: 0.5 · side: K-only · model: Mistral-7B-v0.3 · objective: MSE · datasetOrTask: mean(five tasks) | 53.4% | — | Not assessed |
bpd: 0.5 · side: K-only · model: Mistral-7B-v0.3 · objective: KL · datasetOrTask: mean(five tasks) | 52.5% | — | Not assessed |
bpd: 0.5 · side: K-only · model: Qwen2.5-7B-Instruct · objective: MSE · datasetOrTask: mean(five tasks) | 48.7% | — | Not assessed |
bpd: 0.5 · side: K-only · model: Qwen2.5-7B-Instruct · objective: KL · datasetOrTask: mean(five tasks) | 48.5% | — | Not assessed |